meta-models/Muse-Glimmer-30B-GGUF

Hardware
1× NVIDIA GeForce RTX 5090
Draft model
meta-models/Muse-Glimmer-30B-GGUF
Workload
8192 → 256
Runs
8 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)016920.900run_KBsE3nStXj9Cpvgv: 0.38 req/s, p95 TTFT 16,459 msrun_TxS4jD4cIyfnuhNv: 0.38 req/s, p95 TTFT 16,233.6 msrun_SUZftAPKWzT_kuLE: 0.38 req/s, p95 TTFT 16,920.9 msrun_jvcJi9ATKjORG-x0: 0.36 req/s, p95 TTFT 2,386.4 msrun_p8srQ_PehQYA8j0L: 0.38 req/s, p95 TTFT 15,979.9 msrun_g5t5iqLfYcbR99tN: 0.38 req/s, p95 TTFT 16,524.5 msrun_S7EiD3875wmwwwTZ: 0.38 req/s, p95 TTFT 16,822 msrun_ssxkhaM7u-YmHDj-: 0.36 req/s, p95 TTFT 2,400.8 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
llama.cpp b10689Speculative · Q4_K_Mspeculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384c=898.516,459 ms71.35 ms25,534.6 ms0run_KBsE3nStXj9Cpvgv
llama.cpp b10689Speculative · Q4_K_Mspeculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384c=897.816,233.6 ms72.63 ms25,530 ms0run_TxS4jD4cIyfnuhNv
llama.cpp b10689Speculative · Q4_K_Mspeculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384c=897.116,920.9 ms72.59 ms30,143.8 ms0run_SUZftAPKWzT_kuLE
llama.cpp b10689Speculative · Q4_K_Mspeculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384c=192.62,386.4 ms1.56 ms2,777.9 ms0run_jvcJi9ATKjORG-x0
llama.cpp b10689Speculative · Q4_K_Mspeculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384c=897.515,979.9 ms72.56 ms25,297.6 ms0run_p8srQ_PehQYA8j0L
llama.cpp b10689Speculative · Q4_K_Mspeculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384c=897.816,524.5 ms71.92 ms28,086.7 ms0run_g5t5iqLfYcbR99tN
llama.cpp b10689Speculative · Q4_K_Mspeculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384c=897.416,822 ms72.85 ms28,107.2 ms0run_S7EiD3875wmwwwTZ
llama.cpp b10689Speculative · Q4_K_Mspeculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384c=192.12,400.8 ms1.54 ms2,791.2 ms0run_ssxkhaM7u-YmHDj-
Runtime
llama.cpp b10689Speculative · Q4_K_M draft
Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384
Test load
c=8
Req/s
Output tok/s
98.5
p95 TTFT / TPOT
16,459 ms / 71.35 ms
Runtime
llama.cpp b10689Speculative · Q4_K_M draft
Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384
Test load
c=8
Req/s
Output tok/s
97.8
p95 TTFT / TPOT
16,233.6 ms / 72.63 ms
Runtime
llama.cpp b10689Speculative · Q4_K_M draft
Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384
Test load
c=8
Req/s
Output tok/s
97.1
p95 TTFT / TPOT
16,920.9 ms / 72.59 ms
Runtime
llama.cpp b10689Speculative · Q4_K_M draft
Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384
Test load
c=1
Req/s
Output tok/s
92.6
p95 TTFT / TPOT
2,386.4 ms / 1.56 ms
Runtime
llama.cpp b10689Speculative · Q4_K_M draft
Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384
Test load
c=8
Req/s
Output tok/s
97.5
p95 TTFT / TPOT
15,979.9 ms / 72.56 ms
Runtime
llama.cpp b10689Speculative · Q4_K_M draft
Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384
Test load
c=8
Req/s
Output tok/s
97.8
p95 TTFT / TPOT
16,524.5 ms / 71.92 ms
Runtime
llama.cpp b10689Speculative · Q4_K_M draft
Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384
Test load
c=8
Req/s
Output tok/s
97.4
p95 TTFT / TPOT
16,822 ms / 72.85 ms
Runtime
llama.cpp b10689Speculative · Q4_K_M draft
Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384
Test load
c=1
Req/s
Output tok/s
92.1
p95 TTFT / TPOT
2,400.8 ms / 1.54 ms