meta-models/Muse-Glimmer-30B-GGUF

Hardware
1× NVIDIA GeForce RTX 5090
Draft model
meta-models/Muse-Glimmer-30B-GGUF
Workload
1024 → 256
Runs
8 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)02408.902run_3ipFIwLFOwrmuTaV: 2.1 req/s, p95 TTFT 2,376.9 msrun_0Ngv9Ngq5iFjJAtZ: 1.46 req/s, p95 TTFT 312.7 msrun_OZ046C2bQECSXE-n: 1.46 req/s, p95 TTFT 308.9 msrun_-pmSHKJRhRIryo1f: 1.47 req/s, p95 TTFT 309.6 msrun_6w4n09STcg0arWg9: 2.1 req/s, p95 TTFT 2,408.9 msrun_-nj5RzOKCXMJ2Viu: 1.46 req/s, p95 TTFT 309.3 msrun_SlvPYaK_-LTNjqxw: 1.46 req/s, p95 TTFT 309.6 msrun_7ZuCv8PHg6xQts0P: 1.46 req/s, p95 TTFT 312.8 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
llama.cpp b10689Speculative · Q4_K_Mspeculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384c=8537.92,376.9 ms14.07 ms4,301 ms0run_3ipFIwLFOwrmuTaV
llama.cpp b10689Speculative · Q4_K_Mspeculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384c=1374.4312.7 ms1.56 ms705.8 ms0run_0Ngv9Ngq5iFjJAtZ
llama.cpp b10689Speculative · Q4_K_Mspeculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384c=1374308.9 ms1.55 ms701.6 ms0run_OZ046C2bQECSXE-n
llama.cpp b10689Speculative · Q4_K_Mspeculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384c=1375.8309.6 ms1.54 ms699.3 ms0run_-pmSHKJRhRIryo1f
llama.cpp b10689Speculative · Q4_K_Mspeculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384c=8537.82,408.9 ms13.88 ms4,354 ms0run_6w4n09STcg0arWg9
llama.cpp b10689Speculative · Q4_K_Mspeculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384c=1372.7309.3 ms1.55 ms706.8 ms0run_-nj5RzOKCXMJ2Viu
llama.cpp b10689Speculative · Q4_K_Mspeculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384c=1372.7309.6 ms1.55 ms703.3 ms0run_SlvPYaK_-LTNjqxw
llama.cpp b10689Speculative · Q4_K_Mspeculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384c=1372.9312.8 ms1.54 ms703 ms0run_7ZuCv8PHg6xQts0P
Runtime
llama.cpp b10689Speculative · Q4_K_M draft
Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384
Test load
c=8
Req/s
Output tok/s
537.9
p95 TTFT / TPOT
2,376.9 ms / 14.07 ms
Runtime
llama.cpp b10689Speculative · Q4_K_M draft
Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384
Test load
c=1
Req/s
Output tok/s
374.4
p95 TTFT / TPOT
312.7 ms / 1.56 ms
Runtime
llama.cpp b10689Speculative · Q4_K_M draft
Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384
Test load
c=1
Req/s
Output tok/s
374
p95 TTFT / TPOT
308.9 ms / 1.55 ms
Runtime
llama.cpp b10689Speculative · Q4_K_M draft
Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384
Test load
c=1
Req/s
Output tok/s
375.8
p95 TTFT / TPOT
309.6 ms / 1.54 ms
Runtime
llama.cpp b10689Speculative · Q4_K_M draft
Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384
Test load
c=8
Req/s
Output tok/s
537.8
p95 TTFT / TPOT
2,408.9 ms / 13.88 ms
Runtime
llama.cpp b10689Speculative · Q4_K_M draft
Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384
Test load
c=1
Req/s
Output tok/s
372.7
p95 TTFT / TPOT
309.3 ms / 1.55 ms
Runtime
llama.cpp b10689Speculative · Q4_K_M draft
Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384
Test load
c=1
Req/s
Output tok/s
372.7
p95 TTFT / TPOT
309.6 ms / 1.55 ms
Runtime
llama.cpp b10689Speculative · Q4_K_M draft
Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384
Test load
c=1
Req/s
Output tok/s
372.9
p95 TTFT / TPOT
312.8 ms / 1.54 ms