- Hardware
- 1× NVIDIA GeForce RTX 5090
- Draft model
- meta-models/Muse-Glimmer-30B-GGUF
- Workload
- 8192 → 256
- Runs
- 8 eligible
Efficient frontierOther eligible run
Runs
All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.
| Runtime | Material parameters | Test load | Req/s | Output tok/s | p95 TTFT | p95 TPOT | p95 E2E | Failures | Run |
|---|---|---|---|---|---|---|---|---|---|
llama.cpp b10689Speculative · Q4_K_M | speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384 | c=8 | 98.5 | 16,459 ms | 71.35 ms | 25,534.6 ms | 0 | run_KBsE3nStXj9Cpvgv | |
llama.cpp b10689Speculative · Q4_K_M | speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384 | c=8 | 97.8 | 16,233.6 ms | 72.63 ms | 25,530 ms | 0 | run_TxS4jD4cIyfnuhNv | |
llama.cpp b10689Speculative · Q4_K_M | speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384 | c=8 | 97.1 | 16,920.9 ms | 72.59 ms | 30,143.8 ms | 0 | run_SUZftAPKWzT_kuLE | |
llama.cpp b10689Speculative · Q4_K_M | speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384 | c=1 | 92.6 | 2,386.4 ms | 1.56 ms | 2,777.9 ms | 0 | run_jvcJi9ATKjORG-x0 | |
llama.cpp b10689Speculative · Q4_K_M | speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384 | c=8 | 97.5 | 15,979.9 ms | 72.56 ms | 25,297.6 ms | 0 | run_p8srQ_PehQYA8j0L | |
llama.cpp b10689Speculative · Q4_K_M | speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384 | c=8 | 97.8 | 16,524.5 ms | 71.92 ms | 28,086.7 ms | 0 | run_g5t5iqLfYcbR99tN | |
llama.cpp b10689Speculative · Q4_K_M | speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384 | c=8 | 97.4 | 16,822 ms | 72.85 ms | 28,107.2 ms | 0 | run_S7EiD3875wmwwwTZ | |
llama.cpp b10689Speculative · Q4_K_M | speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384 | c=1 | 92.1 | 2,400.8 ms | 1.54 ms | 2,791.2 ms | 0 | run_ssxkhaM7u-YmHDj- |
- Runtime
- llama.cpp b10689Speculative · Q4_K_M draft
- Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384- Test load
- c=8
- Req/s
- Output tok/s
- 98.5
- p95 TTFT / TPOT
- 16,459 ms / 71.35 ms
- Runtime
- llama.cpp b10689Speculative · Q4_K_M draft
- Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384- Test load
- c=8
- Req/s
- Output tok/s
- 97.8
- p95 TTFT / TPOT
- 16,233.6 ms / 72.63 ms
- Runtime
- llama.cpp b10689Speculative · Q4_K_M draft
- Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384- Test load
- c=8
- Req/s
- Output tok/s
- 97.1
- p95 TTFT / TPOT
- 16,920.9 ms / 72.59 ms
- Runtime
- llama.cpp b10689Speculative · Q4_K_M draft
- Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384- Test load
- c=1
- Req/s
- Output tok/s
- 92.6
- p95 TTFT / TPOT
- 2,386.4 ms / 1.56 ms
- Runtime
- llama.cpp b10689Speculative · Q4_K_M draft
- Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384- Test load
- c=8
- Req/s
- Output tok/s
- 97.5
- p95 TTFT / TPOT
- 15,979.9 ms / 72.56 ms
- Runtime
- llama.cpp b10689Speculative · Q4_K_M draft
- Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384- Test load
- c=8
- Req/s
- Output tok/s
- 97.8
- p95 TTFT / TPOT
- 16,524.5 ms / 71.92 ms
- Runtime
- llama.cpp b10689Speculative · Q4_K_M draft
- Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384- Test load
- c=8
- Req/s
- Output tok/s
- 97.4
- p95 TTFT / TPOT
- 16,822 ms / 72.85 ms
- Runtime
- llama.cpp b10689Speculative · Q4_K_M draft
- Material parameters
speculative.type=draft-dflash speculative.n_max=15 ubatch_size=512 context_size=147456 max_num_seqs=16 flash_attention=on weight_artifact=Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf max_model_len_per_slot=9216 max_num_batched_tokens=16384- Test load
- c=1
- Req/s
- Output tok/s
- 92.1
- p95 TTFT / TPOT
- 2,400.8 ms / 1.54 ms