English Standard Benchmark Release 2026-09-18

DGX Spark GB10 — Local LLM Benchmark

This is the English projection of DevSnack’s current Standard Benchmark release: 32 GGUF model variants run with llama.cpp on an NVIDIA DGX Spark GB10 across 8 suites covering speed, serving, knowledge, coding, tool use, and agent task completion.

Representative results are shown below. MTP and non-MTP rows use different execution paths, so each result must be read together with its variant, quantization, and serving condition.

  • Qwen3.6 35B-A3B TURBO: TG 77.2 t/s, single-slot c1 87.0 t/s.
  • N2.5 Mini Q4_K_M: TG 81.9 t/s under the same public release conditions.
  • North Mini UD-Q4: aggregate c8 throughput 299.8 t/s.

Model families in this release

The model-family pages are currently maintained on the Korean source route. The measurements, JSON release, and protocol are shared across both language surfaces.

Model comparison matrix

Sort
Model / VariantPerformanceServerKnowledgeCodingTool-callExternal tool-evalAgent-singleAgent-multi
Gemma4
NVFP4 · NVFP4
non-MTP
PP 2609.5 t/sTG 37.4 t/s
c1 36.1 t/sc8 175.4 t/s
98%98/100
33.3%4/12
73.3%11/15
100%12/12
70%7/10
Gemma4
Q4_0 · Q4_0
non-MTP
PP 2516.3 t/sTG 51.7 t/s
c1 49.7 t/sc8 193.0 t/s
98%98/100
25%3/12
80%12/15
100%12/12
70%7/10
Laguna S 2.1
APEX-I Balanced · NVFP4
non-MTP
PP 702.2 t/sTG 27.6 t/s
c1 26.8 t/sc8 72.9 t/s
88%88/100
91.7%11/12
86.7%13/15
86 / 10069/69 scored
50%6/12
70%7/10
Laguna XS 2.1
Q4_K_M · Q4_K_M
non-MTP
PP 1915.7 t/sTG 84.2 t/s
c1 79.2 t/sc8 283.0 t/s
85%85/100
100%12/12
86.7%13/15
86 / 10069/69 scored
50%6/12
90%9/10
Laguna XS 2.1
Q5_K_M · Q5_K_M
non-MTP
PP 1838.1 t/sTG 76.8 t/s
c1 72.3 t/sc8 283.9 t/s
88%88/100
100%12/12
86.7%13/15
88 / 10069/69 scored
50%6/12
90%9/10
Laguna XS 2.1
Q6_K_L · Q6_K_L
non-MTP
PP 1675.3 t/sTG 68.2 t/s
c1 64.6 t/sc8 203.3 t/s
88%88/100
100%12/12
86.7%13/15
88 / 10069/69 scored
41.7%5/12
90%9/10
Laguna XS 2.1
Q8_0 · Q8_0
non-MTP
PP 1705.0 t/sTG 65.1 t/s
c1 62.1 t/sc8 177.6 t/s
84%84/100
100%12/12
86.7%13/15
87 / 10069/69 scored
41.7%5/12
90%9/10
Ling 3.0 Flash
Heretic MXFP4 · MXFP4_MOE
MTP
PP 1022.0 t/sTG 28.8 t/s
c1 41.0 t/sc8 101.7 t/s
99%99/100
100%12/12
80%12/15
91.7%11/12
80%8/10
N2 Mini
Q5_K_M · Q5_K_M
non-MTP
PP 1732.2 t/sTG 64.1 t/s
c1 60.5 t/sc8 171.0 t/s
92%92/100
75%9/12
53.3%8/15
82 / 10065/69 scored
50%6/12
60%6/10
N2 Mini
Q6_K · Q6_K
non-MTP
PP 1565.1 t/sTG 58.7 t/s
c1 55.4 t/sc8 157.4 t/s
93%93/100
66.7%8/12
53.3%8/15
82 / 10065/69 scored
66.7%8/12
90%9/10
N2 Mini
UD-Q4 · UD-Q4
non-MTP
PP 1808.5 t/sTG 63.8 t/s
c1 60.0 t/sc8 180.4 t/s
91%91/100
83.3%10/12
60%9/15
85 / 10065/69 scored
58.3%7/12
50%5/10
N2 Mini
UD-Q5XL · UD-Q5XL
non-MTP
PP 1691.7 t/sTG 59.4 t/s
c1 56.5 t/sc8 162.9 t/s
93%93/100
75%9/12
60%9/15
85 / 10065/69 scored
75%9/12
70%7/10
N2.5 Mini
Q4_K_M · Q4_K_M
non-MTP
PP 1862.0 t/sTG 81.9 t/s
c1 76.2 t/sc8 164.9 t/s
93%93/100
91.7%11/12
66.7%10/15
88 / 10065/69 scored
66.7%8/12
90%9/10
N2.5 Mini
Q5_K_M · Q5_K_M
non-MTP
PP 1794.0 t/sTG 74.4 t/s
c1 70.5 t/sc8 165.5 t/s
94%94/100
91.7%11/12
80%12/15
90 / 10065/69 scored
58.3%7/12
100%10/10
N2.5 Mini
Q6_K · Q6_K
non-MTP
PP 1587.2 t/sTG 67.5 t/s
c1 64.0 t/sc8 142.4 t/s
95%95/100
91.7%11/12
80%12/15
91 / 10065/69 scored
66.7%8/12
100%10/10
N2.5 Mini
Q8_0 · Q8_0
non-MTP
PP 1619.9 t/sTG 44.2 t/s
c1 56.0 t/sc8 126.6 t/s
94%94/100
91.7%11/12
80%12/15
90 / 10065/69 scored
66.7%8/12
90%9/10
North Mini
MXFP4 · MXFP4
non-MTP
PP 2454.9 t/sTG 68.8 t/s
c1 69.2 t/sc8 269.9 t/s
88%88/100
91.7%11/12
80%12/15
83.3%10/12
70%7/10
North Mini
UD-Q4 · UD-Q4
non-MTP
PP 2299.1 t/sTG 71.7 t/s
c1 68.5 t/sc8 299.8 t/s
90%90/100
100%12/12
60%9/15
91.7%11/12
70%7/10
North Mini
UD-Q5 · UD-Q5
non-MTP
PP 2127.3 t/sTG 66.0 t/s
c1 63.4 t/sc8 272.3 t/s
93%93/100
91.7%11/12
80%12/15
91.7%11/12
70%7/10
North Mini
UD-Q6 · UD-Q6
non-MTP
PP 1969.3 t/sTG 60.8 t/s
c1 58.3 t/sc8 250.7 t/s
92%92/100
91.7%11/12
73.3%11/15
91.7%11/12
80%8/10
Occamy 1.0
Q4_K_M · Q4_K_M
non-MTP
PP 1953.5 t/sTG 80.6 t/s
c1 76.3 t/sc8 202.6 t/s
94%94/100
100%12/12
73.3%11/15
87 / 10065/69 scored
66.7%8/12
60%6/10
Occamy 1.0
Q5_K_M · Q5_K_M
non-MTP
PP 1854.7 t/sTG 73.7 t/s
c1 68.8 t/sc8 192.5 t/s
95%95/100
91.7%11/12
73.3%11/15
85 / 10065/69 scored
75%9/12
90%9/10
Occamy 1.0
Q6_K · Q6_K
non-MTP
PP 1659.0 t/sTG 66.8 t/s
c1 63.6 t/sc8 161.2 t/s
96%96/100
100%12/12
73.3%11/15
86 / 10065/69 scored
66.7%8/12
100%10/10
Occamy 1.0
Q8_0 · Q8_0
non-MTP
PP 1734.0 t/sTG 58.8 t/s
c1 56.6 t/sc8 152.1 t/s
96%96/100
91.7%11/12
73.3%11/15
86 / 10065/69 scored
75%9/12
90%9/10
Ornith 1.5
Q5_K_M · Q5_K_M
MTP
PP 1634.9 t/sTG 64.5 t/s
c1 68.3 t/sc8 164.1 t/s
92%92/100
75%9/12
73.3%11/15
66.7%8/12
80%8/10
Ornith 1.5
Q6_K · Q6_K
MTP
PP 1432.1 t/sTG 57.9 t/s
c1 62.4 t/sc8 154.5 t/s
94%94/100
75%9/12
73.3%11/15
83.3%10/12
80%8/10
Ornith 1.5
Q8_0 · Q8_0
MTP
PP 1726.7 t/sTG 57.8 t/s
c1 60.3 t/sc8 158.1 t/s
93%93/100
83.3%10/12
73.3%11/15
75%9/12
90%9/10
Qwen3.6 35B-A3B
APEX · APEX
non-MTP
PP 1829.5 t/sTG 68.5 t/s
c1 64.4 t/sc8 195.3 t/s
95%95/100
83.3%10/12
73.3%11/15
75%9/12
90%9/10
Qwen3.6 35B-A3B
HQ · NVFP4 MTP HQ
MTP
PP 2091.6 t/sTG 68.7 t/s
c1 87.1 t/sc8 201.8 t/s
94%94/100
83.3%10/12
73.3%11/15
91.7%11/12
80%8/10
Qwen3.6 35B-A3B
Q8 · Q8
non-MTP
PP 1779.2 t/sTG 58.7 t/s
c1 55.1 t/sc8 182.6 t/s
95%95/100
91.7%11/12
66.7%10/15
75%9/12
90%9/10
Qwen3.6 35B-A3B
TURBO · TURBO
MTP
PP 2241.2 t/sTG 77.2 t/s
c1 87.0 t/sc8 217.3 t/s
94%94/100
91.7%11/12
73.3%11/15
58.3%7/12
80%8/10
Qwen3.8 Flash Next
UD-IQ4_XS · UD-IQ4_XS
non-MTP
PP 280.1 t/sTG 25.8 t/s
c1 26.8 t/sc8 93.7 t/s
98%98/100
83.3%10/12
66.7%10/15
33.3%4/12
40%4/10

What the eight suites measure

Performance

Prompt processing and token generation speed from llama-bench-style measurements.

Server-performance

Aggregate and per-request throughput across concurrent llama-server slots.

Knowledge

Deterministic knowledge, Korea, math, science, and logic questions.

Coding

Executable Python generation evaluated by tests, not explanation quality alone.

Tool-call

Tool selection, arguments, recovery, and final task completion in a fixed simulator.

External tool-eval-bench

A separate 69-scenario deterministic tool-use protocol. Each variant exposes its actual scored/attempted denominator; current N2/N2.5 Mini and Occamy 1.0 runs score 65 after four grammar transport failures, while Laguna S 2.1 and Laguna XS 2.1 runs score all 69.

Agent-single / Agent-multi

Single-agent completion and role handoff under the release’s fixed protocols.

Limitations

  • Results are observations from NVIDIA DGX Spark GB10, llama.cpp, and the versioned public recipes; other hardware, runtimes, or prompt formats may differ.
  • Quantization and MTP mode can change both speed and evaluator outcomes, so there is no universal single “best model” score.
  • Knowledge, coding, tool-call, external tool-eval-bench, and agent suites use bounded protocols. They are not a complete measure of general intelligence or every real-world coding environment.
  • External tool-eval-bench is currently available for eight N2/N2.5 Mini variants, one Laguna S 2.1 variant, and four Laguna XS 2.1 variants; each row shows its actual scored/attempted denominator, and blank cells mean not measured, not zero.
  • Laguna S 2.1’s previous legacy Knowledge v1 result uses 25 questions and remains a historical custom record; the current S result uses Standard Knowledge v1.2 with 100 questions.
  • The release is an immutable public projection. New measurements should update the matrix through a new revision rather than silently rewriting this snapshot.

Data and source

The same machine-readable JSON projection powers the Korean and English benchmark views. Use it to build your own charts or compare model variants without treating the table as a universal leaderboard.