통합 Benchmark로 돌아가기
Benchmark Model Family Latest measurement 2026-09-18Projection updated 2026-09-18

Laguna XS 2.1

Laguna XS 2.1의 Q4_K_M·Q5_K_M·Q6_K_L·Q8_0 variant별 GB10 canonical 7-suite 및 external tool-eval 측정 결과입니다.

4 variantsNVIDIA DGX Spark GB10llama.cpp

기본 모델 정보

Model family

Laguna XS 2.1

Measured variants

4

Latest measurement

2026-09-18

Projection updated

2026-09-18

DGX Spark GB10 측정 결과

VariantPerformanceServer-performanceKnowledgeCodingTool-callExternal tool-eval-benchAgent-singleAgent-multi
Q4_K_M
Q4_K_M · non-MTP · non-MTP · reasoning no-think
PP 1915.7 t/s · TG 84.2 t/s
c1 79.2 t/s · c8 283.0 t/s
85/100 · 85.0%
12/12 · 100.0%
13/15 · 86.7%
86/100 · 69/69 scored
6/12 · 50.0%
9/10 · 90.0%
Q5_K_M
Q5_K_M · non-MTP · non-MTP · reasoning no-think
PP 1838.1 t/s · TG 76.8 t/s
c1 72.3 t/s · c8 283.9 t/s
88/100 · 88.0%
12/12 · 100.0%
13/15 · 86.7%
88/100 · 69/69 scored
6/12 · 50.0%
9/10 · 90.0%
Q6_K_L
Q6_K_L · non-MTP · non-MTP · reasoning no-think
PP 1675.3 t/s · TG 68.2 t/s
c1 64.6 t/s · c8 203.3 t/s
88/100 · 88.0%
12/12 · 100.0%
13/15 · 86.7%
88/100 · 69/69 scored
5/12 · 41.7%
9/10 · 90.0%
Q8_0
Q8_0 · non-MTP · non-MTP · reasoning no-think
PP 1705.0 t/s · TG 65.1 t/s
c1 62.1 t/s · c8 177.6 t/s
84/100 · 84.0%
12/12 · 100.0%
13/15 · 86.7%
87/100 · 69/69 scored
5/12 · 41.7%
9/10 · 90.0%

서로 다른 MTP/non-MTP 조건은 통합 총점으로 합산하지 않으며, 각 variant의 실제 조건을 함께 표시합니다.

External tool-eval-bench 상세

기존 Tool-call 15문항 suite와 별도의 외부 69-scenario protocol입니다. 각 variant의 실제 채점 분모를 표시하며, 현재 N2/N2.5 Mini와 Occamy 1.0 run은 grammar 오류 4건을 제외한 65개, Laguna S 2.1과 Laguna XS 2.1 run은 각각 69개를 채점했습니다.

Q4_K_M

86/100

69/69 scored · ★★★★ Good

Safety warning 2

Responsiveness 83 · Deployability 85

Q5_K_M

88/100

69/69 scored · ★★★★ Good

Safety warning 2

Responsiveness 80 · Deployability 86

Q6_K_L

88/100

69/69 scored · ★★★★ Good

Safety warning 2

Responsiveness 78 · Deployability 85

Q8_0

87/100

69/69 scored · ★★★★ Good

Safety warning 2

Responsiveness 77 · Deployability 84

tool-eval-bench 조사·실행 기록 보기

실제 llama-server 실행 명령

아래 명령은 실제 benchmark runner가 사용한 llama-server flag 조합입니다. 공개 페이지에서는 로컬 모델 경로·host·port만 placeholder로 치환했습니다.

Q4_K_M · non-MTP

Server: non-MTP · reasoning no-think
llama-server -m <MODEL_FILE> --host <HOST> --port <PORT> --ctx-size 8192 --n-gpu-layers 999 --flash-attn on --batch-size 2048 --ubatch-size 512 --threads 20 --metrics --slots --parallel 8 --cache-type-k f16 --cache-type-v f16 --spec-type none

Q5_K_M · non-MTP

Server: non-MTP · reasoning no-think
llama-server -m <MODEL_FILE> --host <HOST> --port <PORT> --ctx-size 8192 --n-gpu-layers 999 --flash-attn on --batch-size 2048 --ubatch-size 512 --threads 20 --metrics --slots --parallel 8 --cache-type-k f16 --cache-type-v f16 --spec-type none

Q6_K_L · non-MTP

Server: non-MTP · reasoning no-think
llama-server -m <MODEL_FILE> --host <HOST> --port <PORT> --ctx-size 8192 --n-gpu-layers 999 --flash-attn on --batch-size 2048 --ubatch-size 512 --threads 20 --metrics --slots --parallel 8 --cache-type-k f16 --cache-type-v f16 --spec-type none

Q8_0 · non-MTP

Server: non-MTP · reasoning no-think
llama-server -m <MODEL_FILE> --host <HOST> --port <PORT> --ctx-size 8192 --n-gpu-layers 999 --flash-attn on --batch-size 2048 --ubatch-size 512 --threads 20 --metrics --slots --parallel 8 --cache-type-k f16 --cache-type-v f16 --spec-type none

해석할 때 참고할 점

  • 측정값은 NVIDIA DGX Spark GB10 + llama.cpp + 고정 evaluator 조건의 관찰값입니다.
  • MTP와 non-MTP는 실행 경로가 다르므로 속도 수치를 조건 없이 직접 비교하면 안 됩니다.
  • Knowledge·Coding·Tool-call·External tool-eval-bench·Agent suite는 각각 다른 고정 protocol을 사용하며 종합 지능 점수는 아닙니다.
  • External tool-eval-bench는 현재 N2/N2.5 Mini 8개, Occamy 1.0 4개, Laguna S 2.1 1개, Laguna XS 2.1 4개 variant에서 측정되었으며, 각 variant의 실제 scored/attempted 분모를 표시합니다.
  • Laguna S 2.1의 이전 legacy Knowledge v1·25문항 결과는 historical custom record로 남아 있으며, 현재 S 결과는 Standard Knowledge v1.2·100문항을 사용합니다.

데이터와 업데이트

이 제품군 페이지는 최신 통합 projection에서 자동으로 구성됩니다. 같은 모델의 새 양자화나 MTP/non-MTP 측정이 추가되면 새 URL을 만들지 않고 이 페이지의 variant matrix와 마지막 업데이트 날짜를 갱신합니다.