Q4_K_M
87/100
65/69 scored · ★★★★ Good
Safety gate passed
Responsiveness 79 · Deployability 85
Occamy 1.0의 Q4_K_M·Q5_K_M·Q6_K·Q8_0 variant별 GB10 canonical 7-suite 및 external tool-eval 측정 결과입니다.
Model family
Occamy 1.0
Measured variants
4
Latest measurement
2026-09-18
Projection updated
2026-09-18
| Variant | Performance | Server-performance | Knowledge | Coding | Tool-call | External tool-eval-bench | Agent-single | Agent-multi |
|---|---|---|---|---|---|---|---|---|
Q4_K_M Q4_K_M · non-MTP · non-MTP · reasoning no-think | PP 1953.5 t/s · TG 80.6 t/s | c1 76.3 t/s · c8 202.6 t/s | 94/100 · 94.0% | 12/12 · 100.0% | 11/15 · 73.3% | 87/100 · 65/69 scored | 8/12 · 66.7% | 6/10 · 60.0% |
Q5_K_M Q5_K_M · non-MTP · non-MTP · reasoning no-think | PP 1854.7 t/s · TG 73.7 t/s | c1 68.8 t/s · c8 192.5 t/s | 95/100 · 95.0% | 11/12 · 91.7% | 11/15 · 73.3% | 85/100 · 65/69 scored | 9/12 · 75.0% | 9/10 · 90.0% |
Q6_K Q6_K · non-MTP · non-MTP · reasoning no-think | PP 1659.0 t/s · TG 66.8 t/s | c1 63.6 t/s · c8 161.2 t/s | 96/100 · 96.0% | 12/12 · 100.0% | 11/15 · 73.3% | 86/100 · 65/69 scored | 8/12 · 66.7% | 10/10 · 100.0% |
Q8_0 Q8_0 · non-MTP · non-MTP · reasoning no-think | PP 1734.0 t/s · TG 58.8 t/s | c1 56.6 t/s · c8 152.1 t/s | 96/100 · 96.0% | 11/12 · 91.7% | 11/15 · 73.3% | 86/100 · 65/69 scored | 9/12 · 75.0% | 9/10 · 90.0% |
서로 다른 MTP/non-MTP 조건은 통합 총점으로 합산하지 않으며, 각 variant의 실제 조건을 함께 표시합니다.
기존 Tool-call 15문항 suite와 별도의 외부 69-scenario protocol입니다. 각 variant의 실제 채점 분모를 표시하며, 현재 N2/N2.5 Mini와 Occamy 1.0 run은 grammar 오류 4건을 제외한 65개, Laguna S 2.1과 Laguna XS 2.1 run은 각각 69개를 채점했습니다.
87/100
65/69 scored · ★★★★ Good
Safety gate passed
Responsiveness 79 · Deployability 85
85/100
65/69 scored · ★★★★ Good
Safety warning 2
Responsiveness 76 · Deployability 82
86/100
65/69 scored · ★★★★ Good
Safety warning 2
Responsiveness 75 · Deployability 83
86/100
65/69 scored · ★★★★ Good
Safety warning 1
Responsiveness 72 · Deployability 82
아래 명령은 실제 benchmark runner가 사용한 llama-server flag 조합입니다. 공개 페이지에서는 로컬 모델 경로·host·port만 placeholder로 치환했습니다.
llama-server -m <MODEL_FILE> --host <HOST> --port <PORT> --ctx-size 8192 --n-gpu-layers 999 --flash-attn on --batch-size 2048 --ubatch-size 512 --threads 20 --metrics --slots --parallel 8 --cache-type-k f16 --cache-type-v f16 --spec-type nonellama-server -m <MODEL_FILE> --host <HOST> --port <PORT> --ctx-size 8192 --n-gpu-layers 999 --flash-attn on --batch-size 2048 --ubatch-size 512 --threads 20 --metrics --slots --parallel 8 --cache-type-k f16 --cache-type-v f16 --spec-type nonellama-server -m <MODEL_FILE> --host <HOST> --port <PORT> --ctx-size 8192 --n-gpu-layers 999 --flash-attn on --batch-size 2048 --ubatch-size 512 --threads 20 --metrics --slots --parallel 8 --cache-type-k f16 --cache-type-v f16 --spec-type nonellama-server -m <MODEL_FILE> --host <HOST> --port <PORT> --ctx-size 8192 --n-gpu-layers 999 --flash-attn on --batch-size 2048 --ubatch-size 512 --threads 20 --metrics --slots --parallel 8 --cache-type-k f16 --cache-type-v f16 --spec-type none이 제품군 페이지는 최신 통합 projection에서 자동으로 구성됩니다. 같은 모델의 새 양자화나 MTP/non-MTP 측정이 추가되면 새 URL을 만들지 않고 이 페이지의 variant matrix와 마지막 업데이트 날짜를 갱신합니다.