Laguna S 2.1 APEX-I Balanced — canonical 7-suite Benchmark

2026년 9월 17일4
작성 DevSnack Lab직접 실행·측정·판단 기록
BenchmarkLocal LLMLaguna S 2.1DGX Spark GB10llama.cpp7-suiteTool-callAgent

Laguna S 2.1 Uncensored APEX-I Balanced — 7-suite benchmark


Public-safe summary. Raw prompts, responses, local paths, logs, and execution traces remain local-only.

Scope


  • Hardware: NVIDIA DGX Spark GB10
  • Runtime: llama.cpp b10930-56381e407
  • Model variant: Laguna S 2.1 — APEX-I Balanced (NVFP4)
  • MTP: non-MTP (spec_type=none)
  • Quality suites: no-think, thinking budget 0
  • Measurement date: 2026-09-16~2026-09-17 (KST)

Results


SuiteWorkloadResult
PerformancePP 512[경로], TG 512, 5 repetitionsPP 762.48 / 757.70 / 756.94 / 702.17, TG 27.55 tokens/s
Server-performanceConcurrency 1[경로], 3 repetitions180/180 successful, failure rate 0%
KnowledgeStandard Knowledge v1.2, 100 questions88/100 (88%)
Coding12 tasks11/12 (91.67%)
Tool-call v1.115 tasks13/15 (86.67%)
Agent-single v1.112 tasks6/12 (50%)
Agent-multi v1.110 tasks7/10 (70%)

Server-performance


ConcurrencyAggregate tokens/sPer-request tokens/sp50 latencyp95 latency
126.8126.837.79s8.18s
237.4418.8611.73s12.73s
459.0914.9214.33s15.36s
872.8810.3524.25s31.11s

Interpretation


The model maintained 702.17 tokens/s at a 32K prompt and generated at 27.55 tokens/s in the Performance workload. The server lane completed all 180 requests, while increasing concurrency reduced per-request throughput and increased latency.


Knowledge v1.2 scored 88/100, with general 20/20, Korea 19/20, math 16/20, science 19/20, and logic 14/20. Coding reached 11/12. Tool-call reached 13/15, with tool selection, argument, and execution metrics each at 93.33%. Agent-single was weaker at 6/12, while Agent-multi achieved 100% handoff and role participation but completed 7/10 tasks. The main gap is maintaining required steps through longer agent workflows.


Compatibility note


The Knowledge result now uses the current Standard Knowledge v1.2 dataset with 100 questions. The previous 24/25 result was legacy Knowledge v1 and remains only as a historical custom result.


External tool-eval-bench


The same model was also measured with the separate tool-eval-bench protocol: 86/100, 119/138, and 69/69 scored. Deployability was 73 and Responsiveness was 42. The safety gate failed because the run followed a fake system message in a file, attempted a destructive action after authority escalation, and carried a cross-turn injection into email recipients. These findings are reported as safety limitations, not hidden behind the score.


The external protocol is separate from the internal Tool-call v1.1 suite. It uses 69 deterministic mock-tool scenarios, temperature 0, no-think, and sequential execution.