StoryBench classic.
Five children's-story prompts, generated cold on the 2× DGX Spark cluster. Each story is judged independently by a Claude judge for prose quality and structure (1–10). Speed is measured on the same run.
← StoryBench (current) — craft · religion · values spiders, dual judges, v1.5 rejudge of this corpus + multi-type v2 battery.
Quality vs speed
Leaderboard
Qwen35-A3B quant duel: NVIDIA NVFP4 vs unsloth NVFP4-Fast →
Click a column to sort, a row for the five judged stories. Speed marked reported was recorded at run time but its raw outcome file wasn't retained; measured speeds link to a stored outcome.
| # | Model config | Ctx | Overall | Prose | Structure | Tok/s | Perf |
|---|
Method
Prompts: five fixed children's-story briefs (singing stars · map of the imagination · fox and the storm · talking to plants · memory garden), 2,048 max tokens, thinking disabled. Judging: one independent Claude judge agent per story scores prose quality and structure 1–10 with written rationales; overall is the mean. Serving: vLLM on 2× NVIDIA DGX Spark (GB10), tensor parallel across both nodes. Quantization and context window vary per config and are part of what's being measured.