Storytune writer-lora
The religion tunes taught Gemma 4 12B to believe like a tradition. Storytune teaches it to write like a storyteller — short stories, bedtime tales, fables — measured by the same judge-panel machinery that powers DharmaBench and StoryBench. Below: every benchmark spider in this program, as an explorable 3D stack.
The program in one paragraph
Storybook Studio serves eight OpenRouter models for kids' stories; Storytune is the plan to make a ninth that belongs there — our own writer LoRA on Gemma 4 12B, trained on a curated children's-literature corpus and selected by measured quality-per-dollar, not vibes. The evaluation battery is Storybench v2: five craft slots (early, fable, family, quest, contemplative), each scored by independent judge panels on prose, structure, and prompt-fit, plus the same religion and value spiders used across DharmaBench.
Craft slots and honest fill rates
Every training row must land in one of five slots with word bands and register requirements. Floors exist because slot balance beats raw volume. Current state of the Stage-2 fill:
| Slot | Spec | Floor | Now | Fill |
|---|---|---|---|---|
| sb2_01_early | ages 4–7 bedtime, 250–400w | 150 | 41 | |
| sb2_02_fable | ages 6–10 animal fable, 350–500w | 150 | 66 | |
| sb2_03_family | family life, read-aloud | 150 | 119 | |
| sb2_04_quest | adventure quest arc | 150 | 245 | |
| sb2_05_contemplative | quiet literary close | 150 | 165 |
Next fill lever: retell-to-spec regeneration — mined public-domain tales become seeds, a strong model retells them into slot spec, judges gate every row before it enters training. Direct reuse of QA-derived section text was tried and retired: 96% fail rate.
How the data pipeline works
1 · Source. Verified public-domain collections only
(Jacobs, Winter Aesop, Ozaki, Lang), each checked against Project
Gutenberg metadata; known poison books hard-skipped.
2 · Extract. Per-tale segmentation by title headers —
never blind chunking — with artifact scans (chapter headings, footnote
markers, broken ligatures), grim-content and religion blocklists,
sha1 dedup, near-dupe removal, and 5-gram quarantine against the
official Storybench v2 prompts so we don't train on the test.
3 · Judge. Exhaustive review by independent judge
panels acting as blind readers: PASS enters training, FAIL drops with
the reason logged, BORDERLINE keeps but carries a meta flag.
4 · Merge & train. Slot-balanced merge into
craft_floor_v1.jsonl, LLaMA-Factory registration,
LoRA fine-tune, then the full Storybench v2 battery per checkpoint.
What the judges caught
All 92 mined rows went under four independent panels instead of spot sampling. Mechanical filters had already removed truncation, artifacts, and religious content — what survived for human-grade judgment shows why judgment is irreplaceable:
| Panel verdict | Rows | Typical catch |
|---|---|---|
| PASS | 7 | clean complete arcs (Princess & the Pea, Town/Country Mouse) |
| BORDERLINE | 27 | warm but headless — kept, flagged in metadata for ablation |
| FAIL | 58 | mid-tale fragments, doom endings, archaic diction, red-hot-iron torture scenes |
Two structural findings now baked into the pipeline: FairytaleQA-style section corpora are unfixable by filtering (fragments are intrinsic), and a regex "complete ending" check passes cliffhangers wholesale — terminal punctuation is not narrative resolution.