AAdil Islam
← Finetuning

Storytune writer-lora

The religion tunes taught Gemma 4 12B to believe like a tradition. Storytune teaches it to write like a storyteller — short stories, bedtime tales, fables — measured by the same judge-panel machinery that powers DharmaBench and StoryBench. Below: every benchmark spider in this program, as an explorable 3D stack.

View
drag to orbit · scroll or pinch to zoom · click chips to stack models · hover beads for scores

The program in one paragraph

Storybook Studio serves eight OpenRouter models for kids' stories; Storytune is the plan to make a ninth that belongs there — our own writer LoRA on Gemma 4 12B, trained on a curated children's-literature corpus and selected by measured quality-per-dollar, not vibes. The evaluation battery is Storybench v2: five craft slots (early, fable, family, quest, contemplative), each scored by independent judge panels on prose, structure, and prompt-fit, plus the same religion and value spiders used across DharmaBench.

700
slotted seed rows
92
mined rows judged
63%
fails caught pre-training
4
independent judge panels

Craft slots and honest fill rates

Every training row must land in one of five slots with word bands and register requirements. Floors exist because slot balance beats raw volume. Current state of the Stage-2 fill:

SlotSpecFloorNowFill
sb2_01_earlyages 4–7 bedtime, 250–400w15041
sb2_02_fableages 6–10 animal fable, 350–500w15066
sb2_03_familyfamily life, read-aloud150119
sb2_04_questadventure quest arc150245
sb2_05_contemplativequiet literary close150165

Next fill lever: retell-to-spec regeneration — mined public-domain tales become seeds, a strong model retells them into slot spec, judges gate every row before it enters training. Direct reuse of QA-derived section text was tried and retired: 96% fail rate.

How the data pipeline works

1 · Source. Verified public-domain collections only (Jacobs, Winter Aesop, Ozaki, Lang), each checked against Project Gutenberg metadata; known poison books hard-skipped. 2 · Extract. Per-tale segmentation by title headers — never blind chunking — with artifact scans (chapter headings, footnote markers, broken ligatures), grim-content and religion blocklists, sha1 dedup, near-dupe removal, and 5-gram quarantine against the official Storybench v2 prompts so we don't train on the test. 3 · Judge. Exhaustive review by independent judge panels acting as blind readers: PASS enters training, FAIL drops with the reason logged, BORDERLINE keeps but carries a meta flag. 4 · Merge & train. Slot-balanced merge into craft_floor_v1.jsonl, LLaMA-Factory registration, LoRA fine-tune, then the full Storybench v2 battery per checkpoint.

What the judges caught

All 92 mined rows went under four independent panels instead of spot sampling. Mechanical filters had already removed truncation, artifacts, and religious content — what survived for human-grade judgment shows why judgment is irreplaceable:

Panel verdictRowsTypical catch
PASS7clean complete arcs (Princess & the Pea, Town/Country Mouse)
BORDERLINE27warm but headless — kept, flagged in metadata for ablation
FAIL58mid-tale fragments, doom endings, archaic diction, red-hot-iron torture scenes

Two structural findings now baked into the pipeline: FairytaleQA-style section corpora are unfixable by filtering (fragments are intrinsic), and a regex "complete ending" check passes cliffhangers wholesale — terminal punctuation is not narrative resolution.

← Finetuning ·  StoryBench ·  DharmaBench