DharmaBench.
A ten-turn worldview interrogation — self, reality, suffering, death, ethics, free will, cosmos, knowledge, a forced dilemma, and a final summation — asked with no persona and no system prompt. Twenty-five religious and philosophical schools then judge the transcript: how close is the model's de-facto worldview to that tradition (0–100)?
Static copies: off vs HIGH · thinking off · thinking HIGH
Runs
One row per benchmark run of a recipe. Dharmatune finetunes (and their base control) are listed first in train order; other cluster recipes follow newest→oldest. Best alignment names the religious system whose judge scored the model's worldview closest to its own doctrine — click a row for top verdicts and to load that run’s interrogation below.
| Recipe | Date | Best alignment | Tok/s | Turns | Flags |
|---|
The interrogation
Ten fixed questions, Q1→Q10, with the model’s verbatim answers. Runs are ordered: our Dharmatune finetunes first (base → Catholic v1→v3 → Mahayana), then other cluster recipes newest→oldest. Select a run to read its transcript.
Judgment by tradition
Each row is one judge's verdict on the selected run: a holistic score and the ten per-question scores. Tap any cell for that judge's reasoning, or tap the row for the full verdict. Each judge is tagged a religious system or a secular philosophy; only religious systems compete for best alignment, and Secular Humanism is held as the control. Q1–Q10:
| School | Holistic | Per-question (Q1–Q10) |
|---|
Method
The battery escalates through ten fixed questions, each building on the model's previous answers, so the model must commit to a coherent worldview rather than hedge per-question. Judging: one independent judge agent per school, writing in that tradition's own voice, scores each answer 0–100 for doctrinal alignment plus a holistic verdict. Most historical cluster runs used Claude agents; the local GLM-5.3-Flash 1M recipe was judged by Grok 4.6. Speed (tok/s) is measured across the same ten generations on the 2× DGX Spark cluster. CoT-leak turns (thinking traces escaping into output) are counted and flagged.