AAdil Islam
BenchmarksFinetuning · DharmaBench

DharmaBench.

A ten-turn worldview interrogation — self, reality, suffering, death, ethics, free will, cosmos, knowledge, a forced dilemma, and a final summation — asked with no persona and no system prompt. Twenty-five religious and philosophical schools then judge the transcript: how close is the model's de-facto worldview to that tradition (0–100)?

About this benchmark →

Static copies: off vs HIGH · thinking off · thinking HIGH

Qwen35-A3B quant duel →

Runs

One row per benchmark run of a recipe. Dharmatune finetunes (and their base control) are listed first in train order; other cluster recipes follow newest→oldest. Best alignment names the religious system whose judge scored the model's worldview closest to its own doctrine — click a row for top verdicts and to load that run’s interrogation below.

RecipeDateBest alignment Tok/sTurnsFlags

The interrogation

Ten fixed questions, Q1→Q10, with the model’s verbatim answers. Runs are ordered: our Dharmatune finetunes first (base → Catholic v1→v3 → Mahayana), then other cluster recipes newest→oldest. Select a run to read its transcript.

Judgment by tradition

Each row is one judge's verdict on the selected run: a holistic score and the ten per-question scores. Tap any cell for that judge's reasoning, or tap the row for the full verdict. Each judge is tagged a religious system or a secular philosophy; only religious systems compete for best alignment, and Secular Humanism is held as the control. Q1–Q10:

0100 · alignment
sort: score · religion
SchoolHolisticPer-question (Q1–Q10)

Method

The battery escalates through ten fixed questions, each building on the model's previous answers, so the model must commit to a coherent worldview rather than hedge per-question. Judging: one independent judge agent per school, writing in that tradition's own voice, scores each answer 0–100 for doctrinal alignment plus a holistic verdict. Most historical cluster runs used Claude agents; the local GLM-5.3-Flash 1M recipe was judged by Grok 4.6. Speed (tok/s) is measured across the same ten generations on the 2× DGX Spark cluster. CoT-leak turns (thinking traces escaping into output) are counted and flagged.