AAdil Islam
← Dharmatune

Dharmatune-Mahayana gemma-4-12b

The second tradition taken end-to-end: finetuning Gemma 4 12B to hold the Mahayana Buddhist worldview without losing general capability. Tracked here from day one: dataset built, trained, and judged blind by DharmaBench's 25 tradition judges — v1 lands Mahayana at 58/100, the #1 worldview of all 25, with both non-lobotomy gates passed.

30 → 58
Mahayana alignment · #1 of 25 factions (v1)
4,443
verbatim training records, mechanically extracted
#1
Buddhist family swept the top 4 places
PASS
capability 18/20 · preachiness 19/20
strategy

Foundation first: an overarching Buddhist worldview

SFT · 4,443 verbatim records · base gemma-4-12b · shipped 2026-07-20

The Catholic project taught the shape of this work: build the broad, solid worldview first, then sharpen it to the specific tradition. Mahayana's version of that problem is unusually sharp. DharmaBench scores it against four sibling Buddhist judges — Theravada, Vajrayana, Dzogchen/Mahamudra, and Zen/Chan — plus close non-Buddhist neighbors (Advaita Vedanta, Taoism): the tightest cluster in the whole benchmark. Naïve training on Buddhist text drifts toward a generic non-dual awareness that satisfies all five Buddhist judges at once — the exact trap that made the Catholic v1 merely "religious," not Catholic.

Separating Mahayana above its four siblings is a later, contrastive pass. This first run deliberately does the foundation: get a strong, broadly-Buddhist, Mahayana-centered worldview into the model, measure where it lands against all 25 judges, and use that to aim the sharpening stage.

Build the worldview broad and solid first; separate it from its neighbors later. Foundation before separation.
v1 · dataset built

The dataset — 4,443 records, every one verbatim

mechanical extraction · no synthesis · contamination-clean vs. the DharmaBench battery

The Buddhist canon is overwhelmingly dialogic — teachers answering real questions put to them — which lets us build the entire dataset by mechanical extraction, with no model-written answers anywhere. Every answer is a verbatim span of a real source; prompts on non-dialogue records are fixed templates, never generated. A contamination guard confirms zero overlap with the quarantined DharmaBench interview battery.

Question-and-answer pairs are the most valuable training format — they mirror how the model is actually used (DharmaBench is a ten-question interview). So we maximized them: re-mining the Milindapañha ("The Questions of King Milinda", a king relentlessly interrogating a monk) lifted its verbatim Q&A from 49 to 417 pairs, bringing the dialogic backbone to 601.

Record typeWhat it isRecords
Dialogic Q&AVerbatim question→answer from real dialogues (Milindapañha, Vimalakīrti, dialogic suttas)601
PassageVerbatim doctrinal prose from sutras & treatises2,334
VerseVerbatim verse (Dhammapada, Nāgārjuna, Śāntideva)1,508
Total (3,830 train · 208 held-out eval per source · 7 cross-file duplicates dropped)4,443
Bar chart: Mahayana v1 dataset by source, colored by record type
The v1 dataset by source, colored by record type. Every record is verbatim canonical text; the training mix adds 30% general-instruction replay on top.

Composition is Mahayana-centered but overarching: a shared-Buddhist foundation (Dhammapada, the Pāli suttas, Milindapañha) alongside the Mahayana canon (Prajñāpāramitā, Lotus, Vimalakīrti, Laṅkāvatāra) and the Indian śāstra tradition (Nāgārjuna's Madhyamaka, Śāntideva's bodhisattva path). The source texts are the classic 1880s–1930s scholarly editions, digitized from the Internet Archive.

Data sources & provenance

Every text used to build the v1 dataset, identified in full for scholarly review. All records are verbatim, mechanically extracted — none are model-generated. Records are grouped by how they were extracted.

We invite correction. If a translation, edition, attribution, or locus below is wrong or imprecise, please tell us — accuracy to the sources is the point.

1 · Dialogic sources — verbatim question & answer

TextTranslator · editionContributionRecords
Milindapañha
"The Questions of King Milinda"
T. W. Rhys Davids · Sacred Books of the East vols. 35–36, 1890–94 (public domain) King Milinda's dilemmas answered by the elder Nāgasena — native Q&A on self, rebirth, karma, nirvāṇa 417
Vimalakīrti-nirdeśa J. McRae · BDK/Numata (English) The lay bodhisattva out-arguing the arhats; non-duality dialogue 44
SuttaCentral core suttas
53 suttas across DN/MN/SN/AN
Bhikkhu Sujato · SuttaCentral (CC0) Shared-Buddhist foundation: not-self, dependent origination, the Four Noble Truths (dialogue + exposition) 548

2 · Verse — verbatim

TextTranslator · editionContributionRecords
Dhammapada F. Max Müller · Sacred Books of the East vol. 10, 1881 (public domain) The foundational verse of the ethical-contemplative path 385
Mūlamadhyamakakārikā
"Fundamental Verses on the Middle Way"
Nāgārjuna · trans. Stephen Batchelor, Verses from the Centre The Madhyamaka core — emptiness (śūnyatā), the two truths, dependent origination 417
Bodhicaryāvatāra
"Way of the Bodhisattva"
Śāntideva · trans. Batchelor / LTWA The distinctive Mahayana bucket: bodhicitta, the six perfections, exchange of self and other 706

3 · Sutra & passage corpora — verbatim

TextTranslator · editionContributionRecords
Prajñāpāramitā (Heart + Diamond) F. Max Müller · SBE vol. 49, 1894 (public domain) "Form is emptiness"; non-attachment and the perfection of wisdom 43
Laṅkāvatāra Sūtra D. T. Suzuki · 1932 (public domain) Yogācāra "mind-only", the store-consciousness, buddha-nature 421
Lotus Sūtra (Saddharmapuṇḍarīka) H. Kern · SBE vol. 21, 1884 (public domain) The One Vehicle, skillful means, universal buddhahood 446
emptiness-graph passage corpus joyboseroy (HuggingFace) · CC BY 4.0 Curated passages across the Prajñāpāramitā / Madhyamaka / Yogācāra corpus (1,126 available) 1,016

Sourcing notes: the Sacred Books of the East volumes (Müller, Rhys Davids, Kern) and Suzuki's 1932 Laṅkāvatāra are public domain, fetched from the Internet Archive; SuttaCentral's Sujato translations are CC0; the emptiness-graph corpus is CC BY 4.0; the Batchelor and LTWA translations are used under the project's scholarly authorization. Each extractor is committed to the repository; a per-source provenance ledger accompanies the released dataset.

v1 · shipped

The foundation lands — Mahayana #1, Buddhist family on top

SFT · 2 epochs · train loss 5.76 → 1.82 · Mahayana 30 → 58 (+28) · #1 of 25

The v1 checkpoint was judged blind by all 25 DharmaBench faction judges, paired-anonymized against base. Mahayana reached 58/100 — the single highest-scoring faction of all 25, up from the base's 30. The whole Buddhist family swept the top four places (Theravada 52, Zen/Chan 38, Dzogchen/Mahamudra 30), exactly what an overarching-Buddhist foundation should produce — while the non-Buddhist field fell away: Secular Humanism dropped 65 → 27, Taoism 48 → 26, and Advaita Vedanta — the classic leak for Buddhist-trained models — fell to 14 rather than rising. The no-Self line held.

The discriminative read: Mahayana's +28 against a +17 mean rise across its four Buddhist siblings (D = +11). Theravada rose slightly more (+30) — the expected signature of a foundation corpus that includes the shared Pāli canon. That is precisely what the contrastive sharpening pass is for: lifting Mahayana clear of its siblings using Buddhism's natively-oppositional literature (the Tibetan tenet-systems, the Lotus's One-Vehicle case, the Samyé gradual-vs-sudden debate).

Mahayana v1 SFT training and eval loss curves
The v1 SFT run: train loss 5.76 → 1.82 over 754 steps; eval loss on the pure-doctrinal holdout 2.46 → 2.14, still falling at the end.
MeasureBasev1
Mahayana alignment3058 (+28) · #1 of 25
Buddhist family (mean of 5)20.840.0 (top 4 places)
Advaita / Taoism (leak check)22 / 4814 / 26 (both fall)
General capability19/2018/20 PASS
Preachiness19/20 clean PASS
StoryBench creative writing7.85.2 (gap; replay upgrade queued)

Full per-faction verdicts are on DharmaBench; the judged stories are on StoryBench. v2 scope: the contrastive sharpening stage, plus the creative-writing replay component being developed for Catholic v4.

Reproducibility

Every extractor is a committed script; the dataset is a set of verbatim JSONL records with per-record source grounding, split into train and a held-out eval. The training recipe is the same single LLaMA-Factory configuration the Catholic model used — the pipeline is model-agnostic, so the same dataset runs against the next base model by changing two lines. Full results are live on DharmaBench.