A structured, machine-readable knowledge base of Sanskrit concepts that resist accurate translation into English, built for AI agents, educators, researchers, and technologists working with dharmic knowledge systems.
13 bundles · 310 concept documents · 132 primary-source reference documents — every file conformant with the Open Knowledge Format v0.2, validated 0-fail.
Interactive bundle visualizers + enrichment waves 1-2 (July 2026): every bundle now ships a self-contained viz.html — a force-directed concept graph with a per-concept detail panel, generated by the repo's own viewer tooling (see the Viz column below; download and open locally, the file is fully self-contained). The corpus's most-consumed terms (karma, yoga, māyā, dharma, mokṣa, saṃsāra, chakra, guru, mūrti, pūjā, ṛta, svadharma, mithyā) now carry a reception_note: presenter guardrail grounded in each file's documented Error Genealogy, and the foundational and Upaniṣadic bundles carry front-page "What This Bundle Does Not Claim" statements. Validator hardened (legacy-body-format recognition, echo-line tokenizer fix); machine-readable not: fields and their human-readable echo lines reconciled so both audiences receive identical bans.
v0.13 sankhya-darshana (July 2026): completes the six classical darśanas. Kapila's dualist, enumerative school — the causal theory (satkāryavāda vs. Nyāya-Vaiśeṣika's ārambhavāda), the plural, nirīśvara puruṣa-prakṛti dualism, and the discriminative (viveka-khyāti) path to kaivalya, contrasted throughout with Yoga-darśana's Īśvara-inclusive path to the same end-state word via nirodha. The bundle deliberately does not re-derive the tattva evolution ladder already published in yoga-darshana, cosmology-creation, and ayurveda-consciousness, and names guṇa's doctrinal source rather than re-treating it as another school-contrast. Recent releases: v0.12 jyotisha-kala (the Jyotiṣa vocabulary of time), v0.11 ayurveda-consciousness (first saṃhitā-sourced bundle).
The Open Knowledge Format (OKF) uses YAML-frontmatter Markdown files. Each concept file does two jobs at once, a dual-action design:
- A boundary — a
not:field listing the English mistranslations an AI must avoid, each paired with why it fails and what to use instead. - A scaffold — positive fields that install the correct understanding: What It Actually Means, Audience Metaphor, Etymology, and Citations.
So OKF does not merely tell a model what a term is not. It clears out the wrong definition and immediately installs the right one, leaving the model with a complete semantic package rather than a severed, "floating" term.
- "Karma = Fate" → Wrong. It is action and consequence. See
okf/dharma-foundation/concepts/karma.md - "Dharma = Religion" → Wrong. It is the natural sustaining order / right conduct. See
okf/dharma-foundation/concepts/dharma.md - "Samadhi = Trance" → Wrong. It is absorption / unified awareness. See
okf/dharma-foundation/concepts/samadhi.md - "Yoga = Exercise" → Wrong. It is the discipline of stilling the mind. See
okf/dharma-foundation/concepts/yoga.md
A non-translatable is school-relative. The same Sanskrit word can be a different technical object in a different darśana, and the corpus preserves the distinction rather than flattening it: karma is moral action-and-consequence in dharma-foundation, the padārtha of motion in Nyāya-Vaiśeṣika, the enjoined ritual act in Mīmāṃsā, and the therapeutic procedure of pañcakarma in Āyurveda; guṇa names the Sāṃkhya triguṇa at its doctrinal source, an unrelated Vaiśeṣika ontological category, and the twenty gurvādi properties of Āyurvedic pharmacology; kaivalya is reached by discrimination alone in Sāṃkhya and by practiced cessation in Yoga. Shared terms carry a school_scope: field, an index contrast note, and reciprocal cross-links — never a single "canonical" definition.
A mistranslation changes the category a model reasons in, which changes its action space and therefore its output. A wellness bot that reads yoga as exercise recommends the squat rack; an advisor that reads karma as fate gives fatalistic, agency-removing advice. The metaphysical error is a logic bug.
See demos/failure-vs-success.md for side-by-side failure-mode vs success-mode transcripts.
The bundles are descriptive vocabulary, and the sensitive domains say so on their front page. The Śākta bundle documents that certain practices are initiation-gated without reproducing anything gated (adhikāra discipline, zero how-to). The Āyurveda bundle is not medical advice, contains no preparations or dosages, and does not adjudicate clinical efficacy. The Jyotiṣa bundle makes no predictive claims and does not endorse astrology while documenting its classical vocabulary. Commercially captured terms carry a reception_note: field naming the capture, and as of the July 2026 enrichment pass the foundational and Upaniṣadic bundles state their own does-not-claim boundaries on their index pages. The corrective work lives in the not: fields, concept by concept, not in polemic.
A not: list in a file is a "do not enter" sign on an unlocked door. The model still has to be told how to read and honor it, and a naive bare prohibition can backfire through negative-prompt leakage. INTEGRATION.md shows how to make the constraint functionally binding: a system-prompt template that injects the positive instead redirect (not a bare ban), a RAG negative-filter pattern, and a lightweight output-check, with honest caveats about what is and is not guaranteed.
| Bundle | bundle_version |
Concepts | Refs | Theme | Viz |
|---|---|---|---|---|---|
okf/dharma-foundation/ |
0.1.4 | 25 | 12 | Foundational Sanskrit non-translatable vocabulary | graph |
okf/yoga-darshana/ |
0.2.2 | 26 | 6 | Patañjali's Yoga Sūtra technical lexicon | graph |
okf/vedanta-epistemology/ |
0.3.3 | 27 | 19 | Pramāṇa-śāstra — Vedantic epistemology | graph |
okf/bhakti-marga/ |
0.4.3 | 15 | 9 | The vocabulary of the devotional path | graph |
okf/dharmic-ethics/ |
0.5.3 | 15 | 6 | Yama–Niyama and the ethical-social vocabulary of dharma | graph |
okf/upanishadic-core/ |
0.6.2 | 26 | 16 | Core Upaniṣadic vocabulary across the Vedānta schools — mahāvākyas, Ātman–Brahman, self-knowledge | graph |
okf/cosmology-creation/ |
0.7.3 | 26 | 12 | Vedic & Purāṇic vocabulary of time, cosmos, and manifestation | graph |
okf/shakta-darshana/ |
0.8.2 | 26 | 12 | Śākta-Tāntric metaphysics of Śakti, Devī, and consciousness-power | graph |
okf/nyaya-vaisheshika/ |
0.9.1 | 27 | 10 | The science of inference and debate — Nyāya apparatus + Vaiśeṣika realist ontology | graph |
okf/mimamsa-dharma/ |
0.10.1 | 25 | 7 | Mīmāṃsā ritual hermeneutics — vidhi, apūrva, and language-as-action | graph |
okf/ayurveda-consciousness/ |
0.11.1 | 26 | 8 | Āyurvedic vocabulary of consciousness, constitution, and health | graph |
okf/jyotisha-kala/ |
0.12.1 | 26 | 8 | Jyotiṣa vocabulary of time — pañcāṅga, kāla-reckoning, and the sidereal celestial frame | graph |
okf/sankhya-darshana/ |
0.13.3 | 20 | 7 | Sāṃkhya's dualist causal theory and discriminative path to kaivalya — completes the six classical darśanas | graph |
| Total | 310 | 132 | 13 bundles spanning the six āstika darśanas + the devotional, ethical, cosmological, medical, and calendrical corpora |
Two consumption surfaces, both first-class, are declared in VERSIONING.md: main is a living vocabulary (concept files are enriched in place — sharper not: fields, added citations, documented genealogies — with bundle_version patch bumps), and release tags are immutable archival snapshots. In-place enrichment waves are logged newest-first in CHANGELOG.md.
Each bundle above carries the tag bundle/<name>/v<bundle_version> — okf/yoga-darshana/ at 0.2.2 is bundle/yoga-darshana/v0.2.2. Twenty-nine older repository-wide vX.Y.Z tags remain valid and unchanged. A tag marks a whole-repository state, not a bundle's files, and the legacy tags do not even name the bundle they were cut for: several share a commit. Cite the bundle name together with a commit SHA — see PROFILE.md §7.1, which documents this in full.
GENEALOGIES.md — "Where the Errors Came From" — documents the histories of the English mistranslations this corpus corrects: not just that a rendering is wrong, but who introduced it, when, and how it propagated into today's training data (e.g., karma-as-fate from Blavatsky 1889; yoga-as-posture via Vivekananda → Krishnamacharya → Singleton 2010; māyā-as-illusion via Schopenhauer 1818). A correction with a genealogy is harder to dismiss than one with only an assertion. The admission bar is strict: a named source, a date, and a documented propagation chain, or it is excluded. The affected concept files carry a matching Error Genealogy section linking back to the canonical entry.
Each concept file contains:
- YAML frontmatter:
type,title,iast,devanagari,description,darshana, a structurednot:(each entryterm/why/instead),related:,tags,okf_version,license(and, where a term recurs across darśanas,school_scope:; where public reception distorts a term,reception_note:) - ## What It Actually Means (and ## Etymology) — the precise positive definition
- ## Audience Metaphor — an accessible analogy engineered for AI and general comprehension
- ## Citations — primary śāstra references, linked into a
references/sub-bundle of first-classtype: Referenceconcepts
Full specification: the base format is Open Knowledge Format v0.2, pinned at commit ad30107. This repository's extensions and stricter rules are documented in PROFILE.md — §2 covers darshana and the structured not: with instead, §2.5 per-claim attribution, §3.2 the required references/ sub-bundle.
Two licences, by scope. Nothing in this repository is licensed by inheritance.
| Scope | Licence | File |
|---|---|---|
The corpus — every document under okf/<bundle>/, and the repository-root prose documents |
CC BY-SA 4.0 (ShareAlike) | LICENSE-CONTENT |
The code — okf/tools/ (the validator) and okf/tests/ (the regression suite) |
Apache 2.0, © Dharma OKF Foundation | LICENSE.md, scope stated in NOTICE |
An earlier revision of this section described the Apache licence as "inherited from the upstream GoogleCloudPlatform fork." That was inaccurate: okf/tools/ has no counterpart upstream, and the validator imports the upstream parser optionally rather than embedding it. The code is this project's own work and is licensed deliberately, not by inheritance.
This is the standalone publication home for the Dharma OKF Foundation. Contributions, corrections, and additions that deepen accuracy are welcome via Pull Request. See CONTRIBUTING.md.