Skip to content

Dharma OKF — Open Knowledge Format for Dharma / Vedic concepts (Sanskrit Non-Translatables)

A structured, machine-readable knowledge base of Sanskrit concepts that resist accurate translation into English, built for AI agents, educators, researchers, and technologists working with dharmic knowledge systems.

13 bundles · 310 concept documents · 132 primary-source reference documents — every file conformant with the Open Knowledge Format v0.2, validated 0-fail.

What's New

Interactive bundle visualizers + enrichment waves 1-2 (July 2026): every bundle now ships a self-contained viz.html — a force-directed concept graph with a per-concept detail panel, generated by the repo's own viewer tooling (see the Viz column below; download and open locally, the file is fully self-contained). The corpus's most-consumed terms (karma, yoga, māyā, dharma, mokṣa, saṃsāra, chakra, guru, mūrti, pūjā, ṛta, svadharma, mithyā) now carry a reception_note: presenter guardrail grounded in each file's documented Error Genealogy, and the foundational and Upaniṣadic bundles carry front-page "What This Bundle Does Not Claim" statements. Validator hardened (legacy-body-format recognition, echo-line tokenizer fix); machine-readable not: fields and their human-readable echo lines reconciled so both audiences receive identical bans.

v0.13 sankhya-darshana (July 2026): completes the six classical darśanas. Kapila's dualist, enumerative school — the causal theory (satkāryavāda vs. Nyāya-Vaiśeṣika's ārambhavāda), the plural, nirīśvara puruṣa-prakṛti dualism, and the discriminative (viveka-khyāti) path to kaivalya, contrasted throughout with Yoga-darśana's Īśvara-inclusive path to the same end-state word via nirodha. The bundle deliberately does not re-derive the tattva evolution ladder already published in yoga-darshana, cosmology-creation, and ayurveda-consciousness, and names guṇa's doctrinal source rather than re-treating it as another school-contrast. Recent releases: v0.12 jyotisha-kala (the Jyotiṣa vocabulary of time), v0.11 ayurveda-consciousness (first saṃhitā-sourced bundle).

What Is OKF?

The Open Knowledge Format (OKF) uses YAML-frontmatter Markdown files. Each concept file does two jobs at once, a dual-action design:

  1. A boundary — a not: field listing the English mistranslations an AI must avoid, each paired with why it fails and what to use instead.
  2. A scaffold — positive fields that install the correct understanding: What It Actually Means, Audience Metaphor, Etymology, and Citations.

So OKF does not merely tell a model what a term is not. It clears out the wrong definition and immediately installs the right one, leaving the model with a complete semantic package rather than a severed, "floating" term.

  • "Karma = Fate" → Wrong. It is action and consequence. See okf/dharma-foundation/concepts/karma.md
  • "Dharma = Religion" → Wrong. It is the natural sustaining order / right conduct. See okf/dharma-foundation/concepts/dharma.md
  • "Samadhi = Trance" → Wrong. It is absorption / unified awareness. See okf/dharma-foundation/concepts/samadhi.md
  • "Yoga = Exercise" → Wrong. It is the discipline of stilling the mind. See okf/dharma-foundation/concepts/yoga.md

A non-translatable is school-relative. The same Sanskrit word can be a different technical object in a different darśana, and the corpus preserves the distinction rather than flattening it: karma is moral action-and-consequence in dharma-foundation, the padārtha of motion in Nyāya-Vaiśeṣika, the enjoined ritual act in Mīmāṃsā, and the therapeutic procedure of pañcakarma in Āyurveda; guṇa names the Sāṃkhya triguṇa at its doctrinal source, an unrelated Vaiśeṣika ontological category, and the twenty gurvādi properties of Āyurvedic pharmacology; kaivalya is reached by discrimination alone in Sāṃkhya and by practiced cessation in Yoga. Shared terms carry a school_scope: field, an index contrast note, and reciprocal cross-links — never a single "canonical" definition.

Why this is an engineering problem, not only a cultural one

A mistranslation changes the category a model reasons in, which changes its action space and therefore its output. A wellness bot that reads yoga as exercise recommends the squat rack; an advisor that reads karma as fate gives fatalistic, agency-removing advice. The metaphysical error is a logic bug.

See demos/failure-vs-success.md for side-by-side failure-mode vs success-mode transcripts.

What the corpus does not claim

The bundles are descriptive vocabulary, and the sensitive domains say so on their front page. The Śākta bundle documents that certain practices are initiation-gated without reproducing anything gated (adhikāra discipline, zero how-to). The Āyurveda bundle is not medical advice, contains no preparations or dosages, and does not adjudicate clinical efficacy. The Jyotiṣa bundle makes no predictive claims and does not endorse astrology while documenting its classical vocabulary. Commercially captured terms carry a reception_note: field naming the capture, and as of the July 2026 enrichment pass the foundational and Upaniṣadic bundles state their own does-not-claim boundaries on their index pages. The corrective work lives in the not: fields, concept by concept, not in polemic.

Integrate it (machine-readable is not the same as obeyed)

A not: list in a file is a "do not enter" sign on an unlocked door. The model still has to be told how to read and honor it, and a naive bare prohibition can backfire through negative-prompt leakage. INTEGRATION.md shows how to make the constraint functionally binding: a system-prompt template that injects the positive instead redirect (not a bare ban), a RAG negative-filter pattern, and a lightweight output-check, with honest caveats about what is and is not guaranteed.

Bundles — thirteen live, all on canonical OKF v0.2

Bundle bundle_version Concepts Refs Theme Viz
okf/dharma-foundation/ 0.1.4 25 12 Foundational Sanskrit non-translatable vocabulary graph
okf/yoga-darshana/ 0.2.2 26 6 Patañjali's Yoga Sūtra technical lexicon graph
okf/vedanta-epistemology/ 0.3.3 27 19 Pramāṇa-śāstra — Vedantic epistemology graph
okf/bhakti-marga/ 0.4.3 15 9 The vocabulary of the devotional path graph
okf/dharmic-ethics/ 0.5.3 15 6 Yama–Niyama and the ethical-social vocabulary of dharma graph
okf/upanishadic-core/ 0.6.2 26 16 Core Upaniṣadic vocabulary across the Vedānta schools — mahāvākyas, Ātman–Brahman, self-knowledge graph
okf/cosmology-creation/ 0.7.3 26 12 Vedic & Purāṇic vocabulary of time, cosmos, and manifestation graph
okf/shakta-darshana/ 0.8.2 26 12 Śākta-Tāntric metaphysics of Śakti, Devī, and consciousness-power graph
okf/nyaya-vaisheshika/ 0.9.1 27 10 The science of inference and debate — Nyāya apparatus + Vaiśeṣika realist ontology graph
okf/mimamsa-dharma/ 0.10.1 25 7 Mīmāṃsā ritual hermeneutics — vidhi, apūrva, and language-as-action graph
okf/ayurveda-consciousness/ 0.11.1 26 8 Āyurvedic vocabulary of consciousness, constitution, and health graph
okf/jyotisha-kala/ 0.12.1 26 8 Jyotiṣa vocabulary of time — pañcāṅga, kāla-reckoning, and the sidereal celestial frame graph
okf/sankhya-darshana/ 0.13.3 20 7 Sāṃkhya's dualist causal theory and discriminative path to kaivalya — completes the six classical darśanas graph
Total 310 132 13 bundles spanning the six āstika darśanas + the devotional, ethical, cosmological, medical, and calendrical corpora

Update contract & documented error genealogies

Two consumption surfaces, both first-class, are declared in VERSIONING.md: main is a living vocabulary (concept files are enriched in place — sharper not: fields, added citations, documented genealogies — with bundle_version patch bumps), and release tags are immutable archival snapshots. In-place enrichment waves are logged newest-first in CHANGELOG.md.

Each bundle above carries the tag bundle/<name>/v<bundle_version>okf/yoga-darshana/ at 0.2.2 is bundle/yoga-darshana/v0.2.2. Twenty-nine older repository-wide vX.Y.Z tags remain valid and unchanged. A tag marks a whole-repository state, not a bundle's files, and the legacy tags do not even name the bundle they were cut for: several share a commit. Cite the bundle name together with a commit SHA — see PROFILE.md §7.1, which documents this in full.

GENEALOGIES.md"Where the Errors Came From" — documents the histories of the English mistranslations this corpus corrects: not just that a rendering is wrong, but who introduced it, when, and how it propagated into today's training data (e.g., karma-as-fate from Blavatsky 1889; yoga-as-posture via Vivekananda → Krishnamacharya → Singleton 2010; māyā-as-illusion via Schopenhauer 1818). A correction with a genealogy is harder to dismiss than one with only an assertion. The admission bar is strict: a named source, a date, and a documented propagation chain, or it is excluded. The affected concept files carry a matching Error Genealogy section linking back to the canonical entry.

Concept file format (OKF v0.2)

Each concept file contains:

  • YAML frontmatter: type, title, iast, devanagari, description, darshana, a structured not: (each entry term / why / instead), related:, tags, okf_version, license (and, where a term recurs across darśanas, school_scope:; where public reception distorts a term, reception_note:)
  • ## What It Actually Means (and ## Etymology) — the precise positive definition
  • ## Audience Metaphor — an accessible analogy engineered for AI and general comprehension
  • ## Citations — primary śāstra references, linked into a references/ sub-bundle of first-class type: Reference concepts

Full specification: the base format is Open Knowledge Format v0.2, pinned at commit ad30107. This repository's extensions and stricter rules are documented in PROFILE.md — §2 covers darshana and the structured not: with instead, §2.5 per-claim attribution, §3.2 the required references/ sub-bundle.

Licensing

Two licences, by scope. Nothing in this repository is licensed by inheritance.

Scope Licence File
The corpus — every document under okf/<bundle>/, and the repository-root prose documents CC BY-SA 4.0 (ShareAlike) LICENSE-CONTENT
The codeokf/tools/ (the validator) and okf/tests/ (the regression suite) Apache 2.0, © Dharma OKF Foundation LICENSE.md, scope stated in NOTICE

An earlier revision of this section described the Apache licence as "inherited from the upstream GoogleCloudPlatform fork." That was inaccurate: okf/tools/ has no counterpart upstream, and the validator imports the upstream parser optionally rather than embedding it. The code is this project's own work and is licensed deliberately, not by inheritance.

Contributing

This is the standalone publication home for the Dharma OKF Foundation. Contributions, corrections, and additions that deepen accuracy are welcome via Pull Request. See CONTRIBUTING.md.

About

The first OKF bundle outside data catalogs — demonstrating the format's power for Dharma knowledge preservation. The key innovation is the `not` field in YAML frontmatter — a list of English mistranslations that AI agents should *avoid* when encountering each concept. No other knowledge format carries this signal for AI Agents / models / LLMs.

Topics

Resources

Code of conduct

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages