Skip to content

Latest commit

 

History

History
19 lines (16 loc) · 1006 Bytes

File metadata and controls

19 lines (16 loc) · 1006 Bytes

MITRA-parallel

Parallel corpora of classical Buddhist literature, mined and released by the Dharmamitra project.

  • v2/ — current release (2026): trilingual Sanskrit–Tibetan, Sanskrit–Chinese, and Chinese–Tibetan corpus of 1,693,730 aligned records (2,338,400 segment pairs), mined with a pivot-free embedding-based span-mining pipeline, decontaminated against the published evaluation suite, and exactly deduplicated. See v2/README.md for statistics, record schema, and cleaning details.
  • v1/ — the previous release (1.74M sentence pairs) together with the multilingual retrieval evaluation benchmarks (v1/eval/).

Fine-tuned translation and embedding models are available in the MITRA Qwen3.5 collection on Hugging Face. The parallel data can be explored interactively at dharmamitra.org/db.