Skip to content

feat: Add Swahili and Yoruba benchmark corpora to evaluate agglutinative morphology #1

Description

@umran666

Description

UniqToken currently evaluates English, Python code, Indic (Hindi/Telugu/Tamil/Bengali), CJK, Arabic, and European agglutinative languages (Turkish, Finnish). We want to expand our empirical benchmark suite to include African agglutinative and morphologically rich languages like Swahili and Yoruba.

Scope of Work

  1. Add raw text evaluation samples for Swahili and Yoruba to BENCHMARK_CORPORA in benchmarks/benchmark_suite.py.
  2. Run python benchmarks/benchmark_suite.py to calculate Bytes/Token and Fertility metrics.
  3. Verify zero fallback regressions.

Getting Started

  • Files to edit: benchmarks/benchmark_suite.py
  • Test command: python benchmarks/benchmark_suite.py

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions