Skip to content

RebuildIndex and migration truncate UTF-8 descriptions by byte, producing mojibake #10

Description

@mashenjun

Problem

Generated OKF indexes can contain replacement characters (�) in Chinese descriptions. The issue is reproducible after compiling/migrating a wiki with CJK text.

Example from a generated catstock wiki:

- [ai-and-tech-sector.md](semantic/ai-and-tech-sector.md) — A股市场存在严重的AI风格抱团。传统消费主题公募基金为追业绩,存在风格�...

The migrated page frontmatter can also preserve broken UTF-8/truncated text:

description: "A股市场存在严重的AI风格抱团。传统消费主题公募基金为追业绩,存在风格偏移——偷偷挪仓位买AI。监管传闻将查处此\xe7\xb1..."

Root Cause

Two code paths truncate strings by byte length instead of rune/UTF-8 character boundaries:

  1. internal/storage/fs.go:572
if len(line) > 100 {
    return line[:100] + "..."
}
  1. internal/okf/migration.go:398
if len(trimmed) > 160 {
    return trimmed[:160] + "..."
}

In Go, len(string) and s[:n] operate on bytes. For multi-byte UTF-8 text, slicing at an arbitrary byte offset can split a character and produce �.

Related Index Quality Issue

RebuildIndex currently extracts descriptions from the first non-empty body line instead of preferring OKF frontmatter description:

  • internal/storage/fs.go:393 calls extractPageDescription(path).
  • extractPageDescription scans raw lines and skips only blanks/comments/headings.

This can produce low-quality index entries, for example when the first body content is a fenced code block:

- [aluminum-supply-chain.md](semantic/commodities/aluminum-supply-chain.md) — ```

Expected Behavior

  1. Description truncation should preserve valid UTF-8.
  2. RebuildIndex should prefer YAML frontmatter description when present.
  3. Fallback body extraction should skip YAML frontmatter and unsuitable lines such as code fence delimiters.
  4. Add regression tests with Chinese text that crosses the truncation boundary.

Suggested Fix

  1. Add a helper such as truncateRunes(s string, max int) string.
  2. Use it in extractPageDescription, migration description extraction, and other user-visible truncation paths.
  3. Teach extractPageDescription to parse or at least read OKF frontmatter description first.
  4. Add tests for CJK text and code-fence-first pages.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions