Problem
Generated OKF indexes can contain replacement characters (�) in Chinese descriptions. The issue is reproducible after compiling/migrating a wiki with CJK text.
Example from a generated catstock wiki:
- [ai-and-tech-sector.md](semantic/ai-and-tech-sector.md) — A股市场存在严重的AI风格抱团。传统消费主题公募基金为追业绩,存在风格�...
The migrated page frontmatter can also preserve broken UTF-8/truncated text:
description: "A股市场存在严重的AI风格抱团。传统消费主题公募基金为追业绩,存在风格偏移——偷偷挪仓位买AI。监管传闻将查处此\xe7\xb1..."
Root Cause
Two code paths truncate strings by byte length instead of rune/UTF-8 character boundaries:
internal/storage/fs.go:572
if len(line) > 100 {
return line[:100] + "..."
}
internal/okf/migration.go:398
if len(trimmed) > 160 {
return trimmed[:160] + "..."
}
In Go, len(string) and s[:n] operate on bytes. For multi-byte UTF-8 text, slicing at an arbitrary byte offset can split a character and produce �.
Related Index Quality Issue
RebuildIndex currently extracts descriptions from the first non-empty body line instead of preferring OKF frontmatter description:
internal/storage/fs.go:393 calls extractPageDescription(path).
extractPageDescription scans raw lines and skips only blanks/comments/headings.
This can produce low-quality index entries, for example when the first body content is a fenced code block:
- [aluminum-supply-chain.md](semantic/commodities/aluminum-supply-chain.md) — ```
Expected Behavior
- Description truncation should preserve valid UTF-8.
RebuildIndex should prefer YAML frontmatter description when present.
- Fallback body extraction should skip YAML frontmatter and unsuitable lines such as code fence delimiters.
- Add regression tests with Chinese text that crosses the truncation boundary.
Suggested Fix
- Add a helper such as
truncateRunes(s string, max int) string.
- Use it in
extractPageDescription, migration description extraction, and other user-visible truncation paths.
- Teach
extractPageDescription to parse or at least read OKF frontmatter description first.
- Add tests for CJK text and code-fence-first pages.
Problem
Generated OKF indexes can contain replacement characters (
�) in Chinese descriptions. The issue is reproducible after compiling/migrating a wiki with CJK text.Example from a generated catstock wiki:
The migrated page frontmatter can also preserve broken UTF-8/truncated text:
Root Cause
Two code paths truncate strings by byte length instead of rune/UTF-8 character boundaries:
internal/storage/fs.go:572internal/okf/migration.go:398In Go,
len(string)ands[:n]operate on bytes. For multi-byte UTF-8 text, slicing at an arbitrary byte offset can split a character and produce�.Related Index Quality Issue
RebuildIndexcurrently extracts descriptions from the first non-empty body line instead of preferring OKF frontmatterdescription:internal/storage/fs.go:393callsextractPageDescription(path).extractPageDescriptionscans raw lines and skips only blanks/comments/headings.This can produce low-quality index entries, for example when the first body content is a fenced code block:
Expected Behavior
RebuildIndexshould prefer YAML frontmatterdescriptionwhen present.Suggested Fix
truncateRunes(s string, max int) string.extractPageDescription, migration description extraction, and other user-visible truncation paths.extractPageDescriptionto parse or at least read OKF frontmatterdescriptionfirst.