feat(analyzer): add Korean-language context terms to KrRrnRecognizer - #2213
Open
juno-junho wants to merge 2 commits into
Open
feat(analyzer): add Korean-language context terms to KrRrnRecognizer#2213juno-junho wants to merge 2 commits into
juno-junho wants to merge 2 commits into
Conversation
juno-junho
force-pushed
the
feat/kr-rrn-korean-context
branch
from
August 18, 2026 02:56
a998981 to
acf4981
Compare
…er placement Per the recognizer instructions (data-privacy-stack#2211): each of the four new Korean terms is shown to raise the score through the enhancer's explicit context argument (text-derived context would need a Korean NLP model the test environment does not ship), and Korean context text before and after the number is covered at the pattern level.
Contributor
Author
|
Updated per the recognizer instructions from #2211: the description now states the detection-behavior change for existing users, and the tests show each of the four Korean terms raising the score through the enhancer, plus placement coverage on both sides of the number. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Change Description
Add Korean-language context terms (
주민등록번호,주민번호,신분증,본인인증) toKrRrnRecognizer.CONTEXT, with a regression test.The context list was English-only, so
LemmaContextAwareEnhancercould never boost KR_RRN confidence on Korean-language text, which is the primary language Korean RRNs appear in. Every other Korean recognizer (driver license, passport, FRN, BRN) already ships Korean context terms. The last two terms are broader identity-related words, so feel free to ask me to drop them if you prefer the conservative set. All four come from a production Korean PII-masking deployment.pytest tests/test_kr_rrn_recognizer.pyandruff checkpass.Issue reference
Fixes #2212
Checklist
Update: aligned with the recognizer instructions from #2211
This changes detection results for existing users: on Korean-language text the context enhancer can now raise KR_RRN scores where it previously never fired. Patterns and base scores are unchanged, so nothing that matched before stops matching.
Tests now show each of the four terms raising the score through the enhancer (0.5 to 0.85 via the explicit
contextargument, the routetest_context_supportuses, since text-derived context needs a Korean NLP model the test environment does not ship), plus Korean context text before and after the number at the pattern level.