Repository navigation
Add JSON Schema constrained decoding for local inference - #564
Conversation
|
Review hardening is pushed at 243e83f. This PR remains draft pending isolated Release app/model/cache proof. SOURCE EVIDENCE: LIVE EVIDENCE: local Release focused engine run Review found and reproduced a donor bypass: Remaining: final app tests/build, real Chat/Responses streaming and nonstreaming, explicit truncation errors, changed-schema cache reuse, ordinary chat/tool regression, and app dependency pin/CI. No merge or release claim. |
|
Qualification update (not a release): SOURCE EVIDENCE: engine LIVE EVIDENCE: isolated Release app (
40 focused engine tests and35 focused app tests passed. Schema live rows reported95.3–102.1tok/s on short outputs. No sustained performance/model-family parity claim. Limits: no qualification for active reasoning, schema+tool envelope, speculative schema generation, other tokenizer families, arbitrary JSON Schema features, media or disk-full eviction. Unsupported requests fail explicitly. This supplies the Architect inference seam, not its workflow graph/UI. Canceled/length-capped partial output is not a success and has no fabricated TPS. Complete raw requests/responses, validator version, source closure, failed first-run evidence and cache receipts are retained locally under |
Final PR564 lint dispositionBASELINE-ONLY FAILURE for the supplied final CI diff. No PR-introduced line violations remain. This classifies the failure; it does not turn the failed GitHub check green.
The previously identified PR-introduced formatting changes have been corrected. Remaining reported transformations concern base content and unrelated files. Do not apply an826-file sweep or claim that this receipt changes repository branch-protection policy; merge disposition belongs to the root/repository owner with this evidence. Machine-readable classification: |
Local inference callers can supply GenerateParameters.jsonSchema to constrain generated tokens to a documented JSON Schema subset. XGrammar masks logits before the existing sampler. No prompt coercion, sampler substitution, closing bias or output repair is used.
Exact ByteLevel tokenizer metadata, request-local grammar state for AR/batch generation, explicit incomplete/cancelled errors and literal JSON preservation are included. Schema requests use AR until speculative matcher rollback is qualified; ordinary requests retain their existing path. Vendored licenses and pinned provenance are retained. See docs/STRUCTURED_OUTPUT.md.
Validation: 40 focused engine tests and35 focused app tests passed. A real isolated Release Osaurus app on local Raptor0.6.1 JANG_6M passed15 API checks: Chat/Responses stream and nonstream, nested schemas, Unicode/literal markers, changed-schema and multi-turn SSD reuse, explicit incomplete/unsupported errors, ordinary chat/tool continuation, disconnect drain and fresh-schema recovery. Raw results were independently validated with jsonschema4.25.1. Responses terminal output equals streamed output. Actual streaming telemetry restored659/662 tokens from disk; paged RAM remained off, KV fp16, TurboQuant0. Visible app model selection/native reasoning toggle/ordinary chat also passed.
Live binary used engine6e41dc75/appdb5e11ab. Final follow-up changes only formatting/documentation and the app pin; runtime Swift non-whitespace text is identical, with a single optional Package array comma. Local syntax checks passed; no unmeasured performance claim. Short schema rows reported95.3–102.1tok/s; this is correctness telemetry, not sustained benchmarking.
Limits: supported ByteLevel tokenizer and documented schema subset only; native reasoning explicitly off. Schema-plus-tools, active reasoning, other tokenizer families and speculative schema decoding reject explicitly. Architect workflow UI/graph is outside this PR. CI remains pending; no release.