Skip to content

Add JSON Schema constrained decoding for local inference - #564

Merged
jjang-ai merged 6 commits into
mainfrom
feat/schema-constrained-decoding-oct6
Oct 7, 2026
Merged

jjang-ai merged 6 commits into
mainfrom
feat/schema-constrained-decoding-oct6

Conversation

@jjang-ai

@jjang-ai jjang-ai commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Local inference callers can supply GenerateParameters.jsonSchema to constrain generated tokens to a documented JSON Schema subset. XGrammar masks logits before the existing sampler. No prompt coercion, sampler substitution, closing bias or output repair is used.

Exact ByteLevel tokenizer metadata, request-local grammar state for AR/batch generation, explicit incomplete/cancelled errors and literal JSON preservation are included. Schema requests use AR until speculative matcher rollback is qualified; ordinary requests retain their existing path. Vendored licenses and pinned provenance are retained. See docs/STRUCTURED_OUTPUT.md.

Validation: 40 focused engine tests and35 focused app tests passed. A real isolated Release Osaurus app on local Raptor0.6.1 JANG_6M passed15 API checks: Chat/Responses stream and nonstream, nested schemas, Unicode/literal markers, changed-schema and multi-turn SSD reuse, explicit incomplete/unsupported errors, ordinary chat/tool continuation, disconnect drain and fresh-schema recovery. Raw results were independently validated with jsonschema4.25.1. Responses terminal output equals streamed output. Actual streaming telemetry restored659/662 tokens from disk; paged RAM remained off, KV fp16, TurboQuant0. Visible app model selection/native reasoning toggle/ordinary chat also passed.

Live binary used engine6e41dc75/appdb5e11ab. Final follow-up changes only formatting/documentation and the app pin; runtime Swift non-whitespace text is identical, with a single optional Package array comma. Local syntax checks passed; no unmeasured performance claim. Short schema rows reported95.3–102.1tok/s; this is correctness telemetry, not sustained benchmarking.

Limits: supported ByteLevel tokenizer and documented schema subset only; native reasoning explicitly off. Schema-plus-tools, active reasoning, other tokenizer families and speculative schema decoding reject explicitly. Architect workflow UI/graph is outside this PR. CI remains pending; no release.

@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Review hardening is pushed at 243e83f. This PR remains draft pending isolated Release app/model/cache proof.

SOURCE EVIDENCE: Libraries/MLXLMCommon/StructuredOutput/JSONSchemaGrammar.swift rejects unsafe numeric spellings and open named-property schemas; Libraries/MLXCXGrammar/xgrammar/cpp/json_schema_converter.cc JSON-escapes property keys before EBNF quoting.

LIVE EVIDENCE: local Release focused engine run engine-proof/build012.log passed 31 XCTest and 7 Swift Testing cases (38 total). These include actual tiny-model generation and BatchEngine tests, not real-model/app proof.

Review found and reproduced a donor bypass: {"answer":true,"answer":null} was accepted for an open object with required boolean answer, while the independent JSON Schema validator rejected it. Named-property schemas now require explicit additionalProperties:false; unsupported semantics are never silently rewritten. Further regressions cover Unicode/control/escaped property names and raw numeric constants before parser rounding.

Remaining: final app tests/build, real Chat/Responses streaming and nonstreaming, explicit truncation errors, changed-schema cache reuse, ordinary chat/tool regression, and app dependency pin/CI. No merge or release claim.

@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Qualification update (not a release):

SOURCE EVIDENCE: engine 648f328ca1c97afc3f51af5da91665fef632c6a5, app f3026dace64588b5d475d184e38d42ae32fc47d4. Runtime implementation is in Libraries/MLXLMCommon/StructuredOutput/, Evaluate.swift, BatchEngine.swift; app admission/wiring in RequestValidator.swift, HTTPHandler.swift, OpenResponsesAPI.swift, ResponseWriters.swift and the MLX adapter. The app pins that engine revision. Final follow-up contains only formatting/docs/pin changes: runtime files have identical non-whitespace text to the tested engine; Package formatting adds one optional trailing comma. Receipt: review/FINAL-FORMATTING-SEMANTIC-EQUIVALENCE.json.

LIVE EVIDENCE: isolated Release app (app-release004.log, executable SHA256 464f076cc87f41bae23ec19205b1285cdec1f220094a2d8311e5e3cb09dd13e0) ran15 cases against a real local Raptor0.6.1 JANG_6M bundle, native reasoning explicitly off and sampler unchanged. live-run002/summary.json has zero failures. Independent raw-byte review is LIVE-RESULTS002.md.

  • Chat/Responses stream and nonstream: nested objects/arrays/enum/const valid under independent jsonschema4.25.1; natural completion. Unicode and literal thinking/tool markers preserved.
  • Responses SSE deltas equal terminal response content; completed items/IDs/order retained.
  • Same-prefix changed schema and growing multi-turn: separate matcher state, valid new output; SSD hits each incremented. Actual streaming event restored659/662 tokens from disk.
  • Ordinary native tool round trip: lookup_fixture({id:alpha}) then answer17. SSD stores26→28 before follow-up, then31; follow-up hit counter10→11.
  • Length1: explicit incomplete-schema stream error, no success. Unsupported pattern/minLength: HTTP400 before SSE.
  • Client disconnect after first content: active slot1→0, pending0; new schema naturally completed afterward.
  • Effective KV fp16;9 ordinary+27 rotating layers; paged RAM disabled; TurboQuant0. No cache precision change.
  • Visible app: selected Raptor through model picker, used native reasoning Off, ordinary chat returned4. Response inspector showed110.6tok/s for1token; this is UI correctness evidence, not a speed benchmark.

40 focused engine tests and35 focused app tests passed. Schema live rows reported95.3–102.1tok/s on short outputs. No sustained performance/model-family parity claim.

Limits: no qualification for active reasoning, schema+tool envelope, speculative schema generation, other tokenizer families, arbitrary JSON Schema features, media or disk-full eviction. Unsupported requests fail explicitly. This supplies the Architect inference seam, not its workflow graph/UI. Canceled/length-capped partial output is not a success and has no fabricated TPS.

Complete raw requests/responses, validator version, source closure, failed first-run evidence and cache receipts are retained locally under /Users/eric/vmlx-private-evidence/structured-output-20261006/. First-run whitespace and Responses receipt failures were fixed and rerun; no output repair or forced closure. CI remains pending; both PRs are draft until merge gates are settled.

@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Final PR564 lint disposition

BASELINE-ONLY FAILURE for the supplied final CI diff. No PR-introduced line violations remain. This classifies the failure; it does not turn the failed GitHub check green.

  • PR head: 648f328ca1c97afc3f51af5da91665fef632c6a5.
  • Base: a46266ddba66667f24c12797c8ecbc829b0266b5.
  • CI merge: dce5027 (engine-ci-lint-final.log188).
  • Tool: swift-format604.0.0 (log323); swift-format and cmake-format both modified files (log555–559), causing exit1 (log148531).
  • Formatter diff spans826files:822 are unchanged by this PR, including CMakeLists.txt.
  • Four PR-touched files also contain inherited formatting debt: BatchEngine.swift615 removed lines, Evaluate.swift495, Tokenizer.swift61, Package.swift11. Zero formatter-removed lines overlap PR-added lines in all four.
  • Independently checked added-only formatter groups: zero insertion anchors at or adjacent to PR-added lines. This closes the addition-only blind spot of removed-line intersection alone.
  • All added schema Swift implementation/tests remain absent from the formatting diff.

The previously identified PR-introduced formatting changes have been corrected. Remaining reported transformations concern base content and unrelated files. Do not apply an826-file sweep or claim that this receipt changes repository branch-protection policy; merge disposition belongs to the root/repository owner with this evidence.

Machine-readable classification: engine-ci-lint-final-classification.json; reproducible parser: classify-lint-final.py. Input log: ../engine-ci-lint-final.log. Assessment is read-only; no source edits, builds or models were run.

@jjang-ai
jjang-ai marked this pull request as ready for review October 7, 2026 02:17
@jjang-ai
jjang-ai merged commit 5fed181 into main Oct 7, 2026
6 of 9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant