Skip to content

Expose bundle-aware speculative defaults before model load - #3029

Draft
jjang-ai wants to merge 26 commits into
mainfrom
review/qwen-spec-defaults-oct7
Draft

jjang-ai wants to merge 26 commits into
mainfrom
review/qwen-spec-defaults-oct7

Conversation

@jjang-ai

@jjang-ai jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Selecting a supported Qwen bundle now exposes speculative On (Adaptive) / Off (AR) before model loading. Flash-Next bundles with real native heads default to Adaptive; Qwen 27B bundles with compatible DFlash2 artifacts use that drafter. The included eval and benchmark routes inherit the same server policy unless explicitly overridden.

Explicit Off disables both native and external drafting. Reset restores the bundle default. Migration changes only provenance-owned untouched defaults, preserving user choices. Drafter weights are released on target unload unless another resident target shares the path; quit preserves its existing no-GPU-free behavior. Guides and picker help describe both draft mechanisms.

Depends on osaurus-ai/vmlx-swift#565; currently pins its reviewed head 78a3539972c881b929e7ecb7e441de87d268e356. The underlying MLX loader dependency is merged in osaurus-ai/mlx#21. Schema-constrained generation remains on its verified AR path.

Evidence so far:

  • Fresh isolated Release app builds succeeded before final fixes.
  • Live intermediate picker: dense 27B JANGH2 Adaptive visible before loading, explicit Off survives reopening, headless Flash 1L has no speculative row. The live test caught Reset writing Off; that route is corrected and its persisted API regression passes.
  • Focused app tests: 8 XCTest and 60 Swift Testing methods passed, including Off/default migration, admission, detection, sampler parity and reset persistence. Additional lifecycle tests and final Release build are running.

Remaining merge gates: final-head engine merge/pin, live generation/tool/cache/media/unload proof, final UI reset/persistence check, included eval/bench activation, and CI review. This stays a draft until those gates are recorded. No release or new model-speed claim is included.

Eric added 21 commits October 6, 2026 07:04
Points the app at vmlx-swift 371f5c40: multi-row bit-exact BF16-affine
verify kernels, lane matmul with safe tiling and the DFlash2 width
chooser, native-MTP copy drafts, and the Qwen4 PLE page cache. Proof
build only; not for release.
Pins vmlx-swift perf/claude-swift-port-oct6 @ a4f99a73 (six pin sites).

Settings > Speculative Decoding now has three modes: Off (AR), Default
and On (Adaptive). Default is the engine's `.familyDefault`: Adaptive
native MTP for Qwen3.8 Flash-Next, and a Qwen 27B bundle's own dflash2/
drafter. A fresh install shows the chat picker's Native MTP row as On
for Flash-Next bundles, and the user can switch it off. The picker and
the phone snapshot show the per-bundle effective state.

ModelRuntime passes the bundle directory to drafter selection, so a
bundled dflash2/ drafter is found without a folder pick.

The Safe Auto materialized-load memory check is advisory: it logs a
warning instead of refusing. It refused Flash-Next JANG_4S on a 128 GB
Mac ("require ~92 GiB, only ~53 GiB available"); the model loads and
runs at 55-115 tok/s. Strict mode keeps its explicit refusal.
Pins vmlx-swift perf/claude-swift-port-oct6 @ 0e26bcb5 (six pin sites).

The engine now loads dense Qwen3.5 JANGH bundles (Qwen3.8-27B-JANGH2:
jangtq2 2-bit MLP banks). Measured in RunBench on max2: AR 30.6 tok/s
prose, DFlash2 28.6 / 57.7 / 135.0 / 141.0 tok/s (prose / code / easy
code / easy prose).
@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Current live qualification found a numerical blocker; this PR remains draft and unmerged.

Source tested: Osaurus ffb03f2, vmlx-swift fd9d925c, MLX fork5aba1efd. The later engine91726781 change is whitespace only; it is not claimed as a new app build.

Live isolated Release app, dense Qwen3.8-27B-JANGH2:

  • Visible composer prose:4117 prompt tokens,302 output tokens,23.0 tok/s; cold prefill12.237s (336.4 prompt tok/s),106 actual DFlash verify cycles, natural completion.
  • Visible follow-up:4507 prompt tokens,26 output tokens,19.2 tok/s, L2 disk hit1/stores2, natural completion.
  • These two UI rows are not a matched AR speed comparison. The attempted UI AR control delegated a helper and is excluded.

Controlled app HTTP diagnostic, same38-token prose, explicit greedy, fresh cache namespace, AR/Adaptive/Adaptive/AR:29.01 /29.39 /27.37 /26.44 tok/s. Clocks differed, so this does not establish a speedup. AR answers matched each other; Adaptive answers differed from AR and from each other.

Holding DFlash width at its trained8 produced identical outputs twice at27.61/27.46 tok/s, but still differed from AR. This is a diagnostic, not a proposed fixed-depth product change.

Causal teacher-forced test on the actual bundle: from the same prefilled snapshot, target row0 of widths5/8/16 differs from a one-token call before acceptance or rollback, both with and without LaneQMM and in eager/staged modes (12 comparisons, max absolute logit difference0.125). Lane installation also changes cold-prefill logits. Next work isolates the first differing operator and tests an arithmetic repair. Changing the controller alone is insufficient.

Local raw artifacts: /Users/eric/vmlx-private-evidence/qwen-spec-review-20261007/{app-live002.log,ui-dense-prose.png,ui-dense-followup.png,dense-prose-abba001,dense-fixed8-001,review/prefill-partition/dense001.json,review/DFLASH-DENSE-ROW-CAUSAL.md}. These local artifacts are not publicly accessible CI attachments. Other quants, complete media/cache/lifecycle coverage, and final app proof remain pending.

@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Correction to the preceding qualification comment: I applied the wrong numerical contract to dense27B DFlash2. The supplied Swift handoff READ-FIRST §5c explicitly documents NAX/Lane verification as speed-first and not bit-identical to AR. The strict AR row-exact gate applies to Flash-Next native MTP.

The measured S1-versus-wide differences are real, but they are not evidence of a new regression or a standalone merge blocker under that documented dense27B contract. No production kernel or adaptive-width change was made. The fixed8 run was diagnostic only.

Review resumes against the correct references: fast QMV vs its replaced S1 kernel; small NAX tile vs its replaced large tile; verification/acceptance and cache commit within the chosen execution path; default selection, live app speeds, cold prefill, tool continuation, and media fallback. Dense27B controlled prose27–29 tok/s is consistent with the handoff28.6 tok/s. The23 tok/s visible-composer run had4117 prompt tokens and is not a matched regression result.

The PR remains unmerged pending the remaining integration/CI proofs, not because dense DFlash differs from plain AR.

Eric added 4 commits October 7, 2026 04:25
A working-set estimate that meets the Safe Auto budget left a 0-byte
allocator headroom, so native-MTP and DFlash 2 generations ran with no
freed-buffer reuse: every decode step re-allocated its intermediates and
paid a Metal residency commit per buffer. The working set already prices
a >= 2 GiB scratch floor for exactly these buffers, so the remaining-budget
clamp may shrink the pool but never below that floor.

Measured live on Qwen3.8 Flash-Next JANG_4S under Safe Auto: 31.4 ms per
target forward with the starved pool vs 21.6 ms with a working pool.
@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Live app proof — osaurus dev build pinned to this engine head

Setup.

  • osaurus review/qwen-spec-defaults-oct7 (53231ce9), pinned to vmlx-swift a5a0689c (docs-only over the proven e40f9fa4).
  • Isolated profile, default Safe Auto memory profile.
  • Plain chat agent (tools and memory off), effort None, bundle sampler defaults.
  • Real composer: three connected turns per bundle (code → pasted image → text follow-up).
  • Speed = wall decode tok/s from the app's step log. Speculation counters come from the engine log.
Bundle T1 code T2 image (answer) T3 follow-up after image (TTFT) Speculation
Qwen3.8-27B JANG_4D 103.0 44.3 (red 7 ✓) 61.3 (0.89 s, 1,202 / 1,407 restored) DFlash 2 trees on all turns
Qwen3.8-27B JANGH2 74.1 40.5 (blue 4 ✓) 35.2 (1.6 s, 801 / 995 restored) DFlash 2 trees on all turns
Flash-Next 4S 71.0 46.6 (green 2 ✓) fully cached (0.76 s) native MTP, 2.1–3.8 / verify
Flash-Next 4M 52.8 56.0 (red 7 ✓) 57.0 (0.65 s) native MTP
Flash-Next 2L 39.0 (cold) 52.2 (blue 4 ✓) 72.0 (0.73 s) native MTP
Flash-Next 6S 37.0 (cold) 39.7 (green 2 ✓) 75.1 (0.84 s) native MTP
Flash-Next 1L 42.6 (cold) 52.9 (red 7) 58.9 (0.64 s) AR (no head, by design)
Allosaurus JANGH2 43.4 (cold) 46.0 (blue 4) 83.8 (0.64 s) native MTP
Flash-Next JANGH4 47.7 (cold) 43.1 (green 2 ✓) 69.7 (1.1 s) native MTP

AR references on the same machine: 27B 4D ~23, 27B JANGH2 ~29, Flash-Next ~50–55 tok/s. "Cold" = the first request
after loading a large bundle (expert pages and kernels warming). Warm turns on the same bundles run 52–84.

App-side defect found and fixed in osaurus (492c66f8)

Under Safe Auto, the remaining-budget clamp could set MLX's freed-buffer pool to 0 bytes for a large model. Every decode
step then re-allocated its intermediates, with a Metal residency commit per buffer.

  • Flash-Next 4S: 31.4 → 20.5 ms per target forward, live, same build and prompt.
  • The clamp now keeps the scratch floor the working-set estimate already admits.

Open, not claimed

  • 27B JANGH2 multi-row verify is ~3× its 1-row cost. The codebook MLP falls to the generic tile above one row, so
    prose speedup on that bundle is small. A dedicated verify kernel is in progress.
  • The first request after a large-bundle load is slow.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant