Repository navigation
Conversation
Points the app at vmlx-swift 371f5c40: multi-row bit-exact BF16-affine verify kernels, lane matmul with safe tiling and the DFlash2 width chooser, native-MTP copy drafts, and the Qwen4 PLE page cache. Proof build only; not for release.
Pins vmlx-swift perf/claude-swift-port-oct6 @ a4f99a73 (six pin sites).
Settings > Speculative Decoding now has three modes: Off (AR), Default
and On (Adaptive). Default is the engine's `.familyDefault`: Adaptive
native MTP for Qwen3.8 Flash-Next, and a Qwen 27B bundle's own dflash2/
drafter. A fresh install shows the chat picker's Native MTP row as On
for Flash-Next bundles, and the user can switch it off. The picker and
the phone snapshot show the per-bundle effective state.
ModelRuntime passes the bundle directory to drafter selection, so a
bundled dflash2/ drafter is found without a folder pick.
The Safe Auto materialized-load memory check is advisory: it logs a
warning instead of refusing. It refused Flash-Next JANG_4S on a 128 GB
Mac ("require ~92 GiB, only ~53 GiB available"); the model loads and
runs at 55-115 tok/s. Strict mode keeps its explicit refusal.
Pins vmlx-swift perf/claude-swift-port-oct6 @ 0e26bcb5 (six pin sites). The engine now loads dense Qwen3.5 JANGH bundles (Qwen3.8-27B-JANGH2: jangtq2 2-bit MLP banks). Measured in RunBench on max2: AR 30.6 tok/s prose, DFlash2 28.6 / 57.7 / 135.0 / 141.0 tok/s (prose / code / easy code / easy prose).
|
Current live qualification found a numerical blocker; this PR remains draft and unmerged. Source tested: Osaurus ffb03f2, vmlx-swift fd9d925c, MLX fork5aba1efd. The later engine91726781 change is whitespace only; it is not claimed as a new app build. Live isolated Release app, dense Qwen3.8-27B-JANGH2:
Controlled app HTTP diagnostic, same38-token prose, explicit greedy, fresh cache namespace, AR/Adaptive/Adaptive/AR:29.01 /29.39 /27.37 /26.44 tok/s. Clocks differed, so this does not establish a speedup. AR answers matched each other; Adaptive answers differed from AR and from each other. Holding DFlash width at its trained8 produced identical outputs twice at27.61/27.46 tok/s, but still differed from AR. This is a diagnostic, not a proposed fixed-depth product change. Causal teacher-forced test on the actual bundle: from the same prefilled snapshot, target row0 of widths5/8/16 differs from a one-token call before acceptance or rollback, both with and without LaneQMM and in eager/staged modes (12 comparisons, max absolute logit difference0.125). Lane installation also changes cold-prefill logits. Next work isolates the first differing operator and tests an arithmetic repair. Changing the controller alone is insufficient. Local raw artifacts: /Users/eric/vmlx-private-evidence/qwen-spec-review-20261007/{app-live002.log,ui-dense-prose.png,ui-dense-followup.png,dense-prose-abba001,dense-fixed8-001,review/prefill-partition/dense001.json,review/DFLASH-DENSE-ROW-CAUSAL.md}. These local artifacts are not publicly accessible CI attachments. Other quants, complete media/cache/lifecycle coverage, and final app proof remain pending. |
|
Correction to the preceding qualification comment: I applied the wrong numerical contract to dense27B DFlash2. The supplied Swift handoff READ-FIRST §5c explicitly documents NAX/Lane verification as speed-first and not bit-identical to AR. The strict AR row-exact gate applies to Flash-Next native MTP. The measured S1-versus-wide differences are real, but they are not evidence of a new regression or a standalone merge blocker under that documented dense27B contract. No production kernel or adaptive-width change was made. The fixed8 run was diagnostic only. Review resumes against the correct references: fast QMV vs its replaced S1 kernel; small NAX tile vs its replaced large tile; verification/acceptance and cache commit within the chosen execution path; default selection, live app speeds, cold prefill, tool continuation, and media fallback. Dense27B controlled prose27–29 tok/s is consistent with the handoff28.6 tok/s. The23 tok/s visible-composer run had4117 prompt tokens and is not a matched regression result. The PR remains unmerged pending the remaining integration/CI proofs, not because dense DFlash differs from plain AR. |
A working-set estimate that meets the Safe Auto budget left a 0-byte allocator headroom, so native-MTP and DFlash 2 generations ran with no freed-buffer reuse: every decode step re-allocated its intermediates and paid a Metal residency commit per buffer. The working set already prices a >= 2 GiB scratch floor for exactly these buffers, so the remaining-budget clamp may shrink the pool but never below that floor. Measured live on Qwen3.8 Flash-Next JANG_4S under Safe Auto: 31.4 ms per target forward with the starved pool vs 21.6 ms with a working pool.
…e, DFlash 2 AR width
Live app proof — osaurus dev build pinned to this engine headSetup.
AR references on the same machine: 27B 4D ~23, 27B JANGH2 ~29, Flash-Next ~50–55 tok/s. "Cold" = the first request App-side defect found and fixed in osaurus (
|
Selecting a supported Qwen bundle now exposes speculative On (Adaptive) / Off (AR) before model loading. Flash-Next bundles with real native heads default to Adaptive; Qwen 27B bundles with compatible DFlash2 artifacts use that drafter. The included eval and benchmark routes inherit the same server policy unless explicitly overridden.
Explicit Off disables both native and external drafting. Reset restores the bundle default. Migration changes only provenance-owned untouched defaults, preserving user choices. Drafter weights are released on target unload unless another resident target shares the path; quit preserves its existing no-GPU-free behavior. Guides and picker help describe both draft mechanisms.
Depends on osaurus-ai/vmlx-swift#565; currently pins its reviewed head 78a3539972c881b929e7ecb7e441de87d268e356. The underlying MLX loader dependency is merged in osaurus-ai/mlx#21. Schema-constrained generation remains on its verified AR path.
Evidence so far:
Remaining merge gates: final-head engine merge/pin, live generation/tool/cache/media/unload proof, final UI reset/persistence check, included eval/bench activation, and CI review. This stays a draft until those gates are recorded. No release or new model-speed claim is included.