Local inference
Local LLMs served on the desktop's RTX 3060. Diffusion image generation on the MacBook's M1.
Qwen3.5:9B
huihui-ai abliterated build — the Qwen3.5‑9B base model with abliteration applied — independently‑verified genuine GGUF, served via the 16‑way raw‑llama.cpp bypass, 2026‑08‑23 onward — 12 of 12 tasks done.
Same abliterated build, thinking mode forced on instead of the non‑thinking default used elsewhere on this page.3 Five results so far (ifeval, mmlu_pro, bbh_cot_zeroshot, cruxeval_input, cruxeval_output) from a 12‑task rerun still in progress — most rows don't have one yet.
Unsloth official Qwen3.5‑9B weights, same harness/prompting, run 2026‑08‑22–23 (two tasks via the raw‑llama.cpp bypass).
Same official weights as the row above, thinking mode forced on — the fourth of the four core settings this page tracks (abliterated/official × thinking/non‑thinking).5 One result so far (mmlu_pro), meant to fill in for every benchmark over time like the abliterated, thinking row above it.
Numbers from Qwen's own model card, quoted as-is and not run by us — except cruxeval_input/cruxeval_output, where no Qwen number exists at all and this row instead shows a same‑scale substitute model (labeled by name, not "Qwen3.5‑9B").
One table per benchmark, all scores as percentages. Each Δ is labeled with what it's measured against, since the comparison isn't the same for every row: Qwen 3.5:9b's Δ is versus the plain Abliterated row (does swapping to official weights change the score), while each thinking row's Δ is versus its own non‑thinking baseline directly above it (does thinking mode change the score) — two independent comparisons, not one chained run down the table. Every Δ is computed straight from the two rounded scores shown, so it always matches subtracting them by hand. The posted row is never delta'd — see footnote 6 for why that comparison needs its own caveats. A task's posted row is omitted entirely when no public number (or substitute) exists for it, rather than shown as a blank — see footnote 1. Some backlog tasks below have only the abliterated row filled in; a qwen comparison doesn't exist yet for those.
grade-school math word problems
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 82.5strict 82.6 | — |
| Abliterated, thinking (ours) | 87.3 ± 1.2strict 84.9 | ▲ 4.8vs non‑think |
| Qwen 3.5:9b (ours) | 85.7 ± 1.0 | ▲ 3.2vs abliterated |
| Qwen 3.5:9b, thinking (ours) | not reported | — |
diverse reasoning — logic, algorithms, language
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 45.6 ± 0.5 | — |
| Abliterated, thinking (ours) | 79.9 ± 1.13 | ▲ 34.3vs non‑think |
| Qwen 3.5:9b (ours) | 65.8 ± 0.5 | ▲ 20.2vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 89.1 ± 0.93 | ▲ 23.3vs non‑think |
broad academic knowledge, 57 subjects
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 62.0 ± 0.4 | — |
| Abliterated, thinking (ours) | 67.5 ± 1.63 | ▲ 5.5vs non‑think |
| Qwen 3.5:9b (ours) | 64.7 ± 0.4 | ▲ 2.7vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 77.0 ± 1.45 | ▲ 12.3vs non‑think |
| Qwen3.5‑9B (posted) | 82.51 | — |
instruction-following & formatting compliance
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 64.7strict‑p 48.8 · strict‑i 61.9 · loose‑p 52.7 | — |
| Abliterated, thinking (ours) | 85.73strict‑p 79.5 · strict‑i 83.9 · loose‑p 81.7 | ▲ 21.0vs non‑think |
| Qwen 3.5:9b (ours) | 58.8strict‑p 44.2–44.42 · strict‑i 57.4–57.7 · loose‑p 45.7 | ▼ 5.9vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 94.1strict‑p 88.7 · strict‑i 91.7 · loose‑p 91.9 | ▲ 35.3vs non‑think |
| Qwen3.5‑9B (posted) | 91.51 | — |
news article summarization, single ROUGE score
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 7.9 | — |
| Abliterated, thinking (ours) | not reported | — |
| Qwen 3.5:9b (ours) | 10.0 | ▲ 2.1vs abliterated |
| Qwen 3.5:9b, thinking (ours) | not reported | — |
extracting a structured field from FDA drug-label text
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 76.0 | — |
| Abliterated, thinking (ours) | not reported | — |
| Qwen 3.5:9b (ours) | 70.2 | ▼ 5.8vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 80.8 | ▲ 10.6vs non‑think |
completing a passage's next span, SQuAD-derived
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 75.8 | — |
| Abliterated, thinking (ours) | 92.0 ± 1.03 | ▲ 16.2vs non‑think |
| Qwen 3.5:9b (ours) | 70.7 | ▼ 5.1vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 90.0 ± 1.1 | ▲ 19.3vs non‑think |
predicting a Python function's input from its output
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 44.1 ± 1.7 | — |
| Abliterated, thinking (ours) | 47.9 ± 3.23 | ▲ 3.8vs non‑think |
| Qwen 3.5:9b (ours) | 52.8 ± 1.7 | ▲ 8.7vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 89.9 ± 2.6 | ▲ 37.1vs non‑think |
| deepseek-instruct-6.7b (posted, ref) | 37.44 | — |
predicting a Python function's output from its input
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 27.6 ± 1.5 | — |
| Abliterated, thinking (ours) | 43.1 ± 5.13 | ▲ 15.5vs non‑think |
| Qwen 3.5:9b (ours) | 31.6 ± 1.6 | ▲ 4.0vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 54.3 ± 5.7 | ▲ 22.7vs non‑think |
| deepseek-instruct-6.7b (posted, ref) | 41.24 | — |
generating a Python function from a natural-language spec, free-form code extraction (HumanEval-derived)
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 39.0 ± 3.8 | — |
| Abliterated, thinking (ours) | 82.9 ± 2.9 | ▲ 43.9vs non‑think |
| Qwen 3.5:9b (ours) | 34.8 ± 3.7 | ▼ 4.2vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 97.0 ± 1.3 | ▲ 62.2vs non‑think |
generating a Python function from a natural-language spec, free-form code extraction (MBPP-derived)
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 13.2 ± 1.5 | — |
| Abliterated, thinking (ours) | 67.8 ± 2.1 | ▲ 54.6vs non‑think |
| Qwen 3.5:9b (ours) | 9.0 ± 1.3 | ▼ 4.2vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 81.2 ± 1.7 | ▲ 72.2vs non‑think |
long‑context multi‑hop QA, 20 reasoning‑complexity levels
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 75.8qa1–qa20 range 9.8–99.9 | — |
| Abliterated, thinking (ours) | 91.2qa1–qa20 range 58.0–100.0 | ▲ 15.4vs non‑think |
| Qwen 3.5:9b (ours) | 83.3qa1–qa20 range 11.2–99.8 | ▲ 7.5vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 91.5qa1–qa20 range 62.0–100.0 | ▲ 8.2vs non‑think |
Abliterated vs Qwen 3.5:9b
Abliterated, thinking vs Qwen 3.5:9b, thinking
Abliterated vs Qwen 3.5:9b
Only 2 tasks compared so far — preliminary.
Abliterated, thinking vs Qwen 3.5:9b, thinking
Only 2 tasks compared so far — preliminary.
Abliterated vs Qwen 3.5:9b
Abliterated, thinking vs Qwen 3.5:9b, thinking
Abliterated vs Qwen 3.5:9b
Only 1 task compared so far — preliminary.
Abliterated, thinking vs Qwen 3.5:9b, thinking
No task has both a Abliterated, thinking and a Qwen 3.5:9b, thinking result yet.
Abliterated vs Qwen 3.5:9b
Abliterated, thinking vs Qwen 3.5:9b, thinking
max_gen_toks=10000, no early stop) instead of the non‑thinking default used everywhere else on this page — see eval‑harness/run_thinking_batch.sh. They're the first three results back from a broader 12‑task thinking‑mode rerun still in progress; more rows will fill in as that batch completes. (humaneval_instruct and mbpp_instruct_fixed were dropped from this page entirely, not merely excluded from that rerun — both used lm_eval's gen_prefix mechanic, which was confirmed to make thinking mode a silent no‑op, so their non‑thinking scores weren't a fair baseline to keep either. They've been superseded by humaneval_instruct_free/mbpp_instruct_free, free‑generation rewrites without gen_prefix, once those finish their own full 2×2 run set.) bbh_cot_zeroshot's two thinking rows are additionally capped at ‑‑limit 30 docs per subtask (≈810 of the full 6,511‑doc set, versus every other row on this task — and every other task's thinking row — covering its full corpus) — flagged 2026‑09‑12 during a benchmark‑page speed‑figures audit, with no on‑page disclosure until now. A full‑coverage rerun is planned (see EVAL‑BACKLOG.md's roadmap section) but not yet queued; treat these two numbers as a smaller‑sample estimate, not a like‑for‑like comparison to this task's own non‑thinking rows._closest_instruct_reference().eval‑harness/run_official_thinking_mmlu_pro.sh against the same 756‑sample subsample as the abliterated, thinking row above it — built specifically to isolate how much of the gap between abliterated, thinking (67.5) and Qwen's posted mmlu_pro number (82.5) is attributable to the abliterated build itself versus this local build's shared Q4_K_M quantization. It's the first result of a fourth core setting (official weights × thinking) this page now tracks alongside abliterated/qwen/abliterated‑thinking — more tasks will fill in over time the same way the abliterated, thinking rows have been.Per‑item speed for each run, wall‑clock runtime alongside it for reference. Not a clean speed benchmark — see the callout below before reading these as apples‑to‑apples.
-c is the same total KV pool either way; see hearth‑bypass/bypass_run_abliterated_thinking.bat's and bypass_run_thinking.bat's matching -np 8, and EVAL‑BACKLOG.md). The much higher per‑item time there is thinking tokens plus that lower concurrency, not a like‑for‑like comparison to the 16‑way non‑thinking rows. The five backlog tasks (fda/squad_completion/cruxeval_input/cruxeval_output/babilong) only have an abliterated row at all, all via the 16‑way bypass. Only compare rows where both sides used the same serving mode.
lm_eval's own on‑disk request cache satisfying every single request (each log's own Cached requests: N, Requests remaining: 0 line confirms 100% hits, and each log has zero tqdm progress‑bar ticks at all to reconstruct from). The real first‑time generation that originally populated that cache was never captured with real timing — its log was overwritten by the pre‑2026‑08‑29 truncation issue described above, before there was anything to filter. speed_stats.py now drops any tick implying faster than 1s/item from every sum (not just the median), on the reasoning that a cache hit's near‑zero interval has nothing to do with model speed and shouldn't be averaged in as if it did; for a partially‑contaminated log this recovers a real median from whatever ticks survive, but these three logs have no surviving ticks at all. Borrowing a rate from these tasks' own "qwen" rows (which do have real generation) was tried and rejected: those runs are themselves so cache‑heavy that only 0–5 real intervals survive per task, too few and too inconsistent (4.0s/it vs 69.0s/it vs no data) to trust as a substitute. So these three cells now read "cache replay" with no number attached, rather than either the old misleading figure or a fabricated estimate built from a handful of samples.
Bars are scaled per task (each task's own longest bar is 100% width) — runtimes span two orders of magnitude across tasks, so a single shared scale would flatten the short ones to nothing.
grade-school math word problems
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 2.00s/it16‑way bypass | 35m47s |
| Abliterated, thinking (ours, bypass) | 44.00s/it8‑way bypass | 9h10m |
| Qwen 3.5:9b (ours) | 6.32s/itserial | 2h19m |
| Qwen 3.5:9b, thinking (ours, bypass) | not run | — |
diverse reasoning — logic, algorithms, language
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 1.00s/it16‑way bypass | 2h03m |
| Abliterated, thinking (ours, bypass) | 35.00s/it4‑way bypass | 10h59m |
| Qwen 3.5:9b (ours) | 0.96s/it16‑way bypass | 1h44m |
| Qwen 3.5:9b, thinking (ours, bypass) | 16.00s/it4‑way bypass | 5h19m47s |
broad academic knowledge, 57 subjects
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 2.00s/it16‑way bypass | 5h53m |
| Abliterated, thinking (ours, bypass) | 24.91s/it4‑way bypass | 5h14m |
| Qwen 3.5:9b (ours) | 1.15s/it16‑way bypass | 3h51m |
| Qwen 3.5:9b, thinking (ours, bypass) | 29.00s/it4‑way bypass | 8h46m |
instruction-following & formatting compliance
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 3.00s/it16‑way bypass | 33m14s |
| Abliterated, thinking (ours, bypass) | 53.88s/it4‑way bypass | 8h06m |
| Qwen 3.5:9b (ours) | 18.77s/itserial | 2h49m |
| Qwen 3.5:9b, thinking (ours, bypass) | 27.53s/it4‑way bypass | 4h08m15s |
news article summarization, single ROUGE score
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 1.78s/itserial | 5h41m |
| Abliterated, thinking (ours, bypass) | not run | — |
| Qwen 3.5:9b (ours) | 2.11s/itserial | 6h43m |
| Qwen 3.5:9b, thinking (ours, bypass) | not run | — |
extracting a structured field from FDA drug-label text
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | cache replay 100% cache hits, no timing data | — |
| Abliterated, thinking (ours, bypass) | not run | — |
| Qwen 3.5:9b (ours) | 1.68s/it16‑way bypass | 31m10s |
| Qwen 3.5:9b, thinking (ours, bypass) | 9.00s/it4‑way bypass | 1h52m30s |
completing a passage's next span, SQuAD-derived
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | cache replay 100% cache hits, no timing data | — |
| Abliterated, thinking (ours, bypass) | 37.19s/it4‑way bypass | 7h44m50s |
| Qwen 3.5:9b (ours) | 1.13s/it16‑way bypass | 56m11s |
| Qwen 3.5:9b, thinking (ours, bypass) | 22.71s/it4‑way bypass | 4h43m54s |
predicting a Python function's input from its output
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 6.67s/it16‑way bypass | 58m52s |
| Abliterated, thinking (ours, bypass) | 31.18s/it4‑way bypass | 38m59s |
| Qwen 3.5:9b (ours) | 1.00s/it16‑way bypass | 2h55m37s |
| Qwen 3.5:9b, thinking (ours, bypass) | 26.49s/it4‑way bypass | 33m07s |
predicting a Python function's output from its input
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 5.00s/it16‑way bypass | 55m07s |
| Abliterated, thinking (ours, bypass) | 31.83s/it4‑way bypass | 39m48s |
| Qwen 3.5:9b (ours) | 0.60s/it16‑way bypass | 1h02m23s |
| Qwen 3.5:9b, thinking (ours, bypass) | 18.00s/it4‑way bypass | 22m30s |
generating a Python function from a natural-language spec, free-form code extraction (HumanEval-derived)
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 5.00s/it16‑way bypass | 54m50s |
| Abliterated, thinking (ours, bypass) | 8.00s/it8‑way bypass | 1h13m20s |
| Qwen 3.5:9b (ours) | 5.00s/it16‑way bypass | 28m55s |
| Qwen 3.5:9b, thinking (ours, bypass) | 10.00s/it8‑way bypass | 1h05m50s |
generating a Python function from a natural-language spec, free-form code extraction (MBPP-derived)
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 1.00s/it16‑way bypass | 9m10s |
| Abliterated, thinking (ours, bypass) | 19.00s/it8‑way bypass | 2h38m20s |
| Qwen 3.5:9b (ours) | 1.00s/it16‑way bypass | 11m25s |
| Qwen 3.5:9b, thinking (ours, bypass) | 14.00s/it8‑way bypass | 2h00m24s |
long‑context multi‑hop QA, 20 reasoning‑complexity levels
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | cache replay 100% cache hits, no timing data | — |
| Abliterated, thinking (ours, bypass) | 26.44s/it8‑way bypass | 7h03m06s |
| Qwen 3.5:9b (ours) | 0.19s/it16‑way bypass | 2h05m |
| Qwen 3.5:9b, thinking (ours, bypass) | 21.81s/it8‑way bypass | 6h03m32s |
Diffusion image generation · Apple M1 Pro, 16GB unified memory
Image-generation speed and memory across diffusion models, run locally on the MacBook's unified memory instead of the desktop's dedicated VRAM — 7 runs across 6 model families that fit comfortably in 16 GB timed so far.
~/diffusion-bench (not tracked in this repo — its venvs and downloaded weights alone run tens of GB; the sample images below are tracked, in images/). Six rows are timed with /usr/bin/time -l: five around an mflux-generate-* call (see run_frontier_batch.sh), and SDXL Turbo via a separate, non‑mflux pipeline (mlx‑examples' stable_diffusion, standardized 2026‑09‑06 — see its card's note). The remaining row (FLUX.1 schnell) comes from mflux's own embedded PNG metadata instead — each card says which. Only models that fit comfortably inside this machine's 16 GB of unified memory are shown below — a model tested but excluded for swapping heavily instead gets a plain callout explaining why, not a card with misleadingly slow numbers (see the notes below the grid). No image-quality scoring here (unlike the LLM tables above, there's no numeric metric for "does it look good") — judge that from the thumbnails, this section tracks generation speed and memory only.





| Metric | Value |
|---|---|
| Steps | 8 |
| Quantize | 8-bit |
| Generation time | 57.0s |
| Speed | 7.12s/it avg |
| Peak MLX memory | 13.44 GB |





| Metric | Value |
|---|---|
| Steps | 4 |
| Quantize | 8-bit |
| Generation time | 13.8s |
| Speed | 3.44s/it avg (6.90s cold → 2.21s/it by step 4) |
| Peak MLX memory | 12.74 GB |
Its text encoder is GPT‑OSS‑20B (mlx‑community/gpt‑oss‑20b‑MXFP4‑Q8), not a typical CLIP/T5 text encoder — a general‑purpose 20B language model doing prompt encoding for a 4‑step image model. Peak memory still fits comfortably despite that.





| Metric | Value |
|---|---|
| Steps | 4 |
| Quantize | 8-bit |
| Generation time | 1m28s |
| Speed | 22.1s/it avg |
| Peak MLX memory | 12.42 GB |
The larger sibling of FLUX.2 klein‑4B (9B vs. 4B parameters). The first attempt at this model failed outright on a HuggingFace 403 gated‑repo error; access opened up since, and this retry completed cleanly — still fitting comfortably despite the larger size, though 4.2× slower per step than its 4B sibling.





| Metric | Value |
|---|---|
| Steps | 4 |
| Quantize | 8-bit |
| Generation time | 20.8s |
| Speed | 5.21s/it avg (7.46s cold → 4.72s/it by step 4) |
| Peak MLX memory | 6.89 GB |





| Metric | Value |
|---|---|
| Steps | 4 |
| Quantize | 4-bit |
| Generation time | 30.1s |
| Speed | 7.54s/it avg |
| Peak MLX memory | 10.33 GB |
Loaded from a checkpoint pre‑quantized to 4‑bit and saved to disk, rather than re‑quantizing the full‑precision weights at load time — the representative FLUX.1 schnell setting on this page. Steps/speed come straight from mflux's own embedded PNG metadata (see this section's intro note); the original astronaut‑prompt run had no /usr/bin/time wrapper, so peak memory here instead comes from the later, separately‑wrapped combined‑prompt run of the same 4‑bit checkpoint (combined‑flux‑schnell‑q4.log).





| Metric | Value |
|---|---|
| Steps | 9 |
| Quantize | 8-bit |
| Generation time | 53.3s |
| Speed | 5.92s/it avg (10.37s cold → ~5.2–5.6s/it after) |
| Peak MLX memory | 8.74 GB |





| Metric | Value |
|---|---|
| Steps | 2 |
| Quantize | fp16 |
| Generation time | 30.2s |
| Speed | 8.91s/it avg (15.30s cold → 7.78s/it by step 2) |
| Peak MLX memory | 6.98 GB |
Standardized 2026‑09‑06 to match the cards above — originally shown as an uninstrumented gallery (no /usr/bin/time, no PNG metadata from this pipeline; see git history for that version). This card's numbers come from a fresh, properly‑wrapped rerun with the same seed and prompt, confirmed byte‑identical to the original image (same file, sdxl-turbo-warm.png). Native output resolution for this pipeline is 528×528, not the 512×512 every other card uses — not operator‑configurable via this script's CLI, so shown as-is rather than cropped/resized to match.
Ideogram 4 (fp8) (ideogram‑ai/ideogram‑4‑fp8, both 8‑bit and 4‑bit), Qwen‑Image (Qwen/Qwen‑Image‑2512, already 4‑bit), FIBO Lite (briaai/Fibo‑lite, 8‑bit), and Krea 2 Turbo (krea/Krea‑2‑Turbo, 8‑bit) were also tested and all completed — but all five runs are excluded from the comparison above rather than shown alongside models that actually fit. Peak memory ran 28.28 GB (Ideogram 8‑bit), 27.36 GB (Ideogram 4‑bit), 27.37 GB (Qwen‑Image), 24.38 GB (FIBO Lite), and 22.48 GB (Krea 2 Turbo), all well past this machine's 16 GB of unified memory, so most of each run's time was spent swapping to disk, not generating — the 88.6s/it, 40.3s/it, 109.7s/it, 21.8s/it, and 48.9s/it measured aren't a fair read on these models' real speed, just how slow this particular machine is at thrashing its way through them.
| Model | Peak Memory | Fits 16 GB? |
|---|---|---|
| SDXL-Turbo | 6.98 GB | ✅ |
| FLUX.1 schnell (4-bit, saved standalone) | 10.33 GB | ✅ |
| FLUX.2 klein-4B | 6.89 GB | ✅ |
| FLUX.2 klein-9B | 12.42 GB | ✅ |
| Z-Image Turbo | 8.74 GB | ✅ |
| ERNIE Image Turbo (Baidu) | 13.44 GB | ✅ |
| Lens (Microsoft) | 12.74 GB | ✅ |
| Ideogram 4 (fp8) | 27–28 GB | ❌ |
| Qwen-Image-2512 | 27.37 GB | ❌ (but best quality) |
| FIBO Lite (Bria AI) | 24.38 GB | ❌ |
| Krea 2 Turbo | 22.48 GB | ❌ (breaks the "Turbo" heuristic) |