Local inference
Local LLMs served on the desktop's RTX 3060. Diffusion image generation on the MacBook's M1.
Qwen3.5:9B
huihui-ai abliterated build — the Qwen3.5‑9B base model with abliteration applied — independently‑verified genuine GGUF, served via the 16‑way raw‑llama.cpp bypass — 12 of 12 tasks done.
Same abliterated build, thinking mode forced on instead of the non‑thinking default used elsewhere on this page.3 Five results so far (ifeval, mmlu_pro, bbh_cot_zeroshot, cruxeval_input, cruxeval_output) from a 12‑task rerun still in progress — most rows don't have one yet.
Unsloth official Qwen3.5‑9B weights, same harness/prompting (two tasks via the raw‑llama.cpp bypass).
Same official weights as the row above, thinking mode forced on — the fourth of the four core settings this page tracks (abliterated/official × thinking/non‑thinking).5 One result so far (mmlu_pro), meant to fill in for every benchmark over time like the abliterated, thinking row above it.
Numbers from Qwen's own model card, quoted as-is and not run by us — except cruxeval_input/cruxeval_output, where no Qwen number exists at all and this row instead shows a same‑scale substitute model (labeled by name, not "Qwen3.5‑9B").
One table per benchmark, all scores as percentages. Each Δ is labeled with what it's measured against, since the comparison isn't the same for every row: Qwen 3.5:9b's Δ is versus the plain Abliterated row (does swapping to official weights change the score), while each thinking row's Δ is versus its own non‑thinking baseline directly above it (does thinking mode change the score) — two independent comparisons, not one chained run down the table. Every Δ is computed straight from the two rounded scores shown, so it always matches subtracting them by hand. The posted row is never delta'd — see footnote 6 for why that comparison needs its own caveats. A task's posted row is omitted entirely when no public number (or substitute) exists for it, rather than shown as a blank — see footnote 1. Some backlog tasks below have only the abliterated row filled in; a qwen comparison doesn't exist yet for those.
grade-school math word problems
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 82.5strict 82.6 | — |
| Abliterated, thinking (ours) | 87.3 ± 1.2strict 84.9 | ▲ 4.8vs non‑think |
| Qwen 3.5:9b (ours) | 85.7 ± 1.0 | ▲ 3.2vs abliterated |
| Qwen 3.5:9b, thinking (ours) | not reported | — |
diverse reasoning — logic, algorithms, language
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 45.6 ± 0.5 | — |
| Abliterated, thinking (ours) | 79.9 ± 1.13 | ▲ 34.3vs non‑think |
| Qwen 3.5:9b (ours) | 65.8 ± 0.5 | ▲ 20.2vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 89.1 ± 0.93 | ▲ 23.3vs non‑think |
broad academic knowledge, 57 subjects
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 62.0 ± 0.4 | — |
| Abliterated, thinking (ours) | 67.5 ± 1.63 | ▲ 5.5vs non‑think |
| Qwen 3.5:9b (ours) | 64.7 ± 0.4 | ▲ 2.7vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 77.0 ± 1.45 | ▲ 12.3vs non‑think |
| Qwen3.5‑9B (posted) | 82.51 | — |
instruction-following & formatting compliance
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 64.7strict‑p 48.8 · strict‑i 61.9 · loose‑p 52.7 | — |
| Abliterated, thinking (ours) | 85.73strict‑p 79.5 · strict‑i 83.9 · loose‑p 81.7 | ▲ 21.0vs non‑think |
| Qwen 3.5:9b (ours) | 58.8strict‑p 44.2–44.42 · strict‑i 57.4–57.7 · loose‑p 45.7 | ▼ 5.9vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 94.1strict‑p 88.7 · strict‑i 91.7 · loose‑p 91.9 | ▲ 35.3vs non‑think |
| Qwen3.5‑9B (posted) | 91.51 | — |
news article summarization, single ROUGE score
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 7.9 | — |
| Abliterated, thinking (ours) | not reported | — |
| Qwen 3.5:9b (ours) | 10.0 | ▲ 2.1vs abliterated |
| Qwen 3.5:9b, thinking (ours) | not reported | — |
extracting a structured field from FDA drug-label text
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 76.0 | — |
| Abliterated, thinking (ours) | not reported | — |
| Qwen 3.5:9b (ours) | 70.2 | ▼ 5.8vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 80.8 | ▲ 10.6vs non‑think |
completing a passage's next span, SQuAD-derived
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 75.8 | — |
| Abliterated, thinking (ours) | 92.0 ± 1.03 | ▲ 16.2vs non‑think |
| Qwen 3.5:9b (ours) | 70.7 | ▼ 5.1vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 90.0 ± 1.1 | ▲ 19.3vs non‑think |
predicting a Python function's input from its output
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 44.1 ± 1.7 | — |
| Abliterated, thinking (ours) | 47.9 ± 3.23 | ▲ 3.8vs non‑think |
| Qwen 3.5:9b (ours) | 52.8 ± 1.7 | ▲ 8.7vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 89.9 ± 2.6 | ▲ 37.1vs non‑think |
| deepseek-instruct-6.7b (posted, ref) | 37.44 | — |
predicting a Python function's output from its input
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 27.6 ± 1.5 | — |
| Abliterated, thinking (ours) | 43.1 ± 5.13 | ▲ 15.5vs non‑think |
| Qwen 3.5:9b (ours) | 31.6 ± 1.6 | ▲ 4.0vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 54.3 ± 5.7 | ▲ 22.7vs non‑think |
| deepseek-instruct-6.7b (posted, ref) | 41.24 | — |
generating a Python function from a natural-language spec, free-form code extraction (HumanEval-derived)
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 39.0 ± 3.8 | — |
| Abliterated, thinking (ours) | 82.9 ± 2.9 | ▲ 43.9vs non‑think |
| Qwen 3.5:9b (ours) | 34.8 ± 3.7 | ▼ 4.2vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 97.0 ± 1.3 | ▲ 62.2vs non‑think |
generating a Python function from a natural-language spec, free-form code extraction (MBPP-derived)
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 13.2 ± 1.5 | — |
| Abliterated, thinking (ours) | 67.8 ± 2.1 | ▲ 54.6vs non‑think |
| Qwen 3.5:9b (ours) | 9.0 ± 1.3 | ▼ 4.2vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 81.2 ± 1.7 | ▲ 72.2vs non‑think |
long‑context multi‑hop QA, 20 reasoning‑complexity levels
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 75.8qa1–qa20 range 9.8–99.9 | — |
| Abliterated, thinking (ours) | 91.2qa1–qa20 range 58.0–100.0 | ▲ 15.4vs non‑think |
| Qwen 3.5:9b (ours) | 83.3qa1–qa20 range 11.2–99.8 | ▲ 7.5vs abliterated |
| Qwen 3.5:9b, thinking (ours) | 91.5qa1–qa20 range 62.0–100.0 | ▲ 8.2vs non‑think |
Abliterated vs Qwen 3.5:9b
Abliterated, thinking vs Qwen 3.5:9b, thinking
Abliterated vs Qwen 3.5:9b
Only 2 tasks compared so far — preliminary.
Abliterated, thinking vs Qwen 3.5:9b, thinking
Only 2 tasks compared so far — preliminary.
Abliterated vs Qwen 3.5:9b
Abliterated, thinking vs Qwen 3.5:9b, thinking
Abliterated vs Qwen 3.5:9b
Only 1 task compared so far — preliminary.
Abliterated, thinking vs Qwen 3.5:9b, thinking
No task has both a Abliterated, thinking and a Qwen 3.5:9b, thinking result yet.
Abliterated vs Qwen 3.5:9b
Abliterated, thinking vs Qwen 3.5:9b, thinking
max_gen_toks=10000, no early stop) instead of the non‑thinking default used everywhere else on this page. They're the first three results back from a broader 12‑task thinking‑mode rerun still in progress; more rows will fill in as that batch completes. (humaneval_instruct and mbpp_instruct_fixed were dropped from this page entirely, not merely excluded from that rerun — both used lm_eval's gen_prefix mechanic, which was confirmed to make thinking mode a silent no‑op, so their non‑thinking scores weren't a fair baseline to keep either. They've been superseded by humaneval_instruct_free/mbpp_instruct_free, free‑generation rewrites without gen_prefix, once those finish their own full 2×2 run set.) bbh_cot_zeroshot's two thinking rows are additionally capped at ‑‑limit 30 docs per subtask (≈810 of the full 6,511‑doc set, versus every other row on this task — and every other task's thinking row — covering its full corpus). A full‑coverage rerun is planned but not yet queued; treat these two numbers as a smaller‑sample estimate, not a like‑for‑like comparison to this task's own non‑thinking rows.Per‑item speed for each run, wall‑clock runtime alongside it for reference. Not a clean speed benchmark — see the callout below before reading these as apples‑to‑apples.
lm_eval's own on‑disk request cache satisfying every single request (each log's own Cached requests: N, Requests remaining: 0 line confirms 100% hits, and each log has zero tqdm progress‑bar ticks at all to reconstruct from). The real first‑time generation that originally populated that cache was never captured with real timing — its log was overwritten by the truncation issue described above, before there was anything to filter. The reconstruction now drops any tick implying faster than 1s/item from every sum (not just the median), on the reasoning that a cache hit's near‑zero interval has nothing to do with model speed and shouldn't be averaged in as if it did; for a partially‑contaminated log this recovers a real median from whatever ticks survive, but these three logs have no surviving ticks at all. A substitute rate from these tasks' own "qwen" rows isn't reliable either: those runs are themselves so cache‑heavy that only 0–5 real intervals survive per task, too inconsistent (4.0s/it vs 69.0s/it vs no data) to use. So these three cells now read "cache replay" with no number attached, rather than either the old misleading figure or a fabricated estimate built from a handful of samples.
Bars are scaled per task (each task's own longest bar is 100% width) — runtimes span two orders of magnitude across tasks, so a single shared scale would flatten the short ones to nothing.
grade-school math word problems
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 2.00s/it16‑way bypass | 35m47s |
| Abliterated, thinking (ours, bypass) | 44.00s/it8‑way bypass | 9h10m |
| Qwen 3.5:9b (ours) | 6.32s/itserial | 2h19m |
| Qwen 3.5:9b, thinking (ours, bypass) | not run | — |
diverse reasoning — logic, algorithms, language
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 1.00s/it16‑way bypass | 2h03m |
| Abliterated, thinking (ours, bypass) | 35.00s/it4‑way bypass | 10h59m |
| Qwen 3.5:9b (ours) | 0.96s/it16‑way bypass | 1h44m |
| Qwen 3.5:9b, thinking (ours, bypass) | 16.00s/it4‑way bypass | 5h19m47s |
broad academic knowledge, 57 subjects
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 2.00s/it16‑way bypass | 5h53m |
| Abliterated, thinking (ours, bypass) | 24.91s/it4‑way bypass | 5h14m |
| Qwen 3.5:9b (ours) | 1.15s/it16‑way bypass | 3h51m |
| Qwen 3.5:9b, thinking (ours, bypass) | 29.00s/it4‑way bypass | 8h46m |
instruction-following & formatting compliance
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 3.00s/it16‑way bypass | 33m14s |
| Abliterated, thinking (ours, bypass) | 53.88s/it4‑way bypass | 8h06m |
| Qwen 3.5:9b (ours) | 18.77s/itserial | 2h49m |
| Qwen 3.5:9b, thinking (ours, bypass) | 27.53s/it4‑way bypass | 4h08m15s |
news article summarization, single ROUGE score
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 1.78s/itserial | 5h41m |
| Abliterated, thinking (ours, bypass) | not run | — |
| Qwen 3.5:9b (ours) | 2.11s/itserial | 6h43m |
| Qwen 3.5:9b, thinking (ours, bypass) | not run | — |
extracting a structured field from FDA drug-label text
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | cache replay 100% cache hits, no timing data | — |
| Abliterated, thinking (ours, bypass) | not run | — |
| Qwen 3.5:9b (ours) | 1.68s/it16‑way bypass | 31m10s |
| Qwen 3.5:9b, thinking (ours, bypass) | 9.00s/it4‑way bypass | 1h52m30s |
completing a passage's next span, SQuAD-derived
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | cache replay 100% cache hits, no timing data | — |
| Abliterated, thinking (ours, bypass) | 37.19s/it4‑way bypass | 7h44m50s |
| Qwen 3.5:9b (ours) | 1.13s/it16‑way bypass | 56m11s |
| Qwen 3.5:9b, thinking (ours, bypass) | 22.71s/it4‑way bypass | 4h43m54s |
predicting a Python function's input from its output
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 6.67s/it16‑way bypass | 58m52s |
| Abliterated, thinking (ours, bypass) | 31.18s/it4‑way bypass | 38m59s |
| Qwen 3.5:9b (ours) | 1.00s/it16‑way bypass | 2h55m37s |
| Qwen 3.5:9b, thinking (ours, bypass) | 26.49s/it4‑way bypass | 33m07s |
predicting a Python function's output from its input
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 5.00s/it16‑way bypass | 55m07s |
| Abliterated, thinking (ours, bypass) | 31.83s/it4‑way bypass | 39m48s |
| Qwen 3.5:9b (ours) | 0.60s/it16‑way bypass | 1h02m23s |
| Qwen 3.5:9b, thinking (ours, bypass) | 18.00s/it4‑way bypass | 22m30s |
generating a Python function from a natural-language spec, free-form code extraction (HumanEval-derived)
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 5.00s/it16‑way bypass | 54m50s |
| Abliterated, thinking (ours, bypass) | 8.00s/it8‑way bypass | 1h13m20s |
| Qwen 3.5:9b (ours) | 5.00s/it16‑way bypass | 28m55s |
| Qwen 3.5:9b, thinking (ours, bypass) | 10.00s/it8‑way bypass | 1h05m50s |
generating a Python function from a natural-language spec, free-form code extraction (MBPP-derived)
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 1.00s/it16‑way bypass | 9m10s |
| Abliterated, thinking (ours, bypass) | 19.00s/it8‑way bypass | 2h38m20s |
| Qwen 3.5:9b (ours) | 1.00s/it16‑way bypass | 11m25s |
| Qwen 3.5:9b, thinking (ours, bypass) | 14.00s/it8‑way bypass | 2h00m24s |
long‑context multi‑hop QA, 20 reasoning‑complexity levels
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | cache replay 100% cache hits, no timing data | — |
| Abliterated, thinking (ours, bypass) | 26.44s/it8‑way bypass | 7h03m06s |
| Qwen 3.5:9b (ours) | 0.19s/it16‑way bypass | 2h05m |
| Qwen 3.5:9b, thinking (ours, bypass) | 21.81s/it8‑way bypass | 6h03m32s |
Diffusion image generation · Apple M1 Pro, 16GB unified memory
Image-generation speed and memory across diffusion models, run locally on the MacBook's unified memory instead of the desktop's dedicated VRAM — 7 runs across 6 model families that fit comfortably in 16 GB timed so far.
/usr/bin/time -l. Only models that fit comfortably inside this machine's 16 GB of unified memory are shown below — a model tested but excluded for swapping heavily instead gets a plain callout explaining why, not a card with misleadingly slow numbers (see the notes below the grid). No image-quality scoring here (unlike the LLM tables above, there's no numeric metric for "does it look good") — judge that from the thumbnails, this section tracks generation speed and memory only.






| Metric | Value |
|---|---|
| Steps | 4 |
| Quantize | 8-bit |
| Generation time | 2m15s |
| Speed | 31.71s/it avg (10.75s cold → 38.32s/it by step 4) |
| Peak MLX memory | 14.77 GB |
10B params, DMD‑distilled to 4 steps, Apache‑2.0. The tightest fit of any card on this page — 14.77 GB peak, within 1.2 GB of the 16 GB ceiling. Unlike every other model here, each step got slower rather than faster (10.75s → 38.32s), the opposite of the usual cold‑start‑then‑settles pattern, plausibly memory pressure this close to the ceiling. --guidance is accepted but ignored (distilled in).






| Metric | Value |
|---|---|
| Steps | 8 |
| Quantize | 8-bit |
| Generation time | 57.0s |
| Speed | 7.12s/it avg |
| Peak MLX memory | 13.44 GB |






| Metric | Value |
|---|---|
| Steps | 4 |
| Quantize | 8-bit |
| Generation time | 13.8s |
| Speed | 3.44s/it avg (6.90s cold → 2.21s/it by step 4) |
| Peak MLX memory | 12.74 GB |
Its text encoder is GPT‑OSS‑20B (mlx‑community/gpt‑oss‑20b‑MXFP4‑Q8), not a typical CLIP/T5 text encoder — a general‑purpose 20B language model doing prompt encoding for a 4‑step image model. Peak memory still fits comfortably despite that.






| Metric | Value |
|---|---|
| Steps | 4 |
| Quantize | 8-bit |
| Generation time | 1m28s |
| Speed | 22.1s/it avg |
| Peak MLX memory | 12.42 GB |
The larger sibling of FLUX.2 klein‑4B (9B vs. 4B parameters) — still fits comfortably despite the larger size, though 4.2× slower per step than its 4B sibling.






| Metric | Value |
|---|---|
| Steps | 4 |
| Quantize | 8-bit |
| Generation time | 20.8s |
| Speed | 5.21s/it avg (7.46s cold → 4.72s/it by step 4) |
| Peak MLX memory | 6.89 GB |






| Metric | Value |
|---|---|
| Steps | 50 |
| Quantize | 8-bit |
| Generation time | 9m13s |
| Speed | 10.83s/it avg (14.85s cold → 10.51s/it by step 50) |
| Peak MLX memory | 8.75 GB |
The full (non‑distilled) Z‑Image release, not the Turbo variant above — using mflux's own documented settings for base Z‑Image (50 steps, guidance 4). Notable because it breaks this page's general pattern: every other large, non‑distilled model tried (Ideogram 4, Qwen‑Image, FIBO Lite, Krea 2 Turbo — see the excluded‑for‑paging note below) overshoots 16 GB regardless of quantization, but base Z‑Image fits at essentially the same peak memory as its own Turbo distillation (8.75 vs. 8.74 GB) — the extra quality costs steps (50 vs. 9) and roughly 10× the time, not memory.






| Metric | Value |
|---|---|
| Steps | 9 |
| Quantize | 8-bit |
| Generation time | 53.3s |
| Speed | 5.92s/it avg (10.37s cold → ~5.2–5.6s/it after) |
| Peak MLX memory | 8.74 GB |
Ideogram 4 (fp8) (ideogram‑ai/ideogram‑4‑fp8, both 8‑bit and 4‑bit), Qwen‑Image (Qwen/Qwen‑Image‑2512, already 4‑bit), FIBO Lite (briaai/Fibo‑lite, 8‑bit), and Krea 2 Turbo (krea/Krea‑2‑Turbo, 8‑bit) were also tested and all completed — but all five runs are excluded from the comparison above rather than shown alongside models that actually fit. Peak memory ran 28.28 GB (Ideogram 8‑bit), 27.36 GB (Ideogram 4‑bit), 27.37 GB (Qwen‑Image), 24.38 GB (FIBO Lite), and 22.48 GB (Krea 2 Turbo), all well past this machine's 16 GB of unified memory, so most of each run's time was spent swapping to disk, not generating — the 88.6s/it, 40.3s/it, 109.7s/it, 21.8s/it, and 48.9s/it measured aren't a fair read on these models' real speed, just how slow this particular machine is at thrashing its way through them.
| Model | Peak Memory | Fits 16 GB? |
|---|---|---|
| FLUX.2 klein-4B | 6.89 GB | ✅ |
| FLUX.2 klein-9B | 12.42 GB | ✅ |
| Z-Image Turbo | 8.74 GB | ✅ |
| Z-Image (full) | 8.75 GB | ✅ |
| Lens (Microsoft) | 12.74 GB | ✅ |
| ERNIE Image Turbo (Baidu) | 13.44 GB | ✅ |
| Boogu Image Turbo | 14.77 GB | ✅ |
| Ideogram 4 (fp8) | 27–28 GB | ❌ |
| Qwen-Image-2512 | 27.37 GB | ❌ (but best quality) |
| FIBO Lite (Bria AI) | 24.38 GB | ❌ |
| Krea 2 Turbo | 22.48 GB | ❌ (breaks the "Turbo" heuristic) |