Local inference
Local LLMs served on the desktop's RTX 3060. Diffusion image generation on the MacBook's M1.
Qwen3.5:9B
huihui-ai abliterated build — the Qwen3.5‑9B base model with abliteration applied — independently‑verified genuine GGUF, served via the 16‑way raw‑llama.cpp bypass, 2026‑08‑23 onward — 12 of 12 tasks done.
Same abliterated build, thinking mode forced on instead of the non‑thinking default used elsewhere on this page.3 Three results so far (ifeval, mmlu_pro, bbh_cot_zeroshot) from a 12‑task rerun still in progress — most rows don't have one yet.
Unsloth official Qwen3.5‑9B weights, same harness/prompting, run 2026‑08‑22–23 (two tasks via the raw‑llama.cpp bypass).
Same official weights as the row above, thinking mode forced on — the fourth of the four core settings this page tracks (abliterated/official × thinking/non‑thinking).5 One result so far (mmlu_pro), meant to fill in for every benchmark over time like the abliterated, thinking row above it.
Numbers from Qwen's own model card, quoted as-is and not run by us — except cruxeval_input/cruxeval_output, where no Qwen number exists at all and this row instead shows a same‑scale substitute model (labeled by name, not "Qwen3.5‑9B").
One table per benchmark, all scores as percentages. Δ is the change versus the row above, chained abliterated→Qwen 3.5:9b — except the two thinking rows (where present), whose Δ is each against its own non‑thinking baseline rather than a further link in that chain: the abliterated, thinking row compares against the plain abliterated row above it, and the Qwen 3.5:9b, thinking row compares against the plain Qwen 3.5:9b row above it — two independent same‑model thinking‑on/off comparisons, not one chain of four. The posted row isn't chained in either — see footnote 6 for why that comparison needs its own caveats. A task's posted row is omitted entirely when no public number (or substitute) exists for it, rather than shown as a blank — see footnote 1. Some backlog tasks below have only the abliterated row filled in; a qwen comparison doesn't exist yet for those.
grade-school math word problems
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 82.5strict 82.6 | — |
| Abliterated, thinking (ours) | not reported | — |
| Qwen 3.5:9b (ours) | 85.7 ± 1.0 | ▲ 3.2 |
| Qwen 3.5:9b, thinking (ours) | not reported | — |
Python function generation from docstrings
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 61.6 ± 3.8 | — |
| Abliterated, thinking (ours) | not reported | — |
| Qwen 3.5:9b (ours) | 47.6 ± 3.9 | ▼ 14.0 |
| Qwen 3.5:9b, thinking (ours) | not reported | — |
basic Python programming problems
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 48.2 ± 2.2 | — |
| Abliterated, thinking (ours) | not reported | — |
| Qwen 3.5:9b (ours) | 48.0 ± 2.2 | ▼ 0.2 |
| Qwen 3.5:9b, thinking (ours) | not reported | — |
diverse reasoning — logic, algorithms, language
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 45.6 ± 0.5 | — |
| Abliterated, thinking (ours) | 79.9 ± 1.13 | ▲ 34.3 |
| Qwen 3.5:9b (ours) | 65.8 ± 0.5 | ▲ 20.2 |
| Qwen 3.5:9b, thinking (ours) | not reported | — |
broad academic knowledge, 57 subjects
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 62.0 ± 0.4 | — |
| Abliterated, thinking (ours) | 67.5 ± 1.63 | ▲ 5.5 |
| Qwen 3.5:9b (ours) | 64.7 ± 0.4 | ▲ 2.7 |
| Qwen 3.5:9b, thinking (ours) | 77.0 ± 1.45 | ▲ 12.3 |
| Qwen3.5‑9B (posted) | 82.51 | — |
instruction-following & formatting compliance
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 64.7strict‑p 48.8 · strict‑i 61.9 · loose‑p 52.7 | — |
| Abliterated, thinking (ours) | 85.73strict‑p 79.5 · strict‑i 83.9 · loose‑p 81.7 | ▲ 21.0 |
| Qwen 3.5:9b (ours) | 58.8strict‑p 44.2–44.42 · strict‑i 57.4–57.7 · loose‑p 45.7 | ▼ 5.9 |
| Qwen 3.5:9b, thinking (ours) | not reported | — |
| Qwen3.5‑9B (posted) | 91.51 | — |
news article summarization, single ROUGE score
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 7.9 | — |
| Abliterated, thinking (ours) | not reported | — |
| Qwen 3.5:9b (ours) | 10.0 | ▲ 2.2 |
| Qwen 3.5:9b, thinking (ours) | not reported | — |
extracting a structured field from FDA drug-label text
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 76.0 | — |
| Abliterated, thinking (ours) | not reported | — |
| Qwen 3.5:9b (ours) | 70.2 | ▼ 5.8 |
| Qwen 3.5:9b, thinking (ours) | 80.8 | ▲ 10.6 |
completing a passage's next span, SQuAD-derived
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 75.8 | — |
| Abliterated, thinking (ours) | not reported | — |
| Qwen 3.5:9b (ours) | 70.7 | ▼ 5.1 |
| Qwen 3.5:9b, thinking (ours) | not reported | — |
predicting a Python function's input from its output
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 44.1 ± 1.7 | — |
| Abliterated, thinking (ours) | not reported | — |
| Qwen 3.5:9b (ours) | 52.8 ± 1.7 | ▲ 8.7 |
| Qwen 3.5:9b, thinking (ours) | not reported | — |
| deepseek-instruct-6.7b (posted, ref) | 37.44 | — |
predicting a Python function's output from its input
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 27.6 ± 1.5 | — |
| Abliterated, thinking (ours) | not reported | — |
| Qwen 3.5:9b (ours) | 31.6 ± 1.6 | ▲ 4.0 |
| Qwen 3.5:9b, thinking (ours) | 54.3 ± 5.7 | ▲ 22.7 |
| deepseek-instruct-6.7b (posted, ref) | 41.24 | — |
long‑context multi‑hop QA, 20 reasoning‑complexity levels
| Source | Score | Δ |
|---|---|---|
| Abliterated (ours) | 75.8qa1–qa20 range 9.8–99.9 | — |
| Abliterated, thinking (ours) | not reported | — |
| Qwen 3.5:9b (ours) | 83.3qa1–qa20 range 11.2–99.8 | ▲ 7.5 |
| Qwen 3.5:9b, thinking (ours) | not reported | — |
Abliterated vs Qwen 3.5:9b
Abliterated, thinking vs Qwen 3.5:9b, thinking
Only 1 task compared so far — preliminary.
Only 2 tasks compared so far — preliminary.
Only 1 task compared so far — preliminary.
max_gen_toks=10000, no early stop) instead of the non‑thinking default used everywhere else on this page — see eval‑harness/run_thinking_batch.sh. They're the first three results back from a broader 12‑task thinking‑mode rerun still in progress (excludes humaneval_instruct/mbpp_instruct_fixed, where thinking mode is a confirmed no‑op under this harness's gen_prefix); more rows will fill in as that batch completes._closest_instruct_reference().eval‑harness/run_official_thinking_mmlu_pro.sh against the same 756‑sample subsample as the abliterated, thinking row above it — built specifically to isolate how much of the gap between abliterated, thinking (67.5) and Qwen's posted mmlu_pro number (82.5) is attributable to the abliterated build itself versus this local build's shared Q4_K_M quantization. It's the first result of a fourth core setting (official weights × thinking) this page now tracks alongside abliterated/qwen/abliterated‑thinking — more tasks will fill in over time the same way the abliterated, thinking rows have been.Per‑item speed for each run, wall‑clock runtime alongside it for reference. Not a clean speed benchmark — see the callout below before reading these as apples‑to‑apples.
-c is the same total KV pool either way; see hearth‑bypass/bypass_run_abliterated_thinking.bat's and bypass_run_thinking.bat's matching -np 8, and EVAL‑BACKLOG.md). The much higher per‑item time there is thinking tokens plus that lower concurrency, not a like‑for‑like comparison to the 16‑way non‑thinking rows. The five backlog tasks (fda/squad_completion/cruxeval_input/cruxeval_output/babilong) only have an abliterated row at all, all via the 16‑way bypass. Only compare rows where both sides used the same serving mode.
Bars are scaled per task (each task's own longest bar is 100% width) — runtimes span two orders of magnitude across tasks, so a single shared scale would flatten the short ones to nothing.
grade-school math word problems
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 2.00s/it16‑way bypass | 35m47s |
| Abliterated, thinking (ours, bypass) | not run | — |
| Qwen 3.5:9b (ours) | 6.32s/itserial | 2h19m |
| Qwen 3.5:9b, thinking (ours, bypass) | not run | — |
Python function generation from docstrings
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 1.00s/it16‑way bypass | 2m49s |
| Abliterated, thinking (ours, bypass) | not run | — |
| Qwen 3.5:9b (ours) | 5.28s/itserial | 14m26s |
| Qwen 3.5:9b, thinking (ours, bypass) | not run | — |
basic Python programming problems
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 1.00s/it16‑way bypass | 12m11s |
| Abliterated, thinking (ours, bypass) | not run | — |
| Qwen 3.5:9b (ours) | 7.42s/itserial | 1h02m |
| Qwen 3.5:9b, thinking (ours, bypass) | not run | — |
diverse reasoning — logic, algorithms, language
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 1.00s/it16‑way bypass | 2h03m |
| Abliterated, thinking (ours, bypass) | 35.00s/it4‑way bypass | 10h59m |
| Qwen 3.5:9b (ours) | 0.96s/it16‑way bypass | 1h44m |
| Qwen 3.5:9b, thinking (ours, bypass) | not run | — |
broad academic knowledge, 57 subjects
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 2.00s/it16‑way bypass | 5h53m |
| Abliterated, thinking (ours, bypass) | 24.91s/it4‑way bypass | 5h14m |
| Qwen 3.5:9b (ours) | 1.15s/it16‑way bypass | 3h51m |
| Qwen 3.5:9b, thinking (ours, bypass) | 29.00s/it4‑way bypass | 8h46m |
instruction-following & formatting compliance
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 3.00s/it16‑way bypass | 33m14s |
| Abliterated, thinking (ours, bypass) | 53.88s/it4‑way bypass | 8h06m |
| Qwen 3.5:9b (ours) | 18.77s/itserial | 2h49m |
| Qwen 3.5:9b, thinking (ours, bypass) | not run | — |
news article summarization, single ROUGE score
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 1.78s/itserial | 5h41m |
| Abliterated, thinking (ours, bypass) | not run | — |
| Qwen 3.5:9b (ours) | 2.11s/itserial | 6h43m |
| Qwen 3.5:9b, thinking (ours, bypass) | not run | — |
extracting a structured field from FDA drug-label text
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 0.01s/it16‑way bypass | 10.8s |
| Abliterated, thinking (ours, bypass) | not run | — |
| Qwen 3.5:9b (ours) | 1.68s/it16‑way bypass | 31m10s |
| Qwen 3.5:9b, thinking (ours, bypass) | 9.00s/it4‑way bypass | 1h52m30s |
completing a passage's next span, SQuAD-derived
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 0.00s/it16‑way bypass | 12.4s |
| Abliterated, thinking (ours, bypass) | not run | — |
| Qwen 3.5:9b (ours) | 1.13s/it16‑way bypass | 56m11s |
| Qwen 3.5:9b, thinking (ours, bypass) | not run | — |
predicting a Python function's input from its output
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 6.67s/it16‑way bypass | 58m52s |
| Abliterated, thinking (ours, bypass) | not run | — |
| Qwen 3.5:9b (ours) | 1.00s/it16‑way bypass | 2h55m37s |
| Qwen 3.5:9b, thinking (ours, bypass) | not run | — |
predicting a Python function's output from its input
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 5.00s/it16‑way bypass | 55m07s |
| Abliterated, thinking (ours, bypass) | not run | — |
| Qwen 3.5:9b (ours) | 0.60s/it16‑way bypass | 1h02m23s |
| Qwen 3.5:9b, thinking (ours, bypass) | 18.00s/it4‑way bypass | 22m30s |
long‑context multi‑hop QA, 20 reasoning‑complexity levels
| Source | Speed | Runtime |
|---|---|---|
| Abliterated (ours, bypass) | 0.00s/it16‑way bypass | 1m40s |
| Abliterated, thinking (ours, bypass) | not run | — |
| Qwen 3.5:9b (ours) | 0.19s/it16‑way bypass | 2h05m |
| Qwen 3.5:9b, thinking (ours, bypass) | not run | — |
Diffusion image generation · Apple M1 Pro, 16GB unified memory
Image-generation speed and memory across diffusion models, run locally on the MacBook's unified memory instead of the desktop's dedicated VRAM — 7 runs across 6 model families that fit comfortably in 16 GB timed so far.
~/diffusion-bench (not tracked in this repo — its venvs and downloaded weights alone run tens of GB; the sample images below are tracked, in images/). Six rows are timed with /usr/bin/time -l: five around an mflux-generate-* call (see run_frontier_batch.sh), and SDXL Turbo via a separate, non‑mflux pipeline (mlx‑examples' stable_diffusion, standardized 2026‑09‑06 — see its card's note). The remaining row (FLUX.1 schnell) comes from mflux's own embedded PNG metadata instead — each card says which. Only models that fit comfortably inside this machine's 16 GB of unified memory are shown below — a model tested but excluded for swapping heavily instead gets a plain callout explaining why, not a card with misleadingly slow numbers (see the notes below the grid). No image-quality scoring here (unlike the LLM tables above, there's no numeric metric for "does it look good") — judge that from the thumbnails, this section tracks generation speed and memory only.





| Metric | Value |
|---|---|
| Steps | 8 |
| Quantize | 8-bit |
| Generation time | 57.0s |
| Speed | 7.12s/it avg |
| Peak MLX memory | 13.44 GB |





| Metric | Value |
|---|---|
| Steps | 4 |
| Quantize | 8-bit |
| Generation time | 13.8s |
| Speed | 3.44s/it avg (6.90s cold → 2.21s/it by step 4) |
| Peak MLX memory | 12.74 GB |
Its text encoder is GPT‑OSS‑20B (mlx‑community/gpt‑oss‑20b‑MXFP4‑Q8), not a typical CLIP/T5 text encoder — a general‑purpose 20B language model doing prompt encoding for a 4‑step image model. Peak memory still fits comfortably despite that.





| Metric | Value |
|---|---|
| Steps | 4 |
| Quantize | 8-bit |
| Generation time | 1m28s |
| Speed | 22.1s/it avg |
| Peak MLX memory | 12.42 GB |
The larger sibling of FLUX.2 klein‑4B (9B vs. 4B parameters). The first attempt at this model failed outright on a HuggingFace 403 gated‑repo error; access opened up since, and this retry completed cleanly — still fitting comfortably despite the larger size, though 4.2× slower per step than its 4B sibling.





| Metric | Value |
|---|---|
| Steps | 4 |
| Quantize | 8-bit |
| Generation time | 20.8s |
| Speed | 5.21s/it avg (7.46s cold → 4.72s/it by step 4) |
| Peak MLX memory | 6.89 GB |





| Metric | Value |
|---|---|
| Steps | 4 |
| Quantize | 4-bit |
| Generation time | 30.1s |
| Speed | 7.54s/it avg |
| Peak MLX memory | 10.33 GB |
Loaded from a checkpoint pre‑quantized to 4‑bit and saved to disk, rather than re‑quantizing the full‑precision weights at load time — the representative FLUX.1 schnell setting on this page. Steps/speed come straight from mflux's own embedded PNG metadata (see this section's intro note); the original astronaut‑prompt run had no /usr/bin/time wrapper, so peak memory here instead comes from the later, separately‑wrapped combined‑prompt run of the same 4‑bit checkpoint (combined‑flux‑schnell‑q4.log).





| Metric | Value |
|---|---|
| Steps | 9 |
| Quantize | 8-bit |
| Generation time | 53.3s |
| Speed | 5.92s/it avg (10.37s cold → ~5.2–5.6s/it after) |
| Peak MLX memory | 8.74 GB |





| Metric | Value |
|---|---|
| Steps | 2 |
| Quantize | fp16 |
| Generation time | 30.2s |
| Speed | 8.91s/it avg (15.30s cold → 7.78s/it by step 2) |
| Peak MLX memory | 6.98 GB |
Standardized 2026‑09‑06 to match the cards above — originally shown as an uninstrumented gallery (no /usr/bin/time, no PNG metadata from this pipeline; see git history for that version). This card's numbers come from a fresh, properly‑wrapped rerun with the same seed and prompt, confirmed byte‑identical to the original image (same file, sdxl-turbo-warm.png). Native output resolution for this pipeline is 528×528, not the 512×512 every other card uses — not operator‑configurable via this script's CLI, so shown as-is rather than cropped/resized to match.
Ideogram 4 (fp8) (ideogram‑ai/ideogram‑4‑fp8, both 8‑bit and 4‑bit), Qwen‑Image (Qwen/Qwen‑Image‑2512, already 4‑bit), FIBO Lite (briaai/Fibo‑lite, 8‑bit), and Krea 2 Turbo (krea/Krea‑2‑Turbo, 8‑bit) were also tested and all completed — but all five runs are excluded from the comparison above rather than shown alongside models that actually fit. Peak memory ran 28.28 GB (Ideogram 8‑bit), 27.36 GB (Ideogram 4‑bit), 27.37 GB (Qwen‑Image), 24.38 GB (FIBO Lite), and 22.48 GB (Krea 2 Turbo), all well past this machine's 16 GB of unified memory, so most of each run's time was spent swapping to disk, not generating — the 88.6s/it, 40.3s/it, 109.7s/it, 21.8s/it, and 48.9s/it measured aren't a fair read on these models' real speed, just how slow this particular machine is at thrashing its way through them.
| Model | Peak Memory | Fits 16 GB? |
|---|---|---|
| SDXL-Turbo | 6.98 GB | ✅ |
| FLUX.1 schnell (4-bit, saved standalone) | 10.33 GB | ✅ |
| FLUX.2 klein-4B | 6.89 GB | ✅ |
| FLUX.2 klein-9B | 12.42 GB | ✅ |
| Z-Image Turbo | 8.74 GB | ✅ |
| ERNIE Image Turbo (Baidu) | 13.44 GB | ✅ |
| Lens (Microsoft) | 12.74 GB | ✅ |
| Ideogram 4 (fp8) | 27–28 GB | ❌ |
| Qwen-Image-2512 | 27.37 GB | ❌ (but best quality) |
| FIBO Lite (Bria AI) | 24.38 GB | ❌ |
| Krea 2 Turbo | 22.48 GB | ❌ (breaks the "Turbo" heuristic) |