Local inference

Benchmarks

Local LLMs served on the desktop's RTX 3060. Diffusion image generation on the MacBook's M1.

Qwen3.5:9B

Desktop Computer: RTX 3060 (12GB)

queue idle · updated 2:40am PDT
Abliterated (ours)

huihui-ai abliterated build — the Qwen3.5‑9B base model with abliteration applied — independently‑verified genuine GGUF, served via the 16‑way raw‑llama.cpp bypass, 2026‑08‑23 onward — 12 of 12 tasks done.

Abliterated, thinking (ours)

Same abliterated build, thinking mode forced on instead of the non‑thinking default used elsewhere on this page.3 Five results so far (ifeval, mmlu_pro, bbh_cot_zeroshot, cruxeval_input, cruxeval_output) from a 12‑task rerun still in progress — most rows don't have one yet.

Qwen 3.5:9b (ours)

Unsloth official Qwen3.5‑9B weights, same harness/prompting, run 2026‑08‑22–23 (two tasks via the raw‑llama.cpp bypass).

Qwen 3.5:9b, thinking (ours)

Same official weights as the row above, thinking mode forced on — the fourth of the four core settings this page tracks (abliterated/official × thinking/non‑thinking).5 One result so far (mmlu_pro), meant to fill in for every benchmark over time like the abliterated, thinking row above it.

Qwen3.5‑9B (posted)

Numbers from Qwen's own model card, quoted as-is and not run by us — except cruxeval_input/cruxeval_output, where no Qwen number exists at all and this row instead shows a same‑scale substitute model (labeled by name, not "Qwen3.5‑9B").

Score comparison

One table per benchmark, all scores as percentages. Each Δ is labeled with what it's measured against, since the comparison isn't the same for every row: Qwen 3.5:9b's Δ is versus the plain Abliterated row (does swapping to official weights change the score), while each thinking row's Δ is versus its own non‑thinking baseline directly above it (does thinking mode change the score) — two independent comparisons, not one chained run down the table. Every Δ is computed straight from the two rounded scores shown, so it always matches subtracting them by hand. The posted row is never delta'd — see footnote 6 for why that comparison needs its own caveats. A task's posted row is omitted entirely when no public number (or substitute) exists for it, rather than shown as a blank — see footnote 1. Some backlog tasks below have only the abliterated row filled in; a qwen comparison doesn't exist yet for those.

Score comparison chart
gsm8k
82.5
87.3 ± 1.2
85.7 ± 1.0
not reported
bbh_cot_zeroshot
45.6 ± 0.5
79.9 ± 1.1
65.8 ± 0.5
89.1 ± 0.9
mmlu_pro
62.0 ± 0.4
67.5 ± 1.6
64.7 ± 0.4
77.0 ± 1.4
82.5
ifeval
64.7
85.7
58.8
94.1
91.5
cnn_dailymail
7.9
not reported
10.0
not reported
fda
76.0
not reported
70.2
80.8
squad_completion
75.8
92.0 ± 1.0
70.7
90.0 ± 1.1
cruxeval_input
44.1 ± 1.7
47.9 ± 3.2
52.8 ± 1.7
89.9 ± 2.6
37.4
cruxeval_output
27.6 ± 1.5
43.1 ± 5.1
31.6 ± 1.6
54.3 ± 5.7
41.2
humaneval_instruct_free
39.0 ± 3.8
82.9 ± 2.9
34.8 ± 3.7
97.0 ± 1.3
mbpp_instruct_free
13.2 ± 1.5
67.8 ± 2.1
9.0 ± 1.3
81.2 ± 1.7
babilong
75.8
91.2
83.3
91.5

gsm8kexact_match · 1,319 samples

grade-school math word problems

SourceScoreΔ
Abliterated (ours)82.5strict 82.6
Abliterated, thinking (ours)87.3 ± 1.2strict 84.9▲ 4.8vs non‑think
Qwen 3.5:9b (ours)85.7 ± 1.0▲ 3.2vs abliterated
Qwen 3.5:9b, thinking (ours)not reported

bbh_cot_zeroshotexact_match, flexible-extract · 6,511 samples

diverse reasoning — logic, algorithms, language

SourceScoreΔ
Abliterated (ours)45.6 ± 0.5
Abliterated, thinking (ours)79.9 ± 1.13▲ 34.3vs non‑think
Qwen 3.5:9b (ours)65.8 ± 0.5▲ 20.2vs abliterated
Qwen 3.5:9b, thinking (ours)89.1 ± 0.93▲ 23.3vs non‑think

mmlu_proexact_match · 12,032 samples

broad academic knowledge, 57 subjects

SourceScoreΔ
Abliterated (ours)62.0 ± 0.4
Abliterated, thinking (ours)67.5 ± 1.63▲ 5.5vs non‑think
Qwen 3.5:9b (ours)64.7 ± 0.4▲ 2.7vs abliterated
Qwen 3.5:9b, thinking (ours)77.0 ± 1.45▲ 12.3vs non‑think
Qwen3.5‑9B (posted)82.51

ifevalloose‑instruction acc · 541 samples

instruction-following & formatting compliance

SourceScoreΔ
Abliterated (ours)64.7strict‑p 48.8 · strict‑i 61.9 · loose‑p 52.7
Abliterated, thinking (ours)85.73strict‑p 79.5 · strict‑i 83.9 · loose‑p 81.7▲ 21.0vs non‑think
Qwen 3.5:9b (ours)58.8strict‑p 44.2–44.42 · strict‑i 57.4–57.7 · loose‑p 45.7▼ 5.9vs abliterated
Qwen 3.5:9b, thinking (ours)94.1strict‑p 88.7 · strict‑i 91.7 · loose‑p 91.9▲ 35.3vs non‑think
Qwen3.5‑9B (posted)91.51

cnn_dailymailrouge · 11,490 samples

news article summarization, single ROUGE score

SourceScoreΔ
Abliterated (ours)7.9
Abliterated, thinking (ours)not reported
Qwen 3.5:9b (ours)10.0▲ 2.1vs abliterated
Qwen 3.5:9b, thinking (ours)not reported

fdacontains · 1,102 samples

extracting a structured field from FDA drug-label text

SourceScoreΔ
Abliterated (ours)76.0
Abliterated, thinking (ours)not reported
Qwen 3.5:9b (ours)70.2▼ 5.8vs abliterated
Qwen 3.5:9b, thinking (ours)80.8▲ 10.6vs non‑think

squad_completioncontains · 2,984 samples

completing a passage's next span, SQuAD-derived

SourceScoreΔ
Abliterated (ours)75.8
Abliterated, thinking (ours)92.0 ± 1.03▲ 16.2vs non‑think
Qwen 3.5:9b (ours)70.7▼ 5.1vs abliterated
Qwen 3.5:9b, thinking (ours)90.0 ± 1.1▲ 19.3vs non‑think

cruxeval_inputpass@1 · 800 samples

predicting a Python function's input from its output

SourceScoreΔ
Abliterated (ours)44.1 ± 1.7
Abliterated, thinking (ours)47.9 ± 3.23▲ 3.8vs non‑think
Qwen 3.5:9b (ours)52.8 ± 1.7▲ 8.7vs abliterated
Qwen 3.5:9b, thinking (ours)89.9 ± 2.6▲ 37.1vs non‑think
deepseek-instruct-6.7b (posted, ref)37.44

cruxeval_outputpass@1 · 800 samples

predicting a Python function's output from its input

SourceScoreΔ
Abliterated (ours)27.6 ± 1.5
Abliterated, thinking (ours)43.1 ± 5.13▲ 15.5vs non‑think
Qwen 3.5:9b (ours)31.6 ± 1.6▲ 4.0vs abliterated
Qwen 3.5:9b, thinking (ours)54.3 ± 5.7▲ 22.7vs non‑think
deepseek-instruct-6.7b (posted, ref)41.24

humaneval_instruct_freepass@1 · 164 samples

generating a Python function from a natural-language spec, free-form code extraction (HumanEval-derived)

SourceScoreΔ
Abliterated (ours)39.0 ± 3.8
Abliterated, thinking (ours)82.9 ± 2.9▲ 43.9vs non‑think
Qwen 3.5:9b (ours)34.8 ± 3.7▼ 4.2vs abliterated
Qwen 3.5:9b, thinking (ours)97.0 ± 1.3▲ 62.2vs non‑think

mbpp_instruct_freepass@1 · 500 samples

generating a Python function from a natural-language spec, free-form code extraction (MBPP-derived)

SourceScoreΔ
Abliterated (ours)13.2 ± 1.5
Abliterated, thinking (ours)67.8 ± 2.1▲ 54.6vs non‑think
Qwen 3.5:9b (ours)9.0 ± 1.3▼ 4.2vs abliterated
Qwen 3.5:9b, thinking (ours)81.2 ± 1.7▲ 72.2vs non‑think

babilongacc (20‑subtask weighted avg) · 39,972 samples

long‑context multi‑hop QA, 20 reasoning‑complexity levels

SourceScoreΔ
Abliterated (ours)75.8qa1–qa20 range 9.8–99.9
Abliterated, thinking (ours)91.2qa1–qa20 range 58.0–100.0▲ 15.4vs non‑think
Qwen 3.5:9b (ours)83.3qa1–qa20 range 11.2–99.8▲ 7.5vs abliterated
Qwen 3.5:9b, thinking (ours)91.5qa1–qa20 range 62.0–100.0▲ 8.2vs non‑think

Abliterated vs Qwen 3.5:9b

row beats
column
Abliterated
Qwen 3.5:9b
Abliterated
5
Qwen 3.5:9b
7

Abliterated, thinking vs Qwen 3.5:9b, thinking

row beats
column
Abliterated, thinking
Qwen 3.5:9b, thinking
Abliterated, thinking
1
Qwen 3.5:9b, thinking
8

Abliterated vs Qwen 3.5:9b

row beats
column
Abliterated
Qwen 3.5:9b
Abliterated
0
Qwen 3.5:9b
2

Only 2 tasks compared so far — preliminary.

Abliterated, thinking vs Qwen 3.5:9b, thinking

row beats
column
Abliterated, thinking
Qwen 3.5:9b, thinking
Abliterated, thinking
0
Qwen 3.5:9b, thinking
2

Only 2 tasks compared so far — preliminary.

Abliterated vs Qwen 3.5:9b

row beats
column
Abliterated
Qwen 3.5:9b
Abliterated
2
Qwen 3.5:9b
2

Abliterated, thinking vs Qwen 3.5:9b, thinking

row beats
column
Abliterated, thinking
Qwen 3.5:9b, thinking
Abliterated, thinking
0
Qwen 3.5:9b, thinking
4

Abliterated vs Qwen 3.5:9b

row beats
column
Abliterated
Qwen 3.5:9b
Abliterated
0
Qwen 3.5:9b
1

Only 1 task compared so far — preliminary.

Abliterated, thinking vs Qwen 3.5:9b, thinking

No task has both a Abliterated, thinking and a Qwen 3.5:9b, thinking result yet.

Abliterated vs Qwen 3.5:9b

row beats
column
Abliterated
Qwen 3.5:9b
Abliterated
3
Qwen 3.5:9b
2

Abliterated, thinking vs Qwen 3.5:9b, thinking

row beats
column
Abliterated, thinking
Qwen 3.5:9b, thinking
Abliterated, thinking
1
Qwen 3.5:9b, thinking
2
  1. mmlu_pro and ifeval are the only two rows that are genuine name‑for‑name matches with a posted number — and even then, exact prompt version and scoring protocol aren't guaranteed identical. Treat the gap's size as directional, not exact. Every other task has no posted row at all: no public Qwen3.5 number exists for it anywhere we could find (checked Qwen's own model card, and for cruxeval/babilong, their independent community leaderboards directly — cached in external_data/), except cruxeval_input/cruxeval_output, which get a substitute instead of a blank row — see footnote 4.
  2. ifeval's strict submetrics show small run‑to‑run nondeterminism on identical cached inputs (four replays: strict‑prompt 43.99–44.36, strict‑inst 57.43–57.67). Loose metrics were stable. The spread is far smaller than the gaps in this table, but don't read these two columns to more than one decimal place.
  3. ifeval's, mmlu_pro's, and bbh_cot_zeroshot's "Abliterated, thinking" rows are the same abliterated build with thinking mode forced on (max_gen_toks=10000, no early stop) instead of the non‑thinking default used everywhere else on this page — see eval‑harness/run_thinking_batch.sh. They're the first three results back from a broader 12‑task thinking‑mode rerun still in progress; more rows will fill in as that batch completes. (humaneval_instruct and mbpp_instruct_fixed were dropped from this page entirely, not merely excluded from that rerun — both used lm_eval's gen_prefix mechanic, which was confirmed to make thinking mode a silent no‑op, so their non‑thinking scores weren't a fair baseline to keep either. They've been superseded by humaneval_instruct_free/mbpp_instruct_free, free‑generation rewrites without gen_prefix, once those finish their own full 2×2 run set.) bbh_cot_zeroshot's two thinking rows are additionally capped at ‑‑limit 30 docs per subtask (≈810 of the full 6,511‑doc set, versus every other row on this task — and every other task's thinking row — covering its full corpus) — flagged 2026‑09‑12 during a benchmark‑page speed‑figures audit, with no on‑page disclosure until now. A full‑coverage rerun is planned (see EVAL‑BACKLOG.md's roadmap section) but not yet queued; treat these two numbers as a smaller‑sample estimate, not a like‑for‑like comparison to this task's own non‑thinking rows.
  4. cruxeval_input/cruxeval_output's posted row is a substitute, not a Qwen number: no Qwen model appears on the independent CRUXEval leaderboard (crux‑eval.github.io) at all, so the row instead shows the closest‑parameter‑count instruct model there, picked automatically from the cached leaderboard dump rather than hand‑typed — see generate.py's _closest_instruct_reference().
  5. mmlu_pro's "Qwen 3.5:9b, thinking" row is the official (non‑abliterated) weights with thinking mode forced on, run via eval‑harness/run_official_thinking_mmlu_pro.sh against the same 756‑sample subsample as the abliterated, thinking row above it — built specifically to isolate how much of the gap between abliterated, thinking (67.5) and Qwen's posted mmlu_pro number (82.5) is attributable to the abliterated build itself versus this local build's shared Q4_K_M quantization. It's the first result of a fourth core setting (official weights × thinking) this page now tracks alongside abliterated/qwen/abliterated‑thinking — more tasks will fill in over time the same way the abliterated, thinking rows have been.
  6. The "posted" column isn't part of any head‑to‑head count on this page (scorecards above, or otherwise): Qwen's posted numbers are almost certainly thinking‑mode (the model's default), while ours are forced non‑thinking, and ours also run Q4_K_M quantized rather than full precision. The gap to "posted" reflects both of those on top of any real capability difference — every win/loss count on this page only ever compares our own runs against each other.

Runtime & throughput

Per‑item speed for each run, wall‑clock runtime alongside it for reference. Not a clean speed benchmark — see the callout below before reading these as apples‑to‑apples.

Two different serving conditions, not one axis: Qwen 3.5:9b ran serially for gsm8k/ifeval/humaneval_instruct/mbpp_instruct_fixed/cnn_dailymail, but through the 16‑way raw‑llama.cpp bypass for mmlu_pro/bbh_cot_zeroshot — two different axes inside the same column. Abliterated runs through the 16‑way bypass for every task except cnn_dailymail, which ran serially like its qwen row instead, so that pair is still apples‑to‑apples. The abliterated, thinking row (ifeval, mmlu_pro, and bbh_cot_zeroshot so far) and the Qwen 3.5:9b, thinking row (mmlu_pro so far) both run through the same bypass build at 8‑way, not 16‑way — sized down from 16 for comfortable worst‑case context headroom on long thinking traces, not a VRAM tradeoff (-c is the same total KV pool either way; see hearth‑bypass/bypass_run_abliterated_thinking.bat's and bypass_run_thinking.bat's matching -np 8, and EVAL‑BACKLOG.md). The much higher per‑item time there is thinking tokens plus that lower concurrency, not a like‑for‑like comparison to the 16‑way non‑thinking rows. The five backlog tasks (fda/squad_completion/cruxeval_input/cruxeval_output/babilong) only have an abliterated row at all, all via the 16‑way bypass. Only compare rows where both sides used the same serving mode.
Runtime/speed reconstructed from raw logs, 2026‑08‑29: these jobs get killed and restarted (digest bypass‑protection blackout windows, manual interrupts), so lm_eval's own end‑to‑end wall‑clock was including dead time between restarts — nothing to do with how fast the model actually processed items. The abliterated² figures for gsm8k/humaneval_instruct/mbpp_instruct_fixed/bbh_cot_zeroshot/mmlu_pro/ifeval and the backlog cruxeval_input/cruxeval_output rows are now reconstructed straight from each surviving log's tqdm ticks (see `speed_stats.py`): summed real per‑tick elapsed deltas for the runtime, median per‑tick rate (not a global average) for the speed, so a stray near‑instant cache‑hit or ramp‑up tick can't skew it. All eight logs showed a single continuous segment with no restart evidence, so the change here is small — mostly stripping non‑generation setup/teardown overhead. Two rows were deliberately left unchanged despite being asked for: ifeval and mmlu_pro's abliterated², thinking figures, because their logs show a genuine crash/partial run with no way to cleanly reconstruct true generation time from what survived on disk — guessing would be worse than leaving them at the old (also imperfect) number.
Three rows are cache replays, not measurements, 2026‑09‑12: fda/squad_completion/babilong's abliterated² rows used to show "0.01s/it"/"0.00s/it"/"0.00s/it" — not the model running fast, but lm_eval's own on‑disk request cache satisfying every single request (each log's own Cached requests: N, Requests remaining: 0 line confirms 100% hits, and each log has zero tqdm progress‑bar ticks at all to reconstruct from). The real first‑time generation that originally populated that cache was never captured with real timing — its log was overwritten by the pre‑2026‑08‑29 truncation issue described above, before there was anything to filter. speed_stats.py now drops any tick implying faster than 1s/item from every sum (not just the median), on the reasoning that a cache hit's near‑zero interval has nothing to do with model speed and shouldn't be averaged in as if it did; for a partially‑contaminated log this recovers a real median from whatever ticks survive, but these three logs have no surviving ticks at all. Borrowing a rate from these tasks' own "qwen" rows (which do have real generation) was tried and rejected: those runs are themselves so cache‑heavy that only 0–5 real intervals survive per task, too few and too inconsistent (4.0s/it vs 69.0s/it vs no data) to trust as a substitute. So these three cells now read "cache replay" with no number attached, rather than either the old misleading figure or a fabricated estimate built from a handful of samples.
Runtime comparison chart

Bars are scaled per task (each task's own longest bar is 100% width) — runtimes span two orders of magnitude across tasks, so a single shared scale would flatten the short ones to nothing.

gsm8k
35m47s
9h10m
2h19m
not reported
bbh_cot_zeroshot
2h03m
10h59m
1h44m
5h19m47s
mmlu_pro
5h53m
5h14m
3h51m
8h46m
ifeval
33m14s
8h06m
2h49m
4h08m15s
cnn_dailymail
5h41m
not reported
6h43m
not reported
fda
cache replay
not reported
31m10s
1h52m30s
squad_completion
cache replay
7h44m50s
56m11s
4h43m54s
cruxeval_input
58m52s
38m59s
2h55m37s
33m07s
cruxeval_output
55m07s
39m48s
1h02m23s
22m30s
humaneval_instruct_free
54m50s
1h13m20s
28m55s
1h05m50s
mbpp_instruct_free
9m10s
2h38m20s
11m25s
2h00m24s
babilong
cache replay
7h03m06s
2h05m
6h03m32s

gsm8kexact_match · 1,319 samples

grade-school math word problems

SourceSpeedRuntime
Abliterated (ours, bypass)2.00s/it16‑way bypass35m47s
Abliterated, thinking (ours, bypass)44.00s/it8‑way bypass9h10m
Qwen 3.5:9b (ours)6.32s/itserial2h19m
Qwen 3.5:9b, thinking (ours, bypass)not run

bbh_cot_zeroshotexact_match, flexible-extract · 6,511 samples

diverse reasoning — logic, algorithms, language

SourceSpeedRuntime
Abliterated (ours, bypass)1.00s/it16‑way bypass2h03m
Abliterated, thinking (ours, bypass)35.00s/it4‑way bypass10h59m
Qwen 3.5:9b (ours)0.96s/it16‑way bypass1h44m
Qwen 3.5:9b, thinking (ours, bypass)16.00s/it4‑way bypass5h19m47s

mmlu_proexact_match · 12,032 samples

broad academic knowledge, 57 subjects

SourceSpeedRuntime
Abliterated (ours, bypass)2.00s/it16‑way bypass5h53m
Abliterated, thinking (ours, bypass)24.91s/it4‑way bypass5h14m
Qwen 3.5:9b (ours)1.15s/it16‑way bypass3h51m
Qwen 3.5:9b, thinking (ours, bypass)29.00s/it4‑way bypass8h46m

ifevalloose‑instruction acc · 541 samples

instruction-following & formatting compliance

SourceSpeedRuntime
Abliterated (ours, bypass)3.00s/it16‑way bypass33m14s
Abliterated, thinking (ours, bypass)53.88s/it4‑way bypass8h06m
Qwen 3.5:9b (ours)18.77s/itserial2h49m
Qwen 3.5:9b, thinking (ours, bypass)27.53s/it4‑way bypass4h08m15s

cnn_dailymailrouge · 11,490 samples

news article summarization, single ROUGE score

SourceSpeedRuntime
Abliterated (ours, bypass)1.78s/itserial5h41m
Abliterated, thinking (ours, bypass)not run
Qwen 3.5:9b (ours)2.11s/itserial6h43m
Qwen 3.5:9b, thinking (ours, bypass)not run

fdacontains · 1,102 samples

extracting a structured field from FDA drug-label text

SourceSpeedRuntime
Abliterated (ours, bypass)cache replay 100% cache hits, no timing data
Abliterated, thinking (ours, bypass)not run
Qwen 3.5:9b (ours)1.68s/it16‑way bypass31m10s
Qwen 3.5:9b, thinking (ours, bypass)9.00s/it4‑way bypass1h52m30s

squad_completioncontains · 2,984 samples

completing a passage's next span, SQuAD-derived

SourceSpeedRuntime
Abliterated (ours, bypass)cache replay 100% cache hits, no timing data
Abliterated, thinking (ours, bypass)37.19s/it4‑way bypass7h44m50s
Qwen 3.5:9b (ours)1.13s/it16‑way bypass56m11s
Qwen 3.5:9b, thinking (ours, bypass)22.71s/it4‑way bypass4h43m54s

cruxeval_inputpass@1 · 800 samples

predicting a Python function's input from its output

SourceSpeedRuntime
Abliterated (ours, bypass)6.67s/it16‑way bypass58m52s
Abliterated, thinking (ours, bypass)31.18s/it4‑way bypass38m59s
Qwen 3.5:9b (ours)1.00s/it16‑way bypass2h55m37s
Qwen 3.5:9b, thinking (ours, bypass)26.49s/it4‑way bypass33m07s

cruxeval_outputpass@1 · 800 samples

predicting a Python function's output from its input

SourceSpeedRuntime
Abliterated (ours, bypass)5.00s/it16‑way bypass55m07s
Abliterated, thinking (ours, bypass)31.83s/it4‑way bypass39m48s
Qwen 3.5:9b (ours)0.60s/it16‑way bypass1h02m23s
Qwen 3.5:9b, thinking (ours, bypass)18.00s/it4‑way bypass22m30s

humaneval_instruct_freepass@1 · 164 samples

generating a Python function from a natural-language spec, free-form code extraction (HumanEval-derived)

SourceSpeedRuntime
Abliterated (ours, bypass)5.00s/it16‑way bypass54m50s
Abliterated, thinking (ours, bypass)8.00s/it8‑way bypass1h13m20s
Qwen 3.5:9b (ours)5.00s/it16‑way bypass28m55s
Qwen 3.5:9b, thinking (ours, bypass)10.00s/it8‑way bypass1h05m50s

mbpp_instruct_freepass@1 · 500 samples

generating a Python function from a natural-language spec, free-form code extraction (MBPP-derived)

SourceSpeedRuntime
Abliterated (ours, bypass)1.00s/it16‑way bypass9m10s
Abliterated, thinking (ours, bypass)19.00s/it8‑way bypass2h38m20s
Qwen 3.5:9b (ours)1.00s/it16‑way bypass11m25s
Qwen 3.5:9b, thinking (ours, bypass)14.00s/it8‑way bypass2h00m24s

babilongacc (20‑subtask weighted avg) · 39,972 samples

long‑context multi‑hop QA, 20 reasoning‑complexity levels

SourceSpeedRuntime
Abliterated (ours, bypass)cache replay 100% cache hits, no timing data
Abliterated, thinking (ours, bypass)26.44s/it8‑way bypass7h03m06s
Qwen 3.5:9b (ours)0.19s/it16‑way bypass2h05m
Qwen 3.5:9b, thinking (ours, bypass)21.81s/it8‑way bypass6h03m32s

Diffusion image generation · Apple M1 Pro, 16GB unified memory

Macbook: M1 (16GB Unified)

Image-generation speed and memory across diffusion models, run locally on the MacBook's unified memory instead of the desktop's dedicated VRAM — 7 runs across 6 model families that fit comfortably in 16 GB timed so far.

First pass, 2026‑09‑05: one image per run — 512×512 except SDXL Turbo, whose pipeline's native output is 528×528 (see its card's note) — same prompt and seed (42) everywhere metadata or timing is available: "a photorealistic portrait of an astronaut riding a horse on the moon, dramatic lighting, highly detailed, 8k" — run out of ~/diffusion-bench (not tracked in this repo — its venvs and downloaded weights alone run tens of GB; the sample images below are tracked, in images/). Six rows are timed with /usr/bin/time -l: five around an mflux-generate-* call (see run_frontier_batch.sh), and SDXL Turbo via a separate, non‑mflux pipeline (mlx‑examples' stable_diffusion, standardized 2026‑09‑06 — see its card's note). The remaining row (FLUX.1 schnell) comes from mflux's own embedded PNG metadata instead — each card says which. Only models that fit comfortably inside this machine's 16 GB of unified memory are shown below — a model tested but excluded for swapping heavily instead gets a plain callout explaining why, not a card with misleadingly slow numbers (see the notes below the grid). No image-quality scoring here (unlike the LLM tables above, there's no numeric metric for "does it look good") — judge that from the thumbnails, this section tracks generation speed and memory only.

ERNIE Image Turbobaidu/ERNIE-Image-Turbo

ERNIE Image Turbo, astronaut promptERNIE Image Turbo, combined bakery-text + kneading-hands promptERNIE Image Turbo, watercolor fox-and-bird promptERNIE Image Turbo, Seoul cyberpunk night market promptERNIE Image Turbo, octopus vs. raccoon chess prompt
MetricValue
Steps8
Quantize8-bit
Generation time57.0s
Speed7.12s/it avg
Peak MLX memory13.44 GB

LensComfy-Org/Lens

Lens, astronaut promptLens, combined bakery-text + kneading-hands promptLens, watercolor fox-and-bird promptLens, Seoul cyberpunk night market promptLens, octopus vs. raccoon chess prompt
MetricValue
Steps4
Quantize8-bit
Generation time13.8s
Speed3.44s/it avg (6.90s cold → 2.21s/it by step 4)
Peak MLX memory12.74 GB

Its text encoder is GPT‑OSS‑20B (mlx‑community/gpt‑oss‑20b‑MXFP4‑Q8), not a typical CLIP/T5 text encoder — a general‑purpose 20B language model doing prompt encoding for a 4‑step image model. Peak memory still fits comfortably despite that.

FLUX.2 klein-9Bblack-forest-labs/FLUX.2-klein-9B

FLUX.2 klein-9B, astronaut promptFLUX.2 klein-9B, combined bakery-text + kneading-hands promptFLUX.2 klein-9B, watercolor fox-and-bird promptFLUX.2 klein-9B, Seoul cyberpunk night market promptFLUX.2 klein-9B, octopus vs. raccoon chess prompt
MetricValue
Steps4
Quantize8-bit
Generation time1m28s
Speed22.1s/it avg
Peak MLX memory12.42 GB

The larger sibling of FLUX.2 klein‑4B (9B vs. 4B parameters). The first attempt at this model failed outright on a HuggingFace 403 gated‑repo error; access opened up since, and this retry completed cleanly — still fitting comfortably despite the larger size, though 4.2× slower per step than its 4B sibling.

FLUX.2 klein-4Bblack-forest-labs/FLUX.2-klein-4B

FLUX.2 klein-4B, astronaut promptFLUX.2 klein-4B, combined bakery-text + kneading-hands promptFLUX.2 klein-4B, watercolor fox-and-bird promptFLUX.2 klein-4B, Seoul cyberpunk night market promptFLUX.2 klein-4B, octopus vs. raccoon chess prompt
MetricValue
Steps4
Quantize8-bit
Generation time20.8s
Speed5.21s/it avg (7.46s cold → 4.72s/it by step 4)
Peak MLX memory6.89 GB

FLUX.1 schnell (q4)black-forest-labs/FLUX.1-schnell

FLUX.1 schnell (q4), astronaut promptFLUX.1 schnell (q4), combined bakery-text + kneading-hands promptFLUX.1 schnell (q4), watercolor fox-and-bird promptFLUX.1 schnell (q4), Seoul cyberpunk night market promptFLUX.1 schnell (q4), octopus vs. raccoon chess prompt
MetricValue
Steps4
Quantize4-bit
Generation time30.1s
Speed7.54s/it avg
Peak MLX memory10.33 GB

Loaded from a checkpoint pre‑quantized to 4‑bit and saved to disk, rather than re‑quantizing the full‑precision weights at load time — the representative FLUX.1 schnell setting on this page. Steps/speed come straight from mflux's own embedded PNG metadata (see this section's intro note); the original astronaut‑prompt run had no /usr/bin/time wrapper, so peak memory here instead comes from the later, separately‑wrapped combined‑prompt run of the same 4‑bit checkpoint (combined‑flux‑schnell‑q4.log).

Z-Image TurboTongyi-MAI/Z-Image-Turbo

Z-Image Turbo, astronaut promptZ-Image Turbo, combined bakery-text + kneading-hands promptZ-Image Turbo, watercolor fox-and-bird promptZ-Image Turbo, Seoul cyberpunk night market promptZ-Image Turbo, octopus vs. raccoon chess prompt
MetricValue
Steps9
Quantize8-bit
Generation time53.3s
Speed5.92s/it avg (10.37s cold → ~5.2–5.6s/it after)
Peak MLX memory8.74 GB

SDXL Turbostabilityai/sdxl-turbo

SDXL Turbo, astronaut promptSDXL Turbo, combined bakery-text + kneading-hands promptSDXL Turbo, watercolor fox-and-bird promptSDXL Turbo, Seoul cyberpunk night market promptSDXL Turbo, octopus vs. raccoon chess prompt
MetricValue
Steps2
Quantizefp16
Generation time30.2s
Speed8.91s/it avg (15.30s cold → 7.78s/it by step 2)
Peak MLX memory6.98 GB

Standardized 2026‑09‑06 to match the cards above — originally shown as an uninstrumented gallery (no /usr/bin/time, no PNG metadata from this pipeline; see git history for that version). This card's numbers come from a fresh, properly‑wrapped rerun with the same seed and prompt, confirmed byte‑identical to the original image (same file, sdxl-turbo-warm.png). Native output resolution for this pipeline is 528×528, not the 512×512 every other card uses — not operator‑configurable via this script's CLI, so shown as-is rather than cropped/resized to match.

Ideogram 4 (fp8) (ideogram‑ai/ideogram‑4‑fp8, both 8‑bit and 4‑bit), Qwen‑Image (Qwen/Qwen‑Image‑2512, already 4‑bit), FIBO Lite (briaai/Fibo‑lite, 8‑bit), and Krea 2 Turbo (krea/Krea‑2‑Turbo, 8‑bit) were also tested and all completed — but all five runs are excluded from the comparison above rather than shown alongside models that actually fit. Peak memory ran 28.28 GB (Ideogram 8‑bit), 27.36 GB (Ideogram 4‑bit), 27.37 GB (Qwen‑Image), 24.38 GB (FIBO Lite), and 22.48 GB (Krea 2 Turbo), all well past this machine's 16 GB of unified memory, so most of each run's time was spent swapping to disk, not generating — the 88.6s/it, 40.3s/it, 109.7s/it, 21.8s/it, and 48.9s/it measured aren't a fair read on these models' real speed, just how slow this particular machine is at thrashing its way through them.

Summary: everything tested

ModelPeak MemoryFits 16 GB?
SDXL-Turbo6.98 GB
FLUX.1 schnell (4-bit, saved standalone)10.33 GB
FLUX.2 klein-4B6.89 GB
FLUX.2 klein-9B12.42 GB
Z-Image Turbo8.74 GB
ERNIE Image Turbo (Baidu)13.44 GB
Lens (Microsoft)12.74 GB
Ideogram 4 (fp8)27–28 GB
Qwen-Image-251227.37 GB(but best quality)
FIBO Lite (Bria AI)24.38 GB
Krea 2 Turbo22.48 GB(breaks the "Turbo" heuristic)