mlx-optiq
Measured · Gemma-4-12B · M3 Max

OptiQ vs Ollama, LM Studio and other 4-bit quants

One model, Gemma-4-12B, served by OptiQ, Ollama and LM Studio on the same M3 Max, then quantized four ways. OptiQ decodes 1.7× faster on prose and up to 2.9× faster on code edits, with identical output, and scores +6.40 over a flat 4-bit quant.

vs Ollama and LM Studio

Decode speed, tokens per second

Each runtime at its own default build of the model. LM Studio is shown twice, with and without its own speculative decoding.

WorkloadOptiQOllama 0.34.4LM Studio 0.4.25LM Studio + MTP
Prose37.021.611.222.3
Code edit5217.919.114.8
Model buildOptiQ mixed 4/8, 8.3 GBQ4_K_M, 7.6 GBMLX 4-bit, 6.8 GBMLX 4-bit, 6.8 GB

On prose OptiQ uses Gemma's own assistant model as a drafter; on code edits it drafts from the prompt with n-gram lookup, because a rewritten file repeats most of its input. Both check every drafted token against the full model and keep only the ones it would have produced, so the output is identical to plain decoding.

A code edit here means asking the model to return a file with one change. Against Ollama that is 2.9× faster, and against LM Studio 2.7×.

What it costs

Download size and memory

The OptiQ build is the largest of the three downloads, 8.3 GB against 7.6 GB and 6.8 GB, and it held the most memory while serving in our runs: 15.9 GB, against 11.0 GB for Ollama and 13.7 GB for LM Studio.

The extra size is where OptiQ puts more bits into the layers that need them. On our Capability Score, a uniform 4-bit quant of this model scores 61.83 and the OptiQ build 68.23. LM Studio's default download for this model is a uniform MLX 4-bit build. The full breakdown is in the sections below.

Method: speed

One M3 Max, one runtime loaded at a time, the same prompts in each, greedy decoding with thinking off, streamed over each runtime's local API. The numbers are decode tokens per second, averaged over repeated rounds. Ollama and LM Studio ran their default builds: gemma4:12b in Ollama and google/gemma-4-12b in LM Studio.

reproducebash
$ pip install -U mlx-optiq
$ optiq serve --model mlx-community/gemma-4-12B-it-OptiQ-4bit --drafter google/gemma-4-12B-it-assistant
$ optiq serve --model mlx-community/gemma-4-12B-it-OptiQ-4bit --ngram-draft 16

The first command serves the prose setup, the second the code-edit setup. The serve docs cover both flags.

vs uniform 4-bit

What the extra 2 GB buys

Both start from google/gemma-4-12B-it. One gives every layer 4 bits. The other measures each layer first and raises the ones that cannot take it.

BenchmarkOptiQ mixed 4/8Uniform 4-bitDelta
MMLU42.6%34.4%+8.3
GSM8K93.4%90.1%+3.3
IFEval73.9%71.2%+2.8
BFCL-V3 simple71.0%71.5%−0.5
HumanEval88.4%76.8%+11.6
HashHop40.0%27.0%+13.0
Capability Score68.2361.83+6.40
On-disk size8.3 GB6.3 GB+2.0

Six points on the mean, for 2.0 GB more on disk. BFCL goes the other way by half a point. At 200 calls that is inside the confidence interval.

vs QAT

How QAT compares

Read this first

QAT is not a switch you flip. Somebody has to run quantization-aware training while the model trains, then publish those weights. Google does it for Gemma-4. Very few others do.

BuildBaseCapability Score
Uniform 4-bitinstruct61.83
OptiQ mixed 4/8instruct68.23
Uniform 4-bitQAT68.27
OptiQ mixed 4/8QAT69.64

Those two middle rows came from separate runs against different baselines, so read them as level. They stack: the sweep on the QAT base reaches 69.64. The margin is smaller there, +1.37 against +6.40. QAT has already taken out most of what the sweep goes looking for.

gemma-4-12B-it-qat-4bit is not a uniform 4-bit quant either. It pins 144 components at 8 bits and weighs 10.2 GB. Comparing scores across the three published repos means comparing three different sizes.

Method: accuracy

Six benchmarks at fixed sample counts, greedy decoding throughout. Each quant is scored against a uniform 4-bit quant of its own base. That isolates the bit allocation from everything else the base brings.

MMLU reads low for this family because its default scoring takes the answer letter by logit argmax. It still separates two quants of the same base. That is all it does here.

reproducebash
$ optiq eval mlx-community/gemma-4-12B-it-OptiQ-4bit --task all --score

Every quant carries the same table on its Hugging Face page, all of them linked from the model list. The eval-framework write-up covers the limits of the Capability Score.