MiniCPM5 on Apple Silicon
OpenBMB's MiniCPM5 family: Llama-architecture models with hybrid reasoning, tool calling and a 128k context, released under Apache-2.0. The mlx-optiq builds are two HF repos, both sensitivity-aware mixed-precision quants: the 1B fits in 875 MB and is the smallest base in the lineup worth fine-tuning; the 2B fits in 1.8 GB and scores level with the 2.8 GB Qwen3.5-4B quant.
The quant
| Model | Size on disk | Capability Score | vs uniform-4 | Best for |
|---|---|---|---|---|
| MiniCPM5-2B-OptiQ-4bit | 1.8 GB | 69.36 | +1.39 | Best small model in the lineup: instruction following, math, code at 1.8 GB |
| MiniCPM5-1B-OptiQ-4bit | 875 MB | 30.28 | +4.44 | Fast on-device assistant, fine-tuning base |
Per-benchmark breakdown: MiniCPM5-2B
| Benchmark | uniform-4 | OptiQ-4 (mixed) | Δ |
|---|---|---|---|
| MMLU (5-shot, 1000) | 56.1% | 59.8% | +3.6 |
| GSM8K (no thinking) | 78.2% | 82.1% | +3.9 |
| IFEval (strict) | 85.4% | 86.7% | +1.3 |
| BFCL V3 (simple AST) | 87.0% | 82.5% | -4.5 |
| HumanEval (pass@1) | 78.0% | 81.1% | +3.0 |
| HashHop (overall) | 23.0% | 24.0% | +1.0 |
| Capability Score | 67.96 | 69.36 | +1.39 |
Per-benchmark breakdown: MiniCPM5-1B
| Benchmark | uniform-4 | OptiQ-4 (mixed) | Δ |
|---|---|---|---|
| MMLU (5-shot, 1000) | 49.0% | 52.4% | +3.4 |
| GSM8K (no thinking) | 1.7% | 2.7% | +1.0 |
| IFEval (strict) | 58.6% | 64.7% | +6.1 |
| BFCL V3 (simple AST) | 0.0% | 0.0% | 0.0 |
| HumanEval (pass@1) | 45.7% | 57.9% | +12.2 |
| HashHop (overall) | 0.0% | 4.0% | +4.0 |
| KL vs bf16 (mean) | 0.350 | 0.136 | 2.6× closer |
Hello world
from mlx_lm import load, generate model, tok = load("mlx-community/MiniCPM5-1B-OptiQ-4bit") prompt = tok.apply_chat_template( [{"role": "user", "content": "Summarize the plot of The Iliad in three sentences."}], tokenize=False, add_generation_prompt=True, enable_thinking=False, ) print(generate(model, tok, prompt=prompt, max_tokens=300))
Hybrid reasoning: think or no-think
MiniCPM5's chat template accepts an enable_thinking flag. With it on, the model emits a <think>...</think> block before the answer, useful for math, multi-step planning, or anywhere you want chain-of-thought. The recipe per the model card:
| Mode | temperature | top_p | Use when |
|---|---|---|---|
| No-think (default) | 0.7 | 0.95 | Fast assistant, rewriting, conversational |
| Think | 0.9 | 0.95 | Math, code, multi-hop reasoning |
Pass the flag via chat_template_kwargs at the OpenAI endpoint level or as a keyword to apply_chat_template directly. optiq serve forwards chat_template_kwargs verbatim, so you can flip thinking per-request from any OpenAI-compatible client.
Serving
$ optiq serve --model mlx-community/MiniCPM5-1B-OptiQ-4bit --port 8000 # From any OpenAI-compatible client: $ curl -s http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model":"mlx-community/MiniCPM5-1B-OptiQ-4bit", "messages":[{"role":"user","content":"What is 17 * 23?"}], "chat_template_kwargs":{"enable_thinking":true}}'
Fine-tuning
The 1B size makes MiniCPM5 the smallest base in the mlx-optiq lineup that's still capable enough to fine-tune for real tasks. On a 24 GB Mac, LoRA training fits comfortably at max_seq_length=2048 with all 7 Unsloth target modules adapted. No num-layers / sequence-length workarounds needed. The sensitivity-aware LoRA overlay reads the optiq/metadata.json sidecar and gives 8-bit layers 2× the adapter rank of 4-bit layers at the same parameter budget.
$ optiq lora train mlx-community/MiniCPM5-1B-OptiQ-4bit \
--data ./my_training_data \
--preset default \
--max-seq-length 2048