mlx-optiq

Changelog

Every version of mlx-optiq published to PyPI, what shipped, and why. The source is the repo's CHANGELOG.md; this page is its public mirror.

v0.5.19

Changed

  • OptiQ Code's system prompt is general, like Claude Code's.
  • optiq serve keeps the prompt cache across agent turns on Qwen3.5/3.6.

Fixed

  • optiq serve --mtp works again.

Added

  • optiq prune-experts supports Qwen3.8-Flash-Next.

v0.5.18

Changed

  • OptiQ Code runs read-only shell commands without asking, like Claude Code.
  • OptiQ Code tells the model its working directory.

Fixed

  • The TUI no longer reports a stall nudge it never sent.

v0.5.17

Changed

  • The OptiQ Code TUI runs like Claude Code; refusals, nudges and build rollback are headless-only.
  • Headless runs can read outside the repository.

Fixed

  • --drafter answers from the whole conversation on follow-up turns, and is now 1.3-1.6x faster than plain decoding on Gemma-4.
  • --drafter and --mtp actually speculate; requests had gone down the batch path.
  • --ngram-draft keeps drafting past Gemma-4's sliding window.
  • --ngram-draft no longer slows turns with nothing to draft.
  • optiq eval --served no longer downloads weights.

Added

  • after_edit_hook, test_nudge_edits and flag_test_edits settings for OptiQ Code.

v0.5.16

Fixed

  • Tool calls are no longer dropped for models whose chat template writes <tool_call><function= without a newline (MiMo-V2.6-Distill-Qwen-9B).
  • OptiQ Code's after-edit check compiles the tests of the Go packages it touched.

Added

  • optiq feedback, /feedback in OptiQ Code and a Send feedback button in the Lab send notes to the OptiQ team; a session or chat is attached only when you choose to.
  • OptiQ Code flags edits to test files that already exist at HEAD.
  • OptiQ Code asks for a test run after eight edits without one.

v0.5.15

Fixed

  • optiq convert with measured sensitivity works from FP8 releases such as Qwen3.8-Flash-Next-FP8.
  • Qwen3.8-Flash-Next's n-gram table is quantized to 4-bit, not 8-bit, in a 4/8 build.
  • Sensitivity sweeps print progress as they go, with seconds per layer and hours left.

v0.5.14

Added

  • Qwen3.8-Flash-Next (qwen4_exp) converts from the FP8 release and serves with its experts and n-gram table streamed off SSD.
  • optiq serve honours a per-request thinking_budget; OptiQ Code sends one to servers that support it.
  • OptiQ Code's exported traces include the pre-compaction conversation and every compaction.

Changed

  • max_turns defaults to 1,000, up from 40.

Fixed

  • The streaming converter quantizes lm_head and embed_tokens, uses group size 32 where 64 does not divide, applies the 2-bit range search, and writes metadata that matches the weights.
  • A run that leaves the build broken restores the last state that built.
  • Repeated read-only bash calls get their output back after compaction; twenty identical refusals end the run.
  • The after-edit build check no longer mistakes the agent's own build break for one already at HEAD.
  • Compaction no longer loses the content it archives.
  • OptiqEngine.generate() works again (broken since 0.5.7).
  • Transport retries are capped at 12 attempts.

Earlier releases

The 14 releases before v0.5.14. Click any version to expand it. Everything before v0.5.0 is in the archive.

v0.5.13

Added

  • Xing4.0 (xing4_0) converts, serves and evaluates.
  • Checkpoints that ship their own config and tokenizer classes load without trust_remote_code.

Fixed

  • Local context sizing no longer underestimates the usable window by 3.8x.
  • Compaction is triggered from the server's reported prompt tokens instead of a character estimate.
  • Compaction also reclaims old tool-call arguments and replayed reasoning.
  • The completion check lists only edited work that was never exercised.
  • MMLU scores tokenizers that encode a leading space as its own token.
  • BFCL parses tool calls in the <param_key>/<param_value> format.
  • HumanEval accepts code indented with three spaces.
  • The GQA kernel falls back cleanly on a quantized KV cache.
  • --ngram-draft works together with --kv-config / --kv-bits.
v0.5.12

Added

  • optiq serve --on-generation-death {restart,exit}.

Changed

  • The GQA decode kernel covers the Qwen3.5/3.6 MoE family and Qwen3.5-9B (1.07-1.15x decoding at 32-64k).
  • OPTIQ_KERNELS, OPTIQ_PREFILL_STEP, OPTIQ_DUMP_REQUESTS and OPTIQ_SERVE_ON_GENERATION_DEATH are documented settings.
  • OPTIQ_KERNELS_DEBUG is removed; a missed kernel explains itself once.
v0.5.11

Changed

  • Lab Chat keeps a sidebar of saved conversations, saves after every reply, and reopens the last one.
  • Lab Chat gives each conversation its own workspace for tools and attached files.
  • The Lab centers on wide screens; Chat uses the full window.
  • OptiQ Code remembers the approval mode per repo and captures the mouse by default (/mouse releases it).
  • optiq serve reports thinking tokens in usage.completion_tokens_details.reasoning_tokens.
  • optiq serve reuses the prompt cache for follow-up messages on Qwen3.5/3.6 hybrids with --ngram-draft or KV quantization.

Fixed

  • A Boost that cannot run falls back to the local model and says why.
  • The tool loop after a Boost turn stays local.
  • Streamed Responses and Anthropic replies report real token usage.
  • Image tokens are counted in usage.prompt_tokens.
  • Images work on REAP-19B and other models whose MTP head does not fit.
  • Structured output stays valid under --ngram-draft.
  • optiq lab picks a free port when 8080 is taken.
  • The Arena waits for loads, answers only for loaded models, can swap model B, and stops with the Lab.
  • Lab chats keep their Deep Research report, model settings and streamed text when stopped; several display fixes.
  • OptiQ Code keeps the cache warm through its own nudges on Qwen3.5/3.6.
  • OptiQ Code counts test runs that print only progress dots.
  • OptiQ Code reports a mid-response server failure as a server error.
  • OptiQ Code logs MCP server stderr to ~/.optiq/logs/.
  • optiq code -p prints the whole reply; -c resumes the last real conversation.
  • optiq --version and the Lab's Account page report the running version.
v0.5.10

Changed

  • optiq cloud login takes effect on a running server at the next /boost.
  • The Lab and optiq run explain how to connect for Boost.
v0.5.9

Added

  • optiq run <agent> -m <model> launches Claude Code, Codex, OpenCode, OpenClaw, Hermes Agent or Mistral Vibe against a local model.
  • Quantized KV cache for latent-attention models and for models with attention sinks (gpt-oss).
  • --ngram-draft and --mtp combine into one speculator.
  • optiq prune-experts drops the MTP head of an MoE model (--keep-mtp keeps it).

Changed

  • --ngram-draft retuned; 16 is the recommended cap (1.62x on coding-agent turns).
  • Faster long-context decoding on Qwen3.6-35B-A3B-REAP-19B (up to 1.31x on 55-60k agent turns).

Fixed

  • Speculation prices verify passes at the request's real context length.
  • MTP heads of Qwen3.5/3.6 MoE quants load their expert weights.
  • OptiQ Code replays reasoning, restoring ~99% cache reuse on hybrids.
  • optiq serve sizes and enforces the prompt cache correctly with KV quantization and --ngram-draft.
  • optiq kv-cache ranks models whose attention depends on the cache.
v0.5.8

Added

  • optiq serve --ngram-draft 8: prompt-lookup speculative decoding for any model (1.3-1.5x on coding-agent turns).

Fixed

  • optiq kv-cache measures in float32 over every position.
  • optiq kv-cache says when an architecture has no quantizable KV cache.
v0.5.7

Added

  • optiq code import <trace.jsonl> continues a saved trace.
  • task_list: a plan whose items record whether anything was run.
  • finish_gate (headless, off by default) refuses a finish while the tree does not build.
  • request_extra: a JSON object merged into every chat request.
  • max_boosts and boost_cooldown settings.
  • optiq code -c -p GOAL continues headlessly.

Changed

  • Requests stay inside the context window: output shrinks first, then compaction.
  • After-edit verification uses per-language syntax checks and the nearest project build.
  • The stall ladder escalates on work outside the repository.
  • The completion check also runs when the agent calls done.
  • Failing tests are labelled pre-existing or new.
  • Long tool results keep their head and tail.
  • The TUI leaves the conversation in scrollback, releases the mouse, and queues typed tasks.

Fixed

  • MTP reuses the prompt cache.
  • optiq serve: health reflects a dead generation thread, missing max_tokens no longer caps at 512, top_k: -1 works.
  • Responses API history, reasoning replay and cached-token counts.
  • Runs that committed their own work or probed outside the repo report their patch.
  • Dropped chat streams are retried.
  • optiq kv-cache calibrates on real text.
  • The MTP head of an expert-pruned quant uses the pruned expert count.
  • gofmt -l output is no longer reported as a syntax error.

Removed

  • /yank (use /copy).
v0.5.6

Changed

  • auto_verify replaces the full test run after every edit with a syntax check (full restores the old behaviour).
  • run_tests takes a target: a path, directory, or k:<expr>.

Added

  • Traces carry model_seconds, turn_latency_p50 and turn_latency_max.
v0.5.5

Removed

  • replace_lines; use edit_file.

Changed

  • git is read-only; writes go through bash.
  • edit_file and write_file require the file to have been read this session.
  • bash takes workdir and timeout.
v0.5.4

Fixed

  • optiq code export includes token accounting.
v0.5.3

Added

  • disable_tools withholds named tools from a run, e.g. OPTIQ_CODE_DISABLE_TOOLS=web_search,web_fetch. They are dropped from what the model is offered and refused at dispatch, so a model that names one anyway is told it is off. For air-gapped runs, and for benchmarks where reaching the open web lets the agent look up the fix instead of deriving it.

Fixed

  • optiq code config and optiq config printed api keys in full. They are now masked to a prefix, a suffix and a length — enough to tell which key is in place, not enough to use.
  • trace_tags is now a real setting (OPTIQ_CODE_TRACE_TAGS, with OPTIQ_TRACE_TAGS still honoured). It was read straight from the environment, so it never appeared in optiq code config and could not be set from a config file.
v0.5.2

Fixed

  • optiq code could not run at all on a clean pip install mlx-optiq: it failed with "MCP servers are configured but the 'mcp' package is not installed" on machines that had no MCP servers configured. The optional import now happens only when a server is actually enabled, and the message names which ones want it.
v0.5.1

Fixed

  • Boost from OpenCode never fired: it sends the prompt JSON-quoted, so the trigger never matched and the local model answered every time, with no error.
  • Boost from OpenClaw never fired either: it stamps the time onto every message, so the trigger no longer started the line. Any bracketed tag a client adds is now allowed in front of it.
  • One /boost could cost two credits: clients send a separate title-generation call carrying the same text, and it opened its own episode. Affected OpenCode and Claude Code.
  • Boost failed outright for clients with rich tool schemas. The cloud model's tool API takes an OpenAPI subset, not JSON Schema, and rejected the whole request over const, patternProperties, exclusiveMinimum and uniqueItems. Schemas are now rewritten into the accepted subset.
  • optiq cloud status and every boosted reply named the backing frontier model. They now report optiq-boost.
  • optiq code told a stalled turn to edit files even when the goal was a question, so asking one could produce edits you never requested. It now offers answering too, and never demands a write in plan mode.
  • Site: the favicon was still the old one (cached at an unchanged URL), and /favicon.ico did not exist.
v0.5.0

Added

  • OptiQ Cloud Boost. /boost (or boost: for clients that eat slash commands) hands one hard turn to a frontier model, then gives control straight back. It lives in optiq serve, so Claude Code, Codex, OpenCode, OpenClaw, Hermes Agent, Mistral Vibe, the Lab and OptiQ Code all get it with no plugin.
  • One Boost is one episode, not one API call: a whole tool loop costs one credit.
  • Works on all three surfaces optiq serve exposes, streaming or not: chat/completions, Anthropic messages, and Responses (including previous_response_id chains).
  • optiq cloud login | status | logout. The token is minted when the CLI collects it, so none is ever stored server-side; logout revokes rather than forgets.
  • OptiQ Code: /boost [task], a bare /boost to retry the last request, and manual / suggest / auto modes. The Lab gets a per-message Boost toggle.
  • 20 free Boosts on sign-up; 200 for $20 as a one-off pack.
  • New pages: OptiQ Cloud and its docs.

Changed

  • New design across the site, the Cloud dashboard and OptiQ Lab. Lighthouse 100 on desktop.
  • optiq serve always sends a content key, null when the model produced only reasoning. Clients reading message["content"] used to hit a KeyError.
  • Boost usage follows OpenAI's convention: reasoning inside completion_tokens, cached inside prompt_tokens. Clients were undercounting a Boost roughly tenfold.
  • optiq code reports cost in Boosts, not dollars: local + 1 Boost · 14 left.

Fixed

  • A failed Boost explains itself and says whether you were charged, instead of printing a Python errno.
  • Benchmark eval tests skip without the datasets extra instead of failing.