Xenon 9 is our ninth generation of models and our strongest yet. It succeeds Xenon 8 at the top of the lineup with three models that are smarter across the board: deeper reasoning, a 1M token context window, and a new system of native effort levels that lets you decide how hard each model thinks before it answers. Every tier improves on its predecessor, and every tier keeps the promises the platform is built on.
Xenon 9 (X9): The flagship
Xenon 9, or X9 for short, is the flagship and the most capable model we have ever shipped. It sits above everything else in the lineup. Compared to Xenon 8, it is plainly smarter: it reasons further before it answers, holds more of the problem in view at once, and gets more of the long, multi-step work right on the first try.
The 1M token context window changes what a single request can contain. Entire repositories, long document sets, or months of conversation history fit in working memory at the same time, and X9 uses that space well: sharper judgment about what matters, fewer confident mistakes, and honest uncertainty when the answer is not in the material.
Use X9 when the problem is hard and the answer matters.
Xenon 9 Code (X9 Code): The engineer
Xenon 9 Code, or X9 Code, is the coding specialist of the family. It takes everything we learned from Xenon 8 Code and pushes it further: it reads your repo as a system instead of a pile of files, plans an approach, writes the change, runs the loop, and checks its own work before handing it back.
It is smarter than its predecessor where it counts: architecture decisions, edge cases, and long agentic runs where a small mistake compounds. With the full context window it can hold large codebases in view at once, so refactors stay consistent across more files and debugging sessions waste fewer steps. It still knows when to ask and when to act.
Use X9 Code when you are shipping something real.
Xenon 9s (X9s): Small, smarter
Xenon 9s is the small model of the family, built for speed without giving up the judgment that makes the bigger models useful. It is a compact, low-latency build of Xenon 9, tuned for rapid tool use and high-frequency loops where every millisecond matters.
X9s shares the improved thinking of its bigger siblings in distilled form: quick plans, fast answers, and the same calibration and personality you already know. For CLI agents, interactive companions, and automation at scale, it is the new default for speed.
Use X9s when latency is the feature.
Native effort levels
Xenon 9 introduces native reasoning effort levels: low, medium, high, xhigh, and max. These are real, provider-managed levels built into the models, not fixed token counts bolted on by the API. Pick low when you want a quick answer, max when the problem deserves everything the model has, and anything in between.
If you do not specify an effort, nothing changes: base model IDs without any controls simply use the provider default, exactly as before.
In OpenCode, the levels appear as native variants on each Xenon 9 model, and the variant selector (default shortcut: Ctrl+T) cycles through them. On the API, each model is available under its base ID, xenon-9, xenon-9s, and xenon-9-code, and the model list also exposes suffixed IDs such as xenon-9-xhigh, xenon-9s-low, and xenon-9-code-max for clients that pick models from a list. You can also set the effort directly on any request:
{
"model": "xenon-9",
"messages": [
{
"role": "user",
"content": "Explain the difference between processes and threads."
}
],
"reasoning_effort": "high",
"stream": true
}
Hidden reasoning, private by design
All three Xenon 9 models reason internally before they answer, and that reasoning stays hidden. Chain of thought is never exposed: not in the chat app, not in API responses, not in streams. During a long hidden reasoning pass the stream sends keep-alive signals so connections stay open, but the thinking itself is never transmitted. You get the benefit of the work without the working.
Tools, streaming, and everything else
Everything you already build on keeps working: tool calling, streaming, multimodal input, and the OpenAI-compatible API are unchanged. Effort levels do not change max_tokens or disable request timeouts, and streaming is still the recommended way to run long reasoning requests. Existing integrations keep working. Just switch the model.
Availability
Xenon 9, Xenon 9s, and Xenon 9 Code are rolling out in the chat app and on the API starting today under the xenon-9 family of model IDs. Xenon 8 remains available for compatibility. If you are building against the API, the models and their effort levels are documented in the API docs, and setup for OpenCode and other tools lives on the integrations page.