Moonshot AI rolled out Kimi K2.8 Preview on September 11, 2026, a new mid-tier model that sits between the company's coding-focused Kimi K2.7 Code and its flagship Kimi K3. The release adds multimodal input, a 1 million token context window for every membership tier, and coding and agent gains that Moonshot describes as close to K3, according to a report from PANews. It went live inside Kimi Code and Kimi Work the same day, with no configuration changes needed for existing users.
Key takeaways
Moonshot AI launched Kimi K2.8 Preview on September 11, 2026, positioning it between Kimi K2.7 Code and the flagship Kimi K3.
The model now accepts text, image, and video input, and every Kimi Code membership tier gets a 1 million token context window, up from the 262,144 tokens on Kimi K2.7 Code.
Kimi K2.8 Preview keeps the same API model ID as K2.7 Code, so developers and third-party tools do not need to change any settings to use it.
Third-party API gateway TokenRa lists pricing around $0.80 per million input tokens and $3.35 per million output tokens for the model.
Independent benchmarks confirming Moonshot's claim that performance is close to K3 have not been published yet, so the comparison currently rests on the company's own description.
What is Kimi K2.8 Preview?
It is Moonshot AI's newest checkpoint in the Kimi K2 line, built to handle everyday coding and agent work without the cost of running the full K3 flagship on every request. Emergent's coverage of the launch describes it as a multimodal model that bridges the gap between K2.7 Code and future flagship iterations, giving developers access to vision and language capability under a proprietary license. That last detail is a departure for Moonshot, which has released K2.7 Code and K3 as open-weight models; K2.8 Preview is closed and available only through Moonshot's own API and apps.
The "preview" label is doing real work here. Moonshot is treating this as an early-access checkpoint, not a finished product, and the company is gathering feedback before deciding how the model fits into its permanent lineup.
What actually changed from Kimi K2.7 Code
The headline change is context length. Every Kimi Code membership tier now gets a 1 million token context window with this release, a jump from the 262,144 tokens available on Kimi K2.7 Code, according to a KuCoin news flash citing Moonshot's own announcement. That kind of headroom matters for anyone feeding an entire codebase, a long document, or hours of transcript into a single prompt without chunking it first.
Reasoning control is the second change worth noting. Kimi K2.8 Preview supports three thinking effort levels, low, high, and max, matching the structure Moonshot introduced with K3, and it runs at max effort by default. The company says thinking efficiency has improved significantly compared with K2.7 Code, alongside broader gains in coding and agent tasks, per the same KuCoin report. Data from llm-stats.com's model card confirms the multimodal input spans text, image, and video, with text-only output, and lists the September 11, 2026 release date.
On pricing, third-party API gateway TokenRa lists Kimi K2.8 Preview at roughly $0.80 per million input tokens, $3.35 per million output tokens, and $0.14 per million cached read tokens. Moonshot has not published first-party API pricing for the preview model as of this writing, so treat gateway pricing as an indicator rather than an official rate card.
How it stacks up against K2.7 Code and K3
Moonshot's own framing places Kimi K2.8 Preview firmly in the middle of its lineup: more capable than K2.7 Code, not a replacement for K3. For context, K3 is Moonshot's 2.8 trillion parameter open-weight flagship, released in July 2026 with its own 1 million token context window and a mixture-of-experts architecture built for long-horizon coding and reasoning work. K2.8 Preview borrows K3's thinking-effort structure and its context length, but Moonshot has not released benchmark scores that would let outside researchers verify the "close to K3" claim directly.
That gap has not gone unnoticed. Reporting from CCTest points out that the claim of near-K3 performance has not yet been supported by public benchmarks or independent testing, and that the model's real value will come down to cost, latency, and reliability across actual coding workloads rather than a marketing line. Until third-party evaluators publish scores on standard suites like LiveCodeBench or SWE-bench, the comparison is Moonshot's word against its own flagship.
Why release a mid-tier model right now
The timing lines up with Moonshot's push to widen usage without running every request through its most expensive model. CCTest's reporting notes that Moonshot has seen rapid revenue growth since K3 shipped and has reportedly begun a Hong Kong IPO process, and frames K2.8 Preview as a way to lower the cost of giving more users long-context, agentic features. Handing every membership tier a 1 million token window on a cheaper model, rather than reserving that context length for flagship-only plans, fits a company trying to scale usage while keeping inference costs predictable.
It also mirrors a pattern other labs have followed this year: ship a flagship, then follow with a lighter, faster sibling that keeps most of the capability at a lower price. We covered a similar dynamic when Kimi K2.5 landed as Moonshot's earlier open-source push, and the same trade-off between raw capability and everyday cost shows up again here.
What this means for creators using AI image and video tools
Kimi K2.8 Preview is a text and code model at heart, not an image or video generator, so it will not directly change how anyone produces visuals. But the broader trend it represents, cheaper access to long-context, multimodal reasoning, is the same trend reshaping creative tooling more generally. Models that can read an image or a video clip alongside a text prompt make it easier to build workflows that plan a shoot, write captions, or reason about a brand's existing assets before generating anything new.
For creators who already lean on AI for visual output, that means the planning and production sides of the workflow are both getting faster, even when the models doing each job are different. If your work involves turning a product photo into a full campaign, tools like MagicShot's image to video generator or its product photo generator handle the visual half, while faster, cheaper coding and agent models like K2.8 Preview are more likely to show up behind the scenes, in the apps and automations that schedule, caption, and route that content. The practical takeaway is not that this model is a creative tool, it is that the cost of running capable AI models keeps dropping, and that pressure eventually reaches every corner of a creator's stack.




