Cognition SWE-2 shipped on September 10, 2026. It scores 50.0% on FrontierCode 1.1 Main against Anthropic Fable 5.1’s 50.9%, and Cognition says it does that for 64% less money. The model is post-trained from Moonshot AI’s 2.8-trillion-parameter Kimi K3. There is no standalone API and no open weights — you can only reach it inside Devin.
That last sentence is the business model. Cognition did not release a model. It released a cheaper engine for a subscription product it already sells, and it built that engine on top of a Chinese open-weights base it did not have to pay to train.
What is Cognition SWE-2?
Cognition SWE-2 is the coding model that now powers Devin, Cognition’s autonomous software engineer. It replaces SWE-1.7 and is available inside Devin Desktop and the Devin CLI first, with Devin Web and Fusion following. Cognition published the release on its own blog on September 10, 2026.
It is not a foundation model. Cognition post-trained it from Kimi K3, the 2.8-trillion-parameter mixture-of-experts model from Beijing-based Moonshot AI. BenchLM lists the context window at 1M tokens.
The headline claim is not raw capability. It is position on the cost-performance curve. Cognition says its post-training added five to six points across many benchmarks while shifting the entire frontier, not just one point on it.
How SWE-2 was trained
The technical claim worth reading is the reinforcement learning method. Cognition describes it as “an RL algorithm that trains all reasoning-effort levels in a single run,” applying a linear cost penalty calibrated to each effort level’s slope on the frontier.
In plain terms: instead of training separate low-, medium- and high-effort models and shipping three artifacts, Cognition trains one model that has internalized the price of its own thinking. The reward function is success minus a weighted cost term.
That matters commercially. Every effort tier trained in one run is one training bill instead of three, and Cognition is a company with roughly $47 billion of valuation to justify — Tech Funding News reported Devin revenue approaching $1 billion as that round came together. We covered the $47 billion valuation and its 52x revenue multiple earlier this month.
How much cheaper is SWE-2 than Fable 5.1?
Cognition’s claim is 64% lower cost at roughly equal FrontierCode performance, and around one-quarter the cost of OpenAI’s GPT-6 Astra. Against its own predecessor SWE-1.7, Cognition reports 81% lower average cost per task and 58% fewer turns on FrontierCode 1.1 Main.
Here is how the published numbers line up across the three models.
| Benchmark | SWE-2 | Fable 5.1 | GPT-6 Astra |
|---|---|---|---|
| FrontierCode 1.1 Main | 50.0% | 50.9% | 53.3% |
| DeepSWE 1.1 | 73.0% | 67.4% | — |
| Terminal-Bench 2.1 | 92.8% | 91.4% | 89.9% |
| Terminal-Bench 4 | 27.3% | 55.8% | 57.9% |
Read the last row twice. On Terminal-Bench 4 — the long-horizon, multi-hour autonomy test — SWE-2 lands at 27.3% against Fable 5.1’s 55.8%. That is not parity. That is less than half.
The cost claim is a curve, not a price list
There is no public per-token rate card for SWE-2. The 64% figure is task-level cost measured inside Cognition’s own harness, against Cognition’s own benchmark selection. explainX, reviewing the launch, put it bluntly: cost claims here “are curves, not SKUs.”
Anyone budgeting against that number is budgeting against a vendor-measured average on a vendor-chosen task mix. Compare that to Fable 5.1’s published $10 input price, which you can put in a spreadsheet.
Why did Cognition build SWE-2 on Kimi K3?
Because training a 2.8-trillion-parameter base from scratch costs more than Cognition wants to spend, and Moonshot gave it away. Kimi K3’s open weights let Cognition skip the most capital-intensive stage of the stack and spend only on post-training.
This is the open-weights trade in its purest form. Chinese labs absorb the pretraining cost; American application companies capture the customer relationship, the subscription revenue and the margin. We saw the same structure in Tencent’s 770B Hy4 drop and in DeepSeek’s 552B V4.1 Flash.
For Moonshot, the payoff is not cash. It is distribution. A $47 billion American company shipping its base model to enterprise developers is the strongest possible endorsement of Kimi K3’s quality — and it costs Moonshot nothing per seat.
Is a Chinese base model a security problem?
It is at least a procurement question, and the timing is awkward. Cognition shipped SWE-2 two days after a federal advisory named Chinese AI firms over industrial-scale distillation of US frontier models — we covered that advisory and the six companies it named on September 9.
Cognition’s structural answer is that weights are not servers. Prompts route through Cognition’s US infrastructure; nothing goes to Moonshot. Tech Times reports Cognition’s internal trustworthiness evaluation passed 145 politically sensitive questions at a 98.0% rate.
The honest caveat: that evaluation is self-reported, and because SWE-2 has no open weights, no third party can audit the finished model. Cognition inherits a Chinese base it can inspect and ships a derivative nobody else can.
Where does SWE-2 fall short?
Four gaps are visible from the launch material alone.
- Long-horizon autonomy. 27.3% on Terminal-Bench 4 against Fable 5.1’s 55.8%. Parity on scoped tasks does not carry over to multi-hour work.
- No independent verification. BenchLM has five verified benchmark rows for SWE-2 and declines to rank it publicly, citing insufficient evidence coverage.
- Harness lock-in. Results are measured inside Devin. With no standalone API, you cannot reproduce them anywhere else or drop the model into your own agent stack.
- No weights to audit. Open base, closed derivative. Security teams get the worst of both: foreign-origin pretraining they can read about, and a final model they cannot inspect.
One behavioral improvement is real and measurable. Cognition reports SWE-2 makes its first code edit at a median of 18 steps versus 48 for SWE-1.7 — a 62.5% cut in how long the agent wanders before doing anything.
Who wins and loses financially?
The money story is not about benchmarks. It is about who pays for the next trillion parameters.
Cognition wins on gross margin. If the 64% cost claim survives contact with real workloads, the company can hold Devin pricing flat while cutting cost of goods sold on every task. At roughly $1 billion of revenue, a structural inference-cost reduction is worth more than a benchmark point.
Anthropic and OpenAI lose a seat, not a war. Cognition was previously a buyer of frontier tokens. It is now, at least partly, a buyer of electricity and a user of free weights. Every application company that makes the same move removes recurring API revenue from the labs’ books. The labs still win Terminal-Bench 4, which is where the enterprise autonomy budget eventually goes.
Moonshot wins mindshare and nothing else. No licensing revenue, no per-seat cut, no customer data. Whether that converts into anything bankable is unproven.
Enterprise buyers get a governance bill. A model with a Beijing-trained base, self-reported safety evaluations and no auditable weights is a compliance review, not a checkbox — however good the price is.
How do you get SWE-2?
Only through Devin. Tech Times reports a one-month free promotion for Pro, Max and Teams subscribers, with Pro starting at $20 per month. There is no standalone API and no weights release.
That distribution choice is deliberate. Cognition is not competing on tokens; it is competing on the agent product. Bundling the model into the subscription makes the cost advantage invisible to the customer and keeps it on Cognition’s income statement.
Frequently asked questions
When did Cognition SWE-2 launch?
September 10, 2026, on Cognition’s blog, initially inside Devin Desktop and the Devin CLI, with Devin Web and Fusion following.
Is SWE-2 better than Claude Fable 5.1?
It is close on scoped coding — 50.0% vs 50.9% on FrontierCode 1.1 Main — and better on DeepSWE 1.1 at 73.0% vs 67.4%. It is much worse on Terminal-Bench 4, at 27.3% vs 55.8%.
What model is SWE-2 built on?
Kimi K3, the 2.8-trillion-parameter open-weights mixture-of-experts model from Moonshot AI in Beijing.
Can I use SWE-2 through an API?
No. Cognition has published no standalone API and no open weights. Access is through Devin only.
How much does SWE-2 cost?
There is no published per-token price. It is included in Devin subscriptions, which Tech Times reports start at $20 per month for Pro.
Is the 64% cost saving independently verified?
No. It is Cognition’s own task-level measurement inside its own harness. BenchLM lists only five verified benchmark rows for SWE-2 and does not rank it publicly.
Does using a Chinese base model create a security risk?
Cognition says prompts route through US infrastructure and never reach Moonshot’s servers, and reports a 98.0% pass rate on 145 politically sensitive questions. That evaluation is self-reported and the finished weights are not public.
The bottom line
Cognition SWE-2 is a margin release dressed as a model release, and that is not an insult — it is the most interesting thing about it. Matching a frontier model’s scoped coding score at 64% less cost, on a base you did not pay to train, is a better business outcome than winning a leaderboard.
But the Terminal-Bench 4 number is the tell. 27.3% against 55.8% says SWE-2 is excellent at bounded tasks and not yet trustworthy across long autonomous runs — which is exactly the capability Cognition’s $47 billion valuation is priced on.
Verdict: a strong buy for teams already paying for Devin and running scoped, well-specified work. A wait for anyone budgeting against the 64% figure or shipping regulated software, until an independent evaluator confirms the cost curve and someone other than Cognition audits the model.
Leave a Reply