MiniMax M3.1 Flash went live on September 27, 2026, dropped straight into the MiniMax Code agent with five reasoning tiers, a 1-million-token context window and a 128K output cap. There is no model card, no benchmark report and no published price. Independent testing clocked it roughly twice as fast as M3 while its pass rate fell from 14 of 15 runs to 10 of 15.
What is MiniMax M3.1 Flash?
MiniMax M3.1 Flash — shipped under the identifier M3.1-Flash-Preview — is a speed-tuned coding and agentic model that appeared inside MiniMax Code on September 27, 2026, with no launch post, no specification sheet and no price list.
Startup Fortune reported the launch the same day and noted the model arrived with “no model card or benchmark report released.”
That is unusual. Chinese labs have spent 2026 competing on published numbers. This one shipped the product and skipped the paperwork.
Where you can actually use it
Access starts inside MiniMax Code, the company’s terminal agent. APIMaster reported zero public routes for the model at launch, against three live routes for the older M3.
One independent tester, writing at AI @ Sulat.com, found the model ID MiniMax-M3.1-Flash-Preview responds on MiniMax’s Anthropic-compatible endpoint anyway — so the wall is softer than the silence suggests.
The five reasoning tiers
AlphaSignal listed five effort levels: low, medium, high, xhigh and max, with max as the default. The Sulat test counted six positions once the unset default is included.
- Thinking cannot be switched off. The Sulat test found reasoning forced on at every tier.
- The ladder is real but short. Low effort produced around 129 reasoning characters; max produced around 365.
- The default is the most expensive setting. Shipping with
maxas the default means the slowest, most token-hungry mode is what a new user gets. - No tier is documented. None of the five appear in published MiniMax docs for this model.
How fast is MiniMax M3.1 Flash?
Roughly twice as fast as M3 on short tasks, according to the only public timing data available. The Sulat tests measured a trivial prompt at 1.02 seconds against M3’s 1.92 seconds, and a dedupe script at 4.26 seconds against 8.11 seconds. Throughput rose from 71 to 94.6 tokens per second.
These are one tester’s numbers, not a lab benchmark. Treat them as directional.
| Task | M3.1 Flash | M3 | M2.7-highspeed |
|---|---|---|---|
| Trivial prompt (latency) | 1.02 sec | 1.92 sec | 2.88 sec |
| FizzBuzz (throughput) | 94.6 tokens/sec | 71 tokens/sec | — |
| Dedupe script (latency) | 4.26 sec | 8.11 sec | 32.97 sec |
| Single-shot pass rate | 10 of 15 | 14 of 15 | — |
The last row is the one that matters commercially. A model that answers in half the time and fails four more times out of fifteen is not a free upgrade. It is a trade.
Where does M3.1 Flash break?
It burns its entire output budget thinking. The Sulat tests ran single-shot tasks with a 16,384-token completion budget. On hard async and concurrency problems, M3.1 Flash consistently spent all 16,384 tokens on reasoning and returned empty code blocks.
That is the worst failure mode for a paid agent. You are billed for the full generation and receive nothing executable.
It also explains the pass-rate drop. M3 cleared 14 of 15. M3.1 Flash cleared 10 of 15 — and the five misses were not wrong answers, they were no answers.
The tester’s conclusion is worth quoting directly: “the model that checks itself against the operating system instead of arguing with the test runner is the one you leave running overnight.”
Why did MiniMax ship it with no price?
Because it is a preview inside a subscription product, not an API product. No price is published because the model is not sold per token yet. MiniMax is running a Sept 28–Oct 7 campaign doubling the daily free quota for M3.1-Flash-Preview, H3 and H3 Max, according to APIMaster — a plan benefit, not a rate card.
There is a quieter story here. Four days before the launch, an anonymous model called stealth/space-bunny-alpha appeared on OpenRouter at $0 per token. Startup Fortune reported developer fingerprinting returned 24/24 and 50/50 token matches to the MiniMax family. MiniMax has not confirmed the link.
Free stealth testing followed by an unpriced in-product preview is a deliberate sequence. It buys usage data before it buys a pricing commitment.
What the architecture reports say — and why to discount them
AlphaSignal relayed unverified claims of a 428-billion-parameter mixture-of-experts base, all-sparse attention replacing M3’s full-attention anchor layers, Q8KV4 cache precision, 4-bit NVFP4 expert weights, and a DSpark speculative decoder shipping as a ~2.3 GB FP8 artifact from a roughly 250 GB gated Hugging Face repository.
The same reports claim one-twentieth the per-token compute at 1M context, 9x faster prompt processing and 15x faster decoding versus M3.
None of it is confirmed by MiniMax. Weights are gated. Until the repo opens, those figures are marketing physics, not measured results.
How much does MiniMax cost compared to US frontier models?
M3.1 Flash has no price. Its predecessor M3 does, and it is the best available anchor: $0.30 per million input tokens and $1.20 per million output, with cached input at $0.06. APIMaster notes the original M3 rate was $0.60/$2.40 — a roughly 50% cut.
Startup Fortune put M3’s total cost at 5–10% of Claude or GPT equivalents.
| Model | Input / 1M | Output / 1M | Context | Status |
|---|---|---|---|---|
| MiniMax M3.1 Flash | Unpublished | Unpublished | 1M | Preview, MiniMax Code |
| MiniMax M3 | $0.30 | $1.20 | 1M (512K guaranteed) | Generally available |
| MiniMax M3 (launch rate) | $0.60 | $2.40 | 1M | Superseded |
On capability, MiniMax’s own page for M3 claims 83.5 on BrowseComp against 79.3 for Claude Opus 4.7, and a #3 overall placing on PostTrainBench at 37.1. Startup Fortune reported M3 at 80.5% on SWE-bench Verified and 59.0% on SWE-Bench Pro.
Those are M3’s numbers. MiniMax has published none for M3.1 Flash.
The pattern is familiar to anyone tracking this market. StepFun shipped Step 5 Preview at $1 per million tokens, and GLM-5.3-FlashX pushed 200 tokens per second on domestic silicon. Speed and price are the axes Chinese labs are competing on. Documentation is not.
Who wins and who loses financially?
Winners: MiniMax shareholders and anyone running high-volume, low-complexity code generation. Losers: US inference margins and any buyer who needs a contract. An unpriced preview cannot be procured, budgeted or audited, which keeps M3.1 Flash out of regulated enterprise pipelines no matter how fast it runs.
The equity story is already loud. Mainland investors poured $1.4 billion into MiniMax stock in August 2026, per Startup Fortune, making it China’s most-bought Hong Kong name that month. Bloomberg reported the same month that MiniMax had emerged as mainland investors’ new favorite stock.
On its first day in Stock Connect the shares rose 17% as southbound buyers took in close to HK$2.7 billion in a single session.
That is the capital context for a launch with no price tag. MiniMax does not need per-token revenue this quarter. It needs usage, benchmark bragging rights and a share price.
The skeptical read
A model that ships without a card, without benchmarks, without a price and without open weights is a product announcement wearing a research announcement’s clothes.
Every hard number in circulation about M3.1 Flash comes from one independent tester or from unverified relays. The one set of numbers that exists shows quality going backwards. Buyers should price that risk, not the 9x decoding claim.
Compare the discipline elsewhere: Xiaomi published a training budget with MiMo V2.6 Pro, and OpenAI published a rate card with GPT-6 Sol at $2 per million tokens. A number you can check is worth more than a tier you cannot.
Frequently asked questions
Is MiniMax M3.1 Flash available through a public API?
Not officially. APIMaster reported zero public routes at launch. One independent tester found the ID MiniMax-M3.1-Flash-Preview responds on MiniMax’s Anthropic-compatible endpoint, but MiniMax has published no API documentation for it.
How much does MiniMax M3.1 Flash cost?
No price has been published. Access runs through MiniMax Code plan quotas, with a double-free-quota campaign running September 28 to October 7, 2026.
What is the context window?
One million tokens, with a 128K maximum output, according to independent testing. M3 guarantees a minimum 512K window on MiniMax’s own spec page.
Are the weights open?
No. Reports describe a gated Hugging Face repository of roughly 250 GB. MiniMax lists a full open-source release for M3 as “coming soon” and has said nothing about M3.1.
Is M3.1 Flash better than M3?
Faster, not better. Independent tests put it at roughly half M3’s latency, and at 10 of 15 single-shot passes against M3’s 14 of 15.
What was space-bunny-alpha?
An anonymous $0 model on OpenRouter from September 23, 2026. Fingerprinting returned 24/24 and 50/50 token matches to MiniMax models. The company has not confirmed the connection.
Does this affect MiniMax stock?
Indirectly. There is no new revenue line attached to an unpriced preview. The stock’s August move — $1.4 billion in mainland inflows — was driven by positioning, not by this model.
The bottom line
MiniMax M3.1 Flash is a real capability release and an unfinished commercial product. The speed gain is measurable and large: half the latency, a third more throughput on short tasks. The quality cost is also measurable, and nobody at MiniMax has acknowledged it.
Until there is a model card, a benchmark report and a rate card, treat it as a fast preview for cheap work and nothing more. The tokens you lose to a model that thinks for 16,384 tokens and returns an empty block are not cheaper than the tokens you would have spent on a model that answers.
For investors, the signal is not the model. It is that a company with $1.4 billion of fresh mainland inflows can ship a frontier-adjacent coding model without charging for it, and take the usage data as payment.
Sources
- Startup Fortune — MiniMax slips a new coding model into its agent tool without a price tag
- APIMaster.AI — MiniMax M3.1 API: Flash Preview is live in MiniMax Code
- AI @ Sulat.com — MiniMax M3.1-Flash-Preview, tested outside MiniMax Code
- AlphaSignal — MiniMax quietly slips M3.1-Flash-Preview into its coding tool
- MiniMax — M3 model page (official specs and benchmarks)
- Bloomberg — MiniMax emerges as mainland China investors’ new favorite stock
Leave a Reply