Mistral Large 4 is Mistralβs new flagship model for multimodal reasoning, coding and tool-driven workflows. The practical distinction is availability: the public-preview API is usable now; downloadable weights are not yet available. This guide separates what a developer can buy today from what Mistral has announced for later.
Availability and model identity
Mistral announced Large 4 on October 6, 2026. Its preview is accessible through Mistral Studio and the API with model ID mistral-large-4. The announcement schedules weights for the end of October, not launch day. A preview can change before the final weight release; do not treat it as a frozen production checkpoint.
| Item | Verified state on October 7 |
|---|---|
| API | Public preview, usable now |
| Model ID | mistral-large-4 |
| Context | 1 million tokens in the model documentation |
| Inputs | Text and images; native multimodal architecture |
| Outputs | Text |
| Architecture | Mixture of experts; approximately 1 trillion total parameters |
| Active parameters | Official sources disagree: 49B in the launch announcement, 52B in live model documentation |
| Weights | Announced for the end of October; not yet downloadable |
Why we do not give one definitive active-parameter number
The launch announcement states 49B active parameters. Both the current Large model route and the versioned Large 4 card returned 52B active, 1.05T total and a 1.6B vision encoder when checked on October 7. Another indexed copy of the versioned card had surfaced 49B. We have not found an official reconciliation. We use the live model documentation for operational API specifications and pricing, but preserve the parameter discrepancy rather than guessing which components account for it.
API pricing: introductory rate versus list rate
The model card displays these USD prices per million tokens:
| Token category | Introductory preview rate | Displayed list rate |
|---|---|---|
| Input | $0.68 | $1.36 |
| Cached input | $0.07 | $0.14 |
| Output | $2.09 | $4.18 |
Mistralβs October 6 changelog says launch pricing is 50% off for two weeks. Budget against the list rate for longer-lived workloads and re-check the offer before sending traffic; the discount is not a permanent price cut. Cached input applies only when the request qualifies for caching.
For illustration, one million uncached input tokens plus 100,000 output tokens costs $0.889 at the introductory rates, versus $1.778 at the displayed list rates. This is arithmetic, not a measured application workload. Tool execution, retries and other services can add costs. The Get API Pulse Mistral calculator lets you compare token budgets, while its provider reference retains earlier models rather than conflating generations.
Capabilities and integration choices
The official card lists structured outputs, function calling and document question answering, including chat-completions and conversations interfaces. That makes the preview relevant to coding assistants, document analysis and agents that must return constrained data or invoke tools. A million-token context is a capacity specification, not a guarantee of reliable retrieval across every token.
Start with an isolated evaluation against your current model: the same tasks, permissions, tool descriptions, error handling and cost accounting. Verify your SDK or gateway passes the documented model ID, inspect usage fields, and test tool-call failures as well as successful responses. Do not silently swap an existing production alias for a changing preview.
Benchmarks: promising, but not our test results
Mistralβs launch material emphasizes agentic coding, business workflows and visual understanding. These are vendor-reported evaluations, not AI Made Tools hands-on benchmarks. Their task sets, scaffolds, refusal policies and scoring differ, so they cannot establish a universal winner or be mixed into an older SWE-bench or HumanEval ranking. We have not independently measured Large 4βs coding reliability, latency or prompt-injection resistance.
For a useful buying decision, evaluate repository tasks, long documents or tool workflows representative of your own application. Compare completed-task cost, not token price alone. Our agent model guide explains the operational dimensions; the Mistral model directory helps distinguish the flagship preview from smaller and older models.
Open weights and local deployment
An announced weight release is not a download. As of this check, we cannot verify a downloadable production checkpoint, final license, quantization support, runtime recipe or serving memory requirements. We therefore do not recommend a GPU configuration or provide a local-install command for Large 4.
Use the open-weight coding guide for models you can actually obtain today. Re-evaluate Large 4 for self-hosting only after official weights, licensing and supported deployment documentation exist. The API preview and the later weight release may also require separate evaluations.
Should you use it now?
- Evaluate the preview if native vision, document reasoning or multi-step tool use matters and you can tolerate preview changes.
- Keep an established production model when stability and predictable behavior are more important than testing a new flagship.
- Wait for the weights when offline operation, a specific license or independently verified self-hosting economics is a hard requirement.
This is a developer and buyer reference, not a claim that we have tested a released weight checkpoint. We will revise the lifecycle and pricing sections when official availability changes.