⏱️ Reading time: 8 min

Mistral AI put into public preview a model that activates less than 5% of its own parameters per token: Mistral Large 4 spreads 49 billion active parameters across a total of 1.05 trillion, an architectural leap from the previous generation of the Large family.

📑 En este artículo
  1. TL;DR
  2. What Mistral Large 4 Is
  3. What Happened
  4. Context and History
  5. Technical Details
  6. Impact and Analysis
  7. What’s Next
  8. Frequently Asked Questions
    1. How many parameters does Mistral AI’s new model have?
    2. What does it mean for Large 4 to be “open-weight”?
    3. How much does it cost to use Mistral’s new MoE via API?
    4. Does Large 3’s successor support images?
    5. Did the context change compared to Mistral’s previous version?
  9. References

The release was documented on October 6, 2026 under the “Public Preview” label and internal version v26.10, with a context window of 1 million tokens and a built-in 1.6 billion-parameter vision encoder to process images alongside text.

TL;DR

  • Mistral AI releases Large 4: an open MoE with 1.05 trillion parameters, 49 billion active.
  • It adds a 1.6 billion-parameter vision encoder to understand text and images together.
  • Context reaches 1 million tokens thanks to granular expert routing (MoE).
  • Pricing ranges from $0.68 to $1.36 per million input tokens and $2.09 to $4.18 for output.
  • Version v26.10 already supports function calling, structured outputs, batching, and agents with built-in tools.

What Mistral Large 4 Is

Mistral Large 4 is Mistral AI’s flagship model, a granular, multimodal, open-weight Mixture-of-Experts system that activates 49 billion parameters out of a total of 1.05 trillion, with a 1.6 billion-parameter vision encoder and a 1-million-token context.

The technical sheet lists it under the identifier mistral-large-4 and marks it as “Open,” the label Mistral reserves for releases that include downloadable weights instead of staying solely behind a closed API.

What Happened

Mistral uploaded the Large 4 sheet to its documentation with the status “Public Preview” and version v26.10. The models page lists the system alongside an additional variant marked “+1” that the available excerpt doesn’t detail, and it specifies support per endpoint: /v1/chat/completions and /v1/conversations cover structured outputs, function calling, Document QnA, and conversation prefixes; /v1/batch enables batch processing; /v1/agents adds agents with built-in tools.

The model takes Large 3’s place at the top of the catalog and keeps the same 1-million-token context window, but it changes the internal compute distribution: instead of activating all of its parameters on every pass, it routes each token to a subset of trained experts.

Context and History

Mistral AI built its Large catalog starting in 2024 as a European alternative to OpenAI and Anthropic models, with the distinctive feature of releasing downloadable weights for much of its lineup instead of staying limited to a closed API. The company had already tested the Mixture-of-Experts architecture before, with Mixtral 8x7B as its first open-weight MoE model in late 2023; Large 4 extends that bet to a much larger scale and with more granular routing.

Mistral releases downloadable weights for much of its Large catalog. Foto de ThisisEngineering en Unsplash

The same documentation page lists Z.ai GLM 5.3, Z.ai GLM 5.2, and Shieldstral 1.0 as “Other Models.” The presence of two consecutive versions of a Chinese model alongside Mistral’s own catalog suggests the company positions Large 4 explicitly against China’s open offerings, not just against the closed models from US labs.

Technical Details

The key difference between Mistral’s new MoE and a dense model lies in how it decides which parts of the network to use. In a Mixture-of-Experts architecture, the model splits its parameters into many small experts instead of a few large ones, and a router learns to send each token to only a subset of them. The architecture being “granular” means, in practice, more and smaller experts than in a traditional MoE, which gives the router more possible combinations to tailor the compute to each token.

That granularity explains the gap between the 49 billion active parameters and the 1.05 trillion total: most of the network stays inactive on any given pass, which lowers inference cost compared to running a dense model of the same total size. The cost of training and storing the full model is still that of 1.05 trillion parameters; what changes is the compute per token at inference time.

The following diagram simplifies that routing: an input token reaches the router, which assigns it to a limited group of experts among the many available, and only those experts take part in the computation for that position.

flowchart TD
A["Input token"] --> B["MoE Router"]
B --> C["Expert 1"]
B --> D["Expert 7"]
B --> E["Expert 23"]
C --> F["Weighted combination"]
D --> F
E --> F
F --> G["Token output"]

In addition, Large 4 integrates a 1.6 billion-parameter vision encoder: this component processes images and translates them into a representation that the rest of the model can combine with text, instead of relying on a separate vision model connected outside the main system.

📌 Note: The original Spanish-language figure uses “1,05 billones,” which follows the Spanish numerical convention where “billón” equals 10^12, the same value as “trillion” in English. This is not a translation error or an exaggeration of Mistral’s original figure.

Mistral Large 4’s pricing sheet lists two cost columns per million tokens, without the available excerpt clarifying whether they correspond to two service modes (the “Speed” and “Performance” tabs appear on the same page) or to two contract tiers:

Token typeTier 1 priceTier 2 price
Input$1.36 /M tokens$0.68 /M tokens
Cached input$0.14 /M tokens$0.07 /M tokens
Output$4.18 /M tokens$2.09 /M tokens

In all three cases, tier 2 costs exactly half of tier 1, a pattern that fits better with a volume discount or a lower-priority mode than with an actual change in which model answers the query.

To try the model, just point the identifier mistral-large-4 at the chat completions endpoint listed in the technical sheet itself:

curl https://api.mistral.ai/v1/chat/completions -H "Authorization: Bearer $MISTRAL_API_KEY" -H "Content-Type: application/json" -d '{"model": "mistral-large-4", "messages": [{"role": "user", "content": "Summarize in one line what a granular MoE is."}]}'

The call follows the same /v1/chat/completions format that Mistral documents for the rest of its catalog: the response arrives as a JSON object with the field choices[0].message.content containing the generated text.

The honest limitation here is that the documentation excerpt doesn’t include any quality benchmarks: there’s no way to confirm, with the available data, whether Large 4’s 49 billion active parameters perform better than a dense model of comparable size. The architecture numbers are confirmed by the technical sheet itself; actual performance remains pending independent evaluations.

Impact and Analysis

The ratio between active and total parameters defines the real inference cost of an MoE. Foto de Trnava University en Unsplash

The model arrives at a moment when the discussion around open models has shifted from “whether they release weights” to “how much compute they activate per token.” The roughly 21-to-1 ratio between total and active parameters puts Mistral in the same conversation as other labs that have been shrinking the active fraction of their MoEs to cut inference costs without giving up the network’s total capacity.

For teams in Latin America already running Mistral models in production, the most concrete change isn’t the model’s size but the API contract: keeping the 1-million-token context and endpoints compatible with the previous generation reduces the migration to changing the model name in the call, without touching the rest of the prompting, function calling, or agent pipeline.

The implicit comparison against Z.ai GLM 5.3, GLM 5.2, and Shieldstral 1.0 on the same catalog page is also a positioning signal: Mistral explicitly competes against Chinese open models, not just against the closed models from US labs.

What’s Next

The Large 4 sheet is still marked “Public Preview,” a status that in Mistral’s history usually precedes general availability weeks or months later, pricing adjustments included. The documentation excerpt mentions an additional variant (“+1”) without detailing it, and lists Shieldstral 1.0 among the “Other Models” on the same page, leaving open the possibility of a safety or moderation model paired with the release.

Try it yourself: request access to the Public Preview at docs.mistral.ai and point a call at /v1/chat/completions with "model": "mistral-large-4" to compare latency against Large 3 for your own use case.

📬 Get new articles by email

We only email about big articles (1-2 a month).

Frequently Asked Questions

How many parameters does Mistral AI’s new model have?

It has 1.05 trillion total parameters, of which it activates 49 billion per token thanks to its granular Mixture-of-Experts architecture.

What does it mean for Large 4 to be “open-weight”?

It means Mistral publishes the trained parameters so anyone can download them and run them on their own infrastructure, unlike a model that’s only accessible through a closed API.

How much does it cost to use Mistral’s new MoE via API?

The documentation lists between $0.68 and $1.36 per million input tokens, between $0.07 and $0.14 for cached input tokens, and between $2.09 and $4.18 per million output tokens, depending on the service tier.

Does Large 3’s successor support images?

Yes: it integrates a 1.6 billion-parameter vision encoder that processes images alongside text in the same conversation.

Did the context change compared to Mistral’s previous version?

No: the context window stays at 1 million tokens, the same value Large 3 had according to Mistral’s documentation.

References

  • Mistral AI Docs: official technical sheet for Mistral Large 4, with pricing, supported features, and preview status.
  • Mistral AI Docs: documentation portal for the API and the full Mistral AI model catalog.
  • Mistral AI: official site of the company that develops the Large model family.

📱 Enjoy this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.

Featured image: Foto de Thomas Peham en Unsplash

Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.

Leave a comment
Categories: Tech News

Andrés Morales

Developer and AI researcher. Writes about language models, frameworks, developer tooling, and open source releases. Covers ML papers, the tech startup ecosystem, and programming trends.

0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *

You can include code inside <code>…</code> or, for several lines, <pre><code>…</code></pre>.

This site uses Akismet to reduce spam. Learn how your comment data is processed.