Source and status check (7 October 2026): Mistral’s dated launch announcement confirms a public API preview, while its live model card lists the current specification and promotional price. The weights are planned for 27 October according to Reuters’ original reporting; they were not downloadable when we checked. TechCrunch, VentureBeat and Le Monde separately reported the launch. The benchmark scores below are Mistral’s claims or its presentation of evaluator data, not a blanket independent endorsement.

Mistral AI has unveiled Mistral Large 4, nicknamed “Le Chonk,” a new 1.05-trillion-parameter multimodal AI model that the French company is positioning as its most capable open-weight model yet. The model entered public preview on October 6, with access available through Mistral Studio, while Mistral says the model weights will be released by the end of October.

Large 4 uses a granular Mixture-of-Experts architecture rather than activating its entire parameter count for every request. Mistral’s launch announcement describes 49 billion active parameters, while the company’s current model documentation lists 52 billion active parameters and 1.05 trillion total parameters. That distinction matters because the trillion-parameter headline describes the total model size, not the amount of computation used for every token.

Key takeaways

  • Mistral Large 4 is a 1.05-trillion-parameter multimodal model.
  • It is based on a Mixture-of-Experts (MoE) architecture.
  • Mistral’s current documentation lists 52 billion active parameters; its launch announcement says 49 billion.
  • The model has a 1-million-token context window.
  • It accepts text and image inputs.
  • Mistral plans to release the model weights by the end of October 2026.
  • The preview API is already available through Mistral Studio.
  • Mistral documentation currently displays promotional preview rates of $0.68 per million input tokens and $2.09 per million output tokens; $1.36/$4.18 are the list rates. Rates can change.
  • Mistral says Large 4 is particularly strong in coding, cybersecurity, agentic workflows, finance, law, science and visual understanding.
  • The model was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s European data centers.

What is Mistral Large 4?

Mistral Large 4 is the latest flagship model from Mistral AI, the Paris-based artificial intelligence company that has built its reputation partly around releasing powerful models with more open access than the industry’s largest closed-model providers.

The company’s latest model represents a significant increase in scale.

Mistral describes Large 4 as a 1-trillion-parameter natively multimodal model. Its current documentation gives the more precise total as 1.05 trillion parameters and describes it as a general-purpose multimodal model using a granular Mixture-of-Experts architecture.

The model can work with text and images and is designed to handle general reasoning as well as specialized workloads.

Mistral is particularly highlighting software engineering, cybersecurity, finance, legal work, manufacturing and other enterprise applications.

The company is also emphasizing agentic use cases, where an AI model does more than answer a single prompt and instead works through multi-step tasks using tools and external information.

Why the 1-trillion-parameter figure matters

A model containing more than one trillion parameters sounds enormous, but parameter count alone does not tell developers how expensive every inference will be.

Large 4 uses a Mixture-of-Experts architecture.

In an MoE model, the full collection of parameters is divided into specialist components, often called experts. The model’s routing mechanism activates only a subset of those experts for a particular token or task.

That means the model can have an enormous total parameter count while requiring substantially less computation per token than a dense model with the same total number of parameters.

This is one reason the distinction between total and active parameters is important.

Mistral’s launch post says Large 4 has 1 trillion parameters with 49 billion active parameters, while its model documentation currently lists 1.05 trillion total and 52 billion active. The underlying message is the same: the model is extremely large overall, but only a fraction is activated for each inference step.

Mistral Large 4 total and active parametersMistral documentation lists 1,050 billion total parameters and 52 billion active parameters. Its launch post said 49 billion active.Total capacity is not compute per token1,050B total52B activeSource: Mistral model documentation; launch post separately reports 49B active.
Mistral’s live model card gives 1.05T total and 52B active parameters; its launch article says 49B active. Neither figure is a serving-cost quote.

A 1-million-token context window

Another major specification is the context length.

Mistral’s documentation lists a 1-million-token context window for Large 4.

A context window determines how much information a model can process within a single interaction.

A million-token window is particularly relevant for enterprise workloads involving large document collections, source-code repositories, financial filings, technical documentation or lengthy research workflows.

For example, a software engineering agent could potentially work with a much larger portion of a repository in one context.

A financial research system could process large quantities of filings and supporting documents.

A legal workflow could potentially analyze extensive case material without repeatedly splitting the information into smaller conversations.

A large context window does not automatically mean perfect comprehension of everything inside it. Retrieval quality, attention allocation, reasoning ability and application architecture still matter.

But it gives developers significantly more room to construct long-context workflows.

Multimodal from the start

Large 4 is not simply a text model with an enormous parameter count.

Mistral describes it as natively multimodal, meaning it can process visual information alongside text.

That opens the model to applications where information exists in formats other than plain text.

Potential workloads include:

  • Technical drawings
  • Charts and graphs
  • Scanned documents
  • Engineering diagrams
  • Financial documents
  • Satellite imagery
  • Photographs
  • Visual inspection
  • PDF analysis

Mistral says its visual capabilities are particularly strong in visual grounding, where a model must identify and reason about specific objects or information inside complex images.

The company reports that Large 4 scored 42% on the Dense 200 visual-grounding benchmark, compared with 41% for GPT-6 Astra in its cited evaluation. That is a Mistral-reported comparison, so it should not be interpreted as a universal ranking across all vision tasks.

Mistral is targeting coding and AI agents

Coding is one of the central areas Mistral is highlighting.

Large 4 reportedly scored:

BenchmarkMistral Large 4 result
DeepSWE v1.161.7%
SWE-Atlas-QnA59.4%
Terminal-Bench 428.3%
Combined Coding Agent Index49.8%

Mistral says the figures place Large 4 ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max on its combined coding-agent comparison. The individual benchmark results are based on figures reported by Artificial Analysis, while the combined index is Mistral’s presentation of those results.

The significance is less about a single benchmark ranking and more about the model’s intended role.

Mistral is building Large 4 for agents that can understand a repository, use a terminal, execute tools and work through multi-step software-engineering tasks.

That puts the model directly into one of the industry’s fastest-growing AI application categories.

Mistral says Large 4 is strong in cybersecurity

Cybersecurity is another major focus.

Mistral says Large 4 ranks among the top five models globally on Artificial Analysis’ Cyber Index and leads open-weight models developed outside China by a significant margin.

The company reports an 82% score on one test involving reproducing and patching a real vulnerability in open-source software.

It also reports that Large 4 solved 93% of Cybench challenges, a set of cybersecurity exercises derived from security competitions.

These capabilities could have legitimate uses in vulnerability research, malware analysis, incident response and defensive security.

They also create obvious dual-use concerns.

Mistral says it is red-teaming the model with cybersecurity leaders, vetted partners and state authorities before the weights are fully released. Those partners receive access to a version with reduced moderation and expanded cyber capabilities.

That testing phase is important because an open-weight model can ultimately be deployed outside the provider’s infrastructure.

Open weights are the bigger story

The most consequential part of the announcement may not be the parameter count.

It is the planned release of the weights.

Mistral says it will release the Large 4 weights by the end of October.

Until then, developers can use the preview API through Mistral Studio, but the model cannot yet be independently self-hosted from the released weights.

Open weights provide developers and enterprises with considerably more control than a conventional hosted API.

A company can potentially run the model in its own infrastructure.

It can control where the model operates.

It can customize deployment.

It can potentially fine-tune or modify the model depending on the final license and release terms.

For organizations with strict data-residency requirements, this can be especially important.

Mistral Large 4 availability timelinePublic preview began October 6. Downloadable weights are planned for October 27, according to Reuters, pending release.A preview is not a weight release6 Oct 2026API preview live27 Oct plannedWeights pendingSources: Mistral announcement; Reuters reporting. Future date remains a company plan.
Teams can test the hosted preview now. Self-hosting depends on the promised weights and licence.

Europe gets its own frontier-model bet

Mistral is also making the model’s infrastructure part of the story.

Large 4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s European data centers. The preview is also being served from that infrastructure.

The company says its European deployment can operate independently of other digital service providers and under European law.

This ties Large 4 to Mistral’s broader argument around European AI sovereignty.

The company wants European organizations to have access to frontier AI without depending entirely on American or Chinese model providers.

That does not mean Mistral operates outside the global AI supply chain. The training infrastructure itself uses NVIDIA hardware.

But owning more of the infrastructure and model stack gives Mistral greater control over deployment.

Mistral is competing with both open and closed models

Large 4 enters an increasingly crowded frontier-model market.

Mistral is competing against closed systems from OpenAI, Anthropic and Google while also facing open-weight models from Chinese and other international developers.

Reuters reported that Mistral is positioning the new model as competitive with leading open-weight models and said CEO Arthur Mensch believes it outperforms several Chinese competitors in particular areas.

Mistral itself makes broader claims, saying Large 4 significantly outperforms open-weight models developed in the US and Europe and is competitive with the strongest open models globally. Those are company claims and should be distinguished from independent evaluations.

The competitive picture will become clearer once the weights are publicly available and independent researchers can run the model under controlled conditions.

The API is already available

Developers do not have to wait until the weight release to experiment with Large 4.

Mistral Studio currently provides a public preview API.

The model’s documentation lists the following standard and promotional preview specifications:

SpecificationMistral Large 4
Total parameters1.05T
Active parameters52B in current documentation
ArchitectureGranular MoE
InputText + image
Context1M tokens
Input price$0.68 promotional / $1.36 list per 1M tokens
Cached input$0.07 promotional / $0.14 list per 1M tokens
Output$2.09 promotional / $4.18 list per 1M tokens
StatusPublic preview

Mistral’s documentation also lists support for structured outputs, function calling, document Q&A, batching, agents and conversations.

These prices are preview API rates and should not be confused with the eventual economics of self-hosting the open-weight model.

The model is multilingual

Mistral says a significant portion of Large 4’s training data was multilingual and covered more than 160 languages, including every official language of the European Union.

That is strategically relevant to Mistral.

Many frontier AI systems are strongest in English, even when they support many languages.

A European model developer has a particular incentive to build strong multilingual capabilities across European languages.

However, training-data language coverage should not be interpreted as proof that the model has equal accuracy across all 160-plus languages.

Independent language-by-language evaluations will be necessary to establish that.

Large 4 also targets finance and legal work

Mistral is positioning Large 4 beyond general chat and coding.

The company says the model performs strongly on professional workflows involving finance and law.

On Harvey’s Legal Agent benchmark, Mistral says Large 4 outperforms other open-source models.

It also reports strong results on Finance Agent v2 and FinWorkBench, which evaluate financial reasoning and spreadsheet-oriented tasks.

These are important workloads because they involve long documents, structured information, numerical reasoning and multi-step processes.

They are also areas where enterprises are increasingly experimenting with AI agents.

Still, benchmark performance is not the same as professional reliability.

A strong score on a legal benchmark does not make the model a substitute for a lawyer, just as a strong finance benchmark does not make it an autonomous financial adviser.

Safety becomes more important once weights are released

Open-weight frontier models create a different safety challenge from hosted models.

A provider can impose restrictions at the API layer.

An organization that downloads the weights can deploy the model under its own infrastructure and policies.

Mistral says Large 4 achieved a 93.3% resistance score on Lakera’s public B3 AI Security Benchmark and reports strong performance against indirect prompt-injection attacks. It also says the model has a relatively high refusal rate on malicious cybersecurity prompts compared with other open models.

Those results are encouraging but do not eliminate the broader risks associated with distributing powerful model weights.

That is why Mistral is conducting additional red-teaming before the complete release.

Mistral says training is still improving

Another unusual part of the announcement is that Mistral does not present Large 4 as a completely finished system.

The company says the reinforcement-learning process behind the preview is still running and that it expects additional improvements as training continues.

Mistral says its training infrastructure can produce roughly 33 billion tokens per day at its current 3,000-GPU scale, including around 16 billion trainable completion tokens after filtering and masking.

The implication is that the October preview may not represent the final capability of Large 4.

Mistral plans to publish more information about the architecture, benchmarks and post-training methodology alongside the weight release.

What the launch means for the open-model race

Large 4 raises the stakes for open-weight AI.

The model demonstrates that open-weight developers are willing to compete at trillion-parameter scale while still using sparse architectures to keep inference manageable.

It also illustrates how the competition is moving beyond raw language performance.

The important battlegrounds increasingly include:

  • Coding agents
  • Cybersecurity
  • Long-context reasoning
  • Multimodal understanding
  • Enterprise workflows
  • Tool use
  • Self-hosting
  • Data sovereignty
  • Inference economics

A model that performs well across those dimensions can be more valuable to enterprises than a model that simply wins a general reasoning benchmark.

What developers should watch next

The biggest milestone now is the weight release.

Until developers can download and run Large 4 themselves, questions remain about actual hardware requirements, inference throughput, quantization options, fine-tuning and the practical cost of self-hosting a model of this scale.

The final license will also matter.

So will independent benchmark results.

Mistral has already provided a substantial set of evaluations, but independent researchers will be able to test the model under their own methodologies only after the weights become available.

The bigger picture

Mistral Large 4 is important not simply because it crosses the one-trillion-parameter threshold.

Its larger significance is that Mistral is attempting to combine frontier-scale capability, open weights, multimodal input, long context and enterprise-focused agentic performance in a single model.

That combination is increasingly becoming the defining competition in open AI.

The model also strengthens Europe’s position in a market dominated by American and Chinese companies. Mistral is using European infrastructure, emphasizing data sovereignty and preparing to release the model weights rather than limiting access to an API.

Whether Large 4 becomes the leading open-weight model will depend on independent evaluations and real-world adoption. But the launch makes one point clear: the open-model race is no longer confined to relatively small research models. Trillion-parameter-class systems are now becoming part of the competition.

Looking Ahead

The next major test arrives when Mistral releases the Large 4 weights later this month. That release will reveal how practical a 1.05T-parameter MoE model is for self-hosting, how much hardware it requires and whether its benchmark performance translates into real developer workloads. It will also allow researchers to independently evaluate the model’s safety and capabilities.

For enterprises, the most interesting question may ultimately be control rather than parameter count. If Large 4 delivers strong coding, multimodal, cybersecurity and knowledge-work performance while allowing organizations to deploy the model under their own infrastructure and policies, Mistral could strengthen the case for open-weight frontier AI as an alternative to purely hosted models.

What this means for Indian AI teams

For Indian developers evaluating Mistral AI, the immediate option is a hosted Large 4 preview. The open-weight promise is a later event, and actual on-premises deployment will depend on the final licence, hardware requirements and independent testing. Teams should compare the promotional API rate with normal pricing, then measure cost per completed task rather than price per token alone. Our earlier coverage of Reflection AI’s pending Beam weights and Moonshot’s Kimi K3 open-weight model provides two distinct availability comparisons. Each model needs its own licence and performance checks.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.