Mistral Large 4 Scores 59.9% on AutomationBench Across Four Workplace Apps

Mistral Large 4's AutomationBench result puts cross-application AI workflows at the center of its public preview, while open weights and EU deployment plans add strategic context.

Mistral Large 4 Scores 59.9% on AutomationBench Across Four Workplace Apps
Mistral Large 4 AutomationBench Score: 59.9%

Mistral AI has introduced Mistral Large 4 (ML4) in public preview, pairing its new multimodal model with a notable result on AutomationBench. According to Mistral AI's official Mistral Large 4 announcement, ML4 scored 59.9% across 657 business workflows spanning simulated Gmail, Google Sheets, Slack, and Salesforce environments. Mistral says the result placed ML4 ahead of Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro.

The result matters because useful workplace AI increasingly depends on work that crosses applications. A workflow might involve reading an email, updating a spreadsheet, notifying a colleague in chat, and recording information in a customer system. AutomationBench is designed to test that kind of cross-application orchestration rather than a single isolated prompt or task.

Mistral is positioning ML4 as an open-weight, multimodal foundation model for specialized successors. The company said it plans to release the model's weights by the end of October 2026. It also described a multi-region deployment strategy, including an EU deployment operated end-to-end under European law.

What Mistral Large 4's AutomationBench result shows

A 59.9% AutomationBench score is a benchmark result, not a guarantee that an AI system can safely run real business processes without oversight. The 657 workflows were performed in simulated workplace applications. Real implementations can involve different permissions, data formats, business rules, exceptions, and consequences when an action is wrong.

Still, the benchmark is useful because it focuses on a practical challenge: coordinating work across the applications where teams already communicate, track information, and manage customers. Mistral describes the result as progress in cross-application workflow orchestration for enterprise use, but the underlying capability is relevant to any organization evaluating AI-assisted operational workflows.

For business teams, the most relevant signal is not simply that a model can draft content or answer questions. It is whether it can maintain context while progressing through a multi-step process that touches several systems. That is the difference between an assistant that suggests the next action and one that can potentially support a broader workflow with appropriate controls.

How ML4 compares with the models named by Mistral

Mistral's announcement identifies three models that ML4 outperformed on AutomationBench. It does not provide their individual benchmark scores in the supplied research, so the comparison should not be read as a full ranking of every workflow automation model or deployment option.

Model AutomationBench result stated in Mistral's announcement Context available from the supplied research
Mistral Large 4 59.9% across 657 business workflows Tested across simulated Gmail, Google Sheets, Slack, and Salesforce environments
Kimi K3 ML4 finished ahead Individual score not supplied
MiMo-V2.6-Pro ML4 finished ahead Individual score not supplied
DeepSeek V4 Pro ML4 finished ahead Individual score not supplied

The comparison is meaningful as evidence that ML4 performed well in this specific evaluation. It does not establish that ML4 will be the best choice for every business process. Model selection still depends on the actual workflow, the systems involved, deployment needs, and the level of human review required.

What open weights and regional deployment could mean

Mistral's stated plan to release ML4's weights is strategically important. Open-weight models can give organizations more flexibility over where and how they run a model than an approach limited to a single hosted service. However, the supplied announcement does not specify implementation requirements, pricing, hardware needs, or the final terms of the planned weights release.

The company also highlighted multilingual training data spanning more than 160 languages and an EU deployment operated under European law. Those details broaden ML4's relevance for organizations that work across languages or place importance on regional deployment. They should not be treated as a substitute for assessing the specific data handling and system setup of a planned workflow.

Mistral also cites agentic workflows, cybersecurity, and knowledge-work tasks among ML4's broader areas of focus, alongside benchmarks such as CyBench, AA Cyber Index, DeepSWE, and Terminal Bench. The AutomationBench result is therefore one component of a wider model launch, but it is particularly concrete for teams exploring automation across familiar business software.

Turning benchmark potential into a workable process

The practical lesson from AutomationBench is that workflow automation should be evaluated as a system, not as a model score alone. Before connecting an AI model to operational tools, teams should identify where automation can create value and where a person must remain the decision-maker.

Useful early candidates often have several characteristics:

  • They involve repetitive steps across email, spreadsheets, messaging, or customer records.
  • The expected outcome can be clearly defined and checked.
  • The workflow has a manageable exception path for unusual cases.
  • A person can review consequential actions before they are finalized.

ML4's benchmark performance suggests improving capability in this category, but it does not demonstrate a ready-made integration with Gmail, Google Sheets, Slack, or Salesforce. The benchmark used simulated applications. A production deployment still needs the relevant system connections, permissions, testing, monitoring, and workflow design.

Cross-application benchmark gains matter only when they translate into dependable production workflows. Businesses considering agentic automation need to identify high-value tasks, connect their existing tools safely, keep people in control where judgment matters, and clarify the implementation path. Scalevise helps teams turn model capabilities into practical workflows that reduce repetitive work without overcomplicating operations. Request an AI automation consultation with Scalevise.

Frequently Asked Questions

What is Mistral Large 4?

Mistral Large 4, or ML4, is a multimodal model introduced by Mistral AI in public preview on October 6, 2026. Mistral describes it as a foundation for specialized open-weight successors.

What did Mistral Large 4 score on AutomationBench?

Mistral Large 4 scored 59.9% on AutomationBench across 657 business workflows involving simulated Gmail, Google Sheets, Slack, and Salesforce environments.

Which models did Mistral Large 4 outperform on AutomationBench?

Mistral said ML4 finished ahead of Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro. The supplied announcement does not provide their individual scores.

Does the AutomationBench result prove ML4 can automate live business systems?

No. The result comes from simulated workplace applications. A live deployment requires actual integrations, permissions, workflow testing, and suitable human oversight.

When will Mistral Large 4 weights be released?

Mistral said it plans to release ML4's weights by the end of October 2026.


Conclusion

Mistral Large 4's 59.9% AutomationBench result gives the public preview a clear practical focus: coordinating work across commonly used business applications. Its lead over the named rivals is notable within this benchmark, while the planned open-weight release and EU deployment add broader deployment context. For businesses, the key next question is how well such capability can be applied to a defined, controlled workflow in real operating conditions.