In short
The new Mistral model is noticeably stronger than its predecessor, but on Artificial Analysis it trails DeepSeek 4.1 Flash, which has fewer total parameters. I examine what this comparison tells us and what we can expect from open weights for now.
Mistral Large 4 looks like a return to the race, but model size alone does not make it a leader. In the Artificial Analysis test, it scored 38 points and fell slightly short of DeepSeek 4.1 Flash, which has 552 billion parameters compared with Mistral’s trillion.
At the same time, only 49 billion of Mistral Large 4’s trillion parameters are active. The model is currently available in preview through the Mistral API. The company promises to release the open weights at the end of the month.
The model has two reasoning levels: “none” and “high.” In Simon Willison’s example, the “high” version handled the pelican drawing better and used fewer output tokens than “none.” An interesting illustration, but not proof that one mode is generally better.
There is also reason for optimism. The previous Mistral Large 3 scored 9 points in the same test, so the progress is significant. But the limitations are also clear: a score of 38 is still below DeepSeek 4.1 Flash, and the source contains no comparisons with other models. The open weights have only been promised so far, so it is not yet possible to evaluate them separately.
Would you choose a model with a larger total number of parameters, or would you focus primarily on its performance on the task you need?
Source: Simon Willison's Weblog