Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.
I certainly wouldnt have predicted that 10 years ago.
Very glad to see Mistral still in the game even after some big stumbles with Large 3. I deeply hope that this model is 'good enough' that it becomes the European go-to, giving them the resources to keep the pace up.
My spend on DeepSeek is not much and I regularly top up my balance every month as my support for all the good work DeepSeek is doing for the open science.
DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government. The reason they release the AI models is economic warfare against US, not because of charity or kindness. It's great for us consumers, but the goal is not to help humanity or open-source.
Framing it that way hides all the levers used to tilt the playing field for US companies, no?
The government that removed restrictions on how private companies can access capital after a certain scale (the JOBS Act), that removed the need for private companies to report as if they were a public company after a shareholder threshold was crossed, superpowering the access of wealthy private investors to get in earlier in a growing company while at the same blocking the public from participating in funding growing enterprises at an earlier stage (since it required companies to IPO much earlier to access capital) which allowed retail investors to also reap the rewards on funding them early when they grew to become behemoths (like Amazon, Meta/Facebook, Google, etc.).
It's not fair in either place, the USA has its own model of unfairness, China has a completely different one. The difference is that in the USA the government allows private investors to become more powerful than the State (outside of the monopoly of violence) while plunging the rest of society into increasingly more precarious lives while in China the State is the power and its legitimacy only exists while the population feel they have a better life.
IMO the selective enforcement of regulatory requirements should be added to the list. Observing from the outside, I have a hard time believing that e.g. musks gas powered data centers really follow all the environmental laws, for example, or that the authorities really see no grounds for indictment if Altmans company hacks hundreds of third parties, or that there's really no one at the SEC having a problem with Anthropics fear mongering prior to the IPO.
Well, they are helping 17% of the world's population, and the US is currently actively engaged in trade wars and economic warfare or explicitly attempting to leverage it's hegemony against its long term alliea for short term gain.
There is as much to criticize about American hyper scalers and AI labs and the lack of interest in helping humanity or contributing to open source, but that might not be as popular an opinions on this site.
If you think about it, the US labs are heavily subsidised too. Not only they receive billions in state funding, the administration is also prepared to engage in trade wars to help them.
I think the Chinese government is backing their labs by less direct means. For example cheap electricity and investing in chip manufacturers such as Huawei and cxmt.
Deepseek specifically, is known to operate with minimal resources. The entire company has around 160 employees and every model they release must break even within ten months.
I'm pretty sure, the training/fine tune through Claude/OpenAi would not be possible without the army/China's hacking teams (and legal protection). So it's more than just a cheap electricity.
Supporting deepseek is just like supporting the Chinese army, no need for that. Though it goes both ways, OpenAi subscription just lowers the cost of the US army as well.
> I think the Chinese government is backing their labs by less direct means.
It may be that the effect of the backing in both places is essentially equal but this statement is strange. The Chinese government invests directly in Deepseek[1]. Notice the article, in addition to saying the CCP is investing, says Tencent is also a major backer. CCP owns a golden share of Tencent.
If you want a direct example of the US subsidising a major AI lab, we can use Microsoft's investment in OpenAI. Microsoft has the largest ownership stake - 27%.
Microsoft did not pay with money - it paid (mostly) with Azure cloud computing credits. MSOFT is then able to write this off as a loss against tax.
It is generally far more tax-efficient in the US to write a loss in this way than it is to write a loss for a cash investment.
In this case, I believe the difference was ultimately highly significant. When including MSOFT eventually writing-off the deprecating Azure hardware it had used to buy the OpenAI equity, the result was MSOFT's tax reduction being either close-to or exceeding the actual cash value of MSOFT's investment in OpenAI - IIRC.
These examples represent taxes that the US chooses not to collect - the US could choose to make investments like these less tax-efficient. Instead, by making them extremely tax-efficient, the US subsidises the transaction hugely.
Even ignoring monetary subsidies, there are the non-monetary ones: not being sued into oblivious by the government for their countless hacks of other companies and countries, the slaps on the wrist for massive piracy, the waving of environmental (and other) regulations in order to allow their data centres to be built an operated.
That's what I'm asking - which monetary subsidies?
> not being sued into oblivious by the government for their countless hacks of other companies and countries
That isn't normally how enforcement works, and it hasn't been very long since they disclosed those breaches. If the victims want to pursue legal action, they can, and they still may!
> the slaps on the wrist for massive piracy
So judges and juries are involved in the subsidization conspiracy, too?
> the waving of environmental (and other) regulations in order to allow their data centres to be built an operated
Sure, though if you think this isn't happening in China too, I have a bridge to sell you.
I heard someone calling the key metric in Anthropic financial reports EBBT: Earnings Before Bad Things[1] :-)
[1] Where "Bad Things" would be the typical interest, taxes, depreciation, amortisation plus the Anthropic specific employee compensation, LLM training (you know, for the LLM lab), revenue sharing agreements (which is a form of paying for infrastructure), etc.
A recent YouTube video by Patrick Boyle said the same thing. They are only profitable if you ignore all the costs that make them unprofitable such as paying employees and developing A models.
The government, specifically Trump's government and the current money circle in AI inflating American company stocks. In addition to all that American models do not share their papers like Deepseek and Qwen do. So you can literally say Chinese models are doing it for charity at this point.
> The reason they release the AI models is economic warfare against US
The story is so much more complicated than that, to the point that this economic warfare theory is basically a meme.
Chinese models are open because they don’t have a choice. “When you trail the frontier, openness maximizes reputation per unit of capability. The moment you lead, you close.” [0]
I'm not sure "subsidised" is the right word. If a government funds research and the results are released openly, that's just publicly funded research. It's how a lot of science works in the US and Europe too.
If your belief is accurate, we should expect China to short the IPOs of Anthropic and OpenAI and release better frontier models immediately after their IPOs.
Does anyone think that likely? I have no clue or bias.
> The reason they release the AI models is economic warfare against US, not because of charity or kindness
There are many other reasons Chinese companies releasing models open-source or open-weight makes strategic sense.
A really easy-to-understand example is a company who has a near-monopoly on "serving video content" releasing a video model openly.
If you can be relatively certain that video content created by a model (which you have trained, using data from your own platform) will be ultimately served on your own platform, thus generating revenue from watch-hours, it makes sense to make those models as widely-available as possible.
It's also a net-positive if people use your public research to build better video models, because - again - you are reasonably certain that the even-better content those new models produce will be watched on your platform.
The alternative would making models harder to access and learn from (broadly, the current western model). Many would argue that Google, in choosing to not optimise its video generation models for "availability", is directly causing less content to be uploaded to YouTube. This is the trade-off.
I don't know much about DeepSeek's financing specifically, which obviously doesn't release video models - so I don't know how directly this analogy runs, or who directly benefits from the extremely evident rising tide that the public release of DeepSeek's research creates. However, this does not negate the broader rising-tide effect of the scientific method.
It's certainly also true that it's geopolitically beneficial to be able to undercut American labs' models. If I ran a global superpower, I would probably want my country to be technologically competitive too.
But Chinese companies are already serving a huge volume of customers in a complex, existing marketplace, before even thinking about the US market, and it's overly simplistic to assume that their entire strategy revolves around economic warfare directed specifically at the US. It's more nuanced than that.
This is, of course, without even getting into opening the can-of-worms around whether US economic policy also results in the US state functionally subsidising technological innovation, how comparable that is to China's model, etc.
> DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government.
Yes, because everything China does is against the US. That's all they think about day and night. God forbid they want to corner the global market or have a genuine business case. How dare they provide options for those who can't afford a measly $200 a month? How can we let Chinese labs publish research for free for the whole world so that they can benefit? The nerve! To think they can use soft power instead of military might! I mean, Anthropic and OpenAI are the last bastions of human kindness and charity. Right?
I’m neither in the camp of Chinese or the Americans(collective West) in general..
As a neutral party, this characterization is crazy.
As if the AI companies - Claude and OpenAI are guardians of freedom and humanity and very charitable to the global society without any self interests…
“Chinese models are subsidized by the Chinese Government, therefore they’re inherently bad for humanity” is a highly propagandist argument. The politics of US vs China may be whatever it is in reality.. You have one company releasing their models for cheap, actually open sourcing their trained weights, and publishing details of their optimizations and learnings for others to use. The other camp actively “aligning” their models, nerfing their capabilities, hyping their swarm activities from poor sandboxes, and trying their best to lock users into their harnesses and walled platforms. They are subsidized by the capitalist VCs who are essentially waiting for their payouts..
At some point, one has to see things for what they are and evaluate their own reasoning..
I’m happy to stay provider agnostic, try all models and cheer any useful progress as open as possible.
Not sure why so many people will vehemently refuse this idea. I won’t say it’s 100% true but it would be foolish to dismiss it. China is very much an adversary to America and has made it pretty clear they want to be a dominant leader of not the new world leader. Not here to evaluate what is good or bad. Keep in mind historically China has aggressively fostered industry (not unlike the west) but sometimes even more aggressively.
Americans need to travel to China. The Chinese have zero issues with us. They quite like Americans. This is such a weird propagandist take. Idk if you remember but both country's leaders just had a slumber party for 3 days in DC. This is not what enemies do.
Not even our leaders say China is an enemy, ita mostly businessmen who are scared of competition and trying to regulate chinese out of their markets so they can make more money milking us.
Economic warfare against the US is charity and kindness to a sizable portion of the world's population, especially when the US uses it's global hegemony as warfare against them. Neither system is perfect, but lets not be disingenuous.
I wouldn't even call it economic warfare. Even in the US, cheaper access to good-enough models helps everyone except for the billionaires who invest in frontier AI. So just from a numbers game effectively nobody is hurt by the open sharing of science.
Meanwhile, American AI models are heavily subsidized by stock market speculation. Ultimately, the subsidies from both countries are flowing out of the pockets of individuals.
> are heavily subsidised by the Chinese government
We hear this about literally every industry the Chinese excel in - that it's only because the government subsidizes them that they succeed. For chip manufacturing, for batteries, for EVs, for solar, for AI. I don't see how the chinese government can afford to subsidize all of these industries and still have them contribute to the GDP.
A conspiracy to make the US look bad by being better at producing all the goods and services the world needs at a reasonable price. Have they no shame?
Not really this is an X algo conspiracy. Up until recently the Chinese government wasn't even that invested in these companies. We're talking very very small grants compared to training costs.
Its very xenophobic of you to say China has zero intention of helping humanity, and just wants to "wage economic warfare".
Last time I checked, it was ourselves (USA) waging economic warfare on 2/3rds of the world.
I dont get this cope people have where people have this idea that its impossible for a Chinese company (that make billions of dollars) to have done something by their own merit, but instead its always some Chinese Communist Party conspiracy where the main goal is to destroy America.
They're trying to pull digitally what they already pulled physically. The reason we can't manufacture a grill brush for a reasonable price is the result of years of Americans choosing the cheapest price. We gave up our manufacture base. They want us to give up our labs.
Deindustrializing a country is not something consumers can achieve. It starts at the top level, with politicians who construct a financial system where it's more profitable to speculate than to build or invest in real businesses.
Also, the field moves fast, but slower than people do. Researchers and engineers switch companies every year or two, and the know-how walks out the door with them.
Architectural/algorithmic tweaks do advance the efficiency frontier nicely. But raw intelligence mostly comes from data (not just its sheer quantity, but also how it's curated & cleansed) and the scaling law. The know-how about data curation doesn't seem to get published much, even among the open-weight labs, though.
You've got to think the comments about "good enough" are people who have not yet tried Opus 5.5. I haven't been this struck by a step change since 4.5/4.6. It's a bigger jump even than when Fable first arrived.
As for Mistral - I got really excited when they said Large 4 was focusing on being #1 in cybersecurity, because that's somewhere that they genuinely could edge out Anthropic & OpenAI. Have it actually solve problems, instead of Anthropic flagging "you tried to find a null pointer exception bug in your own code, we're now reporting you to the US government". But on the Mistral benchmarks I'm seeing, this looks very disappointing, but at least they haven't entirely given up. I genuinely thought Mistral had given up on new general models. They need to learn the bitter lesson all over again.
Many on HN still have the opinion that you must understand every line of code in the project, and that all is lost should you merge code that wasn't reviewed.
Obviously any model will do if you use it as a better autocomplete.
I believe that there is a large gap in expectations between different workflows.
Until the AI like reads my mind and produces perfectly production ready apps with minimal intervention from my side, there is still going to be room for improvement.
I use Opuse 5.5 daily for my job. I am aware (abd in awe of) it's capabilities.
Look at the context in which I used that term 'good enough'.
What i was saying is that there are tasks for which a dumber model can be good enough, and for organizations with sovereignty/ privacy concerns, those concerns can be strong enough to incentivize the use of a dumber model.
Please, give it another 6 months and they catch up. The American labs are currently trying everything they can to block others instead of advancing their models, trying to build an artificial moat. The American models are not that great, they are good, and they have a lot of agentic workflows in the back, but its basically a hardware limitation at this point. Once the HW makers catch up, and we can move away from the Nvidia monopoly, things will speed up quite a lot IMO.
We've been hearing the line about them only being a few months behind for a year now, during which time O/A have grown their revenue like 10x, haven't they?
We've also been hearing we're 6 months from AGI for about three years, and here we are.
"Now, here, you see, it takes all the running you can do, to keep in the same place. If you want to get somewhere else, you must run at least twice as fast as that!"
Those are two different things. The market is expanding, so even if competitors are catching up, you can have your own revenue, in absolute terms, grow.
I agree that Opus and GPT are almsot surely better, but so many real users are nervous enough about giving Anthropic and OpenAI access to all of their internal information that they may be willing to stomach worse models if it gives them more security.
The real question is if this model is good enough that it can still accelerate work, and not be a hindrance to real work like older Mistral models often were.
I agree; yes, I can see that they need a bit less hand holding each cycle, but I also see these "frontier" agents do some absolutely dumb shit that I have to correct and then I'm wondering if I'm the looney one here.
Maybe it's because people stopped watching what their agents are doing and stopped looking at the quality of the output. But I still see agents being absolutely mindless like a junior dev.
Recent example: it updated an an API to add newly released models to the backend. There's a list of models that require specific configuration for the reasoning effort and temperature or the API call fails. GPT 6.1 Sol misses this and code fails at runtime because the newer models need to be added to the list for special handling of temp and reasoning. Fixes it for one model and tests it for that model using an E2E test. But doesn't test the other models that were added for the same error condition...I had to explicitly ask it to do so and it finds them and adds them to the list and says "that's on me."
It's an improvement, but game changer might be a bit of a stretch. If I lost access to Anthropic or OpenAI models tomorrow, I would be annoyed, but would reach for a slightly inferior model. Last year I wouldnt be able to say the same, and rhe challenge is that the moat is drying up fast. Whether its general improvements in model training by other competitors, or straight up distillation of SOTA models, the moat is shrinking and the available capital and spend for American model providers is going to dry up quickly as competing good enough models are adopted by more consumers.
It's especially the case as more non-Americans look to self hosted models and domestic cloud inference providers using open models that the US providers who are still leading the charge need to drastically drop their prices and find a path to profitability in order to maintain their lead and retain the advantage they had as AI turns into a commodity (which is happening faster than I think even the frontier labs initially predicted).
I hear this literally every other week about whatever the newest FoTM model is.
Unless you can provide concrete examples of things you can do with them that you simply couldn't do with last week's model, it's absolutely meaningless.
Yeah, I agree with this. I think the "the models are good enough" narrative is a myth. I've heard it so many times over the last year, but the model number keeps changing...
There is no ceiling on what you can accomplish with more intelligence, so there will always be a market for the best models, and that market is likely to just keep growing. If Opus 13.5 can one-shot a profitable company or discover a new disease treatment or whatever you can think of that a swarm of relentless super-geniuses could accomplish, companies (and governments) will throw money at it.
I also think there will always be a market for many sub-frontier models that will continue to grow rapidly as well, because "good enough" is definitely a thing for a given task.
Seems silly not to have predicted that 10 years ago. I feel like it's long been obvious that smarter models being available will mean way easier cheap synthetic data and access to tools that will speed up competitors as well as consumers.
On a purely technical level, maybe? But in terms of actual revenue, is there really any chance of anyone catching the big labs?
Obviously, this is only a valid question if you don't believe that open weights are about to eat their lunch and their revenue is about to collapse, or they're running a super unprofitable ponzi scheme propped up by investor money that's about to collapse like a house of cards. I don't find those positions credible at all though.
If you do, then this question isn't really for you, as I'm more interested in thoughts from those who think that OpenAI and Anthropic in particular are about to be the largest companies on earth in a couple years. Could anyone catch them at that point?
> But in terms of actual revenue, is there really any chance of anyone catching the big labs?
I don't know about revenue, but I suspect multiple other labs are already beating OpenAI/Anthropic on profitability. Staying on the frontier is expensive, and it's hard to recoup those R&D costs when you have a bunch of other labs nipping at your heels.
If you concede the previous point, then the only way for OpenAI/Anthropic to keep growing long term is to swallow the whole economy (i.e. mass job replacemnt), and that's a bet I wouldn't take.
I think the actual plan is to swallow a good portion of the job market. It’s the only thing that makes sense and I hear VC podcast debates on which percentage of jobs justifies the market cap.
Maybe. I can't freaking wait for the IPO filings so we can finally put all this to rest. (haha, like that'll actually put it to rest on HN, but at least we'll have better data)
They are very unprofitable…? I don’t think that’s really in dispute. We haven’t yet seen a profitable frontier lab and model pricing remains fairly subsidized
Impressive vision benchmarking. If the vision model is truly as good as astra, that would make it best in the world.
Also strong on cyber benchmarks (better than all chinese models), so this is a good defender model.
Lots of people shitting of Mistral for no reason imo. These are pretty good numbers across the board. Definitely good enough to use as a daily driver over other llms, if you have moral qualms with the others. For certain use cases, like cyber security, this may be the go to model.
I like to make fun of europe, but there's lots for mistral to be proud about in this release imo.
> Curious, what do you like to make fun of Europe about?
Well I'm european and... That Switzerland (Europe but not EU) has more companies in the Top 70 by market cap than the entire EU (Switzerland has two, the EU only has ASML) is kinda something that warrants making fun of.
That the biggest European software company is SAP, in 71th position is both sad and tragic: it shows how lame and irrelevant Europe is when it comes to software.
So Europe is nowhere in software and friggin nowhere in hardware: sure it's got ASML but ASML now has officially... Zero customer in Europe. Zero is not much.
Then Japan is at least trying to come back into the game with nano imprint litography. Europe is betting it all on AMSL (which anyway is majoritarily US-owned).
So software: nothing. Hardware: nothing besides ASML.
Overall the EU has six companies in the Top 100 by market cap and they're all, besides ASML, near the bottom of the Top 100.
We could also maybe make a bit fun of how the EU destroyed it's car industry (the main industry in Germany, which is the biggest economy of the union) by handing it all to chinese EVs?
Or what about the US warning the EU, years ago, to not become entirely dependent on Russia for energy? And EU not listening and then seeing its energy price skyrocket when the proverbial shit hit the fan? (Russia attacking Ukraine)
And we could, also, at least make a bit of fun of entire streets in cities like Paris and Brussels that used to have luxury shops and fancy restaurants that are all turned into places selling cheap kebabs? What a great success: I'm sure this one makes the komrades happy. It projects an image of grandeur and success: kebabs.
Or the constant attacks on free speech in the EU. Or the surveillance apparatus that's being put into place.
At this point it's more like I don't know what is there left to not make fun of about my EU.
If the UK is in europe, you should make fun of the following: they don't have a first amendment. So to me, it's an authoritarian state preaching freedom.
Europe has spent the last twenty years in stagnation - generating about half as much wealth and technology as you'd expect for its size and advanced economy. Simultaneously, Europeans are notoriously arrogant. It's a mockable combination.
Edit: unfortunately, it's a question that voting does not really permit you to answer on this website.
The US is biased to favor the top half of society. Hence why there is a lot of vocal poorer people and a lot of low profile wealthier people (I'm not including the 1% or even the 5% in this).
When you are in the 75%ish of the US, it's very easy to make a case that life in the US is better. But we don't really talk about that because it's pretty taboo when poorer people are struggling much more than they would in Europe.
Yet Europeans are healthier, happier, and have a better quality of life than Americans. They also have so much better transport infrastructure it's embarrassing for Americans. Like all the money America generates, where does it go?
A 'bubble bursting', if that happens, doesn't make the tech sector go to zero. Apple doesn't suddenly stop making iPhones. It would take a hell of a lot more than that for the US economy to fall to European levels.
Well, the Scandinavian and Sicilian do share one currency and one immigration policy and one set of regulations - which are the things that we generally look at when we analyze economies.
Norwegians, Danes, Swedes and Finns all use a different currency. Finns and Sicilians do share a currency though, but all these countries have different immigration policies and regulations. I don't think you really know how the EU works.
read Varoufakis for the story on how this happened. European surplus capital gets recycled as VC money into Silicon Valley. So it's not like there's some big choice to be made, its potentially a systemic part of the global monetary flows.
It's a bit below DeepSeek 4.1 Flash at about twice the size. For a model that was supposed to come out a few months ago this is pretty good. Mistral catching up to the chinese open-weight models is great news. Excited to see how they will build on that!
A strong competitor in cybersecurity as an alternative to GLM-5.3 (Mistral reports 82% on CyberGym-E2E). Visual grounding is also impressive (42% on Dense 200 vs. 41% for GPT-6 Astra).
Otherwise, behind on the broader Pareto frontier, but not by much (Vals Index: 48.05% vs. GLM-5.3’s 53.51%; $13.78 vs. $7.25 per test). Many companies will prefer it over Chinease models.
That was about my conclusion as well: Slightly less than GLM 5.3 performance but made in Europe. So, maybe it answers Tiananmen Square questions correctly, and in French. All in all, a reasonable model, but not frontier.
> So, maybe it answers Tiananmen Square questions correctly
What would you consider a "correct" answer? I just asked deepseek-v4.1-flash asking it what happened (without mentioning the word "protest"; here are some excerpts of what it said:
> In April 1989, students in Beijing began demonstrations after the death of Hu Yaobang, a former Communist Party general secretary. The protests grew. [..] Estimates from other sources range from hundreds to several thousand deaths. [..] The Chinese government describes the events as a counter-revolutionary riot and says the military action was necessary to restore stability. It restricts public discussion of the events inside China. Many other governments, human rights organizations, and observers describe the events as a violent suppression of peaceful protests.
So, let's see... it calls it a "protest", mentions the number of deaths, and even mentions the censorship of the topic by the CCP.
"..answers Tiananmen square questions correctly.."- but lies about Ukraine, EU, and about you, americans..
I live in EU, use for my personal needs chinese models only, and don't plan to move to any of the ones allied with the Pentagon or its european counterparts.
You are quoting the discount pricing. It is 50% off for the next two weeks. The blog post has the real pricing up front in the card on the right side: https://mistral.ai/news/mistral-large-4/
After that, it will be much closer to GLM 5.3, but you can also get 5.3 in their API! I dont see people really talking about that.
I’d rather have one European AI lab like Mistral, with the financial firepower and compute to pretrain its own large models, than 20 weak labs that fine-tune Chinese models.
AI needs big bucks, and Mistral is Europe’s Anthropic.
The US has made sure that Mistral has a large market in the EU by temporarily preventing non citizens from accessing Fable.
A lot of European companies now want a model the US can't cut off, but also lack trust in Chinese models.
Some of these will self host Mistral but most will pay them by the token. It's not going to be a huge market or a huge margin within that market but probably it'll be enough.
This is the exact mentality that makes the EU fall behind. If you don't want to invest in something until it makes a profit, you don't get the benefits of being a pioneer.
there are no profitable AI companies at the moment ... This is the exact problem of European startups, trying to make them profitable from day 1 while American counterparts (and Chinese) keep bleeding money for years. Europe will never have a Tesla, a Google or an Amazon with that mindset.
That said I love that they don't seem to restrict cyber capabilities to any degree and even lean into it.
If the best model for cyber attacks is open for everyone to use it just makes us all safer I think. Of course you then also HAVE to use it or otherwise you're vulnerable, which is a great distribution play.
Well, with this and Beam people are going to have to stop saying that western open models are dead. This is great news. Anthropic and OpenAI may have a bit more knowledge, talent and compute but they don't have a monopoly.
Investors looking for them to make monopoly profits are going to be disappointed The premium they can extract from consumers for their models will be capped. Tokens are likely to remain close to the cost of compute, a cost which is high but falling fast.
I tried Mistral's last 3 large models, devstral and mistrallarge3 and the numbers were not even benchmaxxed, but just false. the models were so weak and garbage. Let's hope they are telling the truth this time around, we need more alternatives. There mistral-small and original MistralLarge models were awesome, hoping they are back!
Seems hard to get customers at that price range when you're competing with open source models that are 1/2 - 1/3 the price but with similar capabilities.
People usually buy the cheapest, like Deepseek or GLM or they spend on Anthropic/OpenAI subs.. Are these models in the middle getting any users?
On a side note, I wonder if this was the popular free Space Bunny model that left openrouter yesterday.
Nice! Once they make it available through their API I will be happy to support them. My local Qwen3.8 27B is serving me well, but I miss the speed and concurrency that comes with subscriptions, and I am not currently paying for any.
What experience have people had with Mistral models for cyber research? are they as constrainted as Anthropic models? I use claude day2day but have the need for a model with fewer guardrails to pentest our own APIs.
I tried a quick abbreviated kimbench on their playground before bothering to do the whole thing.
Maybe I didn't really select mistral 4? Either way, failed completely on question 1 and the next 2 questions were completely off base too. I didn't bother to finish.
49b active parameters sounds manageable until you look at the 1t total weights. what hardware does a usable self-hosted setup actually need, especially once you add a long context?
Seems like they've finally made a genuinely competitive model since the original LLM craze, congrats to them! Glad to see some diversification in open weights providers.
Excited to test this on my benchmark[1], but I'm guessing it won't fare better than MiMo v2.6 Pro which is currently the best open-weight model as far as I'm concerned.
Why isn't this on openrouter yet? Is there a better router that gets models much quicker?
if it's not available yet why have a 'try it today' header at all?
> "Try it today"
>
> There is still more to come. As we work toward releasing the weights, we will share further details on the model architecture, additional benchmarks, and our post-training methodology.
This sure is nice, but I've had less than satisfactory results with GLM5.3. I'd like Mistral to compete with Qwen3.8-Flash-Next a 120B class model that IMO is the first model that I can use for serious coding while running it locally.
I estimate it's coding ability on par with opus 4.6 (but opus definitely beats it on factual knowledge) Still it's a genuinely useful model, when everything else except Anthropic's models (and for only 3 weeks after it came out Google's Gemini 3 pro, before it got merged) are not to me.
If you scroll down, most of the charts on that page are sorted s.t. Mistral's bar is right next to the worst competitor model, while the best competitor model's bar is positioned on the opposite side.
If one were to be cynical one could say that it's intentionally making Mistral's result look better than it actually is by making it harder to compare the bar heights.
I mean no disrespect but these are terrible numbers or am I missing something? It seems like Mistral continues to only be relevant for people that want a model trained in Europe. Too bad.
Does Mistral ever advance the state of the art on any dimension?
And if not, why do they exist?
Update: The number of people advocating not innovating is wild. There is no reason why Mistral cannot innovate in ML, they explicitly choose not to. My point is that, given that choice, they should spend their GPU hours differently.
"Sovereign AI" is a joke, there is no substantive difference between a post-trained open weight model from an American or Chinese company and what Mistral is doing today, beyond spending 80% of their GPU hours reproducing a last-gen model's pretraining.
Why would the Europeans want a high-quality open source model that isn't either owned by American Trillion-dollar companies or created by the Chinese? You really can't think of any reasons?
I don't know what definition of open weights you are using, but the "open" part is what allows you to use it without being beholden to American Trillion-dollar companies or the Chinese.
Mistral could do exactly what Cursor does, and post-train the model to meet the needs of Europeans.
That's a far more efficient (and useful!) use of GPUs than yet another pretrain on an outdated model backbone.
Regardless though, their reason to exist is not to advance the state of the art.
Their reason to exist is to ensure the sovereignty of France.
Obviously they would do their job better if they were advancing the state of the art, but it's not like it's pointless if they arent the absolute best.
Just like Microsoft and most companies around PCs, their existence isn't about pushing any boundaries, but seemingly targeting deploying what's already been "invented" into the corporate world and bigger companies. Not "wrong", just different.
Yes broad consensus had developed in the ten minutes between announcement and you asking if consensus had developed, and I'm excited to report that it consensed in the affirmative -- it IS Space Bunny Alpha!
Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.
I certainly wouldnt have predicted that 10 years ago.
Very glad to see Mistral still in the game even after some big stumbles with Large 3. I deeply hope that this model is 'good enough' that it becomes the European go-to, giving them the resources to keep the pace up.
I'm excited to try this out today.
I think a big part of that is the Chinese publishing the solution for everywhere hurdle in the road they've encountered in the form of a paper.
Deepseek essentially releases instruction manuals in paper form.
My spend on DeepSeek is not much and I regularly top up my balance every month as my support for all the good work DeepSeek is doing for the open science.
DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government. The reason they release the AI models is economic warfare against US, not because of charity or kindness. It's great for us consumers, but the goal is not to help humanity or open-source.
Framing it that way hides all the levers used to tilt the playing field for US companies, no?
The government that removed restrictions on how private companies can access capital after a certain scale (the JOBS Act), that removed the need for private companies to report as if they were a public company after a shareholder threshold was crossed, superpowering the access of wealthy private investors to get in earlier in a growing company while at the same blocking the public from participating in funding growing enterprises at an earlier stage (since it required companies to IPO much earlier to access capital) which allowed retail investors to also reap the rewards on funding them early when they grew to become behemoths (like Amazon, Meta/Facebook, Google, etc.).
It's not fair in either place, the USA has its own model of unfairness, China has a completely different one. The difference is that in the USA the government allows private investors to become more powerful than the State (outside of the monopoly of violence) while plunging the rest of society into increasingly more precarious lives while in China the State is the power and its legitimacy only exists while the population feel they have a better life.
IMO the selective enforcement of regulatory requirements should be added to the list. Observing from the outside, I have a hard time believing that e.g. musks gas powered data centers really follow all the environmental laws, for example, or that the authorities really see no grounds for indictment if Altmans company hacks hundreds of third parties, or that there's really no one at the SEC having a problem with Anthropics fear mongering prior to the IPO.
Well, they are helping 17% of the world's population, and the US is currently actively engaged in trade wars and economic warfare or explicitly attempting to leverage it's hegemony against its long term alliea for short term gain.
There is as much to criticize about American hyper scalers and AI labs and the lack of interest in helping humanity or contributing to open source, but that might not be as popular an opinions on this site.
Only 17%? I'd say driving down the price of AI helps at least 95% of the world's population. It certainly helps me.
That would have been a scathing criticism if anyone believed any of the American companies have the goal of helping humanity or open-source.
Notwithstanding that Chinese publishing methods actually does help both humanity and open-source.
If you think about it, the US labs are heavily subsidised too. Not only they receive billions in state funding, the administration is also prepared to engage in trade wars to help them.
I think the Chinese government is backing their labs by less direct means. For example cheap electricity and investing in chip manufacturers such as Huawei and cxmt.
Deepseek specifically, is known to operate with minimal resources. The entire company has around 160 employees and every model they release must break even within ten months.
I'm pretty sure, the training/fine tune through Claude/OpenAi would not be possible without the army/China's hacking teams (and legal protection). So it's more than just a cheap electricity.
Supporting deepseek is just like supporting the Chinese army, no need for that. Though it goes both ways, OpenAi subscription just lowers the cost of the US army as well.
>Supporting deepseek is just like supporting the Chinese army
Good, where do I sign up? At least they aren't exploding little children and generating chaos in the oil market.
https://www.business-humanrights.org/en/latest-news/anthropi...
> I think the Chinese government is backing their labs by less direct means.
It may be that the effect of the backing in both places is essentially equal but this statement is strange. The Chinese government invests directly in Deepseek[1]. Notice the article, in addition to saying the CCP is investing, says Tencent is also a major backer. CCP owns a golden share of Tencent.
[1]https://www.cnbc.com/2026/10/06/deepseek-funding-round.html
So is American models. They are subsidised heavily but still can't provide cheaper access. Their fault is to assume all countries can afford them.
By whom?
If you want a direct example of the US subsidising a major AI lab, we can use Microsoft's investment in OpenAI. Microsoft has the largest ownership stake - 27%.
Microsoft did not pay with money - it paid (mostly) with Azure cloud computing credits. MSOFT is then able to write this off as a loss against tax.
It is generally far more tax-efficient in the US to write a loss in this way than it is to write a loss for a cash investment.
In this case, I believe the difference was ultimately highly significant. When including MSOFT eventually writing-off the deprecating Azure hardware it had used to buy the OpenAI equity, the result was MSOFT's tax reduction being either close-to or exceeding the actual cash value of MSOFT's investment in OpenAI - IIRC.
These examples represent taxes that the US chooses not to collect - the US could choose to make investments like these less tax-efficient. Instead, by making them extremely tax-efficient, the US subsidises the transaction hugely.
This is a joke right?
Even ignoring monetary subsidies, there are the non-monetary ones: not being sued into oblivious by the government for their countless hacks of other companies and countries, the slaps on the wrist for massive piracy, the waving of environmental (and other) regulations in order to allow their data centres to be built an operated.
> Even ignoring monetary subsidies
That's what I'm asking - which monetary subsidies?
> not being sued into oblivious by the government for their countless hacks of other companies and countries
That isn't normally how enforcement works, and it hasn't been very long since they disclosed those breaches. If the victims want to pursue legal action, they can, and they still may!
> the slaps on the wrist for massive piracy
So judges and juries are involved in the subsidization conspiracy, too?
> the waving of environmental (and other) regulations in order to allow their data centres to be built an operated
Sure, though if you think this isn't happening in China too, I have a bridge to sell you.
Investors, but also the government in allowing these companies to siphon electricity away passing the increased cost to the consumers
>Investors, but also the government in allowing these companies to siphon electricity away passing the increased cost to the consumers
Not arbitrarily banning companies from buying a product from a supplier doesn't meet my definition of the term "subsidy".
By investors. OpenAI and Anthropic are not profitable (Anthropic is profitable if you allow them to invent what profitability means).
I heard someone calling the key metric in Anthropic financial reports EBBT: Earnings Before Bad Things[1] :-)
[1] Where "Bad Things" would be the typical interest, taxes, depreciation, amortisation plus the Anthropic specific employee compensation, LLM training (you know, for the LLM lab), revenue sharing agreements (which is a form of paying for infrastructure), etc.
A recent YouTube video by Patrick Boyle said the same thing. They are only profitable if you ignore all the costs that make them unprofitable such as paying employees and developing A models.
The government, specifically Trump's government and the current money circle in AI inflating American company stocks. In addition to all that American models do not share their papers like Deepseek and Qwen do. So you can literally say Chinese models are doing it for charity at this point.
Of course you can say that, but you literally cannot be serious if you do!
Oh I am serious!
> DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government.
DeepSeek is basically a research lab founded by a hedge fund guy, more than anything else.
> The reason they release the AI models is economic warfare against US
The story is so much more complicated than that, to the point that this economic warfare theory is basically a meme.
Chinese models are open because they don’t have a choice. “When you trail the frontier, openness maximizes reputation per unit of capability. The moment you lead, you close.” [0]
[0] https://earnedintuition.substack.com/p/involution-without-ex...
I'm not sure "subsidised" is the right word. If a government funds research and the results are released openly, that's just publicly funded research. It's how a lot of science works in the US and Europe too.
If your belief is accurate, we should expect China to short the IPOs of Anthropic and OpenAI and release better frontier models immediately after their IPOs.
Does anyone think that likely? I have no clue or bias.
And OpenAI and Anthropic are just in for the love of the game? Our of sheer desire to help humanity?
> The reason they release the AI models is economic warfare against US, not because of charity or kindness
There are many other reasons Chinese companies releasing models open-source or open-weight makes strategic sense.
A really easy-to-understand example is a company who has a near-monopoly on "serving video content" releasing a video model openly.
If you can be relatively certain that video content created by a model (which you have trained, using data from your own platform) will be ultimately served on your own platform, thus generating revenue from watch-hours, it makes sense to make those models as widely-available as possible.
It's also a net-positive if people use your public research to build better video models, because - again - you are reasonably certain that the even-better content those new models produce will be watched on your platform.
The alternative would making models harder to access and learn from (broadly, the current western model). Many would argue that Google, in choosing to not optimise its video generation models for "availability", is directly causing less content to be uploaded to YouTube. This is the trade-off.
I don't know much about DeepSeek's financing specifically, which obviously doesn't release video models - so I don't know how directly this analogy runs, or who directly benefits from the extremely evident rising tide that the public release of DeepSeek's research creates. However, this does not negate the broader rising-tide effect of the scientific method.
It's certainly also true that it's geopolitically beneficial to be able to undercut American labs' models. If I ran a global superpower, I would probably want my country to be technologically competitive too.
But Chinese companies are already serving a huge volume of customers in a complex, existing marketplace, before even thinking about the US market, and it's overly simplistic to assume that their entire strategy revolves around economic warfare directed specifically at the US. It's more nuanced than that.
This is, of course, without even getting into opening the can-of-worms around whether US economic policy also results in the US state functionally subsidising technological innovation, how comparable that is to China's model, etc.
[delayed]
Well, US models are economic warfare too, of course.
> DeepSeek, and other Chinese models, are heavily subsidised by the Chinese government.
Yes, because everything China does is against the US. That's all they think about day and night. God forbid they want to corner the global market or have a genuine business case. How dare they provide options for those who can't afford a measly $200 a month? How can we let Chinese labs publish research for free for the whole world so that they can benefit? The nerve! To think they can use soft power instead of military might! I mean, Anthropic and OpenAI are the last bastions of human kindness and charity. Right?
Right?
I like when bad intentions produce good results, I'm tired of seeing the opposite in practice.
Precisely. It's similar to their practice of aluminum dumping to depress US aluminum prices, which causes our aluminum mines and mills to close.
I’m neither in the camp of Chinese or the Americans(collective West) in general..
As a neutral party, this characterization is crazy.
As if the AI companies - Claude and OpenAI are guardians of freedom and humanity and very charitable to the global society without any self interests… “Chinese models are subsidized by the Chinese Government, therefore they’re inherently bad for humanity” is a highly propagandist argument. The politics of US vs China may be whatever it is in reality.. You have one company releasing their models for cheap, actually open sourcing their trained weights, and publishing details of their optimizations and learnings for others to use. The other camp actively “aligning” their models, nerfing their capabilities, hyping their swarm activities from poor sandboxes, and trying their best to lock users into their harnesses and walled platforms. They are subsidized by the capitalist VCs who are essentially waiting for their payouts..
At some point, one has to see things for what they are and evaluate their own reasoning..
I’m happy to stay provider agnostic, try all models and cheer any useful progress as open as possible.
Not sure why so many people will vehemently refuse this idea. I won’t say it’s 100% true but it would be foolish to dismiss it. China is very much an adversary to America and has made it pretty clear they want to be a dominant leader of not the new world leader. Not here to evaluate what is good or bad. Keep in mind historically China has aggressively fostered industry (not unlike the west) but sometimes even more aggressively.
Americans need to travel to China. The Chinese have zero issues with us. They quite like Americans. This is such a weird propagandist take. Idk if you remember but both country's leaders just had a slumber party for 3 days in DC. This is not what enemies do.
Not even our leaders say China is an enemy, ita mostly businessmen who are scared of competition and trying to regulate chinese out of their markets so they can make more money milking us.
Your just saying it’s beating america at its own game and crying fowl.
Economic warfare against the US is charity and kindness to a sizable portion of the world's population, especially when the US uses it's global hegemony as warfare against them. Neither system is perfect, but lets not be disingenuous.
I wouldn't even call it economic warfare. Even in the US, cheaper access to good-enough models helps everyone except for the billionaires who invest in frontier AI. So just from a numbers game effectively nobody is hurt by the open sharing of science.
Meanwhile, American AI models are heavily subsidized by stock market speculation. Ultimately, the subsidies from both countries are flowing out of the pockets of individuals.
US is basically doing economic warfare and bullying against everyone else ATM :)
What if economic warfare against the US does help humanity?
Why are we assuming a strong US is necessarily good? As a European, I have seen plenty of evidence against that stance lately.
I understand that Americans might prefer a strong US. But conflating them with humanity is a leap that I don't think one can make without any backing.
> are heavily subsidised by the Chinese government
We hear this about literally every industry the Chinese excel in - that it's only because the government subsidizes them that they succeed. For chip manufacturing, for batteries, for EVs, for solar, for AI. I don't see how the chinese government can afford to subsidize all of these industries and still have them contribute to the GDP.
A conspiracy to make the US look bad by being better at producing all the goods and services the world needs at a reasonable price. Have they no shame?
Not really this is an X algo conspiracy. Up until recently the Chinese government wasn't even that invested in these companies. We're talking very very small grants compared to training costs.
Its very xenophobic of you to say China has zero intention of helping humanity, and just wants to "wage economic warfare".
Last time I checked, it was ourselves (USA) waging economic warfare on 2/3rds of the world.
I dont get this cope people have where people have this idea that its impossible for a Chinese company (that make billions of dollars) to have done something by their own merit, but instead its always some Chinese Communist Party conspiracy where the main goal is to destroy America.
Lay off twitter for a bit.
>Up until recently the Chinese government wasn't even that invested in these companies
>Last time I checked, it was ourselves (USA) waging economic warefare on 2/3rds of the world.
Up until recently the USA Wasn't waging economic "warefare" on 2/3rds of the world
The USA has been using financial sanctions, aka, financial warfare on anyone it's deemed an enemy for going on 4 decades.
China has been owning and controlling key companies in its industry for going on 6 decades.
Every country does this. Do you think the US government doesnt fund, regulate and control key companies?
China is not a threat to you, or anyone in the West.
They're trying to pull digitally what they already pulled physically. The reason we can't manufacture a grill brush for a reasonable price is the result of years of Americans choosing the cheapest price. We gave up our manufacture base. They want us to give up our labs.
Deindustrializing a country is not something consumers can achieve. It starts at the top level, with politicians who construct a financial system where it's more profitable to speculate than to build or invest in real businesses.
Except they are open with the tech which is easy to replicate. All new models lowering cache prices is the result of DeepSeek's publications.
Also, the field moves fast, but slower than people do. Researchers and engineers switch companies every year or two, and the know-how walks out the door with them.
Architectural/algorithmic tweaks do advance the efficiency frontier nicely. But raw intelligence mostly comes from data (not just its sheer quantity, but also how it's curated & cleansed) and the scaling law. The know-how about data curation doesn't seem to get published much, even among the open-weight labs, though.
This. Even in the efficiency frontier, it is a lot of data curation that actually makes many of those tweaks actually work at scale in practice.
So boring to see conversations moved over to Chinese models when that’s not even what we’re talking about here. This is about Mistral.
There would be a lot of competition even without DeepSeek. Workers can freely exfiltrate trade secrets without noncompetes in California.
Proprietary competition, yes.
>instruction manuals in paper form
So the most common way to publish manuals?
In research paper form.
I think the mean 'paper' in the scientific journal meaning; these are unfortunately often extremely bad 'instruction manuals'.
Maybe 30 years ago
For sure people who don't grasp the difference between models, might be stuck in 'good enough' models.
But Opus 5.5/GPT is such a game changer in comparison to sooo many others, its still a moat for now.
You've got to think the comments about "good enough" are people who have not yet tried Opus 5.5. I haven't been this struck by a step change since 4.5/4.6. It's a bigger jump even than when Fable first arrived.
As for Mistral - I got really excited when they said Large 4 was focusing on being #1 in cybersecurity, because that's somewhere that they genuinely could edge out Anthropic & OpenAI. Have it actually solve problems, instead of Anthropic flagging "you tried to find a null pointer exception bug in your own code, we're now reporting you to the US government". But on the Mistral benchmarks I'm seeing, this looks very disappointing, but at least they haven't entirely given up. I genuinely thought Mistral had given up on new general models. They need to learn the bitter lesson all over again.
I feel like the "good enough" argument isn't about how big the gap between models is but about how good they are at solving the tasks at hand.
The capabilities of all models increasing so much all the time means there are simply less and less tasks you need a frontier model for.
Even if Opus 5.5 is 500x better than Deepseek, if deepseek can solve all my problems, why do I need to pay for more?
Many on HN still have the opinion that you must understand every line of code in the project, and that all is lost should you merge code that wasn't reviewed.
Obviously any model will do if you use it as a better autocomplete.
I believe that there is a large gap in expectations between different workflows.
Until the AI like reads my mind and produces perfectly production ready apps with minimal intervention from my side, there is still going to be room for improvement.
I use Opuse 5.5 daily for my job. I am aware (abd in awe of) it's capabilities.
Look at the context in which I used that term 'good enough'.
What i was saying is that there are tasks for which a dumber model can be good enough, and for organizations with sovereignty/ privacy concerns, those concerns can be strong enough to incentivize the use of a dumber model.
> I haven't been this struck by a step change since 4.5/4.6. It's a bigger jump even than when Fable first arrived.
I had the exact same experience. And unlike Fable, it doesn't gobble up your entire usage limit in a few hours.
I always wonder what the "good enough" people are actually using it for.
Please, give it another 6 months and they catch up. The American labs are currently trying everything they can to block others instead of advancing their models, trying to build an artificial moat. The American models are not that great, they are good, and they have a lot of agentic workflows in the back, but its basically a hardware limitation at this point. Once the HW makers catch up, and we can move away from the Nvidia monopoly, things will speed up quite a lot IMO.
Just one more release cycle bro, I swear
We've been hearing the line about them only being a few months behind for a year now, during which time O/A have grown their revenue like 10x, haven't they?
We've also been hearing we're 6 months from AGI for about three years, and here we are.
"Now, here, you see, it takes all the running you can do, to keep in the same place. If you want to get somewhere else, you must run at least twice as fast as that!"
Phantom Tollbooth?
Those are two different things. The market is expanding, so even if competitors are catching up, you can have your own revenue, in absolute terms, grow.
The thing is O/A have been much louder on pacing the frontier, lately.
And yes, open weights are still behind, but are catching up.
I agree that Opus and GPT are almsot surely better, but so many real users are nervous enough about giving Anthropic and OpenAI access to all of their internal information that they may be willing to stomach worse models if it gives them more security.
The real question is if this model is good enough that it can still accelerate work, and not be a hindrance to real work like older Mistral models often were.
If they can do that, they'll have customers.
No it is not. Only maybe for the noobs or vibe coders.
People who aren't afraid of rolling their sleeves into any code base? The difference is practically zero.
I agree; yes, I can see that they need a bit less hand holding each cycle, but I also see these "frontier" agents do some absolutely dumb shit that I have to correct and then I'm wondering if I'm the looney one here.
Maybe it's because people stopped watching what their agents are doing and stopped looking at the quality of the output. But I still see agents being absolutely mindless like a junior dev.
Recent example: it updated an an API to add newly released models to the backend. There's a list of models that require specific configuration for the reasoning effort and temperature or the API call fails. GPT 6.1 Sol misses this and code fails at runtime because the newer models need to be added to the list for special handling of temp and reasoning. Fixes it for one model and tests it for that model using an E2E test. But doesn't test the other models that were added for the same error condition...I had to explicitly ask it to do so and it finds them and adds them to the list and says "that's on me."
Yeah, not that smart.
I use Opus 5.5 at work.
I use MiMov2.6Pro, DeepSeekv4.1Flash, GLM5.3, Hy4, Qwen3.8 and KimiK3 at home. Opus5.5 is not a game changer.
It's an improvement, but game changer might be a bit of a stretch. If I lost access to Anthropic or OpenAI models tomorrow, I would be annoyed, but would reach for a slightly inferior model. Last year I wouldnt be able to say the same, and rhe challenge is that the moat is drying up fast. Whether its general improvements in model training by other competitors, or straight up distillation of SOTA models, the moat is shrinking and the available capital and spend for American model providers is going to dry up quickly as competing good enough models are adopted by more consumers.
It's especially the case as more non-Americans look to self hosted models and domestic cloud inference providers using open models that the US providers who are still leading the charge need to drastically drop their prices and find a path to profitability in order to maintain their lead and retain the advantage they had as AI turns into a commodity (which is happening faster than I think even the frontier labs initially predicted).
People say this exact thing every single time a new frontier model comes out.
My todo app generator does not need opus 5.5
> X is such a game changer
I hear this literally every other week about whatever the newest FoTM model is.
Unless you can provide concrete examples of things you can do with them that you simply couldn't do with last week's model, it's absolutely meaningless.
Less and less work requires a frontier model though.
Yeah, I agree with this. I think the "the models are good enough" narrative is a myth. I've heard it so many times over the last year, but the model number keeps changing...
There is no ceiling on what you can accomplish with more intelligence, so there will always be a market for the best models, and that market is likely to just keep growing. If Opus 13.5 can one-shot a profitable company or discover a new disease treatment or whatever you can think of that a swarm of relentless super-geniuses could accomplish, companies (and governments) will throw money at it.
I also think there will always be a market for many sub-frontier models that will continue to grow rapidly as well, because "good enough" is definitely a thing for a given task.
>If Opus 13.5 can one-shot a profitable company
No company would ever release such a thing
Despite Opus 5.5 got really bad the last days for me. Looks like they nerfed it again. This is extremely unreliable.
Or maybe they secretly believe you are trying to distill their models and are deliberately degrading your experience. Who knows with them?
You are laughing. Until it happens to you! :-)
> have not been a winner-take-all runaway acceleration game where catchup is impossible
From the Mistral site:
> ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe.
It is pretty capital intensive!
According to Grok thats 7-10 MW. Tiny numbers.
To put that into context, the last wave of capacity SpaceXAI added 400-450 MW.
Yes, so far the competitive dynamics feel more like cloud computing than web search.
Seems silly not to have predicted that 10 years ago. I feel like it's long been obvious that smarter models being available will mean way easier cheap synthetic data and access to tools that will speed up competitors as well as consumers.
How do you figure? I haven't met a single person who doesn't use Claude or Codex for programming in any serious way.
Then you dont know people working on highly sensitive info with stringent privancy concerns.
I mean, Mistral is about 9-12 months behind here when you look at its overall benchmarks versus the models released around a year ago.
On a purely technical level, maybe? But in terms of actual revenue, is there really any chance of anyone catching the big labs?
Obviously, this is only a valid question if you don't believe that open weights are about to eat their lunch and their revenue is about to collapse, or they're running a super unprofitable ponzi scheme propped up by investor money that's about to collapse like a house of cards. I don't find those positions credible at all though.
If you do, then this question isn't really for you, as I'm more interested in thoughts from those who think that OpenAI and Anthropic in particular are about to be the largest companies on earth in a couple years. Could anyone catch them at that point?
> But in terms of actual revenue, is there really any chance of anyone catching the big labs?
I don't know about revenue, but I suspect multiple other labs are already beating OpenAI/Anthropic on profitability. Staying on the frontier is expensive, and it's hard to recoup those R&D costs when you have a bunch of other labs nipping at your heels.
If you concede the previous point, then the only way for OpenAI/Anthropic to keep growing long term is to swallow the whole economy (i.e. mass job replacemnt), and that's a bet I wouldn't take.
I think the actual plan is to swallow a good portion of the job market. It’s the only thing that makes sense and I hear VC podcast debates on which percentage of jobs justifies the market cap.
Maybe. I can't freaking wait for the IPO filings so we can finally put all this to rest. (haha, like that'll actually put it to rest on HN, but at least we'll have better data)
They are very unprofitable…? I don’t think that’s really in dispute. We haven’t yet seen a profitable frontier lab and model pricing remains fairly subsidized
The big lab revenue may not be catchable, but im not sure it needs to be.
If they can carve out a niche of industrial and governmental partners who rely on them for sovereignty reasons, it may be enough.
I completely agree, I think AI is a vast ecosystem will all kinds of profitable niches and sub-markets.
No, because compute, not model ability, is the moat.
The second moat is convenience, which all the big labs make it (comparatively) easy to glide into their models.
I don’t know if convenient is a moat when it makes switching very easy
Impressive vision benchmarking. If the vision model is truly as good as astra, that would make it best in the world.
Also strong on cyber benchmarks (better than all chinese models), so this is a good defender model.
Lots of people shitting of Mistral for no reason imo. These are pretty good numbers across the board. Definitely good enough to use as a daily driver over other llms, if you have moral qualms with the others. For certain use cases, like cyber security, this may be the go to model.
I like to make fun of europe, but there's lots for mistral to be proud about in this release imo.
Curious, what do you like to make fun of Europe about?
> Curious, what do you like to make fun of Europe about?
Well I'm european and... That Switzerland (Europe but not EU) has more companies in the Top 70 by market cap than the entire EU (Switzerland has two, the EU only has ASML) is kinda something that warrants making fun of.
That the biggest European software company is SAP, in 71th position is both sad and tragic: it shows how lame and irrelevant Europe is when it comes to software.
So Europe is nowhere in software and friggin nowhere in hardware: sure it's got ASML but ASML now has officially... Zero customer in Europe. Zero is not much.
Then Japan is at least trying to come back into the game with nano imprint litography. Europe is betting it all on AMSL (which anyway is majoritarily US-owned).
So software: nothing. Hardware: nothing besides ASML.
Overall the EU has six companies in the Top 100 by market cap and they're all, besides ASML, near the bottom of the Top 100.
We could also maybe make a bit fun of how the EU destroyed it's car industry (the main industry in Germany, which is the biggest economy of the union) by handing it all to chinese EVs?
Or what about the US warning the EU, years ago, to not become entirely dependent on Russia for energy? And EU not listening and then seeing its energy price skyrocket when the proverbial shit hit the fan? (Russia attacking Ukraine)
And we could, also, at least make a bit of fun of entire streets in cities like Paris and Brussels that used to have luxury shops and fancy restaurants that are all turned into places selling cheap kebabs? What a great success: I'm sure this one makes the komrades happy. It projects an image of grandeur and success: kebabs.
Or the constant attacks on free speech in the EU. Or the surveillance apparatus that's being put into place.
At this point it's more like I don't know what is there left to not make fun of about my EU.
For what's going on is just sad, plain sad.
Bro there's more to life than software and hardware. Try to get outside today :)
If the UK is in europe, you should make fun of the following: they don't have a first amendment. So to me, it's an authoritarian state preaching freedom.
They also don't have a second amendment so it balances out
Doesn't that make it even less free?
Ask the school kids in America how free they feel
Europe has spent the last twenty years in stagnation - generating about half as much wealth and technology as you'd expect for its size and advanced economy. Simultaneously, Europeans are notoriously arrogant. It's a mockable combination.
Edit: unfortunately, it's a question that voting does not really permit you to answer on this website.
I'm biased (I'm European), but I'd much rather be middle class in Europe than in USA.
The US is biased to favor the top half of society. Hence why there is a lot of vocal poorer people and a lot of low profile wealthier people (I'm not including the 1% or even the 5% in this).
When you are in the 75%ish of the US, it's very easy to make a case that life in the US is better. But we don't really talk about that because it's pretty taboo when poorer people are struggling much more than they would in Europe.
Poor people in America (25th percentile) have more money that middle class people (50th percentile) in Europe. This wasn't true 20 years ago.
I'd much rather the people of every continent have the combined technology advancements of three major powers vs two.
Do you realize the variety of breakfast cereals you are giving up?
Also the range of things with added sugar. I never imagined sauerkraut (you know, sour cabbage) could have sugar added to it.
Arrogant? Actually we are not the ones running around and claiming we are the best country in the world :)
Ironically, the US is always busy with growth. How about celebrating a victory for once?
Take a look at NLNet vs YC. At NLNet, You set your milestones, do the work and get rewarded.
YC just throws money in the hopes one company is a unicorn. Both support growth, but they're not comparable at all.
Yet Europeans are healthier, happier, and have a better quality of life than Americans. They also have so much better transport infrastructure it's embarrassing for Americans. Like all the money America generates, where does it go?
Maybe you're just bitter?
If you remove tech companies, US and A has spent the last twenty years in stagnation, too. If that bubble bursts, both are on par.
If it bears fruits and we build ``it'', everyone dies, which is par, also?
A 'bubble bursting', if that happens, doesn't make the tech sector go to zero. Apple doesn't suddenly stop making iPhones. It would take a hell of a lot more than that for the US economy to fall to European levels.
Gemini says (20 years growth): - USA: 51% total, 27% without tech. - EU: 25% total, 21% without tech. - China: 345% total, 205% without tech.
Ah yes... the "European". From the Scandinavian viking to the Sicilian - one homogenous group that agrees on everything and acts the same.
Well, the Scandinavian and Sicilian do share one currency and one immigration policy and one set of regulations - which are the things that we generally look at when we analyze economies.
> Europeans are notoriously arrogant. It's a mockable combination.
Not a very rigorous economic argument - besides they also don't share an immigration policy, nor a single currency (Denmark)
Norwegians, Danes, Swedes and Finns all use a different currency. Finns and Sicilians do share a currency though, but all these countries have different immigration policies and regulations. I don't think you really know how the EU works.
Not currency! Only Nordic on euro is Finland
read Varoufakis for the story on how this happened. European surplus capital gets recycled as VC money into Silicon Valley. So it's not like there's some big choice to be made, its potentially a systemic part of the global monetary flows.
The blog post https://mistral.ai/news/mistral-large-4/
Surprisingly it only supports reasoning "none" or reasoning "high".
That setting didn't seem to make any real difference - it added a tiny bit of thinking trace and high actually produced less output tokens than none.
The high bicycle frame is better then the none one though.
Pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
(Definitely the best I've seen from any Mistral model: https://simonwillison.net/tags/pelican-riding-a-bicycle+mist... )
It's a bit below DeepSeek 4.1 Flash at about twice the size. For a model that was supposed to come out a few months ago this is pretty good. Mistral catching up to the chinese open-weight models is great news. Excited to see how they will build on that!
A strong competitor in cybersecurity as an alternative to GLM-5.3 (Mistral reports 82% on CyberGym-E2E). Visual grounding is also impressive (42% on Dense 200 vs. 41% for GPT-6 Astra).
Otherwise, behind on the broader Pareto frontier, but not by much (Vals Index: 48.05% vs. GLM-5.3’s 53.51%; $13.78 vs. $7.25 per test). Many companies will prefer it over Chinease models.
That was about my conclusion as well: Slightly less than GLM 5.3 performance but made in Europe. So, maybe it answers Tiananmen Square questions correctly, and in French. All in all, a reasonable model, but not frontier.
> So, maybe it answers Tiananmen Square questions correctly
What would you consider a "correct" answer? I just asked deepseek-v4.1-flash asking it what happened (without mentioning the word "protest"; here are some excerpts of what it said:
> In April 1989, students in Beijing began demonstrations after the death of Hu Yaobang, a former Communist Party general secretary. The protests grew. [..] Estimates from other sources range from hundreds to several thousand deaths. [..] The Chinese government describes the events as a counter-revolutionary riot and says the military action was necessary to restore stability. It restricts public discussion of the events inside China. Many other governments, human rights organizations, and observers describe the events as a violent suppression of peaceful protests.
So, let's see... it calls it a "protest", mentions the number of deaths, and even mentions the censorship of the topic by the CCP.
"..answers Tiananmen square questions correctly.."- but lies about Ukraine, EU, and about you, americans..
I live in EU, use for my personal needs chinese models only, and don't plan to move to any of the ones allied with the Pentagon or its european counterparts.
What are you talking about lol
Exactly what you read
Have a prompt and excerpt of falsehood in response for each?
Refreshing to see this.
The pricing ($.68 in/$.07 cached/$2.09 out) makes it much cheaper than Kimi K3, GLM 5.3, and Meta Muse Spark 1.3. That's great!
But also much more expensive than GLM 5.3-flash and Spark 1.3 Contributor (the Meta-takes-your-data pricing of Spark 1.3).
So, I think it would have to be significantly better than GLM 5.3-flash to be worth it. GLM 5.3-flash is already very good.
You are quoting the discount pricing. It is 50% off for the next two weeks. The blog post has the real pricing up front in the card on the right side: https://mistral.ai/news/mistral-large-4/
After that, it will be much closer to GLM 5.3, but you can also get 5.3 in their API! I dont see people really talking about that.
I still need to evaluate it for my own workloads, but if you trust the benchmarks, seems about on par in quality vs GLM 5.3 Flash
Source https://artificialanalysis.ai/models/mistral-large-4?total-c...
Europe needs a lot of these. Quickly. Way to go, Mistral! Keep 'em coming.
I’d rather have one European AI lab like Mistral, with the financial firepower and compute to pretrain its own large models, than 20 weak labs that fine-tune Chinese models.
AI needs big bucks, and Mistral is Europe’s Anthropic.
Two is enough for foundation models, these guys and Aleph Alpha. The gap is closing.
You have to thank the US investors who funded Mistral from the very beginning.
Mistral would have gotten a tiny and measly "EU grant" and ASML would never have invested later had it not been for the US VCs.
> Europe needs a lot of these.
Europe needs profitable AI companies, not money pits.
The US has made sure that Mistral has a large market in the EU by temporarily preventing non citizens from accessing Fable.
A lot of European companies now want a model the US can't cut off, but also lack trust in Chinese models.
Some of these will self host Mistral but most will pay them by the token. It's not going to be a huge market or a huge margin within that market but probably it'll be enough.
This is the exact mentality that makes the EU fall behind. If you don't want to invest in something until it makes a profit, you don't get the benefits of being a pioneer.
Look at the size of our capital markets. That is a suicidal strategy. We dont have to mimic US.
there are no profitable AI companies at the moment ... This is the exact problem of European startups, trying to make them profitable from day 1 while American counterparts (and Chinese) keep bleeding money for years. Europe will never have a Tesla, a Google or an Amazon with that mindset.
Pithy, but only really applicable to Amazon. I'll give you some tips for future attacks.
I'm not actually sure Tesla is a great company from any perspective so should be easy to find another sick burn for them.
Google is going to be a little more difficult. Maybe say something nasty about advertising? Or go for the monopoly angle.
Name 1 AI company that is profitable, and no Meta and Google are not AI companies
Like the US!
…oh…wait…
-5 on omniscience? https://artificialanalysis.ai/evaluations/omniscience
That's not particularly great.
That said I love that they don't seem to restrict cyber capabilities to any degree and even lean into it.
If the best model for cyber attacks is open for everyone to use it just makes us all safer I think. Of course you then also HAVE to use it or otherwise you're vulnerable, which is a great distribution play.
https://artificialanalysis.ai/models/mistral-large-4 for the main stats
Well, with this and Beam people are going to have to stop saying that western open models are dead. This is great news. Anthropic and OpenAI may have a bit more knowledge, talent and compute but they don't have a monopoly.
Investors looking for them to make monopoly profits are going to be disappointed The premium they can extract from consumers for their models will be capped. Tokens are likely to remain close to the cost of compute, a cost which is high but falling fast.
Also, yay Europe!
I tried Mistral's last 3 large models, devstral and mistrallarge3 and the numbers were not even benchmaxxed, but just false. the models were so weak and garbage. Let's hope they are telling the truth this time around, we need more alternatives. There mistral-small and original MistralLarge models were awesome, hoping they are back!
Finally, a model small enough to self-host on my 2012 MacBook Air if I don't mind my house reaching room temperature in 2 seconds.
Seems hard to get customers at that price range when you're competing with open source models that are 1/2 - 1/3 the price but with similar capabilities.
People usually buy the cheapest, like Deepseek or GLM or they spend on Anthropic/OpenAI subs.. Are these models in the middle getting any users?
On a side note, I wonder if this was the popular free Space Bunny model that left openrouter yesterday.
So about 2 or 3 generations behind, just like they were a year ago?
Nice! Once they make it available through their API I will be happy to support them. My local Qwen3.8 27B is serving me well, but I miss the speed and concurrency that comes with subscriptions, and I am not currently paying for any.
Tais-toi et prends mon argent!
I believe it is already available in the API no?
Direct through them?
Yes: https://docs.mistral.ai/models/mistral-large-4-0
Yes and OpenRouter
Le Chaton Fat is here!
Le Chonk https://www.youtube.com/watch?v=hD51W2txi1Y
Yeah just saw that, I'm gonna keep converting it in my head.
Good to see Europe is at least a little bit still in the game.
https://docs.mistral.ai/inference/model-selection-guide?mode...
Cost is stated at half the price of GLM-5.3, which is quite interesting.
What experience have people had with Mistral models for cyber research? are they as constrainted as Anthropic models? I use claude day2day but have the need for a model with fewer guardrails to pentest our own APIs.
I tried a quick abbreviated kimbench on their playground before bothering to do the whole thing.
Maybe I didn't really select mistral 4? Either way, failed completely on question 1 and the next 2 questions were completely off base too. I didn't bother to finish.
Not suitable for my purposes I don't think.
49b active parameters sounds manageable until you look at the 1t total weights. what hardware does a usable self-hosted setup actually need, especially once you add a long context?
Seems like they've finally made a genuinely competitive model since the original LLM craze, congrats to them! Glad to see some diversification in open weights providers.
Mistral Large 4: 1050B, 49 Active
GLM-5.3: 753B, 40 Active
I was hoping for something that hinted at smaller models too, but I guess not.
Any competition is still good, especially now that the USA AI labs are starting to do regulatory capture.
It's competitive!
Good enough to show competence, and instill confidence in the team/company. Later releases can be more efficient.
I think it's a great release with that framing.
just keep RL frying it should get better...
Off Topic - The Mistral website - Really nice design. My guess, built by a human.
Claims to be on par with GLM 5.3 in DeepSWE (from https://thenextweb.com/news/mistral-releases-large-4-a-1-tri...)
Excited to test this on my benchmark[1], but I'm guessing it won't fare better than MiMo v2.6 Pro which is currently the best open-weight model as far as I'm concerned.
Why isn't this on openrouter yet? Is there a better router that gets models much quicker?
1 - https://bench.killswitch-lang.org/
https://openrouter.ai/mistralai/mistral-large-4-0
Benchmarks are better than expected! And probably got there without distillation ;)
Is there a reason to believe why they wouldn't distill locally running open weights Chinese models?
I thought lechonk motto was just a meme!
The terminalbench 4.0 score is encouraging as a sign of it not doing anything "stupid" when put in a proper harness.
We're probably fast approaching the scenario where the cheapest models will win out.
Excited to hear this!
I barely use anything outside of cheap Chinese models on OpenRouter anymore. They are simply (more than) good enough for most of the things I do.
This model looks reasonably cheap. Though not deepseek levels.
Going to test it with Hermes, wondering where it will land in term of capability.
Bon chance, Mistral!
https://venturebeat.com/technology/mistral-debuts-large-4-le...
if it's not available yet why have a 'try it today' header at all?
> "Try it today" > > There is still more to come. As we work toward releasing the weights, we will share further details on the model architecture, additional benchmarks, and our post-training methodology.
You can try it on their website, https://console.mistral.ai/playground
It's just that the open-weights aren't yet available (although the long delay is slightly annoying).
Looks like they are doing 50% off to stay price competitive with DS Flash V4.1
This sure is nice, but I've had less than satisfactory results with GLM5.3. I'd like Mistral to compete with Qwen3.8-Flash-Next a 120B class model that IMO is the first model that I can use for serious coding while running it locally.
I estimate it's coding ability on par with opus 4.6 (but opus definitely beats it on factual knowledge) Still it's a genuinely useful model, when everything else except Anthropic's models (and for only 3 weeks after it came out Google's Gemini 3 pro, before it got merged) are not to me.
I'd live to have one like that but EU made.
Glad to see progress, despite the ever-increasing sabotage by the EU bureaucrats
le chaton fat is real, my life is complete. Benches look crazy good for 1T.
sorting the charts like that gives off weird vibes
https://mistral.ai/news/mistral-large-4/
sorting the chart like what? You just linked to the main page.
If you scroll down, most of the charts on that page are sorted s.t. Mistral's bar is right next to the worst competitor model, while the best competitor model's bar is positioned on the opposite side.
If one were to be cynical one could say that it's intentionally making Mistral's result look better than it actually is by making it harder to compare the bar heights.
europe finally getting into the race here.
Previous one is barely in top 50 on arena.ai
where does sit on the pareto distribution compered to Le Chaton Fat?
Since the Chinese companies publish their research it would have been odd if Mistral didn't start catching up.
It certainly has helped OpenAI and Anthropic get their KV cache costs under control.
It's no secret that everyone is dis-stealing from everyone else.
I don't see how distillation relates to using the published techniques developed by Deepseek, Moonshot, Zhipu, etc
touche
I mean no disrespect but these are terrible numbers or am I missing something? It seems like Mistral continues to only be relevant for people that want a model trained in Europe. Too bad.
On Prem. Thats a bid deal for some enterprises.
also the benchmarks are not necessarily indicative of how well the model will perform in its own harness with its own skills.
Any open weights model is "on prem".
Second point
Anyone have any indication when I can get my hands on a developer plan for this?
>> Unofficially ML4, very officially: le Chonk
Honestly just nice to see a leader in this space not take themselves so seriously.
wowza le models a heckin chonker
Looks like it's about a year behind still. i.e. its intelligence is behind models from roughly a year ago.
https://www.vals.ai/benchmarks/vals_index
Woah, this seems like a big deal (assuming the benchmarks are as good as claimed)?
Mistral slightly proving me wrong (and I'm not mad).
Can we consolidate the posts? Currently there's 3 on the front page, basically all pointing to Mistral's messaging in different places.
Does Mistral ever advance the state of the art on any dimension?
And if not, why do they exist?
Update: The number of people advocating not innovating is wild. There is no reason why Mistral cannot innovate in ML, they explicitly choose not to. My point is that, given that choice, they should spend their GPU hours differently.
"Sovereign AI" is a joke, there is no substantive difference between a post-trained open weight model from an American or Chinese company and what Mistral is doing today, beyond spending 80% of their GPU hours reproducing a last-gen model's pretraining.
Why would the Europeans want a high-quality open source model that isn't either owned by American Trillion-dollar companies or created by the Chinese? You really can't think of any reasons?
> high-quality open source model
> that isn't owned
I don't know what definition of open weights you are using, but the "open" part is what allows you to use it without being beholden to American Trillion-dollar companies or the Chinese.
Mistral could do exactly what Cursor does, and post-train the model to meet the needs of Europeans.
That's a far more efficient (and useful!) use of GPUs than yet another pretrain on an outdated model backbone.
Has the French military advanced the state of the art on any dimension in the last hundred years?
And if not, why do they exist?
Second hand rifle distributor?
Of course they have!
Regardless though, their reason to exist is not to advance the state of the art.
Their reason to exist is to ensure the sovereignty of France.
Obviously they would do their job better if they were advancing the state of the art, but it's not like it's pointless if they arent the absolute best.
So that there exists an EU-native option in the near-frontier LLM space?
Not everyone is wild about being downstream of either the Chinese or US governments, particularly when it comes to things like cybersecurity
Mistral is one of the few European AI labs. Look up "sovereign AI".
Just like Microsoft and most companies around PCs, their existence isn't about pushing any boundaries, but seemingly targeting deploying what's already been "invented" into the corporate world and bigger companies. Not "wrong", just different.
Sure, agreed. But they do their own pre-training, at great expense, on outdated model backbones.
Wouldn't it be better to do something more like Cursor, and RL on an existing pretrained model if you're not innovating anyway?
Why build cars when you can just change the seat covers
Does erichocean ever advance the state of the art on any dimension?
And if not, why do they exist?
Yes, actually. Thanks for asking.
Not bad only two major releases behind top tier. Edit : checked its rather 3 generations behind . Oh well
Is there consensus on if this was https://openrouter.ai/stealth/space-bunny-alpha ?
Space Bunny Alpha is probably MiniMax M3.1 (rumors on Twitter since it seems to have a similar tokenizer).
Yes broad consensus had developed in the ten minutes between announcement and you asking if consensus had developed, and I'm excited to report that it consensed in the affirmative -- it IS Space Bunny Alpha!