Four Months Ahead for Five Times the Price: Mozilla Does the Math
Mozilla's open-AI report lands this September 15: just 4.4 months now separate the best open-weight models from frontier proprietary ones. Paying the premium only makes sense on a narrow band of tasks.

The report landed on September 15, 2026, and Mozilla handed it to Ars Technica ahead of publication. It’s the second edition of its State of Open Source AI — the first came out on July 14. The whole thing boils down to one number: 4.4 months. That’s the gap between frontier proprietary models and the best open-weight models coming out of China. Four months and change, in a market that keeps selling you a generational technology chasm.
Three points apart, 30 percent of the bill
The number that stings sits on the Artificial Analysis Intelligence Index: Kimi K3, from Moonshot AI, finishes three points behind Anthropic’s Fable 5, at 30 percent of the price. Mozilla’s recommendation is blunt: most organizations should default to open for the bulk of their work.
One distinction matters here. “Open weights” is not open source. You download the model’s main components and run them on your own hardware, but the training data, the data pipeline and the training code all stay locked up at the vendor. You get the record; the master tapes stay in the vault.
The narrow band between eight and twelve hours
To size the gap another way, the report leans on the time horizon defined by the research outfit METR: the length of task — measured by how long a human expert spends on it — that a model completes with a “reliable” 50 percent success rate. That horizon has been doubling at a pace that keeps picking up.
On that measure, the best closed model handles a task 1.7 times longer than the best open one. Mozilla CTO Raffi Krikorian puts it in hours: where open chews through seven hours of work, closed manages twelve — and in four months, open will be doing the twelve while closed is up around twenty.
Which gives you the map you actually care about. Under eight hours, both camps get the job done, so you may as well hand it to the cheaper one. Between eight and twelve, closed clears the bar and open doesn’t, yet. Past twelve, nobody clears it reliably. The proprietary premium, in other words, is a four-hour-wide strip of road.
The harness that inflates the scores
Then there’s the bias marketing departments adore: closed models often show up with their own harness — the software layer that hands them tools and memory so they can act as agents. A harness built in-house by the lab that trained the model can lift its numbers noticeably, while the very same model limps along on somebody else’s.
So the benchmarking firm Vals AI put everyone on the same neutral harness. On Terminal-Bench 2.1, GLM 5.2 from China’s Z.ai (Zhipu AI) lands one point behind Anthropic’s Claude Opus 4.7 and 4.8, at roughly five times less per completed task.
What the invoice still buys
Paying still makes sense, and the report says exactly where: expert professional work, heavy document research, long context. Krikorian frames the choice as a function of the workload, not the organization — the right question isn’t “which model for my company,” it’s “which model for this job.” Closed models also ship the things you can’t train into a model: a service that works without configuration, a compliance wrapper, support, and a named party to hold responsible. Plenty of organizations simply don’t have the staff to run an open model properly.
Some have already made the call task by task: delivery company DoorDash hands its routine work to Kimi and keeps Fable for the jobs that would take a human expert a long time.
Five times the price for a four-month head start is one steep delivery fee. Especially when the head start itself expires in four months.
Sources (1)
Written with AI assistance from the sources cited above, then reviewed and approved before publication by Sébastien Soulier.


