What you need to know

  • Qwen leads every distribution metric in the report. 2,045M downloads in 2026 across repositories with declared parameter counts; 39.6M GGUF pulls a month against Gemma's 20.8M and Llama's 7.5M; 151,448 derivative models on the Hub against Google's 82,506 and Meta's roughly 58,000.
  • The licence split is the decision-relevant fact. Of Chinese releases above 20B parameters, 59% are Apache 2.0 and 22% are MIT. Of American labs' releases, 29% are Apache or MIT, 41% carry custom terms and 30% declare no licence at all.
  • Nobody downloads the frontier. Models under 1B parameters account for 83% of all-time downloads. Models above 100B account for 1%. In 2026 alone, only 3% of download volume came from anything above 70B.
  • The parameter ceiling gap is real but nearly irrelevant to usage. Chinese frontier releases ranged from 754B to 2.78 trillion parameters through 2026; American models stayed under 130B in five of seven months, with NVIDIA's Nemotron 3 Ultra at 561B the notable exception.

The obvious story this week is that NVIDIA now owns the distribution layer for open weights. We have covered the acquisition, the terms and what compute-agnosticism pledges are actually worth separately, and that piece is where to go for the deal mechanics. This one is about a different question, and a more awkward one: what is actually on the hub that changed hands. Because Hugging Face published a detailed census of its own ecosystem shortly before the deal, and the composition it describes is not the one most Western procurement processes are quietly assuming.

The distribution picture

Start with the raw pull counts, because they are the least ambiguous thing in the report. Across repositories with declared parameter counts, Qwen recorded 2,045M downloads in 2026. That is not a narrow win on a niche metric — it repeats across every cut the report offers.

GGUF pulls, the quantised format that llama.cpp, Ollama and LM Studio consume, tell the same story in a different register. Qwen sits at 39.6M downloads a month. Gemma is at 20.8M. Llama is at 7.5M. That puts Qwen at roughly twice Gemma and around five times Llama in the format most closely associated with someone actually running a model on their own machine rather than evaluating it in a notebook.

Derivative models are the third and, we would argue, most informative measure. A download is one person's curiosity. A derivative — a fine-tune, a quantisation, a merge, a distillation — is somebody committing effort, publishing it and inviting maintenance. On that count Qwen has 151,448 derivatives on the Hub, against Google's 82,506 and Meta's approximately 58,000. Qwen's total is about 2.6 times Meta's overall and roughly 4.7 times Llama's specifically.

Family 2026 downloads (declared-param repos) GGUF pulls per month Derivatives on the Hub
Qwen (Alibaba) 2,045M 39.6M 151,448
Google (Gemma family) Not broken out in this cut 20.8M 82,506
Meta (Llama family) Not broken out in this cut 7.5M ~58,000

One number needs handling carefully. Alibaba has said publicly that the Qwen family passed 3 billion downloads in six months. That is the company's own figure and it uses a different counting basis from Hugging Face's 2,045M — different window, different scope, different filter on what counts as a repository. Both can be accurate at once. If you are putting either number in a deck, name the source and do not average them. The discrepancy is a reminder that download counts are marketing artefacts as much as measurements, and that the party doing the counting matters.

Watch out

Every figure in this article measures distribution, not deployment. Nothing here says a single one of those 2,045M pulls ended up serving production traffic, and nothing here says anything about model quality. Treat these as an ecosystem census, not a leaderboard.

The size picture, and why it contradicts itself

Now the part that makes the league table interesting. The report states plainly: "In almost every month of 2026, the largest and most performant open model from a Chinese lab was larger than any model an American lab released." The parameter ceilings back it up. Chinese frontier open releases ranged from 754B to 2.78 trillion parameters across the year. American open releases stayed under 130B in five of seven months, with NVIDIA's own Nemotron 3 Ultra at 561B the notable exception to its side of the ledger.

Hold that next to the download distribution and the tension is immediate. Models under 1B parameters account for 83% of all-time downloads on the Hub. Models above 100B account for 1%. Narrowing to 2026 alone, only 3% of download volume came from models above 70B. Meta, the lab whose brand is most tightly bound to large open models, saw just 9% of its 2026 downloads land above the 70B line.

So the entire public argument — who has the biggest open model, whether a 2.78T open release changes the balance of the field — is being conducted about roughly 1% of what people actually pull. The real open-weight economy is small models, quantised, running close to the product. That is a 4B model doing intent classification on a WhatsApp support flow in Bengaluru. That is a 3B model doing PII redaction inside a Bristol clinic's network because the data cannot leave the building. Neither of those teams cares whether the ceiling this month was 754B or 2.78T.

This should reframe what an open-weight strategy means for a small team. The strategic question is not "can we run the frontier open model" — you almost certainly cannot afford to, and our break-even analysis for self-hosting a large open model on 8×H100 lays out why the arithmetic rarely closes. The question is which sub-30B model, under which licence, with which quantisation ecosystem, you can commit to for the next eighteen months. Our comparison of the 27B-class small open-weight models is a better starting point for that decision than any frontier announcement, and for teams working in Indian languages the emergence of models like Gnani's 30B trained across eleven Indian languages shows the interesting work is happening in exactly that size band.

The licence picture

This is the section that should make it into your next architecture review, because it inverts an assumption that a lot of legal teams are carrying without having checked it.

Among Chinese releases above 20B parameters, 59% carry Apache 2.0 and 22% carry MIT. The report notes that almost none carry non-commercial restrictions, with recent exceptions for Kimi K3 and for Qwen 3.8 at 2.4T — a release whose bespoke terms we unpicked when the weights landed. Among American labs' releases, 29% are Apache or MIT, 41% carry custom terms, and 30% declare no licence at all.

Licence category Chinese releases above 20B params American labs' releases
Apache 2.0 59% Reported only as a combined figure
MIT 22% Reported only as a combined figure
Apache 2.0 or MIT, combined 81% 29%
Custom or bespoke terms Rare; Kimi K3 and Qwen 3.8 (2.4T) noted as recent exceptions 41%
No licence declared Not reported separately 30%

Read that bottom row again. Thirty per cent of American labs' releases in this cut arrive with no declared licence. That is not a permissive default and it is not a restrictive one — it is an absence, and an absence is the worst thing you can inherit. Under both Indian and UK copyright practice, "no licence stated" does not mean "do what you like"; it means the rights position is unresolved and the burden of resolving it sits with you. A custom licence at least tells you what you are arguing about. An undeclared one leaves you with nothing to point at when a customer's procurement team asks the question in eighteen months' time.

None of this is a claim that Chinese models are better. The report measures distribution and licensing, not capability, and we are deliberately not extending it into a quality argument — for that, the independent index coverage of GLM-5.2 is a more honest reference. But "open" has been functioning as a procurement category, a shorthand that means safe to build on, and on this evidence the shorthand has drifted away from where the assumptions in most legal reviews sit. If your organisation's mental model is that American open releases are the permissive ones and Chinese releases are the ones needing scrutiny, the data does not support it in either direction.

Running open weights in production? Put it on your Builder profile.

AI Tech Connect lists AI engineers, founders and researchers across India and the UK — and the people hiring browse it to find them. Adding your profile is free.

Become a Verified Builder →

How to run a licence review that actually holds

Three practices separate a licence review that survives diligence from one that does not.

Read the LICENSE file at the revision you pulled. The licence tag rendered in the Hub's interface is repository metadata. It is set by the publisher, it can be edited after the fact, and it does not necessarily correspond to the text of the LICENSE file sitting in the commit whose weights you downloaded. Pin the revision hash in your build, fetch the LICENSE file from that exact revision, and store it alongside your model artefacts. If your model card and your LICENSE file disagree, the file is the document a court reads. The same discipline applies to provenance more broadly — our guide to verifying model provenance and catching cutoff drift covers the mechanics.

Treat undeclared as a blocker, not a default. Put a hard gate in your pipeline: no model without a resolvable licence text reaches a production branch. This is trivially cheap to implement and it is the single highest-value control on this list, given that 30% figure. For an Indian startup preparing for a Series A, or a UK company selling into NHS or public-sector procurement, "we could not establish the licence" is a finding that stops deals.

Use derivative counts, not download counts, to judge ecosystem support. When you need to know whether a model family will still have working quantisations, serving-engine support and community fixes in a year, the derivative count is the better signal. Downloads measure attention. Derivatives measure sustained effort by people who have to maintain what they published. A family with 151,448 derivatives has a different survival profile from one with 58,000, regardless of who is winning on pulls this month.

Pro tip

Add three fields to your model registry today: the resolved licence identifier, the revision hash the weights came from, and a stored copy of the LICENSE text as fetched. It takes an afternoon. It is the difference between answering a diligence question in ten minutes and spending a fortnight reconstructing what you downloaded eighteen months ago.

The honest counter-argument

We should state the case against our own reading, because it is a decent one.

Download and derivative counts measure distribution, not production deployment. A model can dominate the Hub and appear in almost no revenue-generating systems. Enterprise deployments frequently pull weights once into an internal artefact store and never touch the public hub again, which means the labs with the most conservative enterprise customers are systematically undercounted. GGUF pulls in particular skew heavily towards local and hobbyist use — that 39.6M figure represents a great deal of experimentation on laptops, and experimentation is not adoption.

There is also a limit to what any licence tells you. A clean Apache 2.0 grant from any jurisdiction settles the copyright question and nothing else. It does not answer export-control obligations, which vary by jurisdiction and change without notice. It does not answer your customer's data-residency requirements. It does not answer whether a public-sector buyer in the UK or an enterprise procurement team in India will accept the model's origin, which is a commercial and political question that no LICENSE file has ever resolved. Teams that read the permissive-licence statistic as a green light are making the same category error, in the opposite direction, as the teams currently assuming the opposite.

Our position after weighing that: the report does not tell you which models to use. It tells you that the assumptions underneath a lot of open-weight strategy documents were formed three years ago and have not been re-checked against data since. That is worth an afternoon of anyone's time to correct.

What to do this quarter

Four concrete moves, in order of return on effort.

  1. Audit what you already run. List every open-weight model in your stack with its resolved licence and revision hash. Flag anything undeclared. In our experience most teams find at least one surprise, and it is nearly always a model somebody pulled for a proof of concept that quietly became load-bearing.
  2. Right-size your ambition. If your open-weight plan is built around a frontier-scale model, check the serving arithmetic before the architecture review, not after. The 83%-under-1B statistic is telling you where the working ecosystem is, and small models close to the product is where most teams should be spending their attention. Our guide to self-hosting an open-weight model on vLLM in production is the practical starting point.
  3. Build the abstraction before you need it. The distribution picture in this report will not hold still. Whichever family leads next year, you want a serving layer where swapping the underlying weights is a configuration change rather than a quarter of engineering. That is cheap to build now and expensive to retrofit.
  4. Separate the licence question from the procurement question, explicitly. Write them as two rows in your risk register, with two different owners. Conflating them is how teams end up either blocking a perfectly usable Apache 2.0 model or shipping one their biggest customer will refuse to accept.

The hub changing hands is the news. What is stored on it is the fact. On Hugging Face's own accounting, the open-weight ecosystem's centre of gravity is not where most Western procurement processes assume, and the licence data is the part that should prompt a re-read rather than a reaction.

The full report is published by Hugging Face at huggingface.co/blog, and the model repositories behind every figure quoted here are browsable at huggingface.co/models. We would encourage reading it directly rather than through anyone's summary, including ours.