A reported drop in Hugging Face downloads is raising a harder question: are fewer people running open models themselves, or has the measurement changed?
Nathan Lambert says the top open LLMs tracked by Interconnects have seen download rates fall about 30% for many models. He cautioned that Hugging Face may have changed how downloads are counted slightly, but said the decline has continued. Lambert later said the pattern looks similar when isolating models above 100 billion parameters.
The self-hosting hypothesis
Andrew Trask suggested that fewer people may be trying to host models themselves. Other participants pointed to a possible shift toward providers and gateways. One hypothesis is that developers are accessing large open-weight models through services such as OpenRouter because hosting models above 500 billion parameters is difficult and smaller models from labs are becoming cheaper and faster. Another comment says larger models are increasingly used through providers because agentic inference remains difficult to operate in-house.
Those explanations are not proven causes. The reported decline is not an official Hugging Face statistic, and the possible counting change makes the trend harder to interpret.