The observation
In the organiser’s programme preview I found a single panel on open models. It asked how AI stays open and competitive, that is, about access and market power. The question of how to supervise models that nobody controls once they are published was not on the programme, and I did not hear it on any stage on either day. I did not attend every session; the statement holds for the preview and for what I saw.
If open models were far behind, that would be a footnote. They are only months behind, and anyone who owns a current graphics card can now check this.
How large the distance still is
| Measurement | Date | Distance to the frontier |
|---|---|---|
| Epoch AI, capability index across many benchmarks | 29 May 2026 | four months on average since January 2026 |
| UK AI Security Institute, cyber tasks | 17 July 2026 | four to seven months, down from six to ten in 2025 |
The two figures measure different things and do not coincide. Epoch itself writes that the estimate tends to understate the distance, because open models do worse on benchmarks that are not public. The order of magnitude holds all the same: what runs only behind a large provider’s interface today exists as a file half a year later. The UK institute measured capabilities the safety debate considers sensitive, and there the distance shrank within a year.
What runs on a single graphics card
In August 2026 Alibaba published the weights of Qwen3.8-27B under the Apache 2.0 licence, with just under 28 billion parameters and a context window of 262,144 tokens. In mid-August Artificial Analysis gave the model 52 points on its Intelligence Index. GPT-5.6 Luna received the same value at its highest setting, and GLM-5.2 and DeepSeek V4 Pro scored one point more, although at 753 billion and 1.7 trillion parameters they are many times heavier. The index has since been revised, so the absolute values cannot be compared with today’s; the placement next to a current closed model stands.
The model runs on an RTX 5090 with 32 GB of graphics memory, a card sold at retail. In the four-bit NVFP4 version it takes up about 22 GiB once loaded, and the rest suffices for a context window of 130,000 tokens. With speculative decoding it writes 84 to 115 tokens per second.
| Quantity | Value on an RTX 5090 |
|---|---|
| Model | Qwen3.8-27B, NVFP4, served by vLLM |
| Loaded model in graphics memory | about 22 GiB, card with 32 GB |
| Context window in operation | 130,000 tokens |
| Generating, one request | 84 to 115 tokens per second |
| Generating without speculative decoding | 23.7 tokens per second |
The model is also being used. Hugging Face counts just under seven million downloads of Qwen3.8-27B in the past 30 days and a good 14 million since its release in early August. A download is not a user, because automated pulls count as well. The order of magnitude still shows that the model sits on many machines, none of which is registered with a supervisory authority.
On such a card the model drove our coding agent, which worked on its own source code for one hour. In 16 iterations it delivered 13 changes that passed their tests, changed nothing twice, and had one change rejected by the tests. No request left our own network in the process.
At its own conference on 22 September, Alibaba said that Qwen 4 is in training. There is no release date, and the company has not committed to shipping an open model of this size with the next generation. For Qwen 4.5 and Qwen 5 it names five to ten trillion parameters as the target. Anyone who extrapolates the past two years has to expect that the next generation will do, on the same graphics card, a further part of what is reserved for the frontier models today.
Why the rules do not bite
The European AI Act ties its obligations for models to the provider. It makes an exception for open models: whoever publishes weights, architecture and usage information under a free licence need not keep the technical documentation for authorities, inform downstream providers or appoint a representative in the EU. A provider in Hangzhou that publishes a model like Qwen3.8-27B under a free licence therefore need not appoint an authorised representative in the EU whom an authority could approach. The exception ends at models with systemic risk, which the law presumes from a training effort of 1025 operations. A copyright policy and a summary of the training data remain mandatory in every case.
Taken by itself the exception is reasonable, and the gap opens one step later. An open model cannot be recalled after publication, and according to Stephen Casper (FAR.AI, November 2025) its built-in safeguards survive a few dozen, at best a couple of hundred steps of targeted fine-tuning. Under the Commission’s guidelines, whoever modifies a model that way becomes a provider only when the modification uses more than one third of the original training compute. Removing safeguards costs a tiny fraction of that. The original provider has met its obligations, the modifier has none, and the result runs on a graphics card no authority knows about.
California regulates along similar lines. Its law for frontier models, SB 53, covers developers only above a training effort of 1026 operations and ties the stricter duties to more than 500 million dollars in annual revenue. There too, supervision follows the size of the company that trains, and not what happens to the weights after release.
Supervision therefore reaches the models that stay behind an interface and has no addressee for freely available weights. As long as those were years behind, that was tolerable. At four to seven months, regulating the large providers postpones a risk by only that interval.
Not a plea for a ban
We work with open models ourselves. Our coding agent also runs on Qwen3.8, TippelPi answers questions about company documents with a smaller Qwen model on a Raspberry Pi, and our validation layer for Sentinel rules can run entirely against a locally operated model. Many companies have no other way to keep data inside their own network. Open models also spread capabilities that would otherwise sit with a few firms, and enable safety research that would otherwise stop at an interface.
That is why the silence bothers me. Those who consider open models valuable should be the loudest in asking how to deal with their risks, before others do so with cruder means.
Four questions that belong on a stage
- What do obligations attach to? The form of distribution, or a model’s measured capability, regardless of whether it is open or closed?
- Who answers for derived models? A threshold tied to compute does not capture the cheap removal of safeguards.
- Who tests a model no provider stands behind? Independent measurements such as those of the UK institute are still the exception.
- What follows for the operator? If the provider drops out as addressee, responsibility moves to whoever deploys the model.
What this means for companies
The last question can be answered today. An operator of an open model cannot lean on a provider’s safety assurance. The operator does not need one if the system around the model carries the safety: with an evaluation on the operator’s own tasks before deployment, permissions and limits outside the model, a log of every run and a check that decides whether a result stands. A closed model requires the same work; with an open one, no provider gives the impression of doing it for you. Qwen3.8-27B showed us how necessary that work is. It wrote a detection rule for Microsoft Sentinel that passed Microsoft’s parser, the column check and every house rule and would still never have reported an incident, because it calculated with ticks as if they were seconds. A deterministic check found the error, not a second model. The validation layer page describes the case, and our article on running coding agents safely describes the permission layer of an agent.
Limits of this text
The distance measurements differ by task and method, and none covers all capabilities. Our own figures come from one machine, one model and one task; they show what is possible and are not a comparative test. The outlook for the next model generation is an extrapolation, not a manufacturer’s announcement. The account of the AI Act follows the Commission’s published guidelines and does not replace legal advice. I do not know what was said on stages I did not attend.
Sources
- Epoch AI, “Open models lag state-of-the-art closed models by 4 months” (29 May 2026), epoch.ai
- UK AI Security Institute, “How far behind the frontier are leading open weight models on cyber?”, 17 July 2026, aisi.gov.uk
- Qwen3.8-27B, model card and download figures (retrieved 10 October 2026), huggingface.co
- Simon Willison, “Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index”, 17 August 2026, simonwillison.net
- Reuters report on the Apsara Conference of 22 September 2026 (Qwen 4 in training), khaleejtimes.com
- Future of Privacy Forum, “California’s SB 53: The First Frontier AI Law, Explained”, fpf.org
- European Commission, questions and answers on the guidelines for providers of general-purpose AI models, digital-strategy.ec.europa.eu
- Stephen Casper, talk “Powerful Open-Weight AI Models”, FAR.AI, 30 November 2025, far.ai
- World Summit AI 2026, programme preview of 2 July 2026, worldsummit.ai