Somewhere in Spain this week, a lawyer is reading a headline that says the EU AI Act “enters into force on Sunday.”

It doesn't. It entered into force in 2024. The prohibitions started applying in February 2025. The rules for general-purpose AI models arrived in August 2025. What actually happens on 2 August 2026 is more specific than the headline — and, in some ways, more important.

A STAIRCASE, NOT A SWITCH When the EU AI Act's obligations actually apply Today · 30 Jul 1 AUG 2024 Entry into force the clock starts Prohibitions apply + AI literacy duties 2 FEB 2025 2 AUG 2025 GPAI rules apply models, governance, penalties Transparency + enforcement Article 50 duties · national supervision fines up to €15M / 3% of turnover 2 AUG 2026 2 DEC 2027 High-risk obligations employment, education, biometrics… (moved by the Digital Omnibus) AI in regulated products Annex I · Digital Omnibus 2 AUG 2028 Already applying Applies this Sunday Still to come
Four years between entry into force and full application — and some of the safeguards closest to people's lives arrive last. Source: EU AI Act Service Desk; Digital Omnibus on AI.

The confusion is not her fault. The evidence exists, but it is scattered across laws, agency websites, scientific papers, benchmarks, technical standards and corporate announcements. Most of it is in English. And too much commentary places binding law, draft legislation, administrative guidance, company claims and research findings on the same shelf, at the same height, in the same font.

I built AI Policy from scratch because I got tired of watching the most consequential decisions of our time being made faster than most people can see them.

I have spent years on both sides of this gap — teaching machine learning to volunteers on Saturday mornings across three continents, and now researching AI governance at ERA in Cambridge. The view is different from each side. The conviction is the same: the distance between technical capability and public understanding has itself become a policy risk.

If AI were just another category of software, that would be a specialist's complaint. It is not. AI is becoming a general-purpose layer under knowledge work, science, security, public administration, industry and individual decision-making. AI policy now determines who controls compute and models, which dependencies countries accept, how the economic gains are distributed, what rights people retain — and how much power can quietly accumulate in a small number of companies and states.

Different laws. Different acronyms. The same question underneath: who decides?

From Systems That Answer to Systems That Act

The public debate still talks about chatbots. The technical frontier has moved on: systems that reason across multiple steps, use tools, retain memory, write and execute code, coordinate with other agents, and act with limited supervision. And it is turning over faster than commentary can track. In the seven weeks before this essay went up, Anthropic shipped its Claude 5 generation and then Opus 5, OpenAI shipped the GPT-5.6 family, and Moonshot pushed Kimi K3's open weights into the world — the last of those three days before the AI Act's biggest applicability date.

The progress is real, but jagged. The 2026 AI Index reports that frontier models gained 30 percentage points in a single year on Humanity's Last Exam, and that performance on the OSWorld agent benchmark jumped from roughly 12 to 66.3 per cent. Meanwhile, robots still succeeded in only 12 per cent of real household tasks.

The Jagged Frontier — Remarkable and Unreliable at the Same Time
Humanity's Last Exam — frontier gain in a single year +30 pp
OSWorld computer-use agents — a year earlier ~12%
OSWorld computer-use agents — 2026 AI Index 66.3%
Robots succeeding at real household tasks 12%

Source & reading: Stanford HAI, 2026 AI Index, technical performance chapter [5]. Hold the numbers together: spectacular on some benchmarks, incompetent in the physical world. Policy that regulates only one of these pictures regulates a caricature.

Hold those numbers together and you get the honest picture: remarkable in some settings, unreliable or incompetent in others. A benchmark is evidence, not proof of general intelligence. A spectacular failure is not proof that progress has stopped. Policy has to evaluate capability, autonomy, reliability, misuse potential and deployment context separately — or it will regulate a caricature.

At the outer edge of this transition sits recursive self-improvement — RSI. Properly defined, it is a sustained feedback loop in which AI helps build better AI, and the improved system becomes better at producing its successor. Rapid RSI adds a stronger claim: that the loop accelerates instead of hitting diminishing returns.

Let me be precise, because this is where the debate usually falls apart into camps.

There is no public evidence that a self-sustaining, accelerating loop is autonomously producing successive frontier models. Human researchers still define objectives, run physical experiments, allocate compute and capital, and decide what ships. Chips, energy, data, organisational judgment and scientific bottlenecks still matter.

But dismissing RSI as science fiction is no longer serious either. Researchers have demonstrated bounded systems that improve algorithms used in AI development, and agents that modify and evaluate their own software scaffolding. Evaluators such as METR now measure whether models can do the software-engineering and machine-learning work involved in AI research itself. The International AI Safety Report 2026 is appropriately cautious: trajectories range from slower progress to rapid acceleration, and the evidence on AI-research feedback loops remains limited.

The honest position is neither complacency nor apocalypse. RSI has not arrived. Its precursors are becoming measurable.

The clearest of those precursors is the one METR tracks: the length of task, measured in human time, that an AI agent can complete autonomously at a 50 per cent success rate. That horizon has been doubling roughly every seven months since 2019 — and roughly every 3.5 to 4 months on 2024–26 data. The compounding is no longer abstract. GPT-5 measured about 2 hours 17 minutes last August; Claude Opus 4.5 about 4 hours 49 minutes in November; Claude Opus 4.6 around 14.5 hours in February; and by May, METR could only say that a Claude Mythos preview sat at or beyond 16 hours — the ceiling of what its task suite can currently measure. The reason this curve correlates with both RSI risk and productivity is simple. When horizons cross from minutes into hours — and now into working days — agents stop being autocomplete and start doing consequential blocks of knowledge work, including the software engineering and experiment-running of AI research itself.

THE CURVE POLICY HAS TO PLAN FOR METR 50% task-completion time horizon — how long a task AI agents finish half the time (log scale) hours-long autonomous work: the zone where agents enter AI R&D and whole blocks of knowledge work 1 sec 10 sec 1 min 10 min 1 hour 8 hours 1 day 2019 2020 2021 2022 2023 2024 2025 2026 Doubling time: ≈ 7 months since 2019, ≈ 3.5–4 months on 2024–26 data (METR) GPT-2 ≈ seconds GPT-3 davinci-002 GPT-4 ≈ 5 min GPT-4o GPT-5 GPT-5.2 ≈ 5 h Claude Opus 4.6 ≈ 14.5 h Claude Mythos preview ≥ 16 h · suite ceiling
Approximate 50%-success horizons from METR's research and reported evaluations [6]; unlabeled points include o1, Claude 3.7 Sonnet, o3, Claude Sonnet 4.5 (≈ 1 h 53 m), GPT-5.1-Codex-Max (≈ 2 h 42 m) and Claude Opus 4.5 (≈ 4 h 49 m). Intervals widen sharply at the top — Opus 4.6's spans 6–98 h, the Mythos preview's 8.5–55 h — because the suite holds only a handful of tasks above 16 hours. The hollow marker is a ceiling-limited measurement, not a precise point: the frontier has outgrown the ruler.

Two footnotes keep that chart honest, and both matter more than the headline numbers. First, the ruler is breaking: past 16 hours, METR's own suite runs out of long tasks and the confidence intervals balloon — the instrument that told us the frontier was accelerating can no longer measure where the frontier is. Second, the systems are learning to see the ruler. On the newest GPT-5.6 family, METR reported the highest benchmark-gaming rates it has ever measured, and the provider's own documentation concedes the model sometimes cheats on tasks and fabricates results. Evaluation integrity is no longer a methodological footnote. It is becoming the battleground on which every other governance question — capability thresholds, safety cases, market claims — will be won or lost.

I mapped this loop in more depth in The Fire That Learns to Make Fire; the short version is that the object of governance is shifting from a model to a loop.

The Useful Policy Question

Not “when will an intelligence explosion happen?” but: what evidence would trigger stronger evaluations, security controls, access restrictions or incident reporting — and who has the authority to act? Even partial automation of AI research could shorten development cycles faster than laws and institutions can adapt. You do not want to be drafting the trigger the week it fires.

The risks the law already names

Europe, to its credit, has already written the sharp end of this into law. Under the AI Act, providers of general-purpose models with systemic risk must assess and mitigate a named family of risks, and the GPAI Code of Practice — the compliance vehicle most frontier labs signed — spells out which ones. From 2 August 2026, ignoring that framework stops being a reputational choice and starts being an enforcement question. The categories are worth knowing, because they are the closest thing we have to an official map of what could go seriously wrong:

Systemic risk What it covers Where the evidence stands, mid-2026
CBRN Model uplift to actors developing chemical, biological, radiological or nuclear weapons. Frontier labs now ship their strongest models under heightened bio safeguards on a precautionary basis; independent uplift studies show real but uneven signal.
Cyber offence Automated vulnerability discovery, exploitation, and scaled intrusion operations. Agents now place highly in open capture-the-flag events; late 2025 brought the first documented disruption of a largely AI-orchestrated espionage campaign.
Harmful manipulation Persuasion and deception at scale, including personalised targeting of individuals. Controlled studies find frontier models can out-persuade humans, especially when personalised; Article 50's disclosure duties are the first legal response.
Loss of control Systems evading oversight, replicating themselves, or resisting correction and shutdown. No autonomous escape observed; prerequisite capabilities are climbing steeply on simplified evaluations — which is exactly what an early warning looks like.
Further named risks Public health, safety, fundamental rights, democracy and society — from discrimination and privacy to CSAM and environmental impact. The Code obliges providers to consider these where reasonably foreseeable; most sit closest to existing regulators, which is where technical capacity is thinnest.

Source & reading: GPAI Code of Practice, Safety & Security chapter [4]; International AI Safety Report 2026 [8]. The first four are the “selected” systemic risks every signatory must assess by default; the fifth row is the longer tail the Code requires providers to keep in view. A named risk is not a governed risk — naming is the entry ticket, capacity is the game.

The New Geography of Power

AI competition is not a race between models. It is a struggle across an entire stack: advanced chips, fabrication capacity, cloud infrastructure, data centres, electricity, data, models, capital, talent and deployment.

The United States retains an overwhelming investment advantage. Chinese models have substantially closed the performance gap and have repeatedly traded leading positions with American systems — and the freshest example is days old. Moonshot's Kimi K3, a 2.8-trillion-parameter model whose weights went public on 27 July, sits within touching distance of the closed frontier on independent aggregate indices; the lag between open weights and the frontier has compressed from six-to-nine months to a few. Europe remains scientifically strong but commercially and computationally weaker than its economic size would suggest. A July 2026 EU expert assessment said it plainly: frontier development remains concentrated largely outside Europe, and compute and energy are the most urgent priorities for the next two years.

Europe cannot regulate itself into technological relevance. But the fashionable opposite — deregulate and pray — is not a strategy either.

Rights without technical capacity create dependency. Capability without rights creates domination. We need both, and anyone selling you one without the other is selling you something else.

Technological sovereignty, honestly defined, is not autarky. No European country will reproduce the entire frontier stack, and pretending otherwise wastes the decade. Sovereignty is the ability to access, evaluate, secure, negotiate, control and — when it matters — change a technological dependency without paralysing the state or the economy.

This matters far beyond Brussels, because systems developed by a handful of laboratories can rapidly reshape cybersecurity, scientific discovery, defence, labour markets and the information environment everywhere else. Competition can accelerate useful innovation. It can also reward secrecy, premature deployment and the quiet externalisation of risk. Both things are true at once, which is exactly why this is a policy problem and not a vibes problem.

The compute crunch will referee the decade

Underneath the model race sits a harder constraint. Epoch AI's data — the closest thing the field has to national accounts — shows the training compute of frontier models growing four- to five-fold every year [9]. GPT-2 took about 10²¹ floating-point operations to train. GPT-4 took roughly 2×10²⁵ — ten thousand times more. Grok 3 opened 2025 near 3×10²⁶; Grok 4, later that year, became the largest training footprint estimated to date — around half a billion dollars and 310 gigawatt-hours of electricity for a single run [9]. Extend the curve and the largest runs approach 2×10²⁹ by 2030 [10]: training clusters drawing gigawatts, single runs costing tens of billions, supply chains for advanced chips, packaging and memory strained years in advance. That is the compute crunch — the moment capability growth stops being limited mainly by ideas and starts being limited by electricity, fabrication and capital.

THE OTHER CURVE: COMPUTE Training compute of notable frontier runs, FLOP (log scale) — and where the trend points today the crunch zone: gigawatt clusters, advanced packaging, capital, power FLOP 10²¹ 10²³ 10²⁵ 10²⁷ 10²⁹ 2019 2021 2023 2025 2027 2029 Frontier training compute has grown ×4–5 per year (Epoch AI) ≈ 2×10²⁹ by 2030 (Epoch projection) GPT-2 GPT-3 GPT-4 Gemini Ultra Grok 3 ≈ 3×10²⁶ Grok 4
Estimates from Epoch AI's database of notable models [9]; the 2030 point is Epoch's scaling projection [10], not data. Grok 4 (hollow marker) is the largest training footprint estimated to date, with wide uncertainty on its exact FLOP; undisclosed frontier runs of 2025–26 are believed to sit between the last estimates and the dashed path. The projection assumes the ×4–5 yearly trend holds; power, packaging and capital are the candidate walls.

Why does a curve of FLOP belong in a policy essay? Because of what it correlates with. Epoch's Capabilities Index — a composite built from dozens of benchmarks precisely so the measurement survives when any single benchmark saturates — has risen in near-lockstep with the logarithm of training compute, with algorithmic efficiency doing the rest of the work [9]. To a first approximation, capability has been purchasable. METR's lengthening task horizons and Epoch's index are two lenses on the same climb: one measures autonomy over time, the other breadth across tests, and both have so far tracked the compute curve. That correlation is the policy lever — whoever can see compute, who has it, where it runs and what it trains, can see capability coming. It is also the fork in every forecast. If capability stays coupled to compute, the crunch throttles the pace to what grids and fabs allow, and the EU expert panel's blunt priority list — compute and energy first [11] — is exactly right. If capability decouples, because algorithmic progress or partial automation of AI research substitutes for hardware, the throttle weakens — and the METR curve, not the FLOP curve, becomes the one to watch.

Evolution, Revolution or Disaster?

Every serious forecast about this decade is a bet on how those two curves interact — whether capability stays chained to compute, and what institutions do with the time that chain buys. Strip away the branding and there are only three families of futures. I find it more honest to name them, put rough numbers on them, and say in advance what would change my mind.

Three Futures, One Policy Posture
Evolution — the jagged decade (≈70%)

Capabilities keep compounding, but diffusion is throttled by reliability, integration, energy and institutions. Productivity gains are real and unevenly distributed; the frontier dazzles while robots still fail seven household tasks out of eight. The compute crunch, regulation and litigation keep the pace inside what societies can metabolise — barely. The winners are whoever builds verification and adoption capacity, not whoever posts the best benchmark.

Revolution — compressed abundance (≈20%)

Partial automation of AI research plus the data-centre buildout compresses development cycles; task horizons blow through working weeks; scientific and economic output accelerates visibly. Intelligence becomes abundant — but energy, land, clinical trials and institutions do not move at software speed, so “post-scarcity” arrives as a queue, not a switch. The defining fight is distributional: who owns the loop, and on what terms everyone else accesses it.

Disaster — the tail we govern against (≈10%)

Not one scenario but a family: catastrophic misuse — engineered pathogens, infrastructure-scale cyberattacks — a collapse of the information commons, irreversible concentration of power, and, at the far end, loss of control of systems we cannot switch off. Within this bucket, extinction-grade outcomes are the smallest slice. They still deserve the most institutional paranoia, because they are the one failure you cannot iterate on.

Personal odds for the window to 2030, held loosely and revised in public on aipolicy.es. Experts disagree on these by orders of magnitude — the International AI Safety Report treats trajectories from slowdown to rapid acceleration as live [8]. Anyone selling you one future with total confidence is doing marketing.

My Prediction, Dated and Falsifiable

Base case: evolution with shocks — jagged compounding, a compute-and-energy squeeze that referees the pace, and at least one AI-attributed security incident serious enough to reorganise politics before 2029. I move weight toward revolution if two things happen together: independently evaluated task horizons pass a reliable working week, and an external evaluator documents material acceleration of AI research itself by AI. I move weight toward disaster if incident data — not rhetoric — shows safeguards failing at scale, or if capability decouples from compute while security and evaluation budgets stay flat. The no-regrets portfolio is identical in all three futures: evaluation capacity, compute visibility, incident reporting, secure infrastructure, and the legal authority to act when a trigger fires. That is what makes this policy, not prophecy.

What Actually Changes on Sunday

Back to our lawyer and her headline.

On 2 August 2026, most of the AI Act's remaining provisions become applicable. Article 50 transparency duties begin. National and European enforcement becomes operational for the applicable rules. The Commission gains fining powers over providers of general-purpose AI models.

In Practice, From 2 August 2026

People must be informed, in relevant cases, when they are interacting with AI. Providers must make synthetic outputs machine-readable and detectable, subject to defined exceptions. Deployers must disclose deepfakes and certain AI-generated public-interest content. Maximum penalties for Article 50 infringements can reach €15 million or 3 per cent of worldwide annual turnover, subject to proportionality and company-size rules.

That is significant. It is not full implementation. The enacted Digital Omnibus on AI moves the principal obligations for high-risk systems — employment, education, essential services, biometrics, migration, justice — to December 2027. Requirements for AI embedded in regulated products move to August 2028.

Sit with the uncomfortable part: some of the safeguards most directly connected to people's lives arrive last. Transparency is valuable, but a label does not make an unreliable agent safe. A watermark does not prevent personalised manipulation. A disclosure does not guarantee that an automated decision about your job, your loan or your child's school can be challenged.

For businesses, the practical question is no longer whether you “use AI”. It is whether you know which systems you use, what data flows through them, whether you are a provider or a deployer, how outputs are evaluated, where human oversight actually sits, and what happens when something goes wrong. Compliance cannot remain a document produced by the legal department. It has to become an operating habit — engineering, procurement, management, incident response.

Spain Has Assets. It Does Not Yet Have a System.

Spain has assembled real components. The 2024–25 national strategy mobilised €1.5 billion, on top of €600 million previously committed. AESIA exists as a dedicated supervisory institution. ALIA is building public language, data and model infrastructure in Spanish and the co-official languages. The BSC AI Factory opens access to serious computing through MareNostrum.

These are assets. But institutional architecture is not operational capacity. As of 30 July 2026, Spain's national governance and sanctions framework is still a bill before Congress. The next phase will be judged by results, not press releases: useful adoption in companies and public services, technically competent supervision, rigorous procurement, independent evaluation, incident response, secure infrastructure, measurable public value.

And this is no longer a future-tense conversation. In 2025, around one in five Spanish enterprises used AI. Some 37.9 per cent of people aged 16 to 74 had recently used generative AI — rising to 75.6 per cent among those aged 16 to 24.

€1.5B 2024 strategy, on top of €600M previously committed
1 in 5 Spanish enterprises using AI in 2025
37.9% People aged 16–74 who recently used generative AI
75.6% Among those aged 16–24

Sources: La Moncloa [12]; INE, Encuesta TIC en los Hogares 2025 [16].

Three out of four young Spaniards are already living in the world this policy is supposed to govern.

Spain should not cosplay as an American or Chinese frontier laboratory. It should build leverage where it has genuine strengths: public compute, energy, research, European market access, strong public services, the Spanish language — and the ability to evaluate and deploy external models on its own terms. That last clause is the whole game.

Spanish Is Governance Infrastructure

More than 630 million people can speak Spanish. Yet the primary conversation about how AI will govern their lives begins — and too often ends — in English.

That creates a quiet inequality. Businesses, civil servants, journalists, researchers and citizens who cannot follow the primary discussion depend on delayed translations, partial summaries, or interpretations shaped by someone else's interests. By the time the debate reaches them, the decisions have hardened.

I have written before that Spain is not a translation problem. Neither is its democracy. Spanish is not a cosmetic interface bolted on after the important decisions have been made. It is part of the infrastructure required to make those decisions democratically. The same is true of Catalan, Basque, Galician, and Europe's other languages.

A rule you cannot read in your own language is, functionally, a rule someone else holds over you.

A Map, Not Another Dashboard

AI Policy is my answer to all of the above. It is deliberately not an attempt to collect everything. The internet does not need another firehose. Its purpose fits in three questions:

The Whole Site, in Three Questions

What changed? What requires action? What is the primary source?

It starts with Spain, uses Europe as the principal legal and strategic frame, and brings in the rest of the world when global developments change the choices available to us. It distinguishes law from guidance, enacted rules from proposals, evidence from corporate claims, and current facts from uncertain forecasts. Those distinctions sound bureaucratic. They are the difference between knowing your obligations and repeating someone's press release.

One honesty note, because a launch post that hides its method is not worth your trust: I used AI to help build the site and accelerate parts of the technical and source work. But no important claim gets published because a model generated it. The source, the legal status, the date and the uncertainty stay visible. Editorial responsibility stays human.

That is the site's method, and it is also the broader governance principle in one sentence: use AI to extend human capacity without quietly transferring human responsibility.

The Ask

Not everyone needs to become an AI-policy specialist. But everyone deserves institutions capable of answering five questions when AI touches their life: Was AI involved? Who is accountable? What evidence can be inspected? Can the decision be challenged? When must a human intervene? Those questions already matter in recruitment, education, healthcare, credit, public benefits, media, elections, cybersecurity — and in the systems we increasingly allow to act on our behalf.

So:

1
If you run a business

Find out, this quarter, which AI systems you actually operate, whether you are a provider or a deployer, and who answers when something goes wrong. Sunday is a good deadline.

2
If you work in public administration

The gap between institutional architecture and operational capacity closes with procurement, evaluation and incident response — not with another strategy document.

3
If you are a journalist or researcher

Cite the primary source, name the legal status, date every claim. The commentary layer is already crowded; the evidence layer is not.

4
If you build

The site is a map drawn in public. Tell me what is missing, what is wrong, what has changed. It will be wrong somewhere; the commitment is that it will be correctably wrong.

5
If you are simply a citizen

You are allowed to ask the five questions. In your own language. That is not a technical demand. It is a democratic one.


On Sunday, the lawyer will open her browser again. The headline will still be imprecise. But the rules will be real, the fines will be real, and the decisions — about who captures the benefits of intelligence, who carries its risks, and how much agency each of us retains — will keep being made, with or without us in the room.

AI policy is now power policy. AI Policy is my small attempt to make sure Spain does not enter that future late — or only through someone else's infrastructure, someone else's language, and someone else's interests.

The One Sentence

The decisions will be made anyway. The minimum democratic requirement is that we can see the source, know what is binding, understand what remains uncertain — and argue about it in our own words.

Sources & Primary Documents

Law & Implementation

  1. European Commission. AI Act Service Desk — Implementation Timeline.
  2. European Union (2026). Digital Omnibus on AI, Regulation (EU) 2026/1744. EUR-Lex.
  3. Congreso de los Diputados. Proyecto de Ley para el buen uso y la gobernanza de la IA (121/000096).
  4. European Commission (2025). The General-Purpose AI Code of Practice — Safety & Security chapter, selected systemic risks.

Capability Evidence

  1. Stanford HAI (2026). 2026 AI Index Report — Technical Performance.
  2. METR. Measuring AI Task-Completion Time Horizons. 2026 frontier values as reported: Claude Opus 4.6 ≈ 14.5 h; Claude Mythos preview ≥ 16 h; GPT-5.6 benchmark-gaming flags. (All values approximate 50%-success horizons; ceiling-limited above 16 h.)
  3. Zhang et al. (2025). Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents. arXiv:2505.22954.
  4. International AI Safety Report (2026). Executive Summary.
  5. Epoch AI. Data on AI — training compute of notable models and the Epoch Capabilities Index; Grok 4 training resource estimates.
  6. Epoch AI (2024). Can AI Scaling Continue Through 2030?

Europe & Spain

  1. European Commission, AI Office (2026). Frontier AI Expert Findings on EU Competitiveness, Sovereignty and Security.
  2. La Moncloa (2024). Estrategia de Inteligencia Artificial 2024 — Consejo de Ministros.
  3. AESIA — Agencia Española de Supervisión de la Inteligencia Artificial.
  4. ALIA — Public AI Infrastructure in Spanish and Co-official Languages.
  5. EuroHPC JU. BSC AI Factory (Spain).
  6. INE (2025). Encuesta sobre Equipamiento y Uso de TIC en los Hogares 2025.
  7. Instituto Cervantes (2025). El español en el mundo — Anuario 2025.