From March 2025 to October 2026, I worked on AI at the Office of the Prime Minister of Spain. With a very small team, I led, designed and built three initiatives: PresidencIA, the Office's generative AI adoption program; ServetIA, an AI research platform for civil servants; and CiudadanIA, which processes the letters citizens send to the Prime Minister. All three have now been selected, following OECD review, for publication in the OECD.AI Policy Observatory (PresidencIA, ServetIA, CiudadanIA).
As that work closes, here is what it taught me about adopting AI where mistakes are strategic: where an error is not a bug for the next sprint but a public event, with political, legal and security consequences.
1. Adoption is organizational before it is technical. PresidencIA was organized around community, applications and infrastructure, and it mobilized hundreds of officials. Choosing a model was the easy part. Getting people to use it on their real work was the program.
2. Capability is not deployability. In its third week of testing, CiudadanIA summarized a citizen's letter accurately, then attributed the problem to a ministry that does not exist. The civil servant reviewing the summaries caught it and asked the question that matters: "If this tool invents institutions, how do I trust it with anything?" (the full story). Fluency is not fitness for an institution whose errors make headlines.
3. Design for verification, not trust. The fix was better prompts and entity guardrails, but above all a new workflow: reviewers were no longer asked to trust the model, only to check it. In high-stakes settings, the human in the loop needs a job they can actually do.
4. Make every answer traceable. At the center of government, an answer nobody can source is an answer nobody can use. That is why ServetIA returns answers linked to their references. For a research tool, traceability matters more than fluency.
5. Governance is a mechanism, not a principle. Transparency and human oversight are necessary and never sufficient. What decides whether a system ships is concrete: evaluation criteria, procurement requirements, risk assessments, deployment safeguards. And the institution, not the technology, sets the pace: a very small team could ship applications quickly, while buying eight licenses took six weeks and an eleven-page justification.
What changes with agents
Each of these lessons gets harder as systems become more capable and more agentic. When a system drafts, people review. When it acts, the institution needs evidence, before deployment, of what it can do and how it fails. That is the question I now work on at Cambridge ERA:AI, where my working paper on the Capability–Deployment–Effect gap separates what a model can do, how it is deployed, and what it demonstrably causes.
Build, govern, evaluate: inside government, these were never three jobs but one problem seen from three distances. My thanks to the team that built this with me, and to every official who was willing to try.