Research & Insights

Exploring AI strategy, LLM systems, autonomous agents, and the intersection of technology with public policy and education.

Astra on FORTRESS: One String, Four Scores

I ran GPT-6 Astra against FORTRESS, the national security and public safety benchmark, two days after release. The aggregate risk score is unremarkable. The log is not: the deployed system refuses in two different ways, and the benchmark treats them as opposites — deleting one kind from the sample, grading the other on a harm rubric. The judge then awarded four different scores to nineteen byte-identical refusals, including 6 out of 7 to a refusal that answers nothing. With interactive figures, and what all of it means for measuring manipulation.

AI Policy Is Now Power Policy

The EU AI Act's biggest applicability date lands on Sunday — and most headlines still get it wrong. Why I built aipolicy.es from scratch: what actually changes on 2 August 2026, the systemic risks the GPAI Code of Practice already names, what METR's lengthening task horizons and Epoch's compute curves say about the decade ahead — and a dated, falsifiable prediction: evolution, revolution or disaster. Spain has assets but not yet a system, and Spanish is governance infrastructure, not an interface.

The Fire That Learns to Make Fire: Governing AI at the First Signs of Recursive Self-Improvement

Frontier AI is not yet recursively self-improving in the strong sense — but it is already inside the loop that builds the next system. A five-layer map of RSI, the evidence from METR, AISI, and the lab frameworks, and a seven-point agenda — including an AI Evaluation Corps for Europe and Spain.

Human First, AI Frontier: Why Magnifica Humanitas Is the Rerum Novarum of Artificial Intelligence

Pope Leo XIV is not asking us to choose between technology and humanity. He is asking who governs the infrastructure that is already reshaping work, truth, power, education, war, and freedom. Read from inside AI governance: a Babel-or-Nehemiah civilizational choice, the U.S./China capital asymmetry, and a seven-point Human First agenda for Europe and Spain.

Synthetic Citizens Are Coming to Democracy

Generative AI did not invent fake political participation. It transformed its economics. From the FCC's 18M fake net-neutrality comments to covert AI personas earning karma on r/ChangeMyView, the next wave will produce synthetic citizens that speak, argue and simulate communities. A proposal for democratic provenance — machines can help us listen, but they cannot be counted as the people.

AI Swing States: How Latin America Will Shape US-China Competition and Global AI Safety

Latin America will not decide the AI race by training the largest frontier model. It will decide it by choosing the rules of the ecosystem — procurement, compute, data centers, languages, and safety institutions. A proposal for a Latin American AI Swing States Compact.

Spain Is Not a Translation Problem: Building Frontier AI That Understands the Country It Serves

Why I built spain-reference-personas-frontier — a synthetic reference population of 1M personas, 536K households, and 1,800 benchmark tasks — and how it became open infrastructure for better AI products, services, and public-interest systems for Spain.

Why I Built Acogida.es: A Letter on Welcome, Spain, and the Right to Begin Again

A free, multilingual, login-less tool launching the same week Spain opens a once-in-a-decade regularisation window. Why civic infrastructure — not walls, not rhetorical hugs — is what welcome actually looks like at 11:47 p.m.

Dopamine-Driven Development: When the Machine Keeps Whispering One More Prompt

When the feedback loop between builder and model becomes frictionless, productivity can feel intoxicating. A reflection on vibe coding, supervision fatigue, and the habits needed to build with AI without letting the loop consume the builder.

Intelligence Is No Longer Scarce. Trust in It Is.

Drawing on Catalini, Hui & Wu's MIT Sloan analysis of AGI economics, a three-part exploration of the Measurability Gap, the Missing Junior Loop, and why the cost to verify -- not the cost to automate -- will define the next decade of organizational strategy.

Human First, AI Frontier: Europe's Moment to Lead

Drawing on RAND's AGI preparedness analysis, the Draghi competitiveness report, and Davos 2026, a call for Europe -- and Spain -- to embrace AI with the ingenuity, depth, and humanistic ambition that has defined our civilization at its best.

Building a Spanish Medical Triage AI on a Single GPU

From broken models to balanced solutions: how LoRA fine-tuning, knowledge distillation, and DPO alignment taught me that 350 balanced examples beat 10,000 imbalanced ones — and that the ROJO bias was the most important finding of the project.

The AI Scientist: From Research Assistants to Autonomous Discovery

Mapping the landscape from Elicit and Semantic Scholar to FutureHouse's Robin, Ginkgo's autonomous labs, and the Genesis Mission. How AI is transforming scientific discovery -- with Promethean promise and Pandora's caution.

Governance as Advantage: Open-Weight AI, Agentic Systems, and Compute Diplomacy

The binding constraint for governments is no longer model access -- it is institutional capacity: evaluation, procurement, auditability, and adoption. How open-weight models and compute diplomacy are reshaping policy.

AI and European Democracy: Algorithms, Institutions, and the Fight for Democratic Sovereignty

How AI reshapes European democracy -- from recommender systems and algorithmic governance failures to the EU's four-pillar regulatory architecture and AI-powered civic participation. 30+ academic references.

AI and the Information War: Deepfakes, Democracy, and the Fight for Truth

A technical deep dive into how diffusion models, voice cloning, and LLMs power synthetic disinformation -- and the policy frameworks, provenance standards, and AI-powered countermeasures that can protect democratic institutions.

Building Enterprise Agent Platforms: Lessons from the Field

Practical insights from building GenAI platforms across professional services, infrastructure, finance, and energy sectors. RAG architectures, LLM orchestration patterns, and what actually works in production.

Democratizing AI: Scaling Education Across 12 Countries

The story of Saturdays.AI - from a single bootcamp to an international movement with 30K+ alumni and 500+ AI4Good projects. What we learned about making AI education accessible.

RAG for Industrial Engineering: Lessons from the Energy Sector

Developing RAG systems for Industrial Engineering practice and PMO at a $70Bn energy company. Architecture decisions, chunking strategies, and evaluation frameworks.

AI Use Cases in Defense & Aerospace

Defining and deploying AI use cases to strengthen operations in the defense and aerospace sector. Frameworks for identifying high-impact applications and NATO's evolving AI strategy.

From MVP to Seed Round: Building LLM-Powered Products

Lessons from building an LLM-driven platform for training and talent selection at Tauniqo.AI. How we went from concept to funded startup in the AI gold rush.

Technology in Frontier Markets: Lessons from Microsoft

Reflections on advising Cloud technology product-market fit across Benin, France, and Rwanda. What emerging markets teach us about technology adoption and digital inclusion.

AI in Educational Leadership: Navigating Challenges and Opportunities

Key findings from a scoping review published in the Review of Education (BERA/Wiley), examining how AI is transforming the role of school leaders and reshaping educational decision-making.

Stay Updated

New articles on AI strategy, LLM systems, and digital transformation. Connect to stay in the loop.