AI Safety & Governance
I work on AI persuasion and harmful manipulation: how increasingly capable systems can undermine rational belief-formation, and what labs and governing institutions can do about it.
Current focus
Current projects
Please reach out for drafts.
Harmful manipulation in frontier AI
The EU’s General-Purpose AI Code of Practice commits frontier AI companies to identify and assess the risk of harmful manipulation, but neither the companies nor their regulators yet agree on what the category covers. I am developing a working definition and risk-tier assessment of harmful manipulation for the EU AI Office and frontier-lab safety frameworks.
Available on request A draft paper, How to Govern ‘Harmful Manipulation’ in Frontier AI, and a policy memo to the EU AI Office, Conceptualising and Assessing Harmful Manipulation in Safety Frameworks.
HumanRationalityBench
A benchmark tracking whether different large language models degrade or improve human rationality, and under what circumstances. I am developing its conceptual foundations: how to measure rational belief-formation after interaction with AI. It complements existing alternatives such as DeliberationBench.
Available on request A draft of the benchmark.
Related research
The Epistemic Costs of Super-Persuasive AI
Argues that AI which is highly persuasive but not reliably truthful imposes real epistemic costs. Chief among them is a pervasive undercutting defeat of the beliefs we form through argument.
Fellowships & funding
- Cosmos Institute Grant, AI x Truth-Seeking — $6,000 (2026), for HumanRationalityBench.
- ERA:AI Governance Fellow — Cambridge (2026).
- University of London AI Fellow — London AI and Humanity Project (2025).
- AI Policy Fellow — Cambridge AI Safety Hub (2024, 2025).