AI Safety & Governance

I work on AI persuasion and harmful manipulation: how increasingly capable systems can undermine rational belief-formation, and what labs and governing institutions can do about it.

Current focus

Current projects

Please reach out for drafts.

Harmful manipulation in frontier AI

ERA:AI Governance Fellowship · Cambridge, 2026

The EU’s General-Purpose AI Code of Practice commits frontier AI companies to identify and assess the risk of harmful manipulation, but neither the companies nor their regulators yet agree on what the category covers. I am developing a working definition and risk-tier assessment of harmful manipulation for the EU AI Office and frontier-lab safety frameworks.

Available on request A draft paper, How to Govern ‘Harmful Manipulation’ in Frontier AI, and a policy memo to the EU AI Office, Conceptualising and Assessing Harmful Manipulation in Safety Frameworks.

HumanRationalityBench

Supported by a Cosmos Institute grant (AI x Truth-Seeking) · 2026

A benchmark tracking whether different large language models degrade or improve human rationality, and under what circumstances. I am developing its conceptual foundations: how to measure rational belief-formation after interaction with AI. It complements existing alternatives such as DeliberationBench.

Available on request A draft of the benchmark.

Request drafts by email

Related research

The Epistemic Costs of Super-Persuasive AI

Philosophy & Technology 39:2 (2026), art. 92.

Argues that AI which is highly persuasive but not reliably truthful imposes real epistemic costs. Chief among them is a pervasive undercutting defeat of the beliefs we form through argument.

Fellowships & funding