Research
Most of my prior work has been in metaethics, with a particular focus on nihilism and whether angst about it is warranted. Nihilism is the view that some norm or value (such as moral value) doesn't exist. Many people feel angst when they imagine that nihilism might be true; I argue that there are no good reasons for this. My basic response is that we can often construct alternatives to the kinds of norms and values that nihilism claims do not exist, which can serve the same function in our lives.
Currently, I am developing a research program at the intersection of metaethics and AI safety, exploring what consequences our views on practical reasoning might have for AI alignment.
The first question I'm working on is: what reasons are there to be moral? The literature has largely assumed that the goal is to find reasons to be moral that apply to all rational agents. I think we can make progress by relaxing this assumption, and instead thinking in terms of what kinds of minds have which reasons to behave morally, or equivalently, which assumptions different arguments for morality make about their target audience.
This is related to the alignment problem because traditional arguments for alignment difficulty take the Orthogonality Thesis to show that we can't rely on AIs reasoning their way to morality. But the Orthogonality Thesis only implies that there are no arguments that apply to all minds. Mapping which minds are subject to which reasons would tell us which features to train into AI systems: in effect, characterizing "basins of alignment" in the space of possible minds.
The second question I'm working on is whether we can rationally change our final ends. The received view, instrumentalism, answers "no": practical reasoning concerns only means, not ends. But I think this view is likely mistaken, because it assumes that the line between instrumental and final ends is sharper than it is.
This matters for alignment: traditional arguments for alignment difficulty assume both the Instrumental Convergence Thesis and the Orthogonality Thesis. These both assume a sharp distinction between instrumental and final ends. I doubt this invalidates those arguments entirely, but I suspect they will need reformulating in ways that might make them easier to address.
I am also working on a project about the game theory of the AI race. In The AI Race is Not a Prisoner's Dilemma, I argue that the AI race, contrary to popular belief, is not a Prisoner's Dilemma. Rather, for most reasonable priors on loss of control of powerful AI, it is a Stag Hunt. This matters because it means coordination on a slowdown, both between labs and between countries, is game-theoretically more feasible than is often assumed. And I plan to argue in a followup post that it would be game-theoretically feasible for the frontier labs to pause unilaterally.
- LessWrong · 2026
- Dissertation · The Ohio State University · 2025 · open access via OhioLINK
- [A paper about nihilism]Under review