Peter Wolfendale's work in philosophy at large has been a big influence on my work in AI alignment.
You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
We live in a world in which so-called "artificial agents" are not just commonplace, but growing in number and complexity at an alarming rate. If we are to manage the near-to-long-term consequences of integrating these systems into our work, personal, and civic lives, then it is increasingly important for us to reason about them correctly. Our ability to recognise and understand one another as agents is here a double-edged sword: it provides a useful toolkit of folk-psychological concepts (e.g., hallucination), and it disguises important disanalogies between ourselves and our machines (e.g., jagged competence). We believe that the concept of selfhood is the part of this sword which cuts most deeply.
The canonical theoretical arguments underpinning AI risk discourse make implicit appeals to selfhood in deploying notions like self-preservation, self-improvement, and self-replication to refer to convergent drives among complex goal-directed processes. Yet the extent to which these abstractions presuppose that such "agential" processes possess a concrete self-model has not been thoroughly examined — precisely because it is treated as a presumed feature of agency as such.
By contrast, practical research on the design of natural language based systems (LLMs and their descendants) is appropriating a richer vocabulary of personas, constitutions, and identification from natural language to explain and manipulate holistic patterns of behaviour and their relation to explicit and implicit self-representations. But these remain borrowed folk-psychological concepts in search of adequate abstractions.
The core research clarifies and connects these topics by distinguishing and articulating the relationships between aspects of self-modelling that current debates run together:
agential identity (boundaries) / social identification (roles)
factive self-representation (actual state) / ideal self-image (goal state)
causal autonomy (self-motivation) / normative autonomy (self-legislation)
These distinctions do work beyond alignment narrowly conceived. They give the AI personhood debate a way forward that does not route through intractable questions about phenomenal consciousness, by examining the relation between self-recognition and other-recognition and how it frames questions about oppression and exploitation. And they give the extended agency debate — whether AI-mediated actions are genuinely ours, and whether these systems can serve as prosthetic self-models that enhance rather than diminish autonomy — a framework for questions about the forms of alienation and authenticity engenderd by our institutional frameworks that neither romanticises nor dismisses the technology.
An online conference. The people working on these questions — in alignment organisations, academic philosophy, HCI, and the persona-engineering community — are currently doing so in a disconnected fashion, with incompatible vocabularies (and sometimes values). Our aim is to connect them up: an online conference with an open call across these communities with recorded talks published openly, and the explicit goal of leaving behind a network of channels through which important information about these topics might better flow. This is to say, shared reference points, rather than mere proceedings. For all the advances made in growing the AI research community and building new organisation, the field need institutional interfaces more than it needs another paper. Speaking of which...
A position paper articulating the conceptual framework sketched above, written to be understood and cited across the various subcommunities the conference aims to bring together — the shared reference point that the event revolves around, if only as an excuse for making comrades through comradely criticism. Which engenders...
Public essays and talks that pull apart and analyse the various theoretical threads that the concepts we're exploring bundle together — from blog posts to pre-prints to public lectures and youtube videos — whichever channels keep the signal flowing in the right way. Our aim is to build sustainable interfaces by translating our research framework into the divergent lexicons of the different subcommunities asking these existential questions. Which leads us to...
Peter Wolfendale is a neorationalist philosopher and computer science researcher. His most recent book, The Revenge of Reason (Urbanomic / MIT Press, 2025), develops a computational account of practical rationality (agency, selfhood, value, and freedom) grounded in the conception of theoretical rationality developed by Kant and Hegel, and refactored by Wilfrid Sellars and Robert Brandom. His essay "Geist in the Machine" (Aeon, 2026) is a précis of the research programme this project extends. He previously worked on the DARPA-funded SafeDocs information-security programme with Special Circumstances LLC, and has fifteen years' experience organising, communicating, and teaching interdisciplinary research, from the 'Emancipation as Navigation Summer' school (HKW Berlin, 2014) to a decade of online seminars at The New Centre for Research and Practice. You can find his exo-memory on Twitter (which he refuses to call it X), and his more polished writing on Substack.
Tushita Jha is a scholar of philosophy of AI, currently affiliated with the Mimir Center for Long Term Futures Research. She was previously a research scholar at the Future of Humanity Institute (FHI), a co-founder of PIBBSS, and served as Research Director at the AI Objectives Institute. Her current work spans the foundations of runaway optimization and instrumental convergence, the philosophy of human amplification, and the political economy of cognitive abundance.
Sahil Kulshrestha is an AI safety researcher and the author of the Live Theory research programme, which rethinks alignment through adaptive interfaces and the distribution of meaning. His writing can be found at lesswrong.com/users/sahil-1.
Any funding we receive will support the core team's research time, computing requirements, and all expenses involved with in organising events online (and hoperfully, in person). The potential scope of the project grows the more funding we receive, enabling us organise additional event, and more importantly, to pay other researchers for their time, effort, and the critical feedback and theoretical inspiration they supply. Speaking of which...
Beyond Preferences in AI Alignment (Pete)
Virtue-Ethical Rationality and Training Dynamics (Pete)
The project is developed in communication with ARB and the Autonomy Institute.
Peli Grietzer
14 days ago
Peter Wolfendale's work in philosophy at large has been a big influence on my work in AI alignment.
Pete Wolfendale
14 days ago
@Peli-Grietzer I just want you to know that I wanted to heart react this but that's an emoji in the paid for tier. What is this world we live in.
Callan McGill
19 days ago
sysematic philosophy taking seriously what has happened in the formal sciences in the last 60 years is something we should want to bring about in a world overtaken by hyper-specialized marginal thinking (marginalia)
Ankur Pandey
19 days ago
I have worked with Sahil & have known / interacted with TJ long enough to endorse this. Unique thinkers, and agree with @gleech's "abstract-to-max-abstract" framing.
I'm also interested in the topic because of recent involvement with digital minds literature. Would love to see the online conference, especially.
Oliver Klingefjord
20 days ago
Pete is an extremely underrated philosopher that is severely lacking in the AI alignment world!
Gavin Leech
20 days ago
Sahil, [TJ](https://ai.objectives.institute/blog/the-problem-with-alignment), and Pete have a strong reputation for deep thinking about AI, though on the abstract-to-max-abstract end of the greater AI safety and AI flourishing field. I've been chatting to Pete and others in the cluster for a while and think that the philosophical gap discussed here is actually urgent. I'm hoping (15%?) that along with related work by Grietzer this new subfield soon yields a training method with better alignment properties.
My sole reservation is that the work badly needs a project manager, which is why I've funded a good amount more than the minimum.
Good luck!
CoI: I own Arb, who are providing a little project management / cheerleading pro bono. We won't see any of the money.
Pete Wolfendale
20 days ago
@gleech Thanks for this Gavin. The project would be impossible without your support. And as for badly needing a project manager, no truer word has ever been spoken.