You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
We live in a world in which so-called "artificial agents" are not just commonplace, but growing in number and complexity at an alarming rate. If we are to manage the near-to-long-term consequences of integrating these systems into our work, personal, and civic lives, then it is increasingly important for us to reason about them correctly. Our ability to recognise and understand one another as agents is here a double-edged sword: it provides a useful toolkit of folk-psychological concepts (e.g., hallucination), and it disguises important disanalogies between ourselves and our machines (e.g., jagged competence). We believe that the concept of selfhood is the part of this sword which cuts most deeply.
The canonical theoretical arguments underpinning AI risk discourse make implicit appeals to selfhood in deploying notions like self-preservation, self-improvement, and self-replication to refer to convergent drives among complex goal-directed processes. Yet the extent to which these abstractions presuppose that such "agential" processes possess a concrete self-model has not been thoroughly examined — precisely because it is treated as a presumed feature of agency as such.
By contrast, practical research on the design of natural language based systems (LLMs and their descendants) is appropriating a richer vocabulary of personas, constitutions, and identification from natural language to explain and manipulate holistic patterns of behaviour and their relation to explicit and implicit self-representations. But these remain borrowed folk-psychological concepts in search of adequate abstractions.
The core research clarifies and connects these topics by distinguishing and articulating the relationships between aspects of self-modelling that current debates run together:
agential identity (boundaries) / social identification (roles)
factive self-representation (actual state) / ideal self-image (goal state)
causal autonomy (self-motivation) / normative autonomy (self-legislation)
These distinctions do work beyond alignment narrowly conceived. They give the AI personhood debate a way forward that does not route through intractable questions about phenomenal consciousness, by examining the relation between self-recognition and other-recognition and how it frames questions about oppression and exploitation. And they give the extended agency debate — whether AI-mediated actions are genuinely ours, and whether these systems can serve as prosthetic self-models that enhance rather than diminish autonomy — a framework for questions about the forms of alienation and authenticity engenderd by our institutional frameworks that neither romanticises nor dismisses the technology.
An online conference. The people working on these questions — in alignment organisations, academic philosophy, HCI, and the persona-engineering community — are currently doing so in a disconnected fashion, with incompatible vocabularies (and sometimes values). Our aim is to connect them up: an online conference with an open call across these communities with recorded talks published openly, and the explicit goal of leaving behind a network of channels through which important information about these topics might better flow. This is to say, shared reference points, rather than mere proceedings. For all the advances made in growing the AI research community and building new organisation, the field need institutional interfaces more than it needs another paper. Speaking of which...
A position paper articulating the conceptual framework sketched above, written to be understood and cited across the various subcommunities the conference aims to bring together — the shared reference point that the event revolves around, if only as an excuse for making comrades through comradely criticism. Which engenders...
Public essays and talks that pull apart and analyse the various theoretical threads that the concepts we're exploring bundle together — from blog posts to pre-prints to public lectures and youtube videos — whichever channels keep the signal flowing in the right way. Our aim is to build sustainable interfaces by translating our research framework into the divergent lexicons of the different subcommunities asking these existential questions. Which leads us to...
Peter Wolfendale is a neorationalist philosopher and computer science researcher. His most recent book, The Revenge of Reason (Urbanomic / MIT Press, 2025), develops a computational account of practical rationality (agency, selfhood, value, and freedom) grounded in the conception of theoretical rationality developed by Kant and Hegel, and refactored by Wilfrid Sellars and Robert Brandom. His essay "Geist in the Machine" (Aeon, 2026) is a précis of the research programme this project extends. He previously worked on the DARPA-funded SafeDocs information-security programme with Special Circumstances LLC, and has fifteen years' experience organising, communicating, and teaching interdisciplinary research, from the 'Emancipation as Navigation Summer' school (HKW Berlin, 2014) to a decade of online seminars at The New Centre for Research and Practice. You can find his exo-memory on Twitter (which he refuses to call it X), and his more polished writing on Substack.
Tushita Jha is a scholar of philosophy of AI, currently affiliated with the Mimir Center for Long Term Futures Research. She was previously a research scholar at the Future of Humanity Institute (FHI), a co-founder of PIBBSS, and served as Research Director at the AI Objectives Institute. Her current work spans the foundations of runaway optimization and instrumental convergence, the philosophy of human amplification, and the political economy of cognitive abundance.
Sahil Kulshrestha is an AI safety researcher and the author of the Live Theory research programme, which rethinks alignment through adaptive interfaces and the distribution of meaning. His writing can be found at lesswrong.com/users/sahil-1.
Any funding we receive will support the core team's research time, computing requirements, and all expenses involved with in organising events online (and hoperfully, in person). The potential scope of the project grows the more funding we receive, enabling us organise additional event, and more importantly, to pay other researchers for their time, effort, and the critical feedback and theoretical inspiration they supply. Speaking of which...
Beyond Preferences in AI Alignment (Pete)
Virtue-Ethical Rationality and Training Dynamics (Pete)
The project is developed in communication with ARB and the Autonomy Institute.
There are no bids on this project.