You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Social scientists are now using language models to label texts, code documents, and stand in for survey respondents, mostly with no way to check what happens between the prompt and the output. Mechanistic interpretability has the tools to look inside, but everything written about it assumes a machine learning background. To the best of my knowledge, there is no accessible, hands-on book teaching mechanistic interpretability, with its premises and perils, to social scientists.
I am writing that book in the open, and I want to teach it.
I expect two outputs: one is a short manuscript, about 30,000 words, and the other is a week-long bootcamp that teaches the book and pairs social scientists up to work on a real project.
The book is organized around what you can do to a model. Observation is about what you can describe using probes, the logit lens, sparse autoencoders. Intervention shows what the model actually uses, via activation patching, ablation, steering. Validation is required to check whether the measurement is doing what you think it did. As the field is growing day by day, I planned to design every chapter answering: what does this technique let you claim, and what does it not? That question is the spine of the book. The last section will be a worked case, my own study of political ideology inside open-weight models, starting from the research question through the design to what the result actually licenses.
The draft is public at github.com/tapanyemre/anthropology-of-machines, licensed CC BY-NC-SA for the prose and MIT for the code. The preface is written and the chapters are outlined to the section level with word targets. Chapters appear when they become readable, not when they are finished, so readers can follow the history of my writing and keep me honest, and open an issue when I am wrong.
The hands-on parts, wherever possible, will not require a personal GPU. Instead I will use open-source libraries like nnsight to reach model internals through NDIF's shared machines. That is required if we want the tools accessible to the ones who have the theory and the interest but not the resources.
With funding the manuscript is done in eight months, and the bootcamp follows in summer 2027.
Almost all of it is time. I am in the final year of a PhD, and this money buys months where I can focus on writing the book. The stipend is $3,000 a month in living costs. There will be no computing cost, as the material will be based on Neuronpedia and NDIF, so what a reader needs is a browser.
Ideal: $33,000 covers eight months of stipend ($24,000) plus the bootcamp ($9,000). The bootcamp budget is roughly $5,000 for travel and accommodation so people who are not local can come, $2,500 for food across the week, and $1,500 for a teaching assistant.
Minimum: $12,000 is about one semester part-time. It funds the observation section — the tools for finding what is inside a model — written and published openly. No bootcamp at that level.
In between, $24,000 covers all eight months and the finished manuscript without the bootcamp.
I am the sole author. The draft is public and open to review from anyone who wants to check a claim, and I would rather be reviewed along the way.
I am a PhD candidate in political science (minor in computational social science) at Northeastern. Recently, I've started to work on AI interpretability. In spring 2026 I was the only political scientist invited into David Bau's Neural Mechanics, a full semester of mechanistic interpretability taught from the foundations up. The project I proposed there, on the political ideology of open-weight models, is a working paper now, and it becomes the book's worked case. A second working paper asks whether LLM annotations measure the constructs we ask them for.
My teaching record is longer than the interpretability one. I co-founded SICSS Istanbul in 2019 and co-organized it through 2023 — a two-week summer institute training around twenty early-career social scientists a year — and I went back as a guest lecturer in 2026, where I taught this material to a room that ran the logit lens in their own browsers. In summer 2025 I led the ten-lab programming bootcamp of the AIDE program at Northeastern's Ethics Institute, fifteen philosophy and computer science graduate students, from their first line of Python to working with language models. Last fall I ran two workshops on the mechanics of LLMs on a small NULab grant, which brought an interdisciplinary group of social scientists together around this material.
The most likely cause is time. This is my final year of a PhD, and the dissertation comes first. If the funding is thin I will keep teaching and consulting, the book will slow down, and the bootcamp will be the first thing to go.
There are two smaller risks. The field moves fast, so a book built on a fixed toolkit would date quickly. That is why it teaches what a method assumes and what it lets you claim. And I am one author, new to this field, so I will get some things wrong. Writing in public is how I catch that: mistakes come back as issues instead of in print.
If it fails, not much is lost. Everything is already public and openly licensed, so anyone can read the finished parts. The roadmap shows what has shipped, so you would see me stall while it happened, not at the end.
I believe the real loss would be timing. The literature on validating language models as instruments is being written right now, and the book that teaches social scientists to look inside these models is still missing.
NULab Digital Humanities and Computational Social Science at Northeastern University granted $500 for running two workshops on the mechanics of LLMs in Fall 2025.
There are no bids on this project.