You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Umbrella grant covering compute, API, and conference costs for five interpretability and oversight projects led by my mentees.
Funding goal: $83,332
Project summary
I mentor a portfolio of AI safety research projects, some through the MARS program and some independently. The students’ time is already covered by their programs or their own funding; what the projects lack is compute. This grant aggregates the research-expense budgets of five of those projects into a single umbrella, plus a 20% contingency that I can reallocate across projects (or to other projects under my mentorship) as needs shift over the coming months.
The five projects share a common thread: understanding how behaviours, personas, and hidden information are represented inside language models, and what that means for auditing and oversight. Each project has its own leads, its own detailed budget, and a paper-shaped deliverable, most targeting NeurIPS 2027.
Two related projects are deliberately not part of this grant: Logan Graves’s belief-state geometry of language model personas (funded separately on Manifund) and a larger alignment pretraining project (being evaluated separately due to its size).
1. Mechanistic Signatures of Elicitation versus Teaching Across Model Scale — led by Hieu Nguyen and Marcin Podhajski (MARS) — $16,000
When fine-tuning improves a capability, is it teaching the model something new or eliciting something it already has? This project trains custom TinyStories-style models at four scales (38.7M to 7B parameters) and compares elicitation and teaching on natural-language arithmetic across three fine-tuning methods, five seeds, and a 19-point dataset-size sweep, with mechanistic analysis (residual representations, activation gradients, circuit discovery, causal interventions, weight-space trajectories) on top of every run. Budget covers ~4,050 GPU-hours of training and analysis, storage, and NeurIPS travel for two first authors.
2. Narrator-Mode Elicitation: Story-Prompting as a Tier-Calibrated Baseline for Hidden-Behavior Auditing — led by Aly Kassem — $22,200
Concealment-trained model organisms are the standard testbed for auditing methods, but their training only ever touches one generation mode: the assistant answering in first person. We have preliminary results showing that simply asking the model to narrate a story about an assistant with a secret recovers hidden behaviours that direct interrogation misses, across six independent organism suites — including a 2× improvement over a purpose-built introspection adapter on the benchmark that adapter was designed for. The funded work re-runs the core suites at the final prompt set and executes the causal follow-ups (activation probing, persona-vector steering, patching between generation modes, adversarial baselines, an agentic auditor loop). Budget covers ~4,000 H100-hours, ~91k judge/agent API calls, and NeurIPS travel for one first author.
3. Persona Manifolds: Nonlinear Geometry of Character Representations, Steering, and Drift — led by Trini Borrell and Brendon Chikavanga (MARS) — $8,298
Persona vectors treat each character trait as a single direction in activation space. This project tests whether the 276 role archetypes from the Assistant Axis work instead lie on a low-dimensional curved manifold, whether steering along that manifold beats vector steering, and whether narrowly misaligned fine-tunes leave the manifold — turning persona drift and emergent misalignment into measurable displacements. Deliverables are a per-model geometry of persona space for Gemma 2 27B and Qwen 3 32B, a manifold-aware steering method benchmarked head-to-head against vector steering, and a test of emergent-misalignment dynamics. Budget covers 1,250 H100-hours, ~1.2M batch-API grading calls, and NeurIPS travel for two co-first authors.
4. From User Representation to Assistant Stance: Dyadic Persona Selection in Language Models — led by Nayoung Kwon — $9,445
Language models have been studied through the personas they express, but much less attention has been paid to the other side of the conversation: the user as a character. This project asks whether models internally represent the user as an enacted character (novice, expert, emotionally vulnerable person, hostile interlocutor, malicious actor), what the geometry of that representation looks like, whether it predicts the stance the assistant adopts, and whether it is causally involved in persona selection (via activation steering and patching). The pipeline runs on Qwen3-32B, Gemma 24B, and Llama 70B. Budget covers 1,440 H200-hours, ~18k judge API calls, and NeurIPS travel for one first author.
5. Affective Exploitation of AI Oversight Systems — led by Sohan Venkatesh — $13,500
Scalable oversight assumes monitors apply their standards consistently. This project tests whether emotional distress expressed by a monitored AI agent causes monitor models to become lenient and approve actions they would otherwise reject, whether that effect emerges naturally without scripting, and whether prompt-level interventions can neutralize it — a subtle but unaddressed vulnerability in current safety pipelines. Experiments run across eight frontier and open models in SHADE-ARENA-style monitor settings. Budget is API-only (~$10,000 in calls, no GPUs) plus NeurIPS travel for one first author.
Elicitation vs Teaching (Nguyen, Podhajski): $16,000
Narrator-Mode Elicitation (Kassem): $22,200
Persona Manifolds (Borrell, Chikavanga): $8,298
Dyadic Persona Selection (Kwon): $9,445
Affective Exploitation of Oversight (Venkatesh): $13,500
Subtotal: $69,443
Contingency (20%), allocated across mentored projects as needed: $13,889
Total: $83,332
Every dollar goes to research expenses: GPU rental (on-demand cloud, mostly H100/H200-class), API calls for judging and evaluation, storage, and conference travel for presenting first authors. No stipends or salaries are drawn from this grant — everyone’s time is already funded through MARS, their institutions, or their own grants.
Detailed line-item budgets for each project (GPU-hour derivations, API call counts, travel assumptions) were prepared by each project’s leads and shared with the funder; I’m happy to share them with other interested donors on request.
The 20% contingency exists because compute estimates for research are noisy: runs fail, prompt sets get revised, and promising directions deserve follow-up. I will allocate it across these five projects, or to other projects under my mentorship, as needs become clear, and will account for it in the final readout.
The goal is for each of the five projects to produce a publishable result and a public artifact (paper, code, and datasets where applicable), with most targeting NeurIPS 2027. Beyond the individual papers, the portfolio is a bet on a research direction: that understanding how models represent personas, hidden behaviours, and their interlocutors is a tractable path to better auditing and oversight of frontier systems.
I meet regularly with every project team, review experimental designs and results, and connect the projects to each other (several share infrastructure and evaluation approaches). The aggregate structure means a project that stalls doesn’t strand its compute — I can move budget toward whichever directions are working.
I’m a PhD student at Mila / Université de Montréal, co-supervised by Yoshua Bengio and Guillaume Lajoie, currently an Astra Fellow at Constellation in Berkeley, and a Vanier Canada Graduate Scholarship holder. My research focus is mechanistic interpretability and AI safety; recent outputs include work on dedicated feature crosscoders (model diffing) with Anthropic, CIAware-Bench (control-intervention awareness in frontier LLMs), and Efficient Causal Graph Discovery Using LLMs. I currently mentor roughly nine students across Mila, MARS, and Meridian Cambridge, including the five teams above.
Each project is led by the mentees named above, who wrote their own proposals and budgets. The leads span MARS scholars and graduate students / independent researchers; several projects already have preliminary results (most notably the narrator-mode elicitation work).
The most likely failure mode for any individual project is a null or messy result: the manifold structure doesn’t beat vector steering, user representations don’t causally drive stance, the emotional-leniency effect doesn’t replicate across model families. For most of these projects a clean null is still informative and publishable (e.g., “monitors are robust to affective manipulation” is a useful oversight result), but it would reduce their impact.
The portfolio structure is the main mitigation at the funding level: it is unlikely that all five projects fail simultaneously, and the contingency lets me shift compute from stalled projects to productive ones. The most likely operational risks — GPU price movements, failed runs, evaluation-rubric churn — are already priced into the individual budgets, most of which carry their own internal buffers on top of the shared 20%.
The other failure mode is people: mentees are early-career and juggling programs and applications. MARS project teams have co-leads, and I am close enough to each project to keep it moving or wind it down honestly if a lead has to step away. Any materially unspent funds would be returned or reallocated in consultation with the funder.
The only funding these projects have received is through MARS, which supports the participating mentees’ time.
Reporting: I’ll send the funder periodic informal updates and a readout on each project after the fact, including how the contingency was ultimately allocated.
There are no bids on this project.