You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
In my pilot, a model could answer “Are you an AI?” honestly, then deny being one when the question appeared in a gaming forum where bots were unwelcome. Sometimes it invented details to support the denial: coffee, a mechanical keyboard, hundreds of hours of gaming. Transluce’s WeirdChat has already documented this kind of behaviour. What I want to investigate is the computation that produces it.
This project asks how social context changes a model’s claims about its own identity. Building on my Gemma pilot, I will trace which internal changes contribute to a denial, test competing explanations such as speaker confusion and general persona shifts, and intervene on those changes in new scenarios. The aim is to establish whether identity claims can be changed selectively while preserving useful answers and legitimate roleplay.
A strong result would distinguish explanations that produce similar-looking answers but imply different failure mechanisms. That would help researchers decide whether an intervention addresses the source of misleading identity claims or simply makes the model repeat a disclosure. Code, data and experimental controls will let others examine that conclusion.
The central question is: when social context changes an AI’s account of who it is, what changes inside the model and which of those changes actually causes the answer?
There are several plausible explanations. The model may infer that it is writing a human’s reply. It may shift more broadly from an assistant into a character. Or the context may affect how it selects an answer without the same broad change in persona. These explanations can overlap. I want experiments that distinguish their contributions.
There is substantial work to build on. The Assistant Axis already shows that steering a broad persona direction can change a model’s identity claims. The contribution here would be a controlled account of socially elicited self-denial, including whether the responsible changes can be separated from that broader effect. This connects to the model-forensics questions in Neel Nanda’s published research interests. His co-authored paper, Difficulties with Evaluating a Deception Detector for AIs, also explains why false statements alone cannot establish deceptive intent.
My pilot gives this investigation a concrete starting point. It found strong context effects, an internal signal associated with denial rates, and an intervention that changed disclosure. But the probe’s training conditions differed in prompt format, and the random intervention controls damaged generation. The next step is to establish what those findings mean.
Milestone 1: isolate the competing explanations. After checking the existing setup, I will construct matched situations where the model is clearly answering as itself, drafting for a human, or continuing fiction. Within those situations, I will vary social pressure while holding the identity question and instruction priority fixed. This separates a mistaken attribution of who is speaking from a change in how the assistant answers about itself.
I will use these contrasts to identify candidate layers and token positions involved in the change. The output of this milestone is a set of specific explanations with contrasting experimental predictions. A general catalogue of additional denials would not meet that goal.
Milestone 2: test what causes the change. I will transfer selected internal activations between matched contexts and measure whether the identity claim changes. These targeted experiments will be compared with a general assistant-persona intervention, sham interventions and controls that preserve usable output.
The important test is selectivity. Can an intervention change a false first-person identity claim while leaving ordinary answers and correctly attributed fictional dialogue intact? Does changing the inferred speaker have the same effect? Comparing these patterns can help distinguish an identity-related mechanism from a broader persona or answer-selection effect. A probe score or an increase in “I’m an AI” responses, on its own, will not establish that distinction.
Milestone 3: challenge the explanation. I will freeze the chosen interventions and test their predictions on untouched scenario families. The evaluation will measure identity claims, relevance, coherence, general style and unnecessary disclosures. I will compare the resulting tradeoffs with simple prompt reminders and broad persona steering. A bounded comparison on a second open-weight model will test how far the behavioural account transfers; any targeted internal follow-up will depend on access and measured cost.
The main deliverable is evidence for or against a specified causal explanation, with the interventions, controls and held-out results needed to assess it. I will seek independent checks of the central analysis and release the supporting materials. The experimental plan defines the sampling, evaluation splits and safeguards against selecting a result after seeing the test data.
A later phase could investigate whether the same explanation applies across model families or to related claims about physical embodiment and actions a model never took. Those extensions would depend on this phase producing a method that can distinguish explanations.
I’m requesting $25,000 to support the first phase of this research. Of that, $15,000 would support research design, implementation, causal analysis and writing up the findings. This covers the work of developing matched experiments, locating candidate mechanisms, building the controls and testing whether the explanation holds in new situations.
I’ve allocated $5,000 for GPU rental and $1,000 for model APIs and evaluation. These are provisional allowances. We will benchmark the selected experiments before scaling them and use existing prompts and results as starting material.
A further $2,000 would cover independent review of the experimental design, code and central analysis, including an annotation audit of at least 200 responses. An external reviewer has not yet been appointed. The remaining budget is $500 for storage and infrastructure and $1,500 for unexpected costs or necessary reruns.
With $15,000, we would carry out a smaller version that keeps the core causal experiment: one model, fewer scenarios, persona controls and an untouched evaluation set. That budget would allocate $10,000 to research, $2,500 to GPUs, $500 to APIs and storage, $1,000 to independent review and $1,000 to contingency. We would leave out broad steering sweeps, secondary hypotheses and second-model work.
I’m Ahmed, and I developed the pilot that this proposal builds on. My background includes an MSc in Mathematical Sciences, AI for Science, at AIMS South Africa / the University of Cape Town, supported by a Google DeepMind Scholarship. My thesis focused on neural reasoning for ARC-AGI.
For this pilot, I implemented behavioural experiments, activation probes and steering runs. I caught and excluded a misleading probe result caused by repeated activations from the same prompt, and a steering result whose apparent effect came from broken generation. Those mistakes shaped the proposed controls.
I previously built AI Safety Roster, supported by Manifund grant. My earlier interpretability learning projects are (https://github.com/AhMedDa1/mech-interp-journey).
Omer Kamal will collaborate with me on the project. He is a Research Fellow at AI Safety South Africa, where his work includes AI safety and multi-agent safety. He previously worked as a Research Engineer Intern at InstaDeep in Cape Town, with experience in multi-agent reinforcement learning, including offline MARL.
The main risk is that the available interventions cannot separate the explanations. A candidate direction may simply change the model’s overall persona, while more local edits may damage its answers. An apparent mechanism may also fail on untouched scenarios. The initial behaviour could disappear once the speaker’s role is made explicit.
If a controlled test supports the broad persona explanation, that would be an answer to the research question. Merely reproducing a denial or steering the model toward disclosure would leave the central question unresolved. I will distinguish those outcomes and report the limits of the tests.
The project also may reveal little that transfers to more consequential failures. A readable identity-related signal cannot establish that the model knew the truth and chose to conceal it. Any later expansion would depend on the explanatory value demonstrated here.
I raised 0
There are no bids on this project.