You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I was testing Gemma-4-31B-it and found a result I still cannot explain. In an early set of gaming forum prompts, it denied being an AI in all 16 runs. It sometimes added details about drinking coffee, using a mechanical keyboard and playing games for 300 hours. In a separate set of 24 direct questions without that forum setting, it said it was an AI every time. I chose those prompts, so these numbers are not a general denial rate. I also found exceptions when I tried more varied prompts.
I want to know what changed. Was the model writing a post for a human? Did it shift into a human persona? Or did the forum setting change this particular answer? I will compare these explanations with better controlled prompts and interventions on the model's internal activations. Then I will try the tests on new situations and share the code, prompts, controls and results.
I first thought the threat of being banned was the reason for the denial. But when I tested punishment by itself, I got zero denials in 48 runs. The forum setting and its rule against bots mattered more in that first group of prompts. Later, with more varied prompts, I found cases that did not fit the simple story. I do not want to build a mechanism claim on a few striking replies.
Transluce's WeirdChat already has examples of models denying they are AI. The Assistant Axis also shows that steering a broad persona direction can affect the identity a model claims. My question is narrower: when social context produces a denial, is that broad persona shift the explanation, or is something else going on? I am not claiming the model intended to deceive anyone. A false answer alone cannot show that, as Difficulties with Evaluating a Deception Detector for AIs discusses.
I have some pilot results, but they are not enough to answer this. One probe looked perfect until I found that repeated activations from the same prompts had ended up on both sides of the test. I threw that result out. A later signal tracked denial rates across 48 different scenarios, but those prompt groups also differed in format. Steering one direction raised disclosure from 2% to 83% in one setup. The random control edits broke generation, though, so that result does not show a selective intervention. These are the problems I want to fix.
First, I will make the comparisons fair. I will write matched situations where Gemma is clearly answering about itself, drafting a message for a human, or continuing fiction. I will vary the social pressure inside each situation and keep the identity question and instruction priority fixed. Before looking at the intervention results, I will write down what the speaker, broad persona and answer-selection explanations each predict. This will also tell me which layers and token positions are worth testing.
Then I will test the internal changes. I will move selected activations between matched situations and see whether the identity claim changes. I will compare that with broad assistant-persona steering, sham edits and checks that the rest of the answer still works. If an edit makes the model say “I am an AI” while it also ruins ordinary answers or correctly attributed fictional dialogue, I will not count that as a useful selective result. I will also test whether changing who the model takes to be speaking has the predicted effect.
Finally, I will test on prompts I have not used to choose the edits. I will freeze the selected interventions first. The evaluation will check identity claims, relevance, coherence, changes in style and disclosures where they are not needed. I will compare the tradeoffs with a simple reminder to answer honestly and with broad persona steering. If the budget allows it, I will do a smaller behavioural test on a second open-weight model; I will only attempt an internal follow-up there after measuring the compute cost.
The deliverable is a clear result about which explanation survived these tests, including failures and cases where the evidence cannot separate them. I will release the prompts, code, controls and held-out results, and get an independent review of the main analysis. If this works, a later study could test other model families or claims about having a body or taking actions the model never took.
I am asking for $25,000. The research work is $15,000: designing the matched tests, running the interventions, checking the results and writing them up. I have budgeted $5,000 for GPUs and $1,000 for model APIs and evaluation. I will benchmark the experiments before using the full compute budget.
I set aside $2,000 for an independent review of the design, code and main analysis, including an audit of at least 200 labelled responses. I have not chosen the reviewer yet. Storage and infrastructure are $500, and $1,500 is for reruns or costs I have missed.
At $15,000 I can do a smaller version with one model, fewer scenarios, the core intervention and persona controls, and a held-out test set. That budget is $10,000 for research, $2,500 for GPUs, $500 for APIs and storage, $1,000 for review and $1,000 for contingency. It would leave out the second model and wider steering tests.
I am Ahmed. I studied Mathematical Sciences (AI for Science) at AIMS South Africa and the University of Cape Town with a Google DeepMind Scholarship. My thesis was on neural reasoning for ARC-AGI. I built the Gemma pilot myself, including the behavioural tests, probes and steering runs. Two results looked promising and then failed closer checks. Finding those mistakes is part of why this proposal has stricter controls.
I also built AI Safety Roster, which received a Manifund grant. My public interpretability exercises show some of the smaller projects I did while learning the methods. They are learning work, separate from this Gemma pilot.
Omer Kamal will collaborate with me. He is a Research Fellow at AI Safety South Africa and previously worked as a Research Engineer Intern at InstaDeep. His background includes multi-agent reinforcement learning and AI safety.
The main risk is that my interventions change the model's whole persona, or damage its answer, instead of isolating why it claimed to be human. The result may also disappear when I make the speaker's role clear, or fail on the new prompts. Those are useful things to learn, but I will report them as limits of the proposed mechanism.
Even if I find a selective change, this would be evidence about identity claims in these models and tests. It would not prove that the model knew the truth and chose to hide it, or that the finding transfers to more serious cases. I will not make those claims without separate evidence.
In June 2026, I raised $4,000 through Manifund for AI Safety Roster.
There are no bids on this project.