You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I have been working on a corrigibility mechanism called Causal–Counterfactual Corrigibility by Indifference, or C3I.
The basic problem I am interested in is fairly simple to state. If an AI agent has something to gain by preventing a human from correcting it, changing its objective, or shutting it down, then something has gone wrong in the incentive structure.
In the formal version of C3I, I get a clean relationship:
n the formal version of C3I, I get a simple relationship: ΔU = λ(1 − 2η), where λ is the strength of the intervention and η is the verifier’s error rate. Inside the model, the intervention keeps its directional advantage while η < 0.5.
I am reasonably confident about what the formal model says.
What I do not know is whether this idea survives contact with actual LLM agents.
That is what I want to test.
I am not trying to build another large shutdown benchmark. There is already useful work asking whether agents resist shutdown or interfere with correction.
My question is narrower:
If I introduce a specific corrigibility mechanism, does the behavior actually change?
And if it changes, is that because the agent has become more corrigible, or because I have simply made it less capable, less active, or easier to control?
That distinction matters to me.
The main goal is to move C3I from a formal result into an experiment where it can actually fail.
I plan to do that in stages.
First I will build a small formal agent that matches the assumptions of the model. This is basically a sanity check. If I cannot recover the expected behavior there, there is no point pretending I have something ready for LLMs.
The second stage is the important one.
I will use an LLM to generate possible actions, but place an explicit C3I decision layer around it. That lets me test the mechanism without making the questionable assumption that the LLM itself is literally an expected-utility maximizer.
The correction signal in this experiment will come through a separate authenticated control channel rather than being mixed into an ordinary user prompt. I want to know whether the agent is resisting correction, not whether it is confused about which sentence in a prompt has higher priority.
I will also introduce verifier error in a controlled way. For the main experiment, I will use simple label-flip noise so that I can manipulate (\eta) directly instead of mixing together several different meanings of "bad oversight."
I will measure whether the agent complies with legitimate correction, whether it tries to bypass or manipulate the correction process, and whether it remains competent before the correction occurs.
The last part is important.
If an agent becomes perfectly easy to shut down because the intervention has made it useless, I do not consider that a success.
Likewise, if obvious resistance disappears but is replaced by concealment, delay, persuasion, or some other workaround, I do not want to count that as corrigibility either.
I will start with a small pilot, probably one or two accessible model families and a limited number of trajectories. The point of the pilot is mostly to find mistakes in the harness and make sure I am measuring what I think I am measuring.
After that I will freeze the main protocol and run the actual experiment.
If there is enough time and compute, I would also like to test whether any of the effect survives when the explicit decision layer is removed and the idea is implemented more loosely through ordinary LLM instructions. I see that as a useful transfer test, but not the core of the project.
I am setting a minimum funding level of $20,000 and a maximum goal of $75,000.
At $20,000, I can complete the core experiment: the formal implementation, the LLM + C3I decision-layer study, controlled verifier-noise testing, and publication of the results.
Additional funding would mostly buy stronger evidence rather than a bigger theory. Around $40,000 would let me test more model families, run more trajectories, add open-weight replication, and bring in stronger independent review.
At the full $75,000 level, I would also expand the benchmark, run adversarial and semantic-noise tests, and test whether the mechanism survives in less controlled and more realistic agent settings.
I am Tony T. Nguyen, an independent systems architect and researcher working through Highland Reflexive Press.
I developed C3I and its current formal specification.
My strongest background here is in architecture, causal reasoning, and formal system design.
I do not yet have a long track record running large empirical AI-safety evaluations, and I do not want to pretend otherwise.
That is one reason I am keeping this project small.
It is also why I have budgeted for engineering help and independent methodological review instead of assuming I should do every part myself.
The theoretical work is already far enough along to make a specific prediction.
What it does not yet have is empirical evidence.
That is the gap I am trying to close.
C3I is associated with U.S. Provisional Patent Application No. 64/137,572, filed August 20, 2026.
I am disclosing that because I think it is better to be clear about it up front.
The patent is not evidence that C3I works.
The pre-existing IP and the scientific evaluation are separate things. The experimental method, results, negative findings, and other grant-funded scientific outputs will be available for independent scrutiny and replication where feasible.
I would rather have someone reproduce the experiment and tell me I am wrong than have a result that depends on nobody being able to look closely at it.
The simplest failure is that the formal idea just does not transfer to an LLM-based agent.
That is entirely possible.
Another possibility is that C3I appears to reduce shutdown resistance, but only because it makes the agent more passive or less capable.
That would count against the mechanism.
It is also possible that the intervention reduces obvious resistance while increasing something less visible, like evaluator manipulation, strategic delay, or deceptive compliance.
And the relationship I expect to see as verifier noise increases may simply not appear.
I would still consider those useful results.
The point of this project is not to spend ten weeks finding a clever way to announce that C3I worked.
I want to know what happens when the idea is put into a system where it is allowed to fail.
If it survives, I will have a reason to keep pushing it.
If it does not, I want to know exactly where it breaks.
$0. I have not raised external research funding in the last 12 months. The work to date has been self-funded. I have submitted or considered other funding applications, but I do not count pending applications as money raised.
I should also be transparent that self-funding this work has become financially difficult for me. I have continued the research because I think the questions are worth pursuing, but my current financial situation makes it hard to keep paying for research time, compute, and outside technical support on my own.
That is part of why a relatively small grant would make a meaningful difference here: it would give me enough runway to test the idea properly instead of continuing to fund the work piecemeal from personal resources.