You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I am working on a research question that became more important for me after several of my previous experiments.
The basic problem is simple to explain. An AI system can change internally while still showing similar behavior from outside. If this happens, how much can we really know about what changed?
Most evaluation methods mainly observe outputs, behavior or internal activations. I want to test something more direct: if we actively intervene on the system, can we get information that passive observation cannot give us?
This project will compare passive observation with controlled interventions in artificial systems where I know the hidden change in advance. Then I can test if the method really identifies something, or if it only looks successful.
I am not trying to claim that I already solved interpretability. Actually, some of my earlier results were strong in one setup and then weaker or disappeared in another setup. This is one reason why I want to study the limits more carefully.
My main goal is to understand when intervention gives real extra information about an adaptive AI system.
I will build controlled test environments where I can change things like representation, policy, memory, routing or internal state. Because I know the real hidden transformation, I can compare what the auditing method says with what actually happened.
I want to test three situations:
when the hidden change can actually be recovered,
when we can only reduce uncertainty but not recover one exact answer,
when the correct answer is simply “we cannot identify this from the available evidence.”
I also want to compare this with passive evaluation. If intervention does not improve anything, that is also an important result.
I will use held-out tests, fixed scoring rules and reproducible experiment records. I want negative results to stay in the project, not disappear because they do not support the original idea.
If this works, the final output should be a small open benchmark and experimental framework that other AI safety researchers can test and criticize.
How will this funding be used?
The funding would mainly give me enough time and infrastructure to run this as a serious research project for about one year.
My current plan for the full $50,000 is approximately:
$18,000 for my research time
$12,000 for compute, model/API costs and storage
$8,000 for research engineering and benchmark development
$5,000 for independent methodological or statistical review
$3,000 for replication or adversarial testing
$2,000 for documentation and open research infrastructure
$2,000 for unexpected experimental costs
I may also use part of the research budget for scientific travel if it directly helps the project, for example meeting collaborators, attending a relevant AI safety event, presenting the work or organizing external validation.
If I only receive the minimum funding, I will make the scope smaller. I would focus on the core intervention-versus-passive-observation experiment and do most of the engineering myself.
Who is on your team? What's your track record on similar projects?
At the moment I am leading the project myself.
I am the Founder and Project Director of the Reality & Agency Research Initiative (RARI). My research is mostly about perturbation, identifiability, recoverability and how much we can infer about hidden systems from limited observations.
I have already built several generations of experiments around these questions. Some experiments produced strong results, some produced mixed results, and some failed. I kept the failed cases because they changed how I think about the problem.
This project is partly coming from that experience. I became less interested in trying to prove one large theory, and more interested in asking a narrower question that can be tested directly.
For parts that need extra expertise, I may work with contractors or collaborators for research engineering, statistics or independent replication.
Research overview video:
https://youtu.be/Jqz9QY80SZ8?si=-DsFqlCq1h4UvOfN
ORCID:
https://orcid.org/0009-0002-5884-3660
Website:
https://humanization.tech
LinkedIn:
https://www.linkedin.com/in/emirhan-yildirim-research/
I think there are several realistic ways this project can fail.
The first one is that intervention may not give enough extra information. Maybe passive observation already contains most of the useful signal, or maybe the hidden state is simply not recoverable even after intervention.
Another possibility is that the result works in controlled synthetic systems but does not transfer to more realistic AI agents.
There is also a practical risk. The experiments may become more expensive or engineering-heavy than I currently expect.
I would not consider every negative result a useless failure. If we find that some properties cannot be identified even after reasonable interventions, this is useful information for AI safety. It tells us where an evaluator should stop making strong claims.
For me, the more serious failure would be if the benchmark itself is badly designed — for example if a model succeeds because of information leakage, simulator access or some shortcut instead of actually identifying the hidden change.
In that case I would report it and redesign the experiment rather than present it as a positive result.
I have not yet raised external cash funding specifically for this project.
I was accepted into E2B for Research and received $20,000 in E2B compute and infrastructure credits together with E2B Pro Tier benefits. This is not cash funding, so I do not count it as money raised.
I also have several separate funding applications currently under review. If I receive overlapping funding, I will disclose it and adjust the project budget so the same work is not funded twice.