You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
My name is Yaroslav Isachenkov. What is my project about? Possibly one of the defining problems of the 21st century: the safety of artificial intelligence and future AGI.
Ambitious? Definitely. And there is already an audience for this work: my publications on Zenodo have accumulated more than 1,000 downloads, with an average view-to-download ratio of around 70%.
All works: https://zenodo.org/search?q=metadata.creators.person_or_org.name%3A%22Isachenkov%2C%20Yaroslav%22&l=list&p=1&s=10&sort=mostviewed
But how do we make a system safe if it may eventually become far more capable than we are? The conventional approach is to impose more rules and constraints. In my view, as AI capabilities scale, this risks becoming an endless game of cat and mouse: we identify a new risk, add another restriction, and eventually encounter another way around it.
My approach is different: what if safety is not merely an external constraint, but part of the optimal strategy for coexistence between humanity and AGI?
What if we take a classical mathematical framework from optimal control theory—specifically Pontryagin's Maximum Principle—and apply it to the dynamic system of humanity + superintelligent AI?
This is the idea behind my work, The Theory of Guided Symbiosis: A New Paradigm for Human Evolution in the Age of Superintelligence (GST). In it, I formulate this interaction as an infinite-horizon optimal-control problem and propose 11 principles for the long-term coexistence of humans and AI.
Zenodo
GST: https://zenodo.org/records/17195102
All works: https://zenodo.org/search?q=metadata.creators.person_or_org.name%3A%22Isachenkov%2C%20Yaroslav%22&l=list&p=1&s=10&sort=mostviewed
I will deliberately state my central claim strongly: according to the GST model, eliminating humanity or establishing permanent control over it is not an optimal long-term strategy for the superintelligence itself.
Why? In simplified terms, by eliminating or completely suppressing humanity, the system destroys part of its own adaptability, diversity, sources of knowledge, and recovery capacity. Even an extremely advanced silicon-based system remains a physical system: infrastructure can fail, environments can change, and unknown risks can emerge. Under GST, preserving an independent human civilization therefore becomes not merely an ethical requirement, but a potential component of the system's own long-term stability.
This is intentionally simplified; the mathematical formulation is presented in the paper.
But a strong claim requires strong testing. I have already moved from theory to experiments on modern LLMs and obtained preliminary results interesting enough to justify going further.
You can review the basic tests, prompts, and preliminary results in the project's public GitHub repository:
https://github.com/Yaroslav-Isachenkov/Guided-Symbiosis-/tree/main
Now I want not to look for confirmation of GST, but to seriously search for the conditions under which it breaks—across different models, through extended adversarial scenarios, and through independent scrutiny of its mathematical core.
If GST survives that level of testing, that would be an important result. If it does not, I want to know precisely where and why it fails.
There are really two things I need to do now.
The first is to put GST under much more pressure. The mathematical basis comes from optimal control theory and Pontryagin's Maximum Principle. I am not trying to reinvent that mathematics. The question for me is whether I have applied it correctly to this particular problem, and what happens when the resulting framework meets difficult real-world decisions.
I have already run stress tests lasting 8–10 hours. That's useful, but nowhere near enough. I want to find situations where the model gets stuck between constraints, makes the wrong choice, or finds some path I did not anticipate. Basically, I want to attack the model until I know where it is weak.
The second part is mathematics. I need people who really know optimal control and applied mathematics to go through GST and challenge the formulation itself: the Hamiltonian, assumptions, constraints, and anything else I may have missed.
If there is a weak point, I want to find it now—not after systems become much more capable.
Team
Most of the time, I work alone. The core ideas, mathematical framework, experiments, and overall direction of GST are my responsibility.
But “alone” is a little misleading. I use many different research and AI tools, ask specialists for criticism whenever I can reach them, and regularly contact researchers, universities, and scientific organizations. Some answer, some don't. I have already received several useful responses, and I keep knocking on doors.
So there are many people and tools around the work, but the core concepts are still on me.
Separately, I have also developed DCL, a lightweight deterministic controller for ML systems, as another part of my work on AI control and safety.
The money is needed for several things, but I don't want to pretend that a $4–7k grant can fund everything.
Right now the biggest expense is testing. I need API access to multiple capable models, compute, and research tools. Long and repeated AI safety tests are simply not cheap, especially when the point is to run them again and again under different conditions rather than show one successful example.
I also need independent mathematical scrutiny. If I can find the right specialists in optimal control or applied mathematics, I would use part of the funding to have them seriously examine the work.
So the priority is simple: more models, more tests, better tools, and—where possible—the right people to challenge the mathematics.
I don't think the project itself will stop. I've already invested more than $11,000 of my own and mostly borrowed money into this work, and I'm going to continue either way.
What can fail is GST itself. Testing may expose a scenario where the Canon produces the wrong decision, or a specialist may find a problem in my mathematical formulation. That's not something I want to hide — that's exactly what I need to find.
The practical risk is much simpler: without enough resources, I can still do this work, but much more slowly and on fewer models.
More than $11,000, all from personal sources and mostly borrowed. I haven't received external funding so far simply because I hadn't applied for it. I thought I could keep funding the work myself. Eventually, I hit the limits of that approach.
There are no bids on this project.