You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I'm building GEOSTATE, a geopolitical simulation game where countries operate inside a persistent world with an economy, diplomacy, politics, trade and sanctions.
While building the game, I realized that the same deterministic simulation can also be used to evaluate LLM agents over long periods of time. Instead of a human controlling a country, an LLM can receive information about its country, make decisions, and then experience the consequences of those decisions over many simulated months.
I did not originally set out to build an AI benchmark. This came out of building the game.
The evaluation infrastructure already exists. GEOSTATE currently has a multi-provider model router, 10 agent action types, 10 reproducible scenarios, a six-dimension evaluator, and deterministic replay. So far I have tested it using mocked model calls, not real frontier models.
This project is a small experiment to find out whether GEOSTATE is actually useful as a long-horizon agent evaluation environment.
The main goal is simple: run the existing evaluation system against real models and see what it produces.
Over 4–6 weeks I plan to:
run the existing 10 scenarios against several frontier models;
compare how the models perform across fiscal, economic, social, political, diplomatic and reliability metrics;
check whether the scenarios actually reveal meaningful differences in long-horizon decision-making;
document the methodology and limitations;
publish the results and a small comparison/leaderboard.
The important part for me is that this is an experiment, not a predetermined conclusion. If the results show that GEOSTATE is not a useful benchmark, I will publish that too.
I'm asking for $8,000.
Most of the funding would give me 4–6 weeks of uninterrupted time to run the experiment, handle problems that only appear with real models, analyze the results, and prepare the public write-up.
A smaller part would pay for model API usage, hosting and other costs related to running the evaluation.
The simulation and evaluation infrastructure itself is already built, so the grant would not be paying me to start a new benchmark from scratch.
Approximate budget:
$7,250 — developer/research time
$250 — model API usage
$500 — hosting, retries and contingency
I'm the only person working on the project.
I'm 19, Ukrainian and currently living in Italy. I've been building GEOSTATE solo using AI-assisted development.
I don't have a traditional research background or an AI safety track record. My relevant track record is the system itself: I built the geopolitical simulation, deterministic/versioned world simulation, agent interface, model-provider integration, scenario system, evaluator and replay infrastructure that this experiment will use.
GEOSTATE is already a working software project, not an idea that I would begin building after receiving funding.
The most likely failure is that the environment works technically but does not reveal anything particularly useful about current models.
For example, the scenarios may be too simple, the scoring system may not capture the interesting differences between agents, or the geopolitical setting may turn out to be less useful for evaluation than I expect.
Another possibility is that real models behave in ways that expose weaknesses in the current agent interface or simulation assumptions.
I still think that would be a useful result. The purpose of this grant is to test the idea cheaply before spending months turning it into a separate product.
If the experiment fails, GEOSTATE continues as a game and I will publish what I learned rather than pretending the benchmark worked.
$0 in external funding.
GEOSTATE has not received investment or grant funding so far. I have been building it independently, and this would be the project's first outside funding.