You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
My prototype knowledge engine, CopernicusAI, provides the infrastructure for an AI research pipeline that's been operating for two years. In the most recent iteration, three AI agents (Claude Code, Claude Chat and Cursor) supervised and directed by me do scientific research and write papers. The agent roles include finding, ingesting, indexing and embedding carefully selected specific papers from sources like PubMed and BioRxiv in a research paper database that provides semantic search and retrieval augmented generation. One research line, the Genome Logic Modeling Project, GLMP, is a Claude Project directed at analyzing genomic regulatory structure. The agents generate biological process flowcharts and run "in silico" experiments on virtual cell models using Google Colab. The whole system runs under a written constitution, an enforced behavioral spec, and an explicit division of authority between the AI agents. That governance layer has already caught real mistakes before they became false scientific claims. This project turns that real, dated track record into a documented protocol and a measured benchmark.
This project will generate three concrete outputs:
A documented, general governance protocol including a constitution, behavioral spec, and agent-role authority map, written up so another AI-assisted research pipeline could adopt the pattern, not just read about it.
A fabrication-detection benchmark. I already have real, naturally occurring error catches (a motif-library mismatch that fabricated regulatory "logic gates" from statistical noise; a bug that let a claim outrun its evidence; enforced abstention when evidence is missing). To execute this project, I'll catalog the naturally occurring catches and also build a larger set of deliberately induced fabrication scenarios. Then I'll measure what fraction of the errors the governance layer actually catches, a first quantified detection rate, not just anecdotes.
A published, open-sourced writeup connecting this research to the harder future problem of managing and validating autonomous AI research systems. Those systems will be able to propose questions, make hypotheses, run their own experiments on virtual cell models and write up the results in publishable form. Such systems will need this kind of self-verification built in, vetted and trusted.
My time and compute. Protocol formalization and benchmark construction is the core cost. The rest goes to compute (Colab Pro/Google Cloud costs for running the benchmark) and the costs of publishing and open-sourcing the result. At the minimum funding level of $15K, I'd deliver the protocol and benchmark only; the higher end of $40K adds proper time for the open-source release and a fuller case-study writeup.
Just me, directing three or more AI agents that do the research and writing under my supervision. I'm an independent researcher affiliated with the CUNY Graduate Center New Media Lab. I built and operate CopernicusAI, including the governance layer this project formalizes. That work resulted in a peer-reviewed paper,
"AI-powered knowledge engines as research infrastructure for systematic knowledge discovery" (Discover Artificial Intelligence, Springer Nature, July 2026,
https://doi.org/10.1007/s44163-026-01780-5
My LinkedIn profile:
https://www.linkedin.com/in/garywelz/
My Google Scholar Profile:
https://scholar.google.com/citations?user=3wTcI6EAAAAJ&hl=en
Code and project pages:
https://github.com/garywelz/glmp
https://github.com/garywelz/copernicus-web
https://huggingface.co/spaces/garywelz/glmp
https://huggingface.co/spaces/garywelz/copernicusai
ORCID:
https://orcid.org/0009-0005-7806-0892
The most likely failure mode isn't "the governance layer doesn't work." Instead it's that the benchmark turns out narrower or less generalizable than hoped because induced fabrication scenarios in a genomics pipeline may not transfer cleanly to other domains. This would mean the "general protocol" ends up more GLMP-specific than I'd like. A second risk is scope: as a solo researcher, if the benchmark construction takes longer than planned, the open-source release and writeup could slip past the funding period. If the measured detection rate comes back low, I'd treat and report that as a real finding about the limits of documentation-based governance, rather than not publishing it.
This work has been entirely self-funded — about $2,400 in cash costs over about two years (roughly $200/month for APIs, Google Cloud Services and AI services) from personal savings, no external grants received to date. I have an NSF proposal submitted through CUNY Graduate Center still pending a decision, and a related but smaller/differently-scoped application to EA Funds' Transformative AI Fund currently under evaluation (submitted Aug 11, 2026).