You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I built RIGOR because I kept running into the same problem while writing empirical papers: the criticism that really changes a paper usually arrives before submission, but access to that kind of criticism depends a lot on where you are, who is willing to read your drafts carefully, and sometimes simply whether you can afford to pay someone.
RIGOR is my attempt to make part of that process less scarce. It takes a manuscript and a fixed budget and tries to reconstruct what the paper is claiming, what evidence those claims depend on, and what has to be true for the argument to work. It then checks those pieces separately and, importantly, it does not treat every suspicion as a finding. If something looks wrong but the system cannot establish it, it stays as an unresolved lead.
The system already exists and runs on real papers. What I do not know yet is the harder and more interesting thing: how often is the criticism actually right?
This project is basically an attempt to get that answer rather than relying on a few good-looking reports.
The main goal is simple: when an AI system says an empirical paper has a scientific problem, I want to know how often it is actually correct.
I plan to assemble a few hundred papers in economics, political science, and nearby empirical fields for which some outside criticism already exists, ideally referee reports, replication reports, or documented expert critiques. When possible I will use working-paper versions that predate that criticism, so the system is not being tested on a version that already incorporates the referee's objection.
RIGOR will run on those manuscripts under different compute budgets and, for part of the sample, under different model families. Its findings will then be judged by independent human raters who do not know which experimental condition produced them.
The main things I want to measure are the false-positive rate, how many externally identified problems the system recovers, whether the criticism is actually material, how much raters agree with each other, and whether spending more compute really buys better scientific criticism.
I care especially about false positives. Missing a problem is bad, but confidently telling a researcher that something is wrong when it is not can be worse.
I am seeking $25,000.
Most of the money would go to the part I cannot automate away: independent human judgment.
My current budget is roughly:
$9,000 for blinded expert adjudication
$5,000 for building and annotating the manuscript/criticism corpus
$4,000 for subsidized audits for researchers who otherwise would not pay for them
$3,000 for the public benchmark and scoring infrastructure
$2,000 for model inference not covered by other research-credit programs
$2,000 for project time, quality control, and administration
If the project receives less than the full amount, I would reduce the number of manuscripts rather than cut the human evaluation. A smaller study with credible adjudication is more useful to me than a bigger sample with weak ground truth.
Right now I am the team.
I am an economist and Research Professional at the Center for the Economics of Human Development at the University of Chicago, where I work with James Heckman. Before Chicago I worked at the Inter-American Development Bank, the International Monetary Fund, and Georgetown University. I studied economics at San Marcos in Peru and later at the Harris School.
My research is mostly empirical economics and political economy. I have a coauthored paper forthcoming in the Journal of Labor Economics, and I presented related work at the 2026 NBER Summer Institute in Behavioral Macroeconomics.
But the most relevant track record here is probably that I built RIGOR myself. It is already live and has completed production audits on real empirical papers. One audit of an approximately 80-page economics manuscript used $8.56 of inference from a $10 budget and returned a 24-page structured report in about 38 minutes.
That shows me the system can run. It does not show me that the system is right, and that distinction is basically the reason for this project.
The obvious failure is that RIGOR is simply less reliable than I think.
It may generate too many false positives. It may miss things that human referees catch. It may work quite well on some designs and badly on others. It may also turn out that the underlying model matters much more than the verification architecture I built around it.
Any of those would be disappointing for RIGOR as a product, but scientifically they are still useful results, and I would publish them.
There is another problem, which is that referee reports are not ground truth either. Referees disagree and they miss things. That is why I do not want to evaluate RIGOR by simply asking whether it reproduces a historical report. The project includes independent human adjudication because the old referee record is evidence, not truth.
If the project fails in the strongest sense, meaning the system's criticisms are too unreliable to be useful, the outcome should be a public benchmark showing exactly that and giving other people a better standard than "the review sounded convincing."
$0.
I built and launched RIGOR without outside funding.
I currently have several grant and research-credit applications pending, but none has been awarded yet, so I do not count any of them as money raised.