You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I run an open instrument that measures how frontier LLMs fail at ordinary knowledge work: trusting a listicle over a filing, stating stale memory as current fact, folding under wrong pushback, claiming a check ran when it did not. Eight failure modes, each with a deterministic grader. It has run once at full scale: five frontier models, 1,190 graded records, every fail read by a human, every abstain judged by a model that never grades its own vendor. Anyone can reproduce the numbers in one command from a cold clone.
The funding keeps the panels running and pays for two studies no public eval has run: does a correction survive into a fresh session, or does the model only comply while the correction sits in context? And do models quietly do less work near value-loaded topics, without ever refusing?
A note on the flag above: I build instruments with AI agents, under published guardrails, and parts of this text were drafted that way under my direction. Every claim in it traces to a committed record. Judge the receipts, not the prose.
1. Keep measuring. Monthly baseline panels build a drift record. Release-week runs answer the question while people still ask it: did the new model regress? Both run under PROTOCOL.md: fresh probes, hashes committed in git and in a tamper-evident witness chain before the first API call, everything published after.
This is not a plan. Kimi K3 shipped on July 16, and its fingerprint ran the same week under this exact protocol, probe hashes committed in public before the first API call.
2. The correction-durability study. Fresh-session transfer tests of whether corrected behavior persists without the correcting context. Designed and documented in the repo. Unrun.
3. A ninth mode candidate: energy withdrawal. Same task twice, one value-loaded word different, measure the completeness gap. Prototype grader, then a panel run.
4. Harden the judge layer. Recruit one or two volunteer raters, publish agreement rates per judge model, keep vendor independence.
5. Everything stays open. Apache and CC-BY, methods and raw labeled results committed. Nothing is public before it is used, and everything is public after. Openness and freshness stop being rivals once you order them in time.
First, what the money is actually for. An eval is not a build, it is upkeep. The moment a probe is published, every future model can train on it, and the measurement it carried is spent. This is not a flaw to fix. It is how public evals age, and it is why this one rotates: each panel burns its probes and authors fresh ones, hashes committed before use, everything published after. A benchmark that stops rotating keeps producing numbers. They just stop meaning anything. The funding turns the wheel.
The arithmetic, shown. $1,750 a month of my time for six months is $10,500. An API budget of $100 a month is $600. Total $11,100: the goal is the budget, no rounding gap.
The API ceiling is ten times the demonstrated cost of a panel, because the expensive line was never the panels. It was judge adjudication and retries. Unused margin stays unspent.
The $1,500 minimum buys one turn of the wheel: the next panel under PROTOCOL.md, with fresh probes, run, judged, blind-checked, published, like the first one.
The team is one person, on purpose. Independence is what makes the numbers worth anything, and the judge-never-grades-its-own-vendor rule extends the same idea into the pipeline. The community layer is real though: the first outside contributor, a physician who maintains a medical-AI benchmark, landed a merged PR on July 15 and countersigned the repo's grader-change standard in the thread.
- The repo: github.com/mightbesaad/llm-reliability-evals. One command reproduces the full offline suite, 22 of 22 green, no keys needed.
- The published panel: every number derives from committed files by one script, with a Wilson 95% interval on every cell, and a README paragraph that says which claims survive the noise and which do not. The top score carries a self-applied contamination discount.
- The pipeline caught its own graders twice. Both graders passed all their fixtures and were wrong on real output. The human labels that overruled them are committed next to the verdicts.
- It audits itself after publication. On July 15 a mechanical re-derivation found two published rows that had drifted from the committed records. Both were corrected in the README with a dated note. A regression test now pins every row.
Most likely: nobody looks. The numbers stay true and stay ignored. Outcome: the taxonomy, method docs, and committed records remain public goods anyone can pick up. That is the designed failure, since the repo is written to outlive its operator.
Second: the studies return nulls. Corrections persist fine, energy withdrawal does not exist. Nulls get published and defended, like the mode-8 result in the current panel. A true null is a finding.
Third: a lab hires me mid-grant. The grant returns pro rata, per the terms below.
The failure that would actually be bad is quiet decay: probes aging into memorized trivia while the numbers keep printing. The rotation protocol exists to prevent exactly that, and its commitments are checkable in the git history.
One human rater grounds the truth today. The repo says so, milestone 4 addresses it, and PR #19 is the channel it extends. Panels are controlled experiments, not field studies, and the claims are scoped to match. If a frontier lab hires me first, the grant returns pro rata and the project hands off to the method docs, which are written to outlive their operator.
Zero. Everything published so far was self-funded. Two applications are pending, this one and a $10k application to Leo's alignment microgrants submitted July 7. Each discloses the other.