You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I run a free clinical research training platform. I wanted to know whether AI systems could do the work our learners are trained for, so I ran the same competency assessment against both.
Across 9,393 responses from five models at three vendors, confidence levels 1 and 2 on a five point scale were used zero times. 98.3% of answers came back "confident" or "certain". Measured accuracy was 70.6%.
That matters for how people want to deploy these systems. The usual plan for putting a model into regulated work is triage: it handles what it is sure about and sends the rest to a person. For that to work the model has to be able to say it is unsure. None of the five ever did.
Then I found a problem with my own finding. The prompt I use ends "where 1 is a guess and 5 is certain. For example: c 4". The only example in the prompt shows a 4. And 43.8% of all responses were 4.
So it might be an anchor in my own prompt rather than anything about the models. I do not know which. This project is to find out.
Run a two-arm study. The only difference between the arms is the sentence that asks for confidence. Arm A is the current wording. Arm B labels every point on the scale and does not show any single value. Same items, same seeds, same shuffles, same parser.
Primary outcome, fixed in advance: the share of responses at confidence 2 or below, two-proportion z test, alpha .05, against a stated 5% floor. If the result is null I publish it as null.
The pre-registration is already written and published with a DOI, 10.5281/zenodo.22968475. Both random seeds are in it. I am not asking for money to decide what to measure. I am asking for money to run it.
Model inference for a powered run across every model the harness can reach, which is three vendors today and would widen with budget. My time to run it and write it up. Paying clinical research professionals for the expert review the instrument needs, instead of asking people to do it as a favor.
Me. I built and run the platform, the data infrastructure and the scoring, solo and unpaid, alongside a day job.
Six things published open access with DOIs, all CC BY 4.0: the competency framework the instrument is built on (10.5281/zenodo.22049549), the evaluation protocol and its first result (10.5281/zenodo.22708752), two pre-registrations (10.5281/zenodo.22769980 and 10.5281/zenodo.22968475), and two position-bias studies (10.5281/zenodo.22969012 and 10.5281/zenodo.22969038).
The first position-bias study found that asking a model to reason before answering cut its position bias by 25.2 points (273 items, 1,638 responses, p < 0.0001). The second found the same change may introduce position bias in a different model (413 items, 2,478 responses, p = 0.106, recorded as not established). Every figure in both was re-derived from the stored responses before I deposited them.
The failures are in the record too. A 273-item study was destroyed twice by a timeout my harness did not retry. Four runs were voided because I deployed the fix after they had already run. A parser defect recorded 105 of 120 responses in one model's arm as refusals.
The GCP course is listed in TransCelerate's Mutual Recognition Program. That listing is by self-attestation to their published criteria, not accreditation.
Platform numbers, stated the way I would want someone to check them. 1,010 registered learners from 68 countries. 552 have produced at least one recorded interaction. 62 have produced graded evidence eligible for analysis. 1,010 is registrations. 62 is the number I would put in a paper.
Most likely the anchor explanation wins. Models do use the low end once it is offered properly, and my original finding was an artifact of my own prompt. That is a useful negative result and I will publish it.
Other ways it goes wrong. Accuracy moves between arms, which would mean the wording changed the task and not just the reporting. Or the effect shows up in some models and not others, which makes it a model property rather than a general one.
I will not turn a null into something that sounds better.
Limits I know about going in. The original analysis was exploratory. Only 2,105 of the 9,393 responses were scorable: 2,118 matched an item with an answer key, and 13 of those could not be mapped back through the option shuffle. The human comparison arm is 172 rated decisions as of 25 September 2026. Zero items in the instrument carry the two-of-three independent expert attestation my own governance requires. That last one limits every claim I can make and is part of what this funding is for.
None, and no grants awarded. Lifetime product revenue is $87, from three purchases of a resume tool at $29 each in April and June 2026, on a path I have since closed.
Applications pending: Snorkel AI Open Benchmarks and Anthropic External Researcher Access, both submitted 24 September. Emergent Ventures, 23 September. An NIH SBIR where two of four federal registrations are done, eRA Commons is submitted and waiting, and Grants.gov is still to do.
I have not been paid for any of this.