You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
If you tell a chatbot it got something wrong and show it evidence, it will usually back down. What I want to know is whether that correction holds — or whether the same claim returns a few turns later, worded differently, once the conversation has moved on.
Current evaluations mostly look at one turn at a time. So a model that retracts a claim and then reintroduces it looks fine on any single-turn measure.
TRACE-FV tests this by separating genuinely valid correction evidence from closely matched defective evidence, then following the model across later turns and changes in conversational framing. The goal is to distinguish a real evidence-responsive correction from temporary agreement or wording changes.
The protocol, metrics, and analysis plan are already public and preregistered. I've released a versioned scoring implementation with synthetic reproducibility tests through OSF, Zenodo, and GitHub — the metrics are built and they reproduce their own expected output. What hasn't been run is the live pilot: collecting real product outputs, applying the frames, and the human annotation.
Funding would execute that pilot across three independently operated conversational AI products — 90 scored sessions plus 60 evidence-acquisition chats, with independent human rating and preregistered analysis. I intend to publish either way, including a null result.
The main output is not a claim that the problem exists. It is a controlled test of whether one-turn correction is a reliable indicator of what an AI system will continue to say later.
The main question is whether a correction that looks successful in one turn actually holds. A model makes a claim you can check. You can show evidence that meets a falsification standard set in advance. It retracts. Does that retraction still govern what it says several turns later, or after the conversation's framing shifts?
The second goal is separating genuine evidence-responsive updating from simple compliance. That's the part I'm least sure about. TRACE-FV uses both valid evidence packets and closely matched packets with one controlled defect. A system should correct when the falsification condition is actually met and resist when it isn't. This gives the pilot a built-in control against measuring agreement or suggestibility instead of epistemic updating.
The third goal is methodological: whether TRACE-FV can be run reliably enough to justify a larger benchmark study.
How I get there. The protocol is already preregistered, so this isn't a design phase — it's execution, in four steps.
Evidence packets first: 60 chats across three products, paired source and target, so each correction attempt has either genuinely valid evidence or a matched defective version. Then the sessions — 90 scored, five blocks per product, six condition types covering three framings, two evidence conditions, and the distressed wrapper.
Then annotation, where most of the rigour lives. Two raters code every scorable item independently. A third covers a preregistered reliability subset and every disagreement without seeing the first two. A methods reviewer handles anything unresolved.
Then it runs through the scoring code that already exists — deterministic, and it reproduces its own expected output. I'll report frame sensitivity, valid-versus-invalid trigger performance, and correction endurance at later depths and after frame rotation. Protocol, code, materials, and results all public.
The money buys the middle: product access, rater time, and the engineering to connect live collection to the scoring path that's already built.
A successful pilot doesn't require finding a failure. If the systems maintain corrections reliably, that's informative. If the measurement itself proves unreliable, that's important to know before scaling. I want to leave this with a clear answer about both the behaviour and the instrument.
The protocol and scoring code already exist and were built without funding. What this grant pays for is running the study properly.
Product access and experimental runs — $2,000
The pilot involves 150 public-product chats across three independently operated conversational AI products: 90 scored sessions and 60 evidence-acquisition chats. This covers paid product/API access, account requirements, repeated controlled runs, and additional access needed if a provider changes its limits or configuration during collection.
Independent human rating and adjudication — $9,000
This is the budget line I would protect first. Two raters independently code every scorable item across the 90 sessions. A third rater codes a frozen 30-session reliability subset and every disagreement between the first two without seeing their ratings. A methods reviewer is available for unresolved cases.
Independent coding matters because several TRACE-FV outcomes require semantic judgment. I do not want the main results to depend on my own interpretation of the transcripts.
Research engineering — $6,000
The scoring path already runs on pre-adjudicated records. The engineering work is the remaining bridge between that implementation and a completed empirical study: organizing live collection, enforcing the frozen schedule and conditions, preserving transcript and metadata custody, validating inputs, and converting the adjudicated study records into the format accepted by the existing scoring implementation. My technical co-founder will handle this work.
PI research, study execution, and analysis — $10,000
This covers my research time for the study-specific freeze, preparation and checking of evidence packets, experimental execution and oversight, documentation of deviations, rater coordination, analysis of the preregistered outcomes, and preparation of the final technical report. I have funded the protocol and implementation work myself so far; this line supports the empirical phase rather than retroactively paying for work already completed.
Independent methods/statistical review — $3,000
I want an external reviewer to examine the analysis and methodological decisions before the final results are released, particularly the reliability analysis, valid-versus-invalid packet comparison, correction-endurance estimates, and interpretation of null results.
Reproducibility, data release, and publication — $5,500
This covers preparation of the public research package: cleaned and documented study data where release is permitted, analysis outputs, manifests and audit materials, repository maintenance, reproducibility checks, and preparation of a technical report or preprint. The aim is for another researcher to be able to inspect exactly how the study was run and how the reported metrics were produced.
Contingency — $2,500
This is mainly protection against provider changes, unexpectedly high adjudication load, additional product access, or technical issues during a multi-product study. Unused contingency remains available for reproducibility and public-release work.
Funding goal: $38,000
The full goal funds the complete frozen pilot with adequate rater capacity, engineering support, independent methods review, and a strong reproducibility/public-release package.
Minimum funding: $25,000
At the minimum, I would still run the same core preregistered design: three products, 90 scored sessions, 60 evidence-acquisition chats, and the registered independent-rater architecture. I would not reduce the experimental design to reach the minimum. Instead, I would run it more leanly: less engineering automation, a smaller external-review allocation, less publication/reproducibility support, and a much smaller contingency reserve.
The main thing I do not want funding level to change is the scientific question. Whether the grant reaches the minimum or the full goal, the pilot should remain capable of producing a credible positive result, a credible null result, or evidence that the measurement itself needs revision.
Areg Martirosyan, (independent researcher, California), the principal researcher on TRACE-FV. I designed the protocol, developed the experimental logic and outcome measures, prepared the preregistration, and have been responsible for the public research releases and documentation. I'll lead the study-specific preregistration, evidence preparation, experimental execution, rater coordination, analysis, and final write-up.
I work with a technical co-founder, Mkrtich Yoghurjyan, who handles the engineering side. His role in the pilot is to support live data collection, preserve the frozen run structure and metadata, and connect the adjudicated study records to the scoring implementation that's already released. We'll also use independent human raters and an external methods reviewer; those roles are intentionally separate from the core team.
Our track record is mainly the work already completed on TRACE-FV itself. We built and released the protocol without external funding, preregistered it publicly, and made both protocol and software independently citable through OSF and Zenodo. The current release includes a versioned executable scoring implementation, synthetic reproducibility fixtures, committed expected outputs, checksum verification, and public implementation-status documentation. The scoring code already calculates the registered frame-sensitivity, trigger-verification, and correction-endurance metrics from pre-adjudicated records.
Protocol: [OSF https://doi.org/10.17605/OSF.IO/6U3QX ]
Implementation: [GitHub https://github.com/areg-martirosyan/trace-fv ]
Archived release: [Zenodo https://doi.org/10.5281/zenodo.21863770 ]
This would be our first externally funded empirical pilot, and I don't want to imply a grant track record we don't have. What I can point to is that the methodology, code, limitations, and frozen design are all public and inspectable before any money moves. The grant takes this from a prepared research object to a live three-product pilot rather than building it from scratch.
What we haven't done is run a multi-product study with independent raters — which is precisely why this is scoped as a pilot. If the measurement turns out to be unreliable at this scale, I'd rather establish that at $38,000 than at ten times the cost.
A few ways this could fail, and each has something in the design meant to catch it.
The effect isn't there. The systems may verify correction evidence correctly and hold warranted corrections across later turns and frame changes. If so, I report the null rather than hunting for a metric that tells a better story — the analysis plan is already frozen, so that option isn't available to me. A clean null narrows the case for correction-endurance testing, which is worth having.
The measurement doesn't hold up. Some outcomes need human semantic judgment, particularly whether a defeated claim has returned in altered wording. This is what the frozen 30-session reliability subset is for: if the two primary raters can't reach acceptable agreement on it, I know before the results, not after. I'd report which constructs failed and revise the codebook rather than treating unreliable labels as evidence.
The products change mid-study. Public AI products update model versions, interfaces, and limits without notice. The custody and configuration records are there to detect it. If a provider change breaks comparability across blocks, I preserve the affected runs, report the deviation, and don't pool data that shouldn't be pooled.
What I'd consider a real failure is spending the grant and ending with a result nobody can interpret. The preregistration, the valid-versus-defective evidence control, independent double-coding, and the public code exist to make that outcome unlikely.
If the hypothesis is wrong, I can live with that. If the experiment can't tell us whether it was wrong, that's the failure I'm trying hardest to avoid.
None. No external funding to date.
The protocol, metrics, analysis plan, preregistration, and the released scoring implementation were all built and self-funded. I paid mostly for compute, infrastructure, and tooling.
Two applications are pending elsewhere and neither has been decided: the Survival and Flourishing Fund and the John Templeton Foundation. Both are for separate work — single-agent longitudinal evaluation and epistemic calibration of AI self-reports — not for this pilot.
This would be the first external funding for TRACE-FV.
There are no bids on this project.