You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Most (if not all) of the industry-standard LLM sycophancy metrics right now rely on single-run methodologies. But LLMs are inherently non-deterministic, even at temperature 0, so how do we make sure these methodologies are robust enough to reliably detect sycophancy? My hypothesis is that the answer lies in testing across runs, not just within a run, and seeing if the initial measurements hold up.
In my reproduction of PARROT, I ran 5 models through ~1,300 questions, five times each. Sycophancy rates ranged from ~7% on Sonnet to 71% on GPT-3.5. The models landed in the same rank order every single time, and the spread between runs was smaller than the error you'd get from a single run, so re-running the benchmark didn't move the measurement much.
For PARROT, that means reporting one metric doesn't obscure uncertainty, it's actually more conservative than you'd assume. But this metric happening to be stable doesn't mean we shouldn't check whether other metrics hold up across runs too, which is what this grant funds (finishing the SycEval paper repro and audit, generalizing to 2 or 3 more sycophancy benchmarks, and shipping an arXiv preprint in August).
Beyond my project, this gap matters because sycophancy is a material effect of RLHF. As long as models are trained this way, we should expect it to surface and interfere with our ability to measure alignment. If nobody verifies that these benchmarks are resilient between single-run and across-run testing, sycophantic responses can be missed, and that slows progress toward keeping frontier models aligned with human values.
My goals are to complete the SycEval audit, produce the preprint, and extend the audit harness to additional benchmarks (starting with 2 or 3 more, then ongoing as new methodologies are released) with public replication reports. The harness works for PARROT and is flexible and generalizable, but the work to adapt it to other methodologies still needs to be done.
The gap between the minimum and ideal funding milestones is runway, not scope: more months of protected research time to audit more benchmarks before having to go back to conventional employment or consulting to cover costs.
Minimum ($8k):
$6k partial runway over 4 months to protect independent research time
$2k compute to extend the reliability audit across 2–3 additional safety-eval benchmarks. The sycophancy preprint ships regardless; this floor funds the generalization beyond it.
The sycophancy preprint ships regardless; this floor funds the generalization beyond it.
Ideal ($22k):
$16k for one full protected quarter at a rate that offsets consulting income
$5k compute to run the audit across a broader benchmark set
$1k preprint production (results pre-review and editing)
As mentioned above, the SycEval preprint ships regardless, but the ideal case gives more runway and affords higher generalizability potential than minimum.
I'm an applied ML engineer turned AI safety researcher. I run this project solo, with academic supervision and capstone co-authorship from Dr. Abdulaziz Alharbi (GCU).
Before this, I spent five years as an Applied ML Engineer at Deepgram (Speech AI / ASR platform), building production speech and LLM training and evaluation infrastructure.
My proudest work to date, though, was as a Data Product Management / Engineering fellow in the Civic Digital Fellowship (Coding it Forward): I worked with directors at the NIH National Library of Medicine to discover, evaluate, and implement new Common Data Elements (CDEs) to measure COVID-19's impact on underrepresented, vulnerable populations at the height of the pandemic in 2020.
The most likely cause of a functional failure of this project is running out of runway and having to return to conventional employment rather than continuing with independent research.
The potential outcomes of the study itself might appear as failure on the surface, but generally point to more opportunities for investigation:
If a benchmark's methodology doesn't hold up under audit, that's signal that a pivot into why some metrics are cleanly reproducible and scalable and why others aren't.
If everything replicates and scales cleanly, that would be very surprising (suspicious or exciting), but would still need explaining: models could have known they've been evaluated for longer than previously thought, or there's a common thread across these benchmarks pointing at something about alignment that nobody's identified yet.
Like many institutions, frontier labs have structural bias toward trusting their own published methodologies; they're not well positioned or incentivized to go back and audit their own benchmarks for the fidelity that internal and external stakeholders assume. That is what makes independent empirical research in evaluation invaluable.
So far, I've raised $350 from BlueDot Impact in a rapid grant for the SycEval reproduction section of the research compute on this project.