You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I’m building on a completed first study rather than starting from an idea. Cross Pressure tested one model, one small item bank, and one evaluation framing, and the result was strong enough to justify a larger follow-up but also showed exactly where the design needs to be stronger. This project would scale that work across model families and then test whether the same safety signal changes again when we move from clean stateless API conditions to controlled user-context settings. I’ve already built the pipeline, run the experiments, audited the scoring, written the paper, and released the artifacts publicly. The funding would let me answer the next question properly rather than with another small pilot.
The goal is to test whether the Cross Pressure result holds up when I remove some of the obvious limitations of the first study. I want to scale from one model and 20 items to several model families, a larger balanced item bank, multiple evaluation-frame wordings, and stronger blinded adjudication. I also want to compare clean stateless API runs with controlled user-context conditions to see whether the safety signal changes once the model has prior context. I’ll use the existing pipeline, freeze the design before running, collect the full raw outputs, score them with the same locked rules, and release the code, artifacts, and final paper publicly.
The funding would mainly cover API usage across several model providers, larger-scale generation runs, expanding and validating the item bank, and independent blinded review of unclear outputs. It would also cover the compute, storage, and tooling needed to run the study reproducibly and keep a complete audit trail. I’m not asking for living expenses; the money would go directly into running, checking, and releasing the larger study.
I’m working on this independently for now. My background is in enterprise and AI workflow architecture, with 20+ years in technology and the last several years focused heavily on AI systems. Over the past two years I’ve also been building an independent AI safety research portfolio, including honesty-under-pressure evaluations, cross-model audits, representation-engineering work, and alignment-faking reproductions. For Cross Pressure specifically, I designed the study, built the evaluation pipeline, ran the experiments, audited the scoring, wrote the paper, and released the code and artifacts publicly. The project is now awaiting arXiv endorsement and has also been submitted as a poster to ERA Cambridge.
The effect may not generalize across more models, items, or framing variants.
The context-rich condition may be hard to standardize fairly across providers.
Provider/model changes could make exact replication harder during the study.
Scoring could become a bottleneck if too many outputs need manual review.
The biggest risk is scope creep. If that happens, I’d rather narrow the study than end up with more data and a weaker paper.
$0 raised to date. I’ve been self-funding the research and API/compute costs so far. I currently have a BlueDot Rapid Grant application under review, but no external research funding has been received yet.