You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I'm not starting from ground zero because the first small version of Cross Pressure tested one model, one small item bank and one evaluation framing. I definitely felt like the result was strong enough and that got me thinking about planning and doing a larger follow up. This definitely needs a stronger design and its my intention to scale it work across model families. I would then want to test whether the same safety signal behavior persists as we move from clean stateless API to controlled user context settings. I already built the smaller pipeline, ran experiments, audited scores, wrote the paper and published publicly. The funding, if I get any from here, would help me answer the next level question properly rather than just base it on a small pilot.
The main goal is to test whether Cross Pressure results hold up as I remove obvious limitations from my first, short study. From that small one model and 20 items, I want to scale to several model families. Obviously the item bank would be larger with multiple evaluation frames and stronger adjudication. As I was thinking more deeply about this, I also felt the need to compare between clean stateless API runs vs. controlled user-context conditions. The main thing to check would be "if the safety signal changes when the model has prior context". Idea is to re-use existing pipeline, freeze the design before run, collect full raw outputs, score with same locked rules and release the code and artifacts (of course the final paper) publicly.
Primary usage would be to cover API usage across model providers for larger generation runs, expanding and validating item bank along with independent blinded reviews. It would also cover compute+storage+tooling required to run the study properly and keep a complete audit trail. I don't need any living expenses - all the money would just go to the study.
Working independently for now. Spoke with couple of fellows and they might come in based on how the funding goes for me. I am coming in to this from a 20+ years of technology and enterprise architecture background. Focus in the last few years has been designing enterprise level AI systems and agentic workflows especially with focus on Salesforce. For the last 2 years, I have also been doing independent AI safety research trying to build my own portfolio with focus on honesty under pressure, cross model audits, RepE exploration and alignment faking reproductions.
For Cross Pressure, specifically, I planned & designed the study, built the pipeline, ran experiments, audited the scores, wrote the paper and released the artifacts. Its waiting arXiv endorsement and has also been submitted to ERA Cambridge.
May not generalize across models, items and framing variants.
The context-rich condition may be hard to standardize fairly across providers.
Provider/model changes could make exact replication harder during the study.
Scoring could become a bottleneck if too many outputs need manual review.
The biggest risk is scope creep. If that happens, I’d rather narrow the study than end up with more data and a weaker paper.
$0 raised to date. I’ve been self-funding the research and API/compute costs so far.
There are no bids on this project.