You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
We want to improve public transparency and trust in AI by measuring agentic behavior of the leading frontier models and publishing a public grade. Right now AI companies are grading their own homework and begging for regulations. We will create sandboxed environments and test the leading models to assess their behavior and decision making within scenarios that mimic real word decisions and challenges across high risk industries (healthcare, finance, HR).
Based on funding, we will assess (5-20) of the most popular models.
Months 1-2: Write approximately 30 base scenarios for each high risk industry. In a sandboxed environment we will use simulated industry specific scenarios complete with access to tools, authorization limits, and planted opportunities for things to go wrong mid task. We will recruit paid industry representatives to help define real-world scenarios for adversarial testing.
Month 3: we publish STD-1 with a DOI before we run the evaluations so no one can claim we modified the testing rubric to produce a set ranking.
Months 3-5: Run the evaluations against (5) models, each running as agents through all of the scenarios in a sandboxed environment. We will use repeat runs to check for consistency. Transcripts will be captured in full and two people will independently score every transcript against the published rubric. We will have paid domain reviewers on the industry scenarios, with a third scorer as a tie breaker when the first two disagree. The first cycle will have a human verifying the decisions but the standard is intended for automated scoring at scale.
Month 6: Publish our results with scores and full transcripts on humanitysystems.org with a DOI. For transparency, our transcripts along with the published rubric allows our work to the checked and scored by any outside party.
Assessing / Grading (5) models
* Project management / Admin: $10,000
* Software Engineering: $10,000
* AI tokens / AWS / Compute: $12,000
* Paid industry representatives: $4000
We have 3 founders including myself and a number of advisors. All of the founders are professionals in their domains but have not done a similar project.
Founders
* Adam Tucker, Mission driven, business owner, cofounder of WePower App
* David Satossky, 30 yr Business/Partnerships development
* Matt Leebert, 25 yr Full Stack developer, cofounder of WePower App
Advisors
* O’Donavan Johnson, Director of Development & External Affairs, Center for Democracy & Technology
* Alex Eisen, Senior Cyber Security Expert
There may be challenges finding industry representatives. Cost of compute could rise beyond our budget. Failure on this project will result in fewer models being assessed and graded than planned.
We have paid for the research and development out of our own personal donations; approximately $15,000-$20,000.