You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Ai safety benchmarks determine whether a model is considered safe enough for deployment, however, virtually no one questions whether these benchmarks are actually measuring what the claim to measure. This is the problem AISafetyBenchExplorer aims to address. It is a database-driven platform that documents 200+ safety benchmarks metadata alongside their metric definitions, origins, citations, and known weaknesses so that researchers what constitute the safety claims. It is currently in the pre-launch phase. Creation of a public API and the first-ever longitudinal meta-analysis of benchmark validity are the next two milestones that the funding will help to finalise and deliver.
A paper on my preliminary insights from the project is here: https://tinyurl.com/AiSafetyBenchExplorer-Report
A sample worksheet of benchmark catalogue that is evolving into a web-based (yet to be published) application is here: https://tinyurl.com/AISafetyBenchExplorer-Workbook
Github code access available upon request
The project goals is an independent, transparent evaluation layer for AI safety benchmark. The platform will:
Trace every benchmark to its paper, metric definitions, evidence, repositories, versions, etc.
Validate pluralistically by assessing methodological robustness, safety-construct, representativeness, limitations, maintenance, and expert consensus.
enable independent scrutiny through expert reviews, bug/improvement reports, audit trails, ontology annotation, and reproducible verification.
Deliverables (6 months):
public deployment with API and research-data export
ontology pilot where independent reviewers annotate 30-40 benchmarks, with review trails published
meta-analysis of the catalogue, including which safety construct are tested, and whether evaluation methods are evolving or converging towards proxies.
The platform is ready (PostgreSQL, FastAPI, Next.js, extraction pipelines, review workflow) is already build. The funding covers module completion, validation, and release, not development from scratch
Requested funding will cover 3 areas: compute/api, hosting, and expert reviews.
Compute/API - $1,980 (~$330/month x 6). The extraction pipeline uses three models: two independently extract information from benchmark papers, while the third checks and harmonises their outputs. Frontier closed models substantially outperform open-source models at this task in my experiment. I currently pay these costs personally. A budget of ~$330/month is estimated to cover the cost of API credit for 6-months (Approx. $1,980). Covers continuous cataloguing of new benchmarks, re-extraction when papers are revised, and weekly metadata/citation updates (Semantic Scholar, Crossref, GitHub, and HuggingFace) will run a cron job weekly.
Hosting - $2,000. Production deployment: database, API, frontend, and Redis/Celery infrastructure. Keeps the platform online beyond the grant period.
Expert reviews - $2,000 (4 reviewers x $500). The ontology pipeline marks each annotation "LLM Proposed"; reviewers accept, revise, contest, or reject, with decisions preserved in an audit log. The pilot covers 30-40 benchmarks and it involves evaluating extraction quality and improving ontology framework.
Total $5,980. Funding above this, up to $8,000, is reserved for cost overruns such as API price changes, higher launch-time usage, infrastructure scaling, or re-extraction volume.
Just me running the project on a solo, for now. I am a Researcher and Data Scientist. I engineer applied AI systems across diverse domains including law, finance, and health. I hold a dual Ph.D. in Computer Science (University of Luxembourg) and Legal Informatics (University of Bologna). My research interests intersect AI and Law, especially in high-stakes domains such as criminal justice. I designed and engineered this project so far, including the preliminary paper. If funded, I will onboard PhD-level researchers specifically for independent human-review pilot phase.
Google scholar: https://scholar.google.com/citations?user=xy5v9UAAAAAJ&hl=en
The platform works as of now, so technical fail is expected. But, curation and review coordination may take longer than expected due to my full-time job engagement, and the meta-analysis will be delayed. Although, the catalogue and the platform will still be launched, but the problem the project aims to address simply takes longer to achieve. Also, the experts' onboarding for the pilot annotation phase may suffer a potential fallback. In which case, I will either onboard fewer reviewers on shoulder the burden myself. The project will eventually not die, but its delivery could be slow.
None. At my personal cost so far. An earlier CLI-based version that ran on Ollama utilised a private computing cluster which is no longer sufficient for a multi-model extraction pipeline and verification pipeline. No external funding for AISafetyBenchExplorer in the last 12 months.