Manifund foxManifund
Home
Login
About
People
Categories
Newsletter
HomeAboutPeopleCategoriesLoginCreate
claudiodegenua avatarclaudiodegenua avatar
Claudio De Genua

@claudiodegenua

Independent AI evaluation researcher · ISO 9001/14001 quality auditor · teocentro.com

www.linkedin.com/in/claudiodegenua
$0total balance
$0charity balance
$0cash balance

$0 in pending offers

About Me

I measure how language models handle truths they already know, and I calibrate the judges that do the measuring before trusting their verdicts.

  1. The veil (link to the article) — five models, verifiable claims from ancient religious history, tested at the quiz, in prose, under pressure and on the sources they cite.

  2. precorrect-method (GitHub link, DOI 10.5281/zenodo.22345121) — the judge's public record: blind results next to the baselines that beat them, and a dated ledger of every number I withdrew.

  3. A six-layer parallel Psalter (DOI 10.5281/zenodo.22770451) — 2,469 verses with per-layer measured quality. Before this I worked as a process and quality-systems analyst: ISO 9001/14001 internal and supplier audits, process mapping and standardisation across EMEA. I started in environmental biomonitoring (degree in Environmental Sciences). The habit I bring from there: calibrate the instrument before measuring the part. ORCID 0009-0008-5896-3172.

Projects

Are LLM judges usable at all? A sealed, pre-registered reliability bench

pending admin approval

About ClaudioBy Claudio

Nothing yet.