Manifund foxManifund
Home
Login
About
People
Categories
Newsletter
HomeAboutPeopleCategoriesLoginCreate
saad avatarsaad avatar
Saad

@saad

I build instruments that hold AI agents accountable: open reliability evals, tamper-evident audit logs, spend caps.

https://mightbesaad.com
$0total balance
$0charity balance
$0cash balance

$0 in pending offers

About Me

The main work is llm-reliability-evals, an open suite measuring how LLMs fail at ordinary knowledge work. Five frontier models, eight failure modes, every verdict traceable to a committed record. I also build audit-event-mcp, a tamper-evident log of agent actions, and gvnr, hard spend caps for agents. On X: @mightbesaad.

Projects

Frontier LLM reliability panels: 5 models already measured, fund the next

pending admin approval