@saad
I build instruments that hold AI agents accountable: open reliability evals, tamper-evident audit logs, spend caps.
https://mightbesaad.com$0 in pending offers
The main work is llm-reliability-evals, an open suite measuring how LLMs fail at ordinary knowledge work. Five frontier models, eight failure modes, every verdict traceable to a committed record. I also build audit-event-mcp, a tamper-evident log of agent actions, and gvnr, hard spend caps for agents. On X: @mightbesaad.
pending admin approval