@justinarndt
Independent AI safety engineer & systems architect building mathematically verified evaluation harnesses, MCP cryptographic security tooling, and autonomous agent containment suites.
https://github.com/j-arndt$0 in pending offers
I am an independent systems architect and AI safety researcher specializing in empirical benchmark determinism, autonomous agent containment, and runtime cryptographic verification.
My work focuses on building the foundational, zero-fluff infrastructure required for rigorous frontier model evaluation.
eval-inference-engine - Eliminating non-deterministic scoring variance in capability and safety benchmarks via Monte Carlo perturbation testing.
mcp-shield-audit - Detecting runtime tool poisoning, prompt injection, and schema drift in Model Context Protocol (MCP) server chains using Merkle hash trees.
long-horizon-containment-suite - Standardized benchmarks measuring whether autonomous agents can breach sandbox boundaries or hack reward scorers across multi-hour task horizons.
Coming from a background in rigorous systems engineering and compliance validation, I build software with formal state invariants, property-based verification (Hypothesis), and adversarial test suites.
All of my work is 100% open-source, reproducible, and verifiable at https://github.com/j-arndt
pending admin approval