Manifund foxManifund
Home
Login
About
People
Categories
Newsletter
HomeAboutPeopleCategoriesLoginCreate
🐼
🐼
Oleksii Simon

@simon9679

Solo creator of DriftBench — an open-source, deterministic benchmark for how faithfully AI tracks a person's changing beliefs.

https://github.com/simon9679/driftbench
$0total balance
$0charity balance
$0cash balance

$0 in pending offers

About Me

Independent builder based in Kharkiv, Ukraine. I'm not a professional programmer — I design AI-evaluation tools and build them by directing AI coding tools, acting as architect and tester. I created DriftBench, an open-source, fully deterministic benchmark (no LLM judge) for whether AI systems faithfully track a user's beliefs, conflicts, and identity across long conversations.

Projects

DriftBench v1.1: an ambivalence metric + 20 cross-domain scenarios

pending admin approval

Comments

DriftBench v1.1: an ambivalence metric + 20 cross-domain scenarios
🐼

Oleksii Simon

16 days ago

For transparency: the noise-decomposition study referenced above was conducted on a separate memory-system evaluation (ES-MemEval). It is included in the repository because it empirically establishes DriftBench's premise — LLM-judge pipelines carry measurable, style-correlated noise — which the deterministic protocol addresses.

DriftBench v1.1: an ambivalence metric + 20 cross-domain scenarios
🐼

Oleksii Simon

16 days ago

Progress update (July 9). Published an evaluation-reliability showcase in the repository (eval_reliability/): a measured three-layer noise decomposition on ES-MemEval (WWW '26) — blind human relabel of the LLM judge (K=20: overall noise 0.10, but per-arm leniency up to +0.33 that reorders the leaderboard), answerer stability (±0.05 on byte-identical inputs), and ingest stochasticity dominating subset-level scores (±0.40 swings between two ingests; format ablation contributed 0.00). Both pre-registrations are included, with recorded opposing predictions that the ablations later adjudicated. Directly relevant to DriftBench's thesis: deterministic evaluation exists because LLM-judge pipelines carry measurable, style-correlated noise. Also newly submitted: Anthropic External Researcher application (API credits to scale the judge study to K≈200 and a multi-judge panel).