You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
In May 2026 I backtested a forex strategy that showed +700,000,000,000,000% return. For about an hour I thought I had found the holy grail. Then I re-ran it with a one-bar execution lag — the strategy could only trade on the NEXT bar after its signal, the way real execution works — and the same strategy returned −30%. Nothing in the code was "broken". It was a trade-at-close lookahead artifact, the kind that passes every unit test. I only caught it because I have a hard personal rule: any backtest with Sharpe above 3 must survive lag=1 before I believe it.
I have been trading crypto futures with my own money for three years, and I have paid tuition on backtest bugs like this the whole time. That experience is the basis of this proposal.
Over the last 18 months a cluster of benchmarks has started measuring whether AI agents can trade and forecast: live arenas ranking frontier models by PnL (Alpha Arena, LiveTradeBench), LLM trading-agent papers with public code reporting strong Sharpe ratios (TradingAgents, FinMem), and forecasting benchmarks scoring models against resolved events (ForecastBench). [VERIFY: check từng tên + link trước khi post.] These numbers feed something bigger than academia: UK AISI's RepliBench includes an "acquire money" task family, and METR's autonomy evaluations treat unassisted money-making as an input to resource-acquisition threat models [LINK: RepliBench paper section + METR threat model doc]. If the public benchmarks shaping intuitions about this capability contain lookahead bias or training-data contamination, the numbers are wrong in a known direction: they overstate what agents can do. Nobody has independently checked.
I should be honest about which way my audits cut: they will mostly deflate reported capability numbers, not inflate them. I think that still matters for safety. Decisions need trustworthy evals more than they need scary ones, and a field that tolerates quietly inflated agent-trading numbers is training itself to ignore its own measurements.
Four months, three deliverables:
1. **An open-source (MIT) audit toolkit** — a generalized, documented rewrite of the private robustness suite my own trading has depended on for ~3 years: execution-lag stress tests, point-in-time data checks, deflated Sharpe (Bailey & López de Prado) and bootstrap null models, fee/slippage stress, and LLM-specific contamination checks (do evaluated events predate the model's training cutoff?).
2. **Three public audit reproductions** of benchmarks with fully public code and data — the first collaborative (I contacted the authors before applying: [VERIFY: phải có email trả lời trước khi post]), the other two independent, both with advance notice and a published right of reply.
3. **A failure-mode taxonomy** with minimal reproducible examples, written up publicly.
**Money and conflicts, stated plainly:** zero grant dollars touch any market — this project deploys no capital, paper or real. I trade my own personal account (currently a passive leveraged holding portfolio that takes under 30 minutes a day to monitor), and I have run small copy-trading experiments on Polymarket and Hyperliquid; I sell no trading products. Two binding commitments for this work: for 12 months I will not trade on, or monetize in any form, anything I learn inside an audited benchmark; and every audit write-up will disclose any position I hold in the audited asset class.**Goal:** a harness cheap enough that benchmark authors run it before publication and reviewers start asking for it. The concrete adoption target: by the end of the grant, at least one actively-maintained benchmark (the collaborative audit's team is the first candidate) runs the toolkit as a pre-release check.
**Step 0 — pre-registered protocol (before any target is locked).** I publish the audit protocol first: the exact checks, thresholds, what counts as a finding, and fixed target-selection criteria (public code + public data + reported results that people actually cite). This is deliberate: it means neither my methods nor my target choices can be steered toward "finding something".
**Month 1 — toolkit.** Extract and rewrite my private robustness engine into a documented, pip-installable library with CI. The modules exist and run in production today; the work is generalizing the interfaces and writing the docs, which is why one month is realistic for this part and only this part.
**Month 2 — collaborative audit.** One audit, with the authors of an actively-maintained public benchmark, published with their reply (aiming for their endorsement). This depends on third parties, which is exactly why I contacted them before applying rather than after.
**Month 3 — two independent audits.** Reproductions of two published benchmarks from the pre-registered candidate list, run under point-in-time discipline, with pre-notification and a right-of-reply window; replies are published alongside the audits.
**Month 4 — taxonomy + v1.0.** The failure-mode taxonomy with minimal reproducible examples, toolkit v1.0, and the public write-up (LessWrong / EA Forum).
**Rules I pre-commit to,** because auditing other people's work is socially risky: public-code-and-data targets only; authors always notified before publication with their reply published alongside; results phrased as measurement — "performance under point-in-time discipline" — never as accusation. If all three audits come back clean, I publish that with the same prominence. "These numbers survive strict PIT discipline" would be the first independent confirmation of its kind, and the toolkit stands either way.
**Maintenance:** six months of committed maintenance after v1.0 (versioned releases, CI, issue triage), and I will actively look for a co-maintainer during the grant. I am deliberately not promising more than that as a solo dev; the tool is scoped small — a test harness, not a platform.- **$12,000 — my time.** Four months, full-time. Full-time is real here: my trading currently runs as a passive holding portfolio needing under 30 minutes a day, and my active strategy-mining work is paused.
- **$5,000 — reproduction costs.** Point-in-time-quality historical data is genuinely expensive; this line is sized so the audits don't have to quietly fall back on free survivorship-biased data — and if a specific audit does use free data, the write-up will say so and bound the effect.
- **$3,000 — native-English technical editor + contingency.** English is my second language and the write-ups are the product, so I am budgeting for editing instead of pretending I don't need it.
**If only the minimum ($2,500) is raised:** part-time version — the pre-registered protocol, the toolkit, and the one collaborative audit, over the same four months.Just me. What I can show, and how to check it:
- **Three years of live algorithmic trading** on crypto futures with my own money — not paper. Candidate strategies from my mining pipeline only survive if they pass a robustness suite (lag stress, fee/slippage stress, subperiod stability, bootstrap tests) — the private ancestor of the toolkit in this proposal. Every check was added after a specific loss; the lag test exists because of the May 2026 forex bug above. I'm happy to verify account statements privately with any regrantor considering this project.
- Portfolio: [LINK: https://thanhnguyen.40-160-7-25.sslip.io/#about]. GitHub: [LINK: github.com/Thanh-Van-2001]. Based in Vietnam, working independently. This would be my first grant.- **The social failure mode (most likely).** An audited team reacts badly and the project gets framed as attacks rather than measurement. The mitigations are structural — pre-registered protocol, public-code targets only, pre-notification, right of reply published alongside, collaborative audit first — but the risk isn't zero. If a conflict happens anyway, the toolkit, the protocol, and the collaborative audit still stand on their own.
- **Third-party dependency slips.** The collaborative audit depends on authors' response cycles; that's why contact started before this application. If the first team goes silent, I move to the next candidate on the pre-registered list — the ask is three audits, not three specific audits. Worst case the collaborative one converts to a third independent audit, with the same right-of-reply process.
- **Null results.** All three benchmarks survive PIT discipline. Published as-is, with the same prominence — still the first independent confirmation that these numbers hold up, and still a working toolkit.
- **Solo-maintainer risk.** I under-resource maintenance after the grant. Honest answer: this is why the commitment is six months plus an active search for a co-maintainer, not a promise of forever.$0. I have never raised charitable funding. My trading account is my own savings and is entirely separate from this project.