You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
The conversation-analysis market measures sentiment, topics, and talk-time, but ships no measure of depth — whether an exchange actually moved something or just performed. Over seven years I recorded and transcribed my own conversations across eleven contexts (work, legal, medical, family, facilitated introspection): 50,000+ segments where the producer of the record and the judge of its depth are the same person. On that corpus I validated a five-pattern, five-level depth taxonomy, and the signal turns out to be physical — the deepest segments run slower and more silent (167 vs 200 words per minute; 1.22s vs 0.79s pauses; two pre-registered results, p≈4.5×10⁻⁶ and p≈0.003). This project turns the instrument into a public benchmark: can frontier models tell genuine reasoning from fluent performance?
Three deliverables. First, a human-adjudicated, held-out gold-standard label set built under a pre-registered protocol — the current pipeline labels are AI-applied and quality-gated, and the gold set is what makes every claim testable. Second, published per-model results for frontier models (Claude, GPT, Gemini, open-weights) on the depth-recognition task, including a null result if none beat chance. Third, the taxonomy specification and evaluation harness released openly so others can replicate on their own data. The relevance to safety work is direct: depth discrimination is the same muscle as deception and sycophancy evaluation — telling whether communication is performative.
$5,000 (minimum) funds adjudication tooling, the first gold-set slice, and the harness release. $25,000 funds the full six-month open release: my time writing the specification, a substantially larger adjudicated set, and multi-model evaluation runs beyond the roughly $4,800 in committed compute credits I already hold.
Solo. Before this: initial team at R3 building the Corda blockchain platform (I worked on its first commercial deployment, into the $2.4T international gold market), Head of Financial Services at TradeLens, the Maersk–IBM venture, where the Citi Asia-Pacific implementation enabled $1.2B in paperless trade finance, and ADP Ventures (consumer-permissioned data platform scaled to a $25M projected run rate). Since 2023, alone: 70+ self-running edge services, a public medical-research portal (revasserkernel.com, 10,000+ papers indexed for rare-disease literature synthesis), live commercial products (revasser.nyc, revaddress.com, developer tooling on npm). Seven years collecting the corpus, two years building the instrument.
Three honest risks. The corpus is single-subject — that is precisely what eliminates inter-rater disagreement, but it means transfer beyond one identity is an open question; the benchmark is designed so others can test the taxonomy on their own data, which is the only real answer. The current labels are AI-applied and not yet human-validated — the gold set this funds is the fix, and if adjudication shows the taxonomy discriminates worse than claimed, that result gets published too. And I am one person running real businesses alongside this; the deliverables are scoped so each lands independently.
$0 external to date — everything so far is self-funded from my own products. Pending applications this month, disclosed in full: Emergent Ventures ($25,000, applied 2026-09-06) and the EA Transformative AI Fund ($75,000 over 12 months, applied 2026-09-06 — that application funds my time for the broader research program; this Manifund ask funds the open-release slice, and if both land I will take the union of deliverables and reduce or return the overlap). Also pending: Anthropic AI for Science (API credits) and Google for Startups (~$2,000 infrastructure credits).