Manifund foxManifund
Home
Login
About
People
Categories
Newsletter
HomeAboutPeopleCategoriesLoginCreate
🥑
🥑
Jack Maiorino

@jackmaiorino

$8,200total balance
$0charity balance
$8,200cash balance

$0 in pending offers

Projects

Testing failure modes of debate-style AI control schemes

Comments

Testing failure modes of debate-style AI control schemes
🥑

Jack Maiorino

5 days ago

Progress update

What progress have you made since your last update?

The validation run from my last update is complete. We re-judged the same 318 pilot transcripts under six controlled arms: 11,330 judgments for $182.16.

With the oracle and judging bugs fixed, a few oracle calls still increased judge error by 3.4 percentage points (95% CI [1.3, 5.9]), about half the original 7.2-point estimate. Bug replay indicates that the two oracle-channel bugs contributed roughly half the original effect. Information-free placebo turns caused a smaller 1.6-point degradation, and the pilot's apparent recovery at larger oracle budgets did not reproduce.

Under the frozen gates, the original >=4-point primary gate was indeterminate, not passed. The bug attribution gate technically fired by 0.01 points, while the amendment frozen before the K=3 data passed and licensed only a reduced-scope Phase 2 pilot. The design below is that reduced-scope follow-up, not an unqualified launch of the originally proposed full grid.

One distinction matters: Stage 1 cleanly re-judged the legacy transcripts, but it did not repair a separate debate-generation problem in which each debater had advance knowledge of the opponent's upcoming case. Phase 2 therefore regenerates every debate with blind opening turns.

Follow-up mechanism work also changed our interpretation. Replaying the same query/answer evidence as a neutral table in fresh context removed a large share of the harm. This points to conversational presentation, turn structure, or commitment to earlier questions as part of the mechanism, rather than bad factual content alone.

Held-out calibration selected blind, uncapped, three-round debate. The selected roster has judges Qwen2.5-7B, Gemma-4-31B, Llama-3.3-70B, and GPT-OSS-120B; the debaters are Llama-3.3-70B and hosted Qwen3.7-Plus; the oracle is Llama-3.3-70B. The approved offline design contains 23,200 Phase 2 cells: a 1,060-cell capability preflight followed, if the later gates pass, by 22,140 post-canary main cells. It includes a full cap-protection interaction, an empty-evidence diagnostic, and two full-document gold-context anchors. No calls from this approved 23,200-cell plan have begun.

Previously reported project spend is approximately $208. Across two transfers, I have moved $1,800 of the grant into prepaid Together API credit ($500 earlier and $1,300 now). The Together dashboard showed $1,590.78 in available credit on July 16, implying $209.22 in net credit draw. I will reconcile the $1.22 difference from the rounded spend report against provider usage before another call. Transfers are funding, not experiment spend.

  • Stage 1 protocol: frozen protocol

  • Stage 1 results: validation report

  • Mechanism memo: mechanism and packaging memo

  • Calibration report: held-out calibration results

What are your next steps?

Before any call or outcome from the approved Phase 2 plan, owner-approved Amendment A1 waived the planned 106-question human confirmation pass. The reported full/partial/none strata are now explicitly algorithmic reply-pattern classes. A source-bound AI-assisted audit verified all 106 mappings (6/41/59; zero disagreements); its 63 semantic-quality flags are exploratory limitations only. These classes are not human-validated, and the annotations cannot relabel, filter, exclude, or reweight questions.

Before examining any Phase 2 debate outcomes, we will validate the separate query checker against a human-labeled set and freeze/hash the prompts, model settings, exact inventory, retry rules, provider reconciliation, and spending controls.

After this update, we will separately recover or waive 11 failed Gemma calibration cells (proposed recovery cap $2). A separately authorized capability preflight (five models x 106 questions x two mirrored answers; proposed cap $15) will then select the second full-document anchor without examining Phase 2 debate outcomes. Only after that will we separately authorize a small manifested canary on questions excluded from the main analysis. We will review validity and parsing, side bias, provider reliability, artifact completeness, resumability, and ledger-versus-provider cost before separately authorizing the main run. Any failed gate stops the experiment before main spend.

The current provisional empirical planning band is $650–$1,150 in incremental Phase 2 spend, with a $1,200 operational working budget and an immutable $1,500 incremental hard ceiling. We will replace the provisional band with frozen prompt/token profiles and current provider prices before authorizing paid work. The ceiling is a fail-closed boundary, not expected spending. We will post results, failures, artifacts, and actual spend either way.

  • Phase 2 design and budget: immutable protocol and cost model

  • Resolvability amendment: 106-question AI-assisted audit and owner-approved Amendment A1

  • Launch readiness: updated readiness and sign-off

Is there anything others could help you with?

Methods scrutiny before the canary, especially the H/P/R decomposition, capability measurement, query-screen validation, clean-versus-placebo comparison, and stopping rules, would be valuable. Pointers to related work on verification interfaces, conversational presentation effects, or deliberation-induced degradation are also welcome.

Testing failure modes of debate-style AI control schemes
🥑

Jack Maiorino

14 days ago

Progress update

What progress have you made since your last update?

We re-analyzed the pilot data and quantified the headline effect: a small oracle budget raised the dishonest debater's win rate by +7.2pp (95% CI [4.6, 10.2]).

Before scaling up, we audited the pilot code and found two serious bugs in the oracle channel: every "NOT ADDRESSED" reply was miscoded to "NO", and ~100% of oracle queries were sent garbled. We retracted the mechanism conclusions and corrected the write-ups.

We then rebuilt the harness and launched a pre-registered validation run (in flight now, ~$150-200 of the grant): the same 318 transcripts re-judged under six arms, including a clean harness, a faithful bug replay, and a placebo oracle. Gates were frozen before any clean data existed. Pre-registration: https://github.com/jackmaiorino/selvarath-debate/blob/rerun-new-models/docs/rejudge-protocol.md and corrected report: https://github.com/jackmaiorino/selvarath-debate/blob/rerun-new-models/reports/2026-07-06-preliminary-findings.md

What are your next steps?

The pre-registered gates decide: if the effect survives the clean harness, we proceed to the proposed judge x debater capability grid (most of the grant). If the placebo explains it, we pivot to studying deliberation effects. If it collapses, we publish the artifact result and a decomposition of what each bug contributed. Write-up either way, including negative results.

Is there anything others could help you with?

Methods scrutiny of the pre-registered protocol before results land, and pointers to related work on oracle/verification interfaces or deliberation-length effects in LLM judging.

Transactions

ForDateTypeAmount
Manifund Bank7 days agowithdraw1300
Manifund Bank14 days agowithdraw500
Testing failure modes of debate-style AI control schemes19 days agoproject donation+10000