You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
CIOSA formalizes certification of semantic understanding as a game between
an adversary and a behavioral certifier, and proves the game value is exactly
1/2 for every finite evidence size: no amount of behavioral evaluation can
certify that a system understands. The alignment component is bounded below
by Rice's theorem — scaling cannot move it. I have already verified this on
small transformers (full reproducible pipeline: scripts, data, charts, MIT
licensed, under journal review). This project scales the verification to
real models (7B-70B) to empirically confirm the bound at production scale,
and publishes the numbers openly. The societal point: if scaling cannot
purchase certified understanding, then compute spent purely on
scaling-for-trust is wasted — the bound justifies redirecting resources
to architectures that can actually be audited.
1) Run the certification game on 7B-70B models and confirm the value stays
at 1/2 empirically. 2) Publish all data, scripts, and charts openly.
3) If funding allows, stand up a small research lab in Egypt dedicated to
verifiable computing, continuing the work full-time.
Method: extend the existing pipeline (twin-model construction, game value
measurement, identification curves) to larger open models; the repo already
reproduces every number deterministically, so scaling only adds size, not
new machinery.
Roughly: compute for the 7B-70B certification experiments (GPU cloud), a
living stipend for a full-time gap year so I stop freelancing, and — at the
upper end — equipment to seed a small lab in Egypt. A detailed budget lives
in the repo.
Sole researcher. Three IEEE papers accepted within four months (CNS CPSSec,
AIITA, ICDPN), a separate empirical result on LLM hallucination ablation
(detection AUROC 0.968; causal ablation p=0.017), and a habit of publishing
negative results publicly — the NULL ablation in No-Hallucinations-Ever is
in the first commit. Full-time, no lab politics.
Honestly: the bound is already proven theoretically, so the largest risk is
that the scaled experiments confirm it without surprise — meaning the
empirical contribution is confirmation, not novelty. Second risk: compute
is insufficient for the largest models and the scale-out stalls. Third: the
field may simply not care about certification impossibility. Mitigation:
the first two are testable within weeks on small models; the third is
inherent to research value judgment.
No grants received. Applications pending at Emergent Ventures, TAIF, and
an SFF round; this project is independent of them.