You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
This text was not written by AI despite of the flag by Pangram.
Intro: Presently, most of AI safety relies on "post-hoc" behavioral evaluations (red-teaming, etc.).
AI can game those or pass while suffering internal latent trajectory collapses that lead to unpredictable, and even catastrophic deployment failures. If a transformative AI misaligns or collapses, the downstream impacts on all of us, humans and animals alike, via agricultural, industrial and ecological externalities, etc. may be severe - in extremis, the admonition "If Anyone Builds It, Everyone Dies" - could come true for everyone - including the animals.
The Project: I am building and open-sourcing an interactive hidden state dynamics telemetry framework that allows the operator to monitor the AI's internal geometric state in real-time (using metrics like depth-dependent Friction Delta, Pawula ratios, and Itô Stochastic Differential Equation kinematics), instead of just watching the outputs. See: https://github.com/Ergo-sum-AGI/mycelia_lm_1.6G_params/blob/main/master_telemetry_glossary.pdf
The telemetry translates into a control mechanism of Cache-Fused Kinematic Rail (CFKR). When the system detects on beforehand (ex-ante) an impending AI collapse, the proposed CFKR mechanism intervenes via native KV-cache mutations before "corrupted" tokens are even emitted. See my paper: https://zenodo.org/records/23002427
This project directly addresses the Falcon Fund’s objectives. Ensuring that transformative AI remains structurally stable, legible, interpretable and controllable is the basic requirement for any positive outcome affecting all sentient beings, especially the humans and the animals.
This project hedges against the worst-case scenarios that endanger all life on Earth by an erratic or rogue AI, by enabling the human operator to read the innermost dynamical processes within the AI's Black-Box - the hidden manifold which were, up to now, only sparsely mapped or analyzed as static data, not dynamical events - yet are in fact crucial and determinant for all the subsequent AI outputs.
By implementing weighted corrective mechanisms, instability, collapse and failure can be predicted and prevented.
This goal is achievable, because all the pieces are already in place by now: The 1.62B-parameter custom LM prototype is built and actively trained, and the test results are regularly published on Zenodo.
Once the training is finished, the inference phase implementation and testing can begin. Concomitantly, a rigorous validation protocol shall be executed. See: https://github.com/Ergo-sum-AGI/mycelia_lm_1.6G_params/blob/main/TEST_PROTOCOL.md
The framework shall be tested on larger open-weight models and designed to be model-agnostic.
Thee final goal is to achieve the stated objectives in six months.
This funding shall be used specifically to scale up the empirical validation suite to 7B+ open-weight models, and to finalize the Triton/CUDA kernels for vLLM integration.
It shall cover any extra compute, APIs, HW upgrades, telecommunication costs and it shall also provide a modest stipend for the unaffiliated Principal Investigator, allowing a full-time focus on this research without further financial distractions.
I am Daniel Solis (solis@dubito-ergo.com), the sole team member. The track record is the project itself, as it is at an advanced stage that can vouch for itself. See: https://scholar.google.com/citations?user=fOI5z-cAAAAJ
This project is a win-win game. Obviously, it should succeed since the foundations have been already laid and seem robust. However, there are various scenarios how it could fail:
a) Scale/Architecture Non-Generalization
The geometric diagnostics were developed and validated on a specific 1.62B-parameter, 24-layer dense transformer and the metrics do not generalize to frontier scales (70B+) or alternative architectures (e.g. MoE, where the concept of a unified "layer variance" breaks down because routing mechanisms dynamically split the active parameters. If the metrics do not reliably predict hallucination in a 70B MoE model, the framework fails as a universal diagnostics).
Or,
b) Pragmatic Deployment Failure
As the CFKR requires writing complex, Triton/CUDA kernels to mutate the KV-cache in virtual LLM with paged attention, it introduces non-trivial math overhead and latency. If a simpler, existing method (e.g., standard speculative decoding with a small verifier model, or logit-biasing) achieves the exact same reduction in hallucinations or error-rates for lower compute, then CFKR fails the pragmatic deployment test and become more something of a theoretical curiosity, rather than a sought after safety tool.
Or,
c) Diagnostic Evasion (the Goodhart’s Law)
The model learns to "game" the geometric telemetry when it learns to maintain a "healthy" metrics while still schemes or generates deceptive or hallucinatory outputs. The diagnostic would flash a fatal blind spot.
Even in such case, the failure still makes a scientific contribution by showing exactly what does not work and which path is a dead end.
Nonetheless, the geometric telemetry does not become completely useless either. The framework would still serve purely to facilitate training-time observability and early stopping indicator, saving a lot of money in unnecessary training compute.
The project would in either case map the boundaries of where white-box geometric interpretability works and where it doesn't.
I have raised $100,000 in compute credits from AWS via NVIDIA inception program.
There are no bids on this project.