You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
When you fine-tune a model with RLHF, safety degrades by 40 to 60 percent in full fine-tuning runs. You find out from evals, after training ends. The information about when it happened, which training step, which layer, which gradient conflict did it, is gone. Training is a black box and nobody has a mechanism to see inside it while it runs.
I built ORMAS to fix this at the architectural level. The core move: restrict each node's local gradient chain to four operations instead of entangling everything into one global backward pass. This makes every node independently measurable during training its gradient conflict, activation health, and correction history logged at every step as a structural property of the forward and backward pass, not reconstructed afterward.
When a node's health drops below its own baseline, ORMAS fires a bounded correction and logs what happened and why. All corrections are bounded by a formal Input-to-State Stability argument, so training still converges even while repairs run.
Key result: I destroyed a fully converged convolutional layer at epoch 100. ORMAS recovered to 80.3% accuracy through 85 bounded, causally attributed corrections. A standard network collapsed permanently to 10.0% across all three seeds. The gap is 70.3 percentage points.
This was accepted at DeepMath 2026 after double-blind peer review. NeurIPS 2026 AI for Good workshop reviewers called the results striking, described the framework as one that would help with AI safety, and specifically said the next step is testing at transformer scale. That is exactly what this grant funds.
Preprint: https://zenodo.org/records/21730363
Code: https://anonymous.4open.science/r/ormas-EB73/
Site: https://raadh.me
The goal is to find out whether ORMAS works at GPT-2 scale, where alignment research actually happens. Everything currently proven is on CIFAR-10 and ResNet-18. Safety-critical neurons exist in transformer attention heads and MLP blocks, not image classifiers. The ISS stability bound needs to be validated on transformer architectures before this is useful to anyone working on RLHF safety.
Three specific experiments:
First: train ORMAS on GPT-2-small (117M parameters) on OpenWebText and measure whether correction frequency decays from high to near-zero as training stabilizes. This is the foundational validation does the stability bound hold for attention mechanisms? Reviewers said this is the test the paper needs to be publishable at a main conference.
Second: fine-tune that GPT-2-small model on Anthropic's public HH-RLHF dataset with ORMAS monitoring active. Log which nodes show the highest gradient conflict during safety fine-tuning versus capability fine-tuning. This produces the first per-node, timestamped conflict log during RLHF ever recorded. If safety behaviors live in specific circuits (which Neuron-Level Safety Realignment research suggests), those circuits will show measurably higher conflict when safety training modifies them.
Third: destroy one attention head at epoch 50 and measure whether ORMAS recovers it. At CNN scale the recovery gap was 70.3 percentage points. Does this hold for transformer attention heads, which fail differently? If yes, this is publishable as evidence that training-time structural repair works at GPT-2 scale. If no, the mechanism needs rethinking for transformers, which is equally important to know.
I achieve this by running the experiments on AWS GPU instances over 60 days and releasing a preprint and the full RLHF conflict dataset publicly.
AWS p3.2xlarge, 1x V100, 8 hours/day for 60 days (GPT-2 training and RLHF fine-tuning experiments): $2,880
AWS p3.8xlarge, 4x V100, 20 hours total (lesion recovery at GPT-2-XL scale): $1,220
S3 storage for experiment logs, correction records, telemetry archives for 60 days: $180
Weights & Biases Teams for reproducible experiment tracking and public sharing, 3 months: $240
Contingency for spot instance interruptions and re-runs on failed seeds: $280
Total: $4,800
Everything produced comes out publicly. A preprint covering all three experiments, released regardless of whether results are positive or negative. The per-node RLHF gradient conflict dataset from the second experiment is released as a standalone artifact because it is useful to anyone working on surgical alignment or plasticity preservation, independent of whether ORMAS turns out to be the right long-term architecture.
Just me. Rokib Al Dhin Raadh, 18 years old, Dhaka, Bangladesh. Self-taught, no university, no lab, no advisor.
For ORMAS specifically: I wrote all 16,316 lines of the codebase, derived the ISS stability proof, designed the experimental protocol, and ran all 383 controlled experiments myself on a personal RTX 3090. The codebase runs from a single script and has a full reproducibility checklist.
For track record on larger systems: I built OXIMO, a 40,933-line autonomous operating system that ran a live UK e-commerce company (Black Bloxie LTD, incorporated in England and Wales) for 12 months without human input. Its causal impact was verified through a controlled ablation study removing the system cut output by 91%, restoring it produced a 3.3x recovery overshoot. This is the same methodology I use for ORMAS: controlled removal of components to measure causal contribution.
Current validation: DeepMath 2026 accepted the ISS theoretical contribution after double-blind peer review. NeurIPS 2026 AI for Good workshop reviewers called the results striking and the framework one that would help with AI safety. University of Chicago Assistant Professor Mina Lee (Stanford PhD, MIT Technology Review Innovators Under 35) called the idea super interesting and impressive in a direct conversation. I review papers for the NeurIPS 2026 Trustworthy AI for Good workshop. I have a talk on the mathematical foundations at Cohere Labs on November 2.
Two realistic failure modes:
The more likely one: the 4-operation local gradient chain constraint does not translate cleanly to transformer attention mechanisms. Convolutional layers have spatial locality that makes the bounded local chain tractable. Attention heads do not have the same structure query-key-value interactions are global by design. It is possible that ORMAS's per-node health monitoring adds so much computational overhead at attention scale that training becomes impractically slow, or that the correction mechanism interferes with attention's ability to learn long-range dependencies. If this happens, the experiments will show it clearly. The output is still a published negative result that tells the field why this approach doesn't scale to transformers, which is useful information.
The less likely but more frustrating one: the experiments work but I fail to execute them cleanly. Spot instance interruptions, bugs in the transformer-adapted codebase, seed failures that require more compute than budgeted. The $280 contingency covers some of this but not all of it. If the budget runs out before clean results are in, the experiments would need to be run at smaller scale or with fewer seeds, reducing statistical confidence in the results.
In both cases I commit to publishing whatever results exist at the end of the grant period, positive or negative, with full logs. The RLHF conflict dataset gets released regardless because even partial data from that experiment is useful.
Nothing for this research specifically. ORMAS was built entirely on personal compute (RTX 3090 I own) with no external funding.
Separately, Black Bloxie LTD (the UK e-commerce company OXIMO runs) generated $6,691.68 in verified revenue over 12 months, reconciled against Shopify export records. That is passive income from the deployed system, not a grant or investment, and none of it went toward ORMAS research.
This is the first funding application for ORMAS-specific compute.