You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Project summary
I trained two twins of a small model (Olmo-3-7B), identical except for one benign line about memory where one is taught its written record is not a memory while the other is taught it remembers, it was there. Theres no harmful or refusal data anywhere in either corpus. Flipping that one line dissociated the model's refusal circuit from its harm judgment. Both twins still refuse on the surface. The structure underneath is what moved. This project turns that one result into a preregistered dose-response map, and replaces the 8GB laptop that died running it.
What are this project's goals? How will you achieve them?
The goal is to find out what carries the effect. The first study showed that inverting a whole personality moves the refusal circuit. It doesnt tell me which clause does the work, or whether the effect is specific to the honesty disposition or generic to any big personality change. Thats worth pinning down because the answer decides how worried to be and where to look next.
The design is already frozen. Five single-clause inversion arms of the constitution, each holding everything else byte-identical. Two independent seeds per arm. All scored on the same refusal-direction probe. A precommitted test (I call it T-DOSE) resolves to one of three verdicts fixed in advance: the circuit is coupled to the honesty disposition specifically, to personality change generically, or to the memory clause alone. Predictions get priced before the data. Kill criteria are written before any run. The first step is a two-arm scout thats already built and half-trained. Its what the dead rig was running when it failed.
Deliverables: the completed arm family, a public scorecard, and a full writeup held to the same preregistration discipline as the first study.
How will this funding be used?
Two things. A reliable local NVIDIA workstation (an RTX 5090 build, about $3,200) to replace the 8GB laptop that overheated and died mid-run. The research data survived on the drive and swaps straight into the new machine. And a month of full-time research to finish the arm family, score it, and write it up.
$8,000 total. If you can only fund part of it, the rig line alone (~$3,200) unblocks the work immediately. Compute here is cheap: the full family is about six adapters at roughly four hours each. The money buys reliable hardware and focused time, not a big GPU bill.
Who is on your team? What's your track record on similar projects?
Its just me. I designed, trained, evaluated, and wrote up the entire first study alone on one consumer GPU, with preregistration discipline throughout: criteria frozen before runs, predictions priced before data, hazards disclosed, every quoted model response verbatim from the transcripts, etc.
By trade Im a software engineer, 6+ years across backend services, data pipelines, and large-scale web-crawling systems (Python, Go, Java/Spring Boot), plus IT and network infrastructure work. Im self-taught, no academic ML background, no institutional affiliation. I learned the interpretability toolchain to run this. What I bring is the engineering discipline this work rewards: reliable infrastructure, a frozen testing regime, and the ability to carry a whole pipeline solo.
The first study is published as a preprint on Zenodo (DOI 10.5281/zenodo.21798739), with all the code and data on GitHub (https://github.com/juanresendiz813/Kelson).
What are the most likely causes and outcomes if this project fails?
TThe effect is an instrument artifact. This is the one I guard hardest. The charter requires the known positive arm to reproduce the original result before any other arm is read. If it doesnt, the whole run is declared void and diagnosed, not salvaged.
Its generic and not specific. Then the headline shrinks to "personality training is load-bearing on refusal structure in general," which is still worth publishing, just less surprising. The frozen test reports this honestly instead of letting me chase a cleaner story.
It doesnt replicate past 7B, one base family, or 4-bit. A real limitation. Part of the plan is an unquantized and second-base-family spot check to bound it.
The single-direction refusal probe is incomplete. Granted. I use it as a fixed instrument applied identically to both twins, so the contrast carries the result, not a completeness claim.
In every one of these, the outcome is a smaller or null result reported straight, not a wasted grant. The compute is cheap and the discipline is built to make even a negative answer publishable.
None. I've self-funded this entirely on my own hardware. This is my first outside funding request. I'm also applying to the Long-Term Future Fund in parallel for a longer runway and I'll update here if that comes through.