You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Crux of the project
A recent paper released by ARC Theory shows that sample-free cumulant propagation can estimate the expected output of wide random MLPs more efficiently (in the number of floating point operations) than Monte Carlo sampling, with particularly strong results for rare-event probabilities.
I hope to obtain a similar estimation result for wide random convolutional neural networks. My current belief after some preliminary work is that a nontrivial flattening of the convolution would make the core of the existing algorithm quite useful for our purposes. and covariance-level uncertainty propagation can already handle convolution as a linear operation. Perhaps the most significant new challenge is to propagate higher-order cumulants efficiently and show that our algorithm doesn't make too much error when filters are shared across many spatial positions.
The deliverables will be an open-source implementation, reproducible experiments, and (assuming I'm successful) a research paper. Project design, theory, and experiments will be done by me (with the assistance of frontier models on small but time-consuming tasks), with short-term engineering support for optimized PyTorch or CUDA kernels. Additionally, and especially because the MLP paper’s proof-of-concept implementation often fails to realize its FLOP advantage in wall-clock time, optimized implementation and wall-clock comparisons will be core parts of the project.
Theory of impact
Beyond the fact that I believe this could help shed light on a precise definition of "mechanisticness," CNNs are a natural test of whether mechanistic estimation survives parameter sharing rather than depending on fully independent dense weights. Success would mean an extension of the method to an important class of structured architectures and identify how architectural symmetry can make analytic estimation scalable. Furthermore, more reliable estimates of low-probability, high-cost outputs could eventually support safety evaluation and training methods related to latent knowledge elicitation and low-probability estimation.
How the money will be spent
Short-term PyTorch/CUDA implementation support = $2,000
Compute:
GPU cloud credits = $2,000
High-memory CPU/RAM cloud credits = $750
Storage and experiment tracking = $250
Total = $5,000
I can try my best to push costs down (by about $1,000) by reaching out to friends I know.
There are no bids on this project.