You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Transformer models consume far too much GPU energy. My goal in this 1-year project is to fix that. Based on my recent IEEE TIT paper ("Rank-Adaptive Local Empirical Processes in Low-Rank Attention", (under review)DOI: 10.5281/zenodo.23147548), I build lightweight training algorithms that cut GPU power and FLOPs by up to 90% without sacrificing accuracy.
During training, Self-Attention weights naturally collapse into low-rank subvarieties M_{≦s}. I exploit this geometry to dynamically allocate rank and prune weights on the fly. The math is already done in my paper. Now I am turning it into working PyTorch code for LLaMA and Mistral.
I have three concrete technical deliverables for this 1-year project:
Sub-Root Adaptive LoRA: Dynamic rank allocation during fine-tuning based on my sub-root fast rate bound r^* = O(s(d_out + d_in) log s)/(λ_min n). Cuts GPU power during fine-tuning.
Provable Zero-Shot Pruning: One-shot model compression using my SVD prediction-distortion bound E_P [1/2 ||WX - ∏_s(W)X||_2^2] ≤ (λ_max(Σ_X)/2) (Σ_{j=s+1} σ_{j}(W)^2). Guaranteed zero distortion under exact rank collapse.
Rank-Collapsing Regularization: A new objective function using nuclear-norm penalties to force early rank collapse during training, lowering GPU load as training proceeds.
Timeline:
Months 1–4: Extend my TIT theorems to multi-head PyTorch attention layers and design the adaptive rank selection algorithm.
Months 5–8: Implement in PyTorch. Measure actual GPU power draw (Joules/token, W/epoch) on LLaMA/Mistral models using cloud GPUs.
Months 9–12: Release the open-source PyTorch library on GitHub and publish the findings.
PI Stipend ($45,000): Covers my living expenses for 1 year so I can work 100% full-time on this research without taking outside contracting work.
GPU Cloud Compute ($8,000): RunPod / Lambda Labs H100/A100 instances for GPU power and rank dynamics benchmarks.
Workstation & Tooling ($3,000): Development environment, API costs, and verification tooling.
Publication & Outreach ($4,000): Open-access publication fees and documentation for the GitHub repository.
I am Hideki Ishiyama, an independent researcher specializing in mathematical learning theory and algebraic geometry. I design the core theory, math, and algorithms myself. I use AI tools (like Gemini Notebook) as interactive assistants for coding and verification.
Track Record: I proved and published the foundational theory single-handedly in my IEEE TIT paper and Zenodo preprints. In a 4096 x 4096 attention projection layer, my theory proves a 99.22% parameter space reduction (from 16.77M to 130.8K dimensions) under rank s=16 with shrinking local Rademacher complexity.
Risk: In extremely deep models or long-context inputs, the minimum eigenvalue of the input covariance might degenerate (λ_{min}->0). This expands the local radius ρ_r = \sqrt{2r / λ_min } and could slightly degrade zero-shot pruning accuracy.
Mitigation / Fallback: My rank-collapsing regularization explicitly prevents spectral degeneration. Even if extreme compression fails on certain huge models, I will release the open-source PyTorch profiling tools and attention spectrum datasets. This provides value to the AI Safety and DevInterp research communities regardless.DevInterp, and Green AI (energy-efficient AI).
$0 USD. I self-funded 100% of my research so far. This grant is my first external funding application.