You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Project description
What this project is
Figure 5 of Algorithmic Progress in Language Models illustrates that there was a 410^10× scale-up of LMs from 2012 to 2023, of which 1.710^7× was from physical compute scaling and 2.2*10^4× was from algorithmic progress (including algorithms, optimizers, architectures, and training data quality improvements in this second category). The goal is to isolate the effects of data quality improvements and estimate the data-only compute-equivalent gain, by doing experiments such as the following:
1. Fixed recipe, swap only the data
1. Hold constant: architecture, tokenizer, optimizer, hyperparameters, training code, numeric precision, evaluation procedure
2. Vary dataset: raw web, deduplicated web, filtered web, curated or enriched datasets
2. Test at several scales
3. Data time-leap
1. {old, new} data x {old, new} algorithms+architectures
2. Then do a Shapley decomposition to find compute savings from each.
4. Ablations on modern data pipeline
5. Longitudinal experiment: same recipe; dataset from each year 2014-2026
6. Generate datasets from models of different intelligence; measure student model quality
Theory of impact
Informs x-risk reduction plans, especially those calling for slowdown of AI capabilities. Specifically, it affects plans that involve restrictions on algorithmic progress or AI R&D. If data-only CEG is large, then data governance is more important.
How the money will be spent
Compute for synthetic data generation, model training and fine-tuning, and evaluations.
There are no bids on this project.