You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Current post-training learns by interpreting human feedback as a reward to be maximized. However, human usually form preferences by comparing options over the long run, not by chasing a myopic, short score. Hence, treating feedback as a reward maximization does not match how people actually make judgements [1]. This project builds a novel technique of learning from human feedback beyond reward maximization that stays closer to real human judgment, and efficient even when human feedback is limited and costly to collect. Our approach has two parts:
The first is reading human feedback the way people actually produce it, which we call cognitive alignment. In our previous work PPL [2] and RePO [3], we model a preference as regret, meaning a human judgment of how much better a different choice would have been. This gives a more accurate picture of human decision-making, and in our experiments, it improved both alignment and efficiency at the same time. PPL began with robotic manipulation and RePO with LLMs, and this project extends the idea to agentic AI, where a system takes many steps and mixes human and automated feedback, tested on tool use benchmarks [4, 5].
The second part is keeping the learning provably correct when training runs at large scale asynchronously. Frontier post-training relies on asynchronous RL, which decouples rollout generation from policy updates to gain throughput across large clusters. This naturally leads to off-policy, because rollouts become stale and the offline data is heavily filtered and reused. In this regime, the estimators (e.g., importance ratios, divergence term) become unreliable and the policy gradient grows biased and high variance as staleness accumulates. Method such as GRPO [6] recover stability by normalization and hyperparameter tuning, yet no principled criterion tells us which of these corrections remain permissible \[7]. This project characterizes how far the learned reward under ad-hoc off-policy corrections departs from the ground truth.
Today's post-training increasingly uses LLM judges and RLVR to replace human feedback on verifiable tasks such as code and math [1, 2]. But the settings where advanced AI is actually being deployed are dominated by non-verifiable interaction (e.g., service quality in agentic AI, and control and safety in physical/embodied AI) involve an enormous number of human-AI interactions whose correctness cannot be checked automatically. Some work proposes giving even these a process-level automatic signal, yet quantifying the quality of a process is inherently ambiguous, so automated judges cannot fully replace human feedback in these domains [3]. The human role therefore does not disappear as models scale, rather it concentrates on exactly these non-verifiable, process-level judgments, and that irreducible human role is itself a safeguard against AI we can no longer directly supervise.
For that safeguard to hold at scale, human feedback must be interpreted correctly and learned efficiently, but today it is neither. Standard post-training treats a human preference as a reward to maximize. However, a growing line of human-AI alignment work, including PPL, RePO, and regret-minimization formulations [4, 5, 6], shows this is misaligned with how people actually judge. Recently, reward-maximization objectives are provably misspecified estimators whose implicit reward drifts from human intent [7, 8]. Furthermore, reward hacking learned this way generalizes into broad emergent misalignment on unrelated tasks [9, 10, 11]. Misreading the very signal meant to encode human values means that scaling only widens the gap.
This project's core bet is that correctly modeling the human cognitive mechanism [12] is not merely more faithful but a powerful inductive bias that makes learning from scarce human feedback dramatically more sample-efficient. We have demonstrated in two top-AI conference papers (ICML spotlight 2025 and 2026) [4, 5]. This approach is decisive because the common alternative typically fixes one metric while regressing others [8]. That alternative involves letting a model learn a proxy and then trying to correct its behavior after the fact. By contrast, interpreting the cognitive mechanism correctly lets the model internalize the structure of human judgment itself rather than patch its symptoms.
Put together, making human process-level feedback learnable in a sample-efficient, cognitively-aligned way keeps a meaningful human role precisely where automated judges cannot reach. This refers to the non-verifiable interactions that will govern agentic and physical AI. Preserving that existential human role in the loop, rather than designing it out, is how this work contributes to reducing x-risk. All methods and tools will be released open-source so any lab can adopt and verify them.
Personnel top-up (min $13,000, ideal $20,000): dedicated-time supplements for the project lead and a collaborating student, on top of base support already covered by existing funding
LLM-judge / RLVR API (min $8,000, ideal $12,000): external API for generating feedback data and running evaluations across models and trigger conditions (GPU compute itself is covered in-kind by the co-PI's group)
International conference travel (2 people, 1 trip, min $7,000, ideal $9,000): presenting results at a top ML / AI conference (registration + flights + lodging + per diem)
Human-preference & process-level feedback data (min $7,000, ideal $9,000): collecting and annotating the non-verifiable human feedback the project depends on, plus storage and experiment tracking
Taehyun Cho: Vector Distinguished Postdoctoral Fellow and Fields Institute Postdoctoral Fellow at the Vector Institute for Artificial Intelligence. He completed his Ph.D. at the Cognitive Machine Learning Lab, Seoul National University (advised by Prof. Jungwoo Lee), working on reinforcement learning from human feedback and cognitively-aligned reward learning. He is the first author of RePO (ICML 2026) and PPL (ICML 2025), the two papers this project builds on.
Tim G. J. Rudner (co-PI, supervisor): Assistant Professor of Statistics and Computer Science at the University of Toronto, a Canada CIFAR AI Chair and Faculty Member at the Vector Institute, and Chief Scientist at Vijil. His research spans the statistical foundations of machine learning and scalable oversight of frontier AI. He provides supervision and in-kind group resources, including compute and student collaboration, and students from his group will work on the project.
Track record: PPL [2] (robotic manipulation) and RePO [3] (LLMs) are exactly the methods this project extends. Both were accepted at ICML (2025 and 2026) as a spotlight and showed that modeling preferences as regret improves alignment and sample efficiency at the same time.
When the data mixes human and automated feedback in a hybrid way, it becomes harder to verify and analyze, and the outcome could end up worse than relying on either source alone.
For the off-policy correction, if compute is abundant the statistical inefficiency can simply be overcome by collecting more data, which can make the correction terms less meaningful. Minimizing variance to stabilize training also does not always improve actual performance, so when such a trade-off arises we may need to define a sound evaluation criterion for judging which corrections are worthwhile.
Fields Institute – Principles of Intelligence Postdoctoral Fellowship (Jul 2026–Jun 2027): a secondment with no salary from Fields or PI, plus up to CAD $10,000 for travel from Prof. Rudner's grant.
This grantmaking.ai award ($35,000), now being set up through Manifund. (https://app.grantmaking.ai/projects/9fe5e9f7-98d2-4129-a7b7-95a92906e645)
There are no bids on this project.