·reading

Reward is Enough

Silver, Singh, Precup, and Sutton's maximalist hypothesis — that every ability we call intelligence falls out of maximizing a single scalar reward.

paper · finished
David Silver, Satinder Singh, Doina Precup, Richard S. Sutton
Artificial Intelligence 299, 103535 (2021)
source ↗

The hypothesis is stated without hedging: knowledge, learning, perception, social intelligence, language, generalization, and imitation all arise in the service of maximizing reward, and a sufficiently powerful agent maximizing a sufficiently rich reward in a sufficiently rich environment will develop them on the way.

I hold this one at arm's length and value it anyway. The argument is genuinely strong for a large class of abilities and genuinely load-bearing on those three "sufficientlys" — which do a great deal of unexamined work. My own affect-based RL work started from the suspicion that a scalar signal is the wrong shape for what actually drives agents, so I read this as the strongest available statement of the position I keep pushing against.

That is what makes it a favorite. A thesis this clean is falsifiable, and you learn more from where it strains than from a hedged version that never commits.

Neighborhood

Related

Grokking: Generalization Beyond OverfittingGrokking: Generalizatio...Language Models are Few-Shot LearnersLanguage Models are Few...On the Measure of IntelligenceOn the Measure of Intel...Rhythms of the BrainRhythms of the BrainThe Bitter LessonThe Bitter LessonThe Brain from Inside OutThe Brain from Inside O...The free-energy principle: a rough guide to the brain?The free-energy princip...The Unreasonable Effectiveness of *The Unreasonable Effect...Research and technical writingResearch and technical ...Broaden and Build Conference 2021Broaden and Build Confe...Broadening and building beyond classical reinforcement learningBroadening and building...Reaching for the IntangibleReaching for the Intang...Broadening and Building Beyond Classical Reinforcement LearningBroadening and Building Bey...What is intelligence?What is intelligence?Predictive General IntelligencePredictive General Intellig...Reward is Enough