<aside>

AI agents learn from their interactions with the world

TL;DR

DeepMind's 2024 paper introducing "Experience Replay" for AI agents. Unlike traditional human-centric RLHF, this approach grounds rewards in real-world signals (health metrics, exam results, environmental data) rather than human preferences. The key insight: agents that learn from environmental consequences can surpass human knowledge and discover novel strategies.


<aside>

image.png

image.png

image.png

ReAct : Observations and Actions

Ultimately, experiential data will eclipse the scale and quality of human generated data.

Rewards

<aside>

What if experiential agents could learn from external events and signals, and not just human preferences?

</aside>