<aside>
DeepMind's 2024 paper introducing "Experience Replay" for AI agents. Unlike traditional human-centric RLHF, this approach grounds rewards in real-world signals (health metrics, exam results, environmental data) rather than human preferences. The key insight: agents that learn from environmental consequences can surpass human knowledge and discover novel strategies.
<aside>



Ultimately, experiential data will eclipse the scale and quality of human generated data.
<aside>
What if experiential agents could learn from external events and signals, and not just human preferences?
</aside>