Maybe we will run out of good RL environments so the models won’t get better. Like everything will Be too easy. We will run out of hard environments that make models smarter. You won’t get things like alpha go where you had self play. (Dwarkesh, Noam Brown)

This is again the case for maximizing learnable novelty or interesting structure

back