The ultimate goal is an embodied learner that starts with little priors. But, by through self play, is able to generalize over the physical world. Policy updates require feedback, and when rewards are sparse, how does the learner know what to do?

back