Non-Parametric Rehearsal Learning via Conditional Mean Embeddings
Nanjing University
Hosted by RIKEN Center for Advanced Intelligence Project
Abstract
How can learning systems recommend actions that prevent an undesirable predicted outcome? Rehearsal learning approaches this problem through influence relations, but existing parametric assumptions can limit its use in complex systems. This talk presents a non-parametric method based on conditional mean embeddings. It expresses the objective in a reproducing kernel Hilbert space, replaces a discontinuous desirability indicator with a smooth Probit surrogate, and uses nested kernel ridge regression to estimate outcome distributions conditional on actions. The resulting estimator can be identified from observational data and has finite-sample error bounds and consistency guarantees. Synthetic and semi-synthetic experiments assess its effectiveness and flexibility.