Publications

Overview of the language-critique imitation learning framework: an agent receives both a scalar score and structured natural-language critique describing and labeling its action.
arXiv preprint

Language-Critique Imitation Learning from Suboptimal Demonstrations

We propose a language-critique framework for imitation learning from suboptimal demonstrations that instead leverages natural language as a structured supervision signal, avoiding the collapse of expressive feedback into scalars.

Overview of the Hierarchical Programmatic Option framework (HIPO): a high-level policy selects human-readable programmatic options that execute in the environment.
NeurIPS 2024

Hierarchical Programmatic Option Framework

Yu-An Lin*, Chen-Tao Lee*, Chih-Han Yang*, Guan-Ting Liu*, Shao-Hua Sun

Deep reinforcement learning aims to learn deep neural network policies to solve large-scale decision-making problems. However, approximating policies using deep neural networks makes it difficult to interpret the learned decision-making process.

Awards