news

Jul 02, 2026 Our paper LUMINA: Long-horizon Understanding for Multi-turn Interactive Agents appeared in Findings of ACL 2026.
Dec 02, 2025 Our paper Majority of the Bests: Improving Best-of-N via Bootstrapping appeared at NeurIPS 2025. Code and data are available.
May 05, 2025 Deflated Dynamics Value Iteration was published in TMLR.
Aug 09, 2024 Our paper PID Accelerated Temporal Difference Algorithms appeared at RLC 2024.
May 07, 2024 Our paper Maximum Entropy Model Correction in Reinforcement Learning appeared at ICLR 2024.