My research focuses on making intelligent systems that can reason and act robustly over long horizons. My work in reinforcement learning (RL) spans accelerated planning and learning algorithms, model-based RL with imperfect models, and adversarial attacks. More recently, I have focused on LLM reasoning and AI agents, developing efficient inference-time algorithms and evaluation frameworks for targeted post-training, partly during internships at Autodesk AI Lab and Qualcomm AI Research.
Education
University of Toronto — PhD in Computer Science September 2020–October 2026 (expected) · Toronto, Canada Supervisor: Amir-massoud Farahmand
Sharif University of Technology — BSc in Computer Engineering September 2020 · Tehran, Iran
Research Experience
University of Toronto and Vector Institute — PhD Researcher September 2020–Present · Toronto, Canada
Qualcomm AI Research — Research Intern June–September 2025 · Amsterdam, Netherlands
Autodesk AI Lab — Research Intern February–May 2025 · Toronto, Canada
Max Planck Institute for Software Systems (MPI-SWS) — Research Intern and Collaborator July 2019–January 2021 · Saarbrücken, Germany
Chinese University of Hong Kong (CUHK) — Research Intern and Collaborator July 2018–January 2019 · Hong Kong
Selected Honors and Awards
Silver Medal, International Mathematical Olympiad (IMO) — July 2016
Borealis AI Global Fellowship — May 2022
Three University of Toronto Research Awards — 2020, 2021, 2023
Iran’s National Elite Foundation Fellowship — 2015–2020
Research Topics
AI Agents Model Evaluation
What capabilities are needed to solve long-horizon agentic tasks, and which of them is the bottleneck for our model and task?
With LUMINA, we use a POMDP formulation of agentic tasks to identify a set of critical capabilities and develop an evaluation suite that measures the importance of each capability for a model–task pair. The resulting diagnoses guide targeted model post-training and agentic system design.
@inproceedings{rakhsha2026lumina,title={{LUMINA}: Long-horizon Understanding for Multi-turn Interactive Agents},author={Rakhsha, Amin and Hehn, Thomas and Mazzaglia, Pietro and Massoli, Fabio Valerio and Behboodi, Arash and Orekondy, Tribhuvanesh},booktitle={Findings of the Association for Computational Linguistics: ACL 2026},month=jul,year={2026},address={San Diego, California, United States},publisher={Association for Computational Linguistics},pages={3913--3926},doi={10.18653/v1/2026.findings-acl.190},}
Parallel Test-time Scaling of LLM Reasoning
How should we utilize parallel inference to scale LLM reasoning when no ground-truth evaluation is available?
We propose Majority-of-the-Bests (MoB), an algorithm for selecting among independently generated LLM outputs. MoB uses bootstrapping to become more robust to noisy reward models compared to Best-of-N.
@inproceedings{rakhsha2025majority,title={Majority of the Bests: Improving {Best-of-N} via Bootstrapping},author={Rakhsha, Amin and Madan, Kanika and Zhang, Tianyu and Farahmand, Amir-massoud and Khasahmadi, Amir},booktitle={Advances in Neural Information Processing Systems},volume={38},pages={37844--37877},year={2025},publisher={Curran Associates, Inc.},doi={10.52202/085713-1268},}
Model-based Reinforcement Learning with Imperfect Models
How can an RL agent utilize an erroneous model of the environment?
In many applications, an approximate model of the environment is available that is not completely accurate: a robotic simulator, a pretrained foundation model, or a misspecified learned model. We develop specialized RL algorithms that can benefit from these models while remaining robust to the model’s error.
@inproceedings{rakhsha2024maximum,title={Maximum Entropy Model Correction in Reinforcement Learning},author={Rakhsha, Amin and Kemertas, Mete and Ghavamzadeh, Mohammad and Farahmand, Amir-massoud},booktitle={The Twelfth International Conference on Learning Representations},year={2024},}
@inproceedings{rakhsha2022operator,title={Operator Splitting Value Iteration},author={Rakhsha, Amin and Wang, Andrew and Ghavamzadeh, Mohammad and Farahmand, Amir-massoud},booktitle={Advances in Neural Information Processing Systems},volume={35},pages={38373--38385},year={2022},publisher={Curran Associates, Inc.},doi={10.52202/068431-2780},}
Accelerated Reinforcement Learning
How can we design general and scalable acceleration methods for iterative RL algorithms, analogous to those in optimization?
Standard iterative RL algorithms can converge slowly as the task horizon grows. We develop accelerated methods with improved convergence rates while retaining comparable per-iteration cost.
@article{lee2025deflated,title={Deflated Dynamics Value Iteration},author={Lee, Jongmin and Rakhsha, Amin and Ryu, Ernest K. and Farahmand, Amir-massoud},journal={Transactions on Machine Learning Research},year={2025},}
@article{bedaywi2024pid,title={{PID} Accelerated Temporal Difference Algorithms},author={Bedaywi, Mark and Rakhsha, Amin and Farahmand, Amir-massoud},journal={Reinforcement Learning Journal},volume={5},pages={2071--2095},year={2024},}
@inproceedings{rakhsha2024maximum,title={Maximum Entropy Model Correction in Reinforcement Learning},author={Rakhsha, Amin and Kemertas, Mete and Ghavamzadeh, Mohammad and Farahmand, Amir-massoud},booktitle={The Twelfth International Conference on Learning Representations},year={2024},}
@inproceedings{rakhsha2022operator,title={Operator Splitting Value Iteration},author={Rakhsha, Amin and Wang, Andrew and Ghavamzadeh, Mohammad and Farahmand, Amir-massoud},booktitle={Advances in Neural Information Processing Systems},volume={35},pages={38373--38385},year={2022},publisher={Curran Associates, Inc.},doi={10.52202/068431-2780},}
Adversarial Attacks in Reinforcement Learning
How vulnerable are RL agents to adversarial attacks that alter the state transitions or rewards?
We study the security of RL systems, including training-time attacks that manipulate rewards or transition dynamics. Our work characterizes when an agent can be steered toward an adversarially chosen policy and how costly such attacks must be.
@inproceedings{rakhsha2021reward,title={Reward Poisoning in Reinforcement Learning: Attacks Against Unknown Learners in Unknown Environments},author={Rakhsha, Amin and Zhang, Xuezhou and Zhu, Xiaojin and Singla, Adish},booktitle={NeurIPS Workshop on Learning and Decision-Making with Strategic Feedback},year={2021},}
@article{rakhsha2021policy,title={Policy Teaching in Reinforcement Learning via Environment Poisoning Attacks},author={Rakhsha, Amin and Radanovic, Goran and Devidze, Rati and Zhu, Xiaojin and Singla, Adish},journal={Journal of Machine Learning Research},volume={22},number={210},pages={1--45},year={2021},}
@inproceedings{rakhsha2020policy,title={Policy Teaching via Environment Poisoning: Training-time Adversarial Attacks against Reinforcement Learning},author={Rakhsha, Amin and Radanovic, Goran and Devidze, Rati and Zhu, Xiaojin and Singla, Adish},booktitle={Proceedings of the 37th International Conference on Machine Learning},editor={Daum\'{e} III, Hal and Singh, Aarti},volume={119},series={Proceedings of Machine Learning Research},pages={7974--7984},month={13--18 Jul},year={2020},publisher={PMLR},}