Amin Rakhsha

amin-rakhsha.webp

Toronto, Canada

I am a PhD candidate in Computer Science at the University of Toronto, advised by Amir-massoud Farahmand, and a PhD researcher at the Vector Institute

My research focuses on making intelligent systems that can reason and act robustly over long horizons. My work in reinforcement learning (RL) spans accelerated planning and learning algorithms, model-based RL with imperfect models, and adversarial attacks. More recently, I have focused on LLM reasoning and AI agents, developing efficient inference-time algorithms and evaluation frameworks for targeted post-training, partly during internships at Autodesk AI Lab and Qualcomm AI Research.

Education


University of TorontoPhD in Computer Science
September 2020–October 2026 (expected) · Toronto, Canada
Supervisor: Amir-massoud Farahmand

Sharif University of TechnologyBSc in Computer Engineering
September 2020 · Tehran, Iran

Research Experience


University of Toronto and Vector InstitutePhD Researcher
September 2020–Present · Toronto, Canada

Qualcomm AI ResearchResearch Intern
June–September 2025 · Amsterdam, Netherlands

Autodesk AI LabResearch Intern
February–May 2025 · Toronto, Canada

Max Planck Institute for Software Systems (MPI-SWS)Research Intern and Collaborator
July 2019–January 2021 · Saarbrücken, Germany

Chinese University of Hong Kong (CUHK)Research Intern and Collaborator
July 2018–January 2019 · Hong Kong

Selected Honors and Awards


  • Silver Medal, International Mathematical Olympiad (IMO) — July 2016
  • Borealis AI Global Fellowship — May 2022
  • Three University of Toronto Research Awards — 2020, 2021, 2023
  • Iran’s National Elite Foundation Fellowship — 2015–2020

Research Topics


AI Agents Model Evaluation

Comparison of conventional and specialized model-based planners

What capabilities are needed to solve long-horizon agentic tasks, and which of them is the bottleneck for our model and task?

With LUMINA, we use a POMDP formulation of agentic tasks to identify a set of critical capabilities and develop an evaluation suite that measures the importance of each capability for a model–task pair. The resulting diagnoses guide targeted model post-training and agentic system design.

  1. LUMINA: Long-horizon Understanding for Multi-turn Interactive Agents
    Amin Rakhsha, Thomas Hehn, Pietro Mazzaglia, Fabio Valerio Massoli, Arash Behboodi, and Tribhuvanesh Orekondy
    In Findings of the Association for Computational Linguistics: ACL 2026, Jul 2026

Parallel Test-time Scaling of LLM Reasoning

Preview of the developed agent evaluation framework

How should we utilize parallel inference to scale LLM reasoning when no ground-truth evaluation is available?

We propose Majority-of-the-Bests (MoB), an algorithm for selecting among independently generated LLM outputs. MoB uses bootstrapping to become more robust to noisy reward models compared to Best-of-N.

  1. Majority of the Bests: Improving Best-of-N via Bootstrapping
    Amin Rakhsha, Kanika Madan, Tianyu Zhang, Amir-massoud Farahmand, and Amir Khasahmadi
    In Advances in Neural Information Processing Systems, 2025

Model-based Reinforcement Learning with Imperfect Models

Comparison of conventional and specialized model-based planners

How can an RL agent utilize an erroneous model of the environment?

In many applications, an approximate model of the environment is available that is not completely accurate: a robotic simulator, a pretrained foundation model, or a misspecified learned model. We develop specialized RL algorithms that can benefit from these models while remaining robust to the model’s error.

  1. Maximum Entropy Model Correction in Reinforcement Learning
    Amin Rakhsha, Mete Kemertas, Mohammad Ghavamzadeh, and Amir-massoud Farahmand
    In The Twelfth International Conference on Learning Representations, 2024
  2. Operator Splitting Value Iteration
    Amin Rakhsha, Andrew Wang, Mohammad Ghavamzadeh, and Amir-massoud Farahmand
    In Advances in Neural Information Processing Systems, 2022

Accelerated Reinforcement Learning

Comparison of classical and accelerated reinforcement-learning updates toward the optimal value

How can we design general and scalable acceleration methods for iterative RL algorithms, analogous to those in optimization?

Standard iterative RL algorithms can converge slowly as the task horizon grows. We develop accelerated methods with improved convergence rates while retaining comparable per-iteration cost.

  1. Deflated Dynamics Value Iteration
    Jongmin Lee, Amin Rakhsha, Ernest K. Ryu, and Amir-massoud Farahmand
    Transactions on Machine Learning Research, 2025
  2. RLC
    PID Accelerated Temporal Difference Algorithms
    Mark Bedaywi*†, Amin Rakhsha*, and Amir-massoud Farahmand
    Reinforcement Learning Journal, 2024
  3. Maximum Entropy Model Correction in Reinforcement Learning
    Amin Rakhsha, Mete Kemertas, Mohammad Ghavamzadeh, and Amir-massoud Farahmand
    In The Twelfth International Conference on Learning Representations, 2024
  4. Operator Splitting Value Iteration
    Amin Rakhsha, Andrew Wang, Mohammad Ghavamzadeh, and Amir-massoud Farahmand
    In Advances in Neural Information Processing Systems, 2022

Adversarial Attacks in Reinforcement Learning

Adversarial manipulation of next-state and reward feedback in reinforcement learning

How vulnerable are RL agents to adversarial attacks that alter the state transitions or rewards?

We study the security of RL systems, including training-time attacks that manipulate rewards or transition dynamics. Our work characterizes when an agent can be steered toward an adversarially chosen policy and how costly such attacks must be.

  1. Reward Poisoning in Reinforcement Learning: Attacks Against Unknown Learners in Unknown Environments
    Amin Rakhsha*, Xuezhou Zhang*, Xiaojin Zhu, and Adish Singla
    In NeurIPS Workshop on Learning and Decision-Making with Strategic Feedback, 2021
  2. Policy Teaching in Reinforcement Learning via Environment Poisoning Attacks
    Amin Rakhsha, Goran Radanovic, Rati Devidze, Xiaojin Zhu, and Adish Singla
    Journal of Machine Learning Research, 2021
  3. Policy Teaching via Environment Poisoning: Training-time Adversarial Attacks against Reinforcement Learning
    Amin Rakhsha, Goran Radanovic, Rati Devidze, Xiaojin Zhu, and Adish Singla
    In Proceedings of the 37th International Conference on Machine Learning, 13–18 jul 2020