Arya Ebrahimi

I am a Master's student at the University of Alberta, supervised by Dr. Jun Jin. I work at the intersection of robotics and machine learning, with a focus on continual reinforcement learning. My goal is to build robots that keep learning throughout their lifetime, using the broad priors of foundation models as a starting point and reinforcement learning to keep improving from their own experience.


Education
  • University of Alberta
    University of Alberta
    Department of Electrical and Computer Engineering
    M.Sc. in Computer Engineering
    May. 2025 - present
  • Ferdowsi University of Mashhad
    Ferdowsi University of Mashhad
    B.Sc. in Computer Engineering
    Sep. 2019 - Feb. 2024
News
2025
I am attending the ICML 2025 in Vancouver!
Jul 01
Selected Publications (view all )
Retrospective and Structurally Informed Exploration via Cross-task Successor Feature Similarity
Retrospective and Structurally Informed Exploration via Cross-task Successor Feature Similarity

Xirui Shi, Arya Ebrahimi, Yi Hu, Jun Jin

arXiv 2026

Diffusion models are increasingly used for robot learning, but current designs face a clear trade-off. Action-chunking diffusion policies like ManiCM are fast to run, yet they only predict short segments of motion. This makes them reactive, but unable to capture time-dependent motion primitives, such as following a spring-damper-like behavior with built-in dynamic profiles of acceleration and deceleration. Recently, Movement Primitive Diffusion (MPD) partially addresses this limitation by parameterizing full trajectories using Probabilistic Dynamic Movement Primitives (ProDMPs), thereby enabling the generation of temporally structured motions. Nevertheless, MPD integrates the motion decoder directly into a multi-step diffusion process, resulting in prohibitively high inference latency that limits its applicability in real-time control settings. We propose FODMP (Fast One-step Diffusion of Movement Primitives), a new framework that distills diffusion models into the ProDMPs trajectory parameter space and generates motion using a single-step decoder. FODMP retains the temporal structure of movement primitives while eliminating the inference bottleneck through single-step consistency distillation. This enables robots to execute time-dependent primitives at high inference speed, suitable for closed-loop vision-based control. On standard manipulation benchmarks (MetaWorld, ManiSkill), FODMP runs up to 10 times faster than MPD and 7 times faster than action-chunking diffusion policies, while matching or exceeding their success rates. Beyond speed, by generating fast acceleration-deceleration motion primitives, FODMP allows the robot to intercept and securely catch a fast-flying ball,whereas action-chunking diffusion policy and MPD respond too slowly for real-time interception.

Retrospective and Structurally Informed Exploration via Cross-task Successor Feature Similarity

Xirui Shi, Arya Ebrahimi, Yi Hu, Jun Jin

arXiv 2026

Diffusion models are increasingly used for robot learning, but current designs face a clear trade-off. Action-chunking diffusion policies like ManiCM are fast to run, yet they only predict short segments of motion. This makes them reactive, but unable to capture time-dependent motion primitives, such as following a spring-damper-like behavior with built-in dynamic profiles of acceleration and deceleration. Recently, Movement Primitive Diffusion (MPD) partially addresses this limitation by parameterizing full trajectories using Probabilistic Dynamic Movement Primitives (ProDMPs), thereby enabling the generation of temporally structured motions. Nevertheless, MPD integrates the motion decoder directly into a multi-step diffusion process, resulting in prohibitively high inference latency that limits its applicability in real-time control settings. We propose FODMP (Fast One-step Diffusion of Movement Primitives), a new framework that distills diffusion models into the ProDMPs trajectory parameter space and generates motion using a single-step decoder. FODMP retains the temporal structure of movement primitives while eliminating the inference bottleneck through single-step consistency distillation. This enables robots to execute time-dependent primitives at high inference speed, suitable for closed-loop vision-based control. On standard manipulation benchmarks (MetaWorld, ManiSkill), FODMP runs up to 10 times faster than MPD and 7 times faster than action-chunking diffusion policies, while matching or exceeding their success rates. Beyond speed, by generating fast acceleration-deceleration motion primitives, FODMP allows the robot to intercept and securely catch a fast-flying ball,whereas action-chunking diffusion policy and MPD respond too slowly for real-time interception.

Retrospective and Structurally Informed Exploration via Cross-task Successor Feature Similarity
Retrospective and Structurally Informed Exploration via Cross-task Successor Feature Similarity

Arya Ebrahimi, Jun Jin

The Exploration in AI Today Workshop at ICML 2025

We introduce Cross-task Successor Feature Similarity Exploration (C-SFSE), a novel intrinsic reward mechanism that leverages retrospective similarities in task-conditioned successor features to prioritize exploration of semantically meaningful states. C-SFSE constructs a cross-task similarity signal from previously learned policies, identifying regions, such as bottlenecks or reusable subgoals, that consistently support goal-directed behavior. This enables the agent to focus its exploration on state space areas that are not only novel but informative across tasks.

Retrospective and Structurally Informed Exploration via Cross-task Successor Feature Similarity

Arya Ebrahimi, Jun Jin

The Exploration in AI Today Workshop at ICML 2025

We introduce Cross-task Successor Feature Similarity Exploration (C-SFSE), a novel intrinsic reward mechanism that leverages retrospective similarities in task-conditioned successor features to prioritize exploration of semantically meaningful states. C-SFSE constructs a cross-task similarity signal from previously learned policies, identifying regions, such as bottlenecks or reusable subgoals, that consistently support goal-directed behavior. This enables the agent to focus its exploration on state space areas that are not only novel but informative across tasks.

All publications