Open to full-time robotics, reinforcement learning & simulation roles

Robotics & Reinforcement Learning Engineer

PhD in Mechanical Engineering

I develop reinforcement-learning controllers, physics-based simulation environments, and sim-to-real workflows for physical systems.

My work focuses on continuous-control reinforcement learning, policy regularization, robotics, and adaptive optics. I developed State-Adaptive Proportional Policy Smoothing (SAPPS), implemented PPO, SAC, and DDPG from scratch in PyTorch, built custom Gymnasium and MuJoCo environments, and developed real-hardware evaluation workflows.

  • Ottawa, ON, Canada
  • Authorized to work in Canada
  • Open to relocation across Ontario
Payam Parvizi

Technical Skills

Learning-based control for physical systems

Reinforcement Learning

PPO, SAC, DDPG, actor-critic methods, continuous control, policy regularization, reward design, and multi-seed evaluation.

Simulation & Robotics

Gymnasium, MuJoCo, custom physics-based environments, Crazyflie hardware integration, sim-to-real workflows, and robotic skill learning.

ML & Research Computing

Python, PyTorch, Tianshou, NumPy, Linux, HPC/SLURM, Weights & Biases, experiment design, analysis, and visualization.

Control & Dynamics

Feedback control, stability analysis, state-space modelling, system dynamics, MATLAB/Simulink, and adaptive-optics control.

Projects

Methods, research software, and physical systems

A focused selection of projects demonstrating original method development, scientific software engineering, simulation, and real-hardware evaluation.

Wavefront-sensorless adaptive-optics control system for an optical satellite communication downlink

02 · RL FOR ADAPTIVE OPTICS

Adaptive Optics Gym

A configurable Gymnasium-compatible environment for reinforcement-learning control of wavefront-sensorless adaptive optics in optical satellite communication downlinks.

The environment models atmospheric turbulence, optical propagation, deformable-mirror control, focal-plane sensing, and single-mode-fiber coupling. It includes training and evaluation workflows for PPO, SAC, and DDPG implemented from scratch in PyTorch.

  • Adaptive Optics
  • Gymnasium
  • PyTorch
  • PPO / SAC / DDPG

03 · SIM-TO-REAL ROBOTICS

Crazyflie Reinforcement Learning

An end-to-end simulation-to-hardware PPO workflow for Crazyflie 2.1 altitude control.

Multi-seed simulation training was followed by 100 real-hardware episodes with real-time telemetry, safety constraints, and automatic landing. The trained policy achieved an inference time of approximately 312 microseconds at a 10 Hz control rate.

  • PPO
  • PyTorch
  • Crazyflie
  • cflib
  • Hardware Integration

04 · PHYSICS-BASED SIMULATION

MuJoCo Truck-Packing Simulator & PPO Agent

A completed Gymnasium-compatible MuJoCo research prototype for sequential 3D truck packing, developed as the simulation and reinforcement-learning component of a team project for the 2026 Dexterity Foresight API Challenge.

The environment models randomized box dimensions and masses, MuJoCo contact physics, truck-boundary and axis-aligned overlap checks, displacement-based instability, and three projected occupancy grids. I designed and implemented the simulator and customized PPO training workflow, including continuous placement outputs, four axis-aligned orientation choices, and custom neural-network encoders. The team’s REST API integration was handled separately.

  • MuJoCo
  • Gymnasium
  • Tianshou
  • PPO
  • CNN
Experimental robotic deburring setup with a haptic device and workpiece

05 · LEARNING FROM DEMONSTRATION

Robotic Deburring Skill Learning

Dynamic Movement Primitive (DMP) models of expert deburring motions derived from human demonstrations recorded with 6-DOF and 1-DOF haptic devices.

I implemented particle swarm optimization for DMP parameter identification, generated trajectories adaptable to new start and goal conditions, developed force-trajectory DMP methods, and conducted experiments using haptic devices and a hexapod platform.

  • DMPs
  • Haptics
  • PSO
  • MATLAB
  • Hexapod

Publications

Selected publications

Full record on Google Scholar
  1. 2026

    Action-Regularized Reinforcement Learning for Adaptive Optics in Optical Satellite Communication

    Payam Parvizi, Colin Bellinger, Ross Cheriton, Abhishek Naik, Davide Spinello

    Journal of the Optical Society of America B · Early posting, Jul 20, 2026

  2. 2026

    Adaptive Policy Regularization for Smooth Control in Reinforcement Learning

    Payam Parvizi, Abhishek Naik, Colin Bellinger, Ross Cheriton, Davide Spinello

    TechRxiv preprint · Manuscript under revision

  3. 2023

    Reinforcement Learning Environment for Wavefront Sensorless Adaptive Optics in Single-Mode Fiber Coupled Optical Satellite Communications Downlinks

    Payam Parvizi, Runnan Zou, Colin Bellinger, Ross Cheriton, Davide Spinello

    Photonics · 10(12), 1371

  4. 2017

    Parametrization of Robotic Deburring Process with Motor Skills from Motion Primitives of Human Skill Model

    Payam Parvizi, Musab Çağrı Uğurlu, Kemal Açıkgöz, Erhan İlhan Konukseven

    22nd International Conference on Methods and Models in Automation and Robotics (MMAR) · IEEE · pp. 373–378

  5. 2017

    Dynamic Movement Primitives and Force Feedback: Teleoperation in Precision Grinding Process

    Kemal Açıkgöz, Payam Parvizi, Abdulhamit Dönder, Musab Çağrı Uğurlu, Erhan İlhan Konukseven

    10th International Conference on Electrical and Electronics Engineering (ELECO) · IEEE · pp. 722–726

EXPERIENCE

Experience & Education

  1. Nov 2025 – Mar 2026

    Research Associate (Postdoctoral Fellow)

    University of Ottawa

    Advanced SAPPS and led cross-domain implementation, experiment design, multi-environment and multi-seed evaluation, analysis, visualization, and manuscript development using Linux, HPC/SLURM, and Weights & Biases.

  2. Sep 2022 – Aug 2025

    Research Assistant

    University of Ottawa · Collaborative project with the National Research Council of Canada (NRC)

    Designed and implemented Adaptive Optics Gym; implemented PPO, SAC, and DDPG from scratch in PyTorch; and led software development, experiments, analysis, figures, repository maintenance, and primary manuscript writing.

  3. Jan 2019 – Apr 2022

    Teaching Assistant

    University of Ottawa

    Served as a Teaching Assistant across seven academic terms in Control Systems, System Dynamics, and Biomedical System Dynamics; nominated for the 2022 University of Ottawa Excellence Award for Teaching Assistants.

  4. Nov 2015 – Nov 2017

    Project Assistant

    Middle East Technical University · TÜBİTAK Grant 114E274

    Developed DMP-based robotic deburring skill models, implemented particle swarm optimization and force-trajectory methods, and conducted haptic-device and hexapod experiments.

Selected presentations

  • 2025 Photonics North poster · Action-Regularized RL for Adaptive Optics
  • 2025 University of Ottawa MCG seminar · Wavefront-Sensorless Adaptive Optics
  • 2023 NRC AI for Design seminar · RL-Based Adaptive Optics
  • 2017 MMAR oral presentation · Robotic Deburring Skill Learning

Contact

Let’s connect

I am seeking full-time opportunities in robotics, reinforcement learning, simulation, and applied machine learning for physical systems.