Reinforcement Learning
PPO, SAC, DDPG, actor-critic methods, continuous control, policy regularization, reward design, and multi-seed evaluation.
Open to full-time robotics, reinforcement learning & simulation roles
PhD in Mechanical Engineering
I develop reinforcement-learning controllers, physics-based simulation environments, and sim-to-real workflows for physical systems.
My work focuses on continuous-control reinforcement learning, policy regularization, robotics, and adaptive optics. I developed State-Adaptive Proportional Policy Smoothing (SAPPS), implemented PPO, SAC, and DDPG from scratch in PyTorch, built custom Gymnasium and MuJoCo environments, and developed real-hardware evaluation workflows.
Technical Skills
PPO, SAC, DDPG, actor-critic methods, continuous control, policy regularization, reward design, and multi-seed evaluation.
Gymnasium, MuJoCo, custom physics-based environments, Crazyflie hardware integration, sim-to-real workflows, and robotic skill learning.
Python, PyTorch, Tianshou, NumPy, Linux, HPC/SLURM, Weights & Biases, experiment design, analysis, and visualization.
Feedback control, stability analysis, state-space modelling, system dynamics, MATLAB/Simulink, and adaptive-optics control.
Projects
A focused selection of projects demonstrating original method development, scientific software engineering, simulation, and real-hardware evaluation.
01 · CONTINUOUS-CONTROL RL
A trajectory-local policy-regularization method for continuous-control reinforcement learning that scales the penalty on consecutive action changes according to changes in observed state.
Implemented with PPO and evaluated across MuJoCo continuous-control benchmarks, simulated wavefront-sensorless adaptive optics, and real Crazyflie altitude-control experiments through multi-environment, multi-seed studies.
02 · RL FOR ADAPTIVE OPTICS
A configurable Gymnasium-compatible environment for reinforcement-learning control of wavefront-sensorless adaptive optics in optical satellite communication downlinks.
The environment models atmospheric turbulence, optical propagation, deformable-mirror control, focal-plane sensing, and single-mode-fiber coupling. It includes training and evaluation workflows for PPO, SAC, and DDPG implemented from scratch in PyTorch.
03 · SIM-TO-REAL ROBOTICS
An end-to-end simulation-to-hardware PPO workflow for Crazyflie 2.1 altitude control.
Multi-seed simulation training was followed by 100 real-hardware episodes with real-time telemetry, safety constraints, and automatic landing. The trained policy achieved an inference time of approximately 312 microseconds at a 10 Hz control rate.
04 · PHYSICS-BASED SIMULATION
A completed Gymnasium-compatible MuJoCo research prototype for sequential 3D truck packing, developed as the simulation and reinforcement-learning component of a team project for the 2026 Dexterity Foresight API Challenge.
The environment models randomized box dimensions and masses, MuJoCo contact physics, truck-boundary and axis-aligned overlap checks, displacement-based instability, and three projected occupancy grids. I designed and implemented the simulator and customized PPO training workflow, including continuous placement outputs, four axis-aligned orientation choices, and custom neural-network encoders. The team’s REST API integration was handled separately.
05 · LEARNING FROM DEMONSTRATION
Dynamic Movement Primitive (DMP) models of expert deburring motions derived from human demonstrations recorded with 6-DOF and 1-DOF haptic devices.
I implemented particle swarm optimization for DMP parameter identification, generated trajectories adaptable to new start and goal conditions, developed force-trajectory DMP methods, and conducted experiments using haptic devices and a hexapod platform.
Publications
Payam Parvizi, Colin Bellinger, Ross Cheriton, Abhishek Naik, Davide Spinello
Journal of the Optical Society of America B · Early posting, Jul 20, 2026
Payam Parvizi, Abhishek Naik, Colin Bellinger, Ross Cheriton, Davide Spinello
TechRxiv preprint · Manuscript under revision
Payam Parvizi, Runnan Zou, Colin Bellinger, Ross Cheriton, Davide Spinello
Photonics · 10(12), 1371
Payam Parvizi, Musab Çağrı Uğurlu, Kemal Açıkgöz, Erhan İlhan Konukseven
22nd International Conference on Methods and Models in Automation and Robotics (MMAR) · IEEE · pp. 373–378
Kemal Açıkgöz, Payam Parvizi, Abdulhamit Dönder, Musab Çağrı Uğurlu, Erhan İlhan Konukseven
10th International Conference on Electrical and Electronics Engineering (ELECO) · IEEE · pp. 722–726
PhD thesis · 2025
University of Ottawa
MSc thesis · 2018
Middle East Technical University
EXPERIENCE
University of Ottawa
Advanced SAPPS and led cross-domain implementation, experiment design, multi-environment and multi-seed evaluation, analysis, visualization, and manuscript development using Linux, HPC/SLURM, and Weights & Biases.
University of Ottawa · Collaborative project with the National Research Council of Canada (NRC)
Designed and implemented Adaptive Optics Gym; implemented PPO, SAC, and DDPG from scratch in PyTorch; and led software development, experiments, analysis, figures, repository maintenance, and primary manuscript writing.
University of Ottawa
Served as a Teaching Assistant across seven academic terms in Control Systems, System Dynamics, and Biomedical System Dynamics; nominated for the 2022 University of Ottawa Excellence Award for Teaching Assistants.
Middle East Technical University · TÜBİTAK Grant 114E274
Developed DMP-based robotic deburring skill models, implemented particle swarm optimization and force-trajectory methods, and conducted haptic-device and hexapod experiments.
Contact
I am seeking full-time opportunities in robotics, reinforcement learning, simulation, and applied machine learning for physical systems.