Cevahir Koprulu

I am a Ph.D. student at the Department of Electrical and Computer Engineering, University of Texas at Austin. I am advised by Prof. Ufuk Topcu. I received my B.Sc. degree from the Department of Electrical and Electronics Engineering at Bilkent University, where I worked with Assoc. Prof. Yildiray Yildiz.

Email  /  CV  /  Google Scholar  /  Github  /  LinkedIn Twitter  /  Center for Autonomy

profile photo
Taken at UT Austin, 2024.

News

Research

My research focuses on reinforcement learning (RL): which tasks an agent trains on and what signal it learns from. My work proposes automated curriculum generation methods for heavy tailed task distributions, temporal specifications, and cost constraints, and for the scale of batched simulators. In addition, it develops methods for reward synthesis from task agnostic experience with a few demonstrations, and from physics informed dynamics models with distance aware uncertainty estimates. My publications can be found on my Google Scholar page, and some are listed with details below.

My Ph.D. includes internships at Waymo, on learning from unsupervised driving data within the data flywheel for onboard planner models; at Bosch Center for Artificial Intelligence, on CL4AD, the first integration of CL into batched autonomous driving simulators, accelerating training by 77%; and at Honda Research Institute, on Gen2Spec, an action advising framework that distills knowledge from generalist agents to specialists in a continual learning setting.

Ongoing Work

  • Curricula for Composing Traffic Behaviors: My work guides synthetic traffic scenario generation in batched driving simulators through curricula that compose driving behaviors.
  • Curricula for Reinforcement Unlearning: My work escalates task difficulty during RL to make large language models forget targeted knowledge.

What Agents Train On: Automated Curriculum Generation

A curriculum sequences the tasks an agent trains on and adapts as the agent improves. Standard curriculum learning assumes target task distributions with bounded tails, Markovian rewards, and no safety constraints. My work asks what a curriculum should do when these assumptions fail, and proposes risk aware curricula for heavy tailed task distributions, specification aware curricula guided by reward machines, and safety prioritizing curricula for constrained RL. It also scales curriculum learning and unsupervised environment design to batched autonomous driving simulators that train on real traffic scenarios.

Publications

Scaling Curriculum Learning for Autonomous Driving
Cevahir Koprulu, David Paz, Feng Tao, Yuliang Guo, Xinyu Huang, Ufuk Topcu, Liu Ren
ArXiv, 2026

Safety-Prioritizing Curricula for Constrained Reinforcement Learning
Cevahir Koprulu, Thiago D. Simão, Nils Jansen, Ufuk Topcu
International Conference on Learning Representations (ICLR), 2025

Risk-aware Curriculum Generation for Heavy-tailed Task Distributions
Cevahir Koprulu, Thiago D. Simão, Nils Jansen, Ufuk Topcu
Conference on Uncertainty in Artificial Intelligence (UAI), 2023

Reward-Machine-Guided, Self-Paced Reinforcement Learning
Cevahir Koprulu, Ufuk Topcu
Conference on Uncertainty in Artificial Intelligence (UAI), 2023

Reward-Machine-Guided, Self-Paced Reinforcement Learning
Cevahir Koprulu, Ufuk Topcu
International Conference on Autonomous Agents and Multiagent Systems (AAMAS), 2023 (accepted as Extended Abstract)

Workshops

Prioritizing Safety via Curriculum Learning
Cevahir Koprulu, Thiago D. Simão, Nils Jansen, Ufuk Topcu
RLBRew and RLSW at Reinforcement Learning Conference, 2024

What Agents Learn From: Reward Synthesis from Offline Data and Prior Knowledge

Sparse rewards and imperfect offline data limit how much an agent learns from each transition. My work synthesizes rewards from offline data and prior knowledge. For long horizon tasks with sparse rewards, it combines task agnostic prior experience with a few demonstrations to construct dense rewards through potential based reward shaping. For offline model based RL, it learns dynamics as neural stochastic differential equations that carry prior physics knowledge, and penalizes synthetic rollouts by their uncertainty estimates.

Publications

Dense Dynamics-Aware Reward Synthesis: Integrating Prior Experience with Demonstrations
Cevahir Koprulu, Po-han Li, Tianyu Qiu, Ruihan Zhao, Tyler Westenbroek, David Fridovich-Keil, Sandeep Chinchali, Ufuk Topcu
Learning for Dynamics and Control (L4DC) Conference, 2025

Neural Stochastic Differential Equations for Uncertainty-Aware Offline RL
Cevahir Koprulu, Franck Djeumou, Ufuk Topcu
International Conference on Learning Representations (ICLR), 2025

Other Work

Publications

Joint Learning of Reward Machines and Policies in Environments with Partially Known Semantics
Christos Verginis, Cevahir Koprulu, Sandeep Chinchali, Ufuk Topcu
Artificial Intelligence, 2024

Act to Reason: A Dynamic Game Theoretical Driving Model for Highway Merging Applications
Cevahir Koprulu, Yildiray Yildiz
IEEE Conference on Control Technology and Applications (CCTA), 2021


This website is based on Jon Barron's source code. His website can be found here.