| ||
News
|
||
ResearchMy research focuses on reinforcement learning (RL): which tasks an agent trains on and what signal it learns from. My work proposes automated curriculum generation methods for heavy tailed task distributions, temporal specifications, and cost constraints, and for the scale of batched simulators. In addition, it develops methods for reward synthesis from task agnostic experience with a few demonstrations, and from physics informed dynamics models with distance aware uncertainty estimates. My publications can be found on my Google Scholar page, and some are listed with details below. My Ph.D. includes internships at Waymo, on learning from unsupervised driving data within the data flywheel for onboard planner models; at Bosch Center for Artificial Intelligence, on CL4AD, the first integration of CL into batched autonomous driving simulators, accelerating training by 77%; and at Honda Research Institute, on Gen2Spec, an action advising framework that distills knowledge from generalist agents to specialists in a continual learning setting. Ongoing Work
What Agents Train On: Automated Curriculum GenerationA curriculum sequences the tasks an agent trains on and adapts as the agent improves. Standard curriculum learning assumes target task distributions with bounded tails, Markovian rewards, and no safety constraints. My work asks what a curriculum should do when these assumptions fail, and proposes risk aware curricula for heavy tailed task distributions, specification aware curricula guided by reward machines, and safety prioritizing curricula for constrained RL. It also scales curriculum learning and unsupervised environment design to batched autonomous driving simulators that train on real traffic scenarios.PublicationsCevahir Koprulu, David Paz, Feng Tao, Yuliang Guo, Xinyu Huang, Ufuk Topcu, Liu Ren ArXiv, 2026 Cevahir Koprulu, Thiago D. Simão, Nils Jansen, Ufuk Topcu International Conference on Learning Representations (ICLR), 2025 Cevahir Koprulu, Thiago D. Simão, Nils Jansen, Ufuk Topcu Conference on Uncertainty in Artificial Intelligence (UAI), 2023 Cevahir Koprulu, Ufuk Topcu Conference on Uncertainty in Artificial Intelligence (UAI), 2023 Cevahir Koprulu, Ufuk Topcu International Conference on Autonomous Agents and Multiagent Systems (AAMAS), 2023 (accepted as Extended Abstract) WorkshopsCevahir Koprulu, Thiago D. Simão, Nils Jansen, Ufuk Topcu RLBRew and RLSW at Reinforcement Learning Conference, 2024 What Agents Learn From: Reward Synthesis from Offline Data and Prior KnowledgeSparse rewards and imperfect offline data limit how much an agent learns from each transition. My work synthesizes rewards from offline data and prior knowledge. For long horizon tasks with sparse rewards, it combines task agnostic prior experience with a few demonstrations to construct dense rewards through potential based reward shaping. For offline model based RL, it learns dynamics as neural stochastic differential equations that carry prior physics knowledge, and penalizes synthetic rollouts by their uncertainty estimates.PublicationsCevahir Koprulu, Po-han Li, Tianyu Qiu, Ruihan Zhao, Tyler Westenbroek, David Fridovich-Keil, Sandeep Chinchali, Ufuk Topcu Learning for Dynamics and Control (L4DC) Conference, 2025 Cevahir Koprulu, Franck Djeumou, Ufuk Topcu International Conference on Learning Representations (ICLR), 2025 Other WorkPublicationsChristos Verginis, Cevahir Koprulu, Sandeep Chinchali, Ufuk Topcu Artificial Intelligence, 2024 Cevahir Koprulu, Yildiray Yildiz IEEE Conference on Control Technology and Applications (CCTA), 2021 |
|
This website is based on Jon Barron's source code.
His website can be found here.
|