|
Strategic Games: Having peaked at server rank #9 in Teamfight Tactics (TFT) during a competitive season, I treat strategic games as experimental laboratories for studying decision-making under uncertainty. Competitive multi-agent environments like TFT require continuous adaptation, opponent modeling, and risk-sensitive planning. By analyzing my own high-level gameplay, I study how strategies are formed and adjusted in response to stochastic dynamics and strategic interaction. I am particularly interested in formalizing how these strategies emerge in competitive settings and translating these principles into reinforcement learning agents.
Behavioral Economics and Cognitive Science: Behavioral economics and cognitive science play a central role in how I think about human decision-making. Many reinforcement learning formulations focus primarily on maximizing expected return, abstracting away inference-time constraints and bounded rationality. I am interested in understanding how human biases, risk perception, and cognitive limitations shape actual decisions. This perspective motivates my research on modeling human-aligned decision processes that go beyond expectation maximization and instead account for uncertainty, regret, and subjective evaluation of outcomes.
Philosophy: Since 2020, I have been part of a regular philosophy gathering called "Sunday Salon." Our dialogues, initially inspired by French "Baccalauréat" topics, have expanded to address AI and its societal implications. These discussions have shaped my thinking on human uniqueness, responsibility, and subjective fulfillment. I am focused on understanding which aspects of human reasoning (such as regret, narrative self-reflection, and emotional responses) remain difficult to formalize, and how they might inform the development of human-aligned AI systems.
|
|
My research focuses on sequential decision-making under uncertainty, particularly in the context of human feedback.
I have extensively studied distributional reinforcement learning (DistRL), reinforcement learning from human feedback (RLHF), and regret analysis,
aiming to bridge theory and practice.
My long-term goal is to build human-aligned, socially aware agents — systems that reason about people, navigate the dynamics of human interaction, and act on their preferences, values, and decisions.
I draw inspiration from how humans make decisions. By seeking to understand the cognitive mechanisms underlying human choice and mathematically modeling the structure that governs interaction,
I aim to develop both theoretical insights and practical algorithms for robust decision-making under uncertainty.
Currently, I am focusing on (i) a regret minimization framework for agentic AI,
and (ii) the theoretical foundations of off-policy post-training.
Feel free to reach out if you are interested!
|
|
A Distributional Perspective on Human-Aligned Decision Making under Uncertainty
Taehyun Cho
Department of Electrical and Computer Engineering, Seoul National University Distinguished Dissertation Award
paper /
|
|
The Unique Off-Policy Permissibility of KL-divergence in Direct Preference Optimization
Taehyun Cho, Amirabbas Afzali, Suhwan Kim, Tim G. J. Rudner
Work In Progress
|
|
An Axiomatization of Process Score Model: Your Process-level Feedback is Not a Reward
Taehyun Cho
Work In Progress
|
|
Off-Policy Trust Region Preference Optimization
Seungyub Han*, Taehyun Cho*, Suhwan Kim, Seokhun Ju, Dohyeong Kim, Kyungjae Lee, Jungwoo Lee
Work In Progress
|
|
When to Truncate Traces: Stochastic Truncation for Multi-Step Off-Policy RL
Seungyub Han, Taehyun Cho, Dohyeong Kim, Kyungjae Lee, Jungwoo Lee
Under submission
|
|
Probabilistic Smoothing with Ratio-Monotone Transforms for Global Optimization
Kukyoung Jang, Taehyun Cho, Junrui Zhang, Ping Xu, Kyungjae Lee
Under submission
|
|
A Regret Minimization Framework on Preference Learning in Large Language Models
Suhwan Kim*, Taehyun Cho*, Geonhyeong Kim, Yujin Kim, Youngsoo Jang, Moontae Lee, Jungwoo Lee
ICML 2026 Spotlight (Top 2.2%)
paper /
project page /
slides /
|
|
MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint Reconstruction
Jung Min Lee, Dohyeok Lee, Seokhun Ju, Taehyun Cho, Jin Woo Koo, Li Zhao, Sangwoo Hong, Jungwoo Lee
ICML 2026
paper /
project page /
|
|
Pareto Optimal Risk-Agnostic Distributional Bandits with Heavy-Tail Rewards
Kyungjae Lee, Dohyeong Kim, Taehyun Cho, Chaeyeon Kim, Yunkyung Ko, Seungyub Han, Seokhun Ju, Dohyeok Lee, Sungbin Lim
NeurIPS 2025
paper /
|
|
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
Taehyun Cho*, Seokhun Ju*, Seungyub Han, Dohyeong Kim, Kyungjae Lee, Jungwoo Lee
ICML 2025 Spotlight (Top 2.6%)
paper /
arxiv /
|
|
Bellman Unbiasedness: Toward Provably Efficient Distributional Reinforcement Learning with General Value Function Approximation
Taehyun Cho, Seungyub Han, Seokhun Ju, Dohyeong Kim, Kyungjae Lee, Jungwoo Lee
ICML 2025
paper /
arxiv /
|
|
Spectral-Risk Safe Reinforcement Learning with Convergence Guarantees
Dohyeong Kim, Taehyun Cho, Seungyub Han, Hojun Chung, Kyungjae Lee, Songhwai Oh
NeurIPS 2024
paper /
arxiv /
|
|
Pitfall of Optimism: Distributional Reinforcement Learning by Randomizing Risk Criterion
Taehyun Cho, Seungyub Han, Heesoo Lee, Kyungjae Lee, Jungwoo Lee
NeurIPS 2023
paper /
arxiv /
|
|
SPQR: Controlling Q-ensemble Independence with Spiked Random Model for Reinforcement Learning
Dohyeok Lee, Seungyub Han, Taehyun Cho, Jungwoo Lee
NeurIPS 2023
paper /
arxiv /
code /
|
|
On the Convergence of Continual Learning with Adaptive Methods
Seungyub Han, Yeongmo Kim, Taehyun Cho, Jungwoo Lee
UAI 2023
paper /
arxiv /
|
|
Adaptive Methods for Nonconvex Continual Learning
Seungyub Han, Yeongmo Kim, Taehyun Cho, Jungwoo Lee
NeurIPS 2022 Optimization for Machine Learning Workshop
paper /
|
|
Perturbed Quantile Regression for Distributional Reinforcement Learning
Taehyun Cho, Seungyub Han, Heesoo Lee, Kyungjae Lee, Jungwoo Lee
NeurIPS 2022 Deep RL Workshop
paper /
|
|
Chebyshev polynomial codes: Task entanglement-based coding for distributed matrix multiplication
Sangwoo Hong, Heecheol Yang, Youngseok Yoon, Taehyun Cho, Jungwoo Lee
ICML 2021
paper /
arxiv /
|
|
Learning Graph Based Individual Intrinsic Reward For Multi-Agent Reinforcement Learning
Seokhun Ju, Seungyub Han, Taehyun Cho, Jungwoo Lee, Taeyoung Lee, Minkyoung Kim, Jinho Ahn
ICT Express
paper /
|
|
Optimized shallow neural networks for sum-rate maximization in energy harvesting downlink multiuser NOMA systems
Heasung Kim, Taehyun Cho, Jungwoo Lee, Wonjae Shin, H. Vincent Poor
IEEE Journal on Selected Areas in Communications
paper /
arxiv /
|
|
An Efficient Neural Network Architecture for Rate Maximization in Energy Harvesting Downlink Channels
Heasung Kim, Taehyun Cho, Jungwoo Lee, Wonjae Shin, H. Vincent Poor
2020 IEEE International Symposium on Information Theory (ISIT)
paper /
arxiv /
|
|
Grantmaking.ai Research Grant
Faithful Preference Learning: Cognitively-Aligned Post-Training
Funded by Grantmaking.ai
(
2026.09 - 2027.09
)
Total funding: USD 50K
|
|
Fields Institute - Principles of Intelligence Postdoctoral Fellowship
Funded by Fields Institute
(
2026.08 - 2027.08
)
|
|
Sejong Science Fellowship
Distributional Regret Analysis for Human-Aligned Interactive Agentic AI under High Uncertainty
Funded by National Research Foundation of Korea
(
2026.03 - 2031.02
)
Total funding: KRW 600M
|
|
Vector Institute
Vector Distinguished Postdoctoral Fellow
2026.08 - 2027.08 · Toronto, Canada
|
|
Fields Institute
Principles of Intelligence Postdoctoral Fellow
2026.08 - 2027.08 · Toronto, Canada
|
|
Seoul National University
Postdoctoral Fellow, Cognitive Machine Learning Lab
2026.03 - Present · Seoul, South Korea
|
|
LG AI Research
Research Intern, Superintelligence Lab
2024.12 - 2025.05
|
|
Seoul National University
Ph.D./M.S. in Electrical and Computer Engineering
2020.03 - 2026.02
|
|
Korea University
B.S. in Mathematics
2013.03 - 2020.02
|
Sep 2026 – Sep 2027 — Research Grant (USD 50K)
Grantmaking.ai
|
Aug 2026 – Aug 2027 — Vector Distinguished Postdoctoral Fellowship
Vector Institute
|
Aug 2026 – Aug 2027 — Principles of Intelligence Postdoctoral Fellowship
Fields Institute
|
Jul 2026 — SK Inc. AX Best Research Talk Award
Cortiq Summit 2026: Agentic AI
|
May 2026 — INMC Young Researcher Award
Institute of New Media and Communications (INMC)
|
May 2026 — Gold Reviewer Award (Top 25%)
International Conference on Machine Learning (ICML)
|
Mar 2026 – Feb 2031 — Sejong Science Fellowship
National Research Foundation of Korea (NRF)
|
Feb 2026 — Distinguished Dissertation Award
Seoul National University (SNU)
|
Jan 2023 — Certificate of Commendation
Center for Applied Research in Artificial Intelligence (CARAI)
|
Mar 2020 – Feb 2026 — Brain Korea 21 Plus Scholarship
Seoul National University (SNU)
|
Jul 2026 — RIKEN-AIP · Vector · NAIRL Workshop, KAIST Seocho AICT
A Regret Minimization Framework on Preference Learning in Large Language Models
|
Jul 2026 — Cortiq Summit, Seoul COEX
A Regret Minimization Framework on Preference Learning in Large Language Models
|
Apr 2026 — Vector Visitor Research Talk, Vector Institute
Rethinking Human Feedback: A Regret Minimization Perspective on Preference Learning
|
Apr 2026 — SNU AI Summit, Seoul National University
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
|
May 2025 — LG AI Research Seminar, LG AI Research
Policy Optimization with Process Score in LRMs
|
Dec 2024 — LG AI Research Seminar, LG AI Research
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
|
Aug 2023 — LG Tech Talk, LG AI Research
Pitfall of Optimism: Distributional Reinforcement Learning with Randomized Risk Criterion
|
May 2023 — AIIS Spring Retreat Program, Seoul National University
Pitfall of Optimism: Distributional Reinforcement Learning with Randomized Risk Criterion
|
Conference Reviewer
- International Conference on Machine Learning (ICML) 2022 – 2026
- Gold Reviewer Award (Top 25%)
- Neural Information Processing Systems (NeurIPS) 2022 – 2025
- International Conference on Learning Representations (ICLR) 2023 – 2025
- AAAI Conference on Artificial Intelligence (AAAI) 2023
Workshop Reviewer
- RLxF: Reinforcement Learning from World Feedback @ ICML 2026
Journal Reviewer
- Transactions on Machine Learning Research (TMLR) 2026
- The Journal of Korean Institute of Communications and Information Sciences (JKICS) 2026
|
|