|
Mitsuhiko Nakamoto
Hi, I'm Mitsuhiko! I am a final-year PhD candidate in the Berkeley Artificial Intelligence Research (BAIR) Lab at UC
Berkeley, advised by Professor Sergey
Levine.
My research goal is to develop algorithms that empower robots with human-level dexterity and
reasoning capabilities, and to integrate them into everyday life.
To this end, my research focuses on data-driven approaches to solving real-world robotic
tasks, including offline reinforcement learning (RL), online RL fine-tuning, and
imitation learning.
Recently, I have been particularly interested in using RL to post-train generalist robot
policies to improve their low-level dexterity and high-level reasoning capabilities.
I had a great opportunity to spend time as a research intern at TRI in Cambridge, MA
(2025.06-2025.08) and at Google DeepMind in London, UK (2025.09-2025.12). Prior to Berkeley, I
received
my bachelor's degree from
University of Tokyo in March 2022.
Email /
Google Scholar
/
GitHub /
Twitter /
Talks
|
|
|
Research
* denotes equal contribution, † denotes core contributor
|
Conference Papers and Pre-prints
|
|
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies
Sarvesh Patil*, Mitsuhiko Nakamoto*, Manan Agarwal, Shashwat Saxena, Jesse Zhang,
Giri Anantharaman, Cleah Winston, Chaoyi Pan, Douglas Chen, Nai-Chieh Huang, Zeynep Temel, Oliver
Kroemer, Sergey Levine, Abhishek Gupta, Hongkai Dai, Paarth Shah, Max Simchowitz.
ICML 2026
arXiv link /
project page /
code
We propose OGPO (Off-Policy Generative Policy Optimization), a sample-efficient algorithm for full
fine-tuning of diffusion- and flow-based control policies. OGPO combines off-policy Q-learning with
GRPO-style policy updates to enable fast & stable fine-tuning.
|
|
|
Steering Your Diffusion Policy with Latent Space Reinforcement Learning
Andrew Wagenmaker†, Mitsuhiko Nakamoto†, Yunchu
Zhang†, Seohong Park, Waleed Yagoub,
Anusha Nagabandi, Abhishek Gupta†, Sergey Levine†.
CoRL 2025 (Oral Presentation, Best Paper Award
Nomination)
arXiv link /
project page /
code
We propose DSRL (Diffusion Steering via Reinforcement Learning), a sample-efficient and lightweight
approach for improving pre-trained diffusion- or flow-based policies by steering the initial noise
input through RL.
We successfully demonstrate real-world RL-based fine-tuning of Pi-0, a pre-trained generalist policy
from Physical Intelligence.
|
Journal Papers
Deep Learning Algorithm to Detect Cardiac Sarcoidosis From Echocardiographic Movies
S.Katsushika, S.Kodera, M.Nakamoto, K.Ninomiya, N.Kakuda, H.Shinohara, R.Matsuoka,
H.Ieki, M.Uehara, Y.Higashikuni, K.Nakanishi, T.Nakao, N.Takeda, K.Fujiu, M.Daimon, J.Ando,
H.Akazawa,
H.Morita, I.Komuro.
In Circulation Journal, 2021
paper
link
|
Deep learning model to detect significant aortic regurgitation using
electrocardiography
S.Sawano, S.Kodera, S.Katsushika, M.Nakamoto, K.Ninomiya, H.Shinohara,
Y.Higashikuni,
K.Nakanishi, T.Nakao, T.Seki, N.Takeda, K.Fujiu, M.Daimon, H.Akazawa, H.Morita, I.Komuro.
In Journal of Cardiology, 2021
paper
link
|
Automatic detection of vessel structure by deep learning using intravascular
ultrasound
images of the coronary arteries
H.Shinohara, S.Kodera, K.Ninomiya, M.Nakamoto, KS.Katsushika, A.Saito, S.Minatsuki,
H.Kikuchi, A.Kiyosue, Y.Higashikuni, N.Takeda, K.Fujiu, J.Ando, H.Akazawa, H.Morita, I.Komuro.
In PLOS ONE, 2021
paper
link
|
|
|
Teleoperated Robot
- Technologies: CoppeliaSim, Arduino, Servo Motor, 3D CAD design, 3D Printer, potentiometers
- [Demo
videos]
|
|
|
Wireless Remote Switch Toggler
- A device for turning on/off the electric switch remotely using a smartphone.
- Technologies: ESP32, Servo Motor, BLE socket programming, electronic design (EAGLE)
- [Demo
videos]
|
|
Teaching
|
Services
- Reviews: NeurIPS (2024, 2025), ICLR (2024, 2025, 2026), ICML (2025, 2026), ICRA (2025, 2026),
CoRL (2025, 2026),
RA-L
|
I am from Yokohama, Japan. My name in Japanese is 中本光彦.
During my free time, I enjoy playing and watching sports, especially soccer. I also enjoy playing
Japanese chess (aka Shogi).
|
|