Mitsuhiko Nakamoto

Hi, I'm Mitsuhiko! I am a final-year PhD candidate in the Berkeley Artificial Intelligence Research (BAIR) Lab at UC Berkeley, advised by Professor Sergey Levine.

My research goal is to develop algorithms that empower robots with human-level dexterity and reasoning capabilities, and to integrate them into everyday life. To this end, my research focuses on data-driven approaches to solving real-world robotic tasks, including offline reinforcement learning (RL), online RL fine-tuning, and imitation learning.

Recently, I have been particularly interested in using RL to post-train generalist robot policies to improve their low-level dexterity and high-level reasoning capabilities.

I had a great opportunity to spend time as a research intern at TRI in Cambridge, MA (2025.06-2025.08) and at Google DeepMind in London, UK (2025.09-2025.12). Prior to Berkeley, I received my bachelor's degree from University of Tokyo in March 2022.

Email  /  Google Scholar  /  GitHub  /  Twitter  /  Talks

profile photo
News
Research

* denotes equal contribution, denotes core contributor

Conference Papers and Pre-prints
OGPO method overview
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies
Sarvesh Patil*, Mitsuhiko Nakamoto*, Manan Agarwal, Shashwat Saxena, Jesse Zhang, Giri Anantharaman, Cleah Winston, Chaoyi Pan, Douglas Chen, Nai-Chieh Huang, Zeynep Temel, Oliver Kroemer, Sergey Levine, Abhishek Gupta, Hongkai Dai, Paarth Shah, Max Simchowitz.
ICML 2026
arXiv link / project page / code
We propose OGPO (Off-Policy Generative Policy Optimization), a sample-efficient algorithm for full fine-tuning of diffusion- and flow-based control policies. OGPO combines off-policy Q-learning with GRPO-style policy updates to enable fast & stable fine-tuning.
Steering Your Diffusion Policy with Latent Space Reinforcement Learning
Andrew Wagenmaker, Mitsuhiko Nakamoto, Yunchu Zhang, Seohong Park, Waleed Yagoub, Anusha Nagabandi, Abhishek Gupta, Sergey Levine.
CoRL 2025 (Oral Presentation, Best Paper Award Nomination)
arXiv link / project page / code
We propose DSRL (Diffusion Steering via Reinforcement Learning), a sample-efficient and lightweight approach for improving pre-trained diffusion- or flow-based policies by steering the initial noise input through RL. We successfully demonstrate real-world RL-based fine-tuning of Pi-0, a pre-trained generalist policy from Physical Intelligence.
Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance
Mitsuhiko Nakamoto, Oier Mees, Aviral Kumar, Sergey Levine.
CoRL 2024
arXiv link / project page / video / code
We propose V-GPS (Value-Guided Policy Steering), a general and broadly applicable approach that enhances the performance of pre-trained generalist robot policies at deployment time by re-ranking their actions according to a value function learned via offline RL.
SuSIE: Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models
Kevin Black*, Mitsuhiko Nakamoto*, Pranav Atreya, Homer Walke, Chelsea Finn, Aviral Kumar, Sergey Levine.
ICLR 2024
arXiv link / project page / code
We propose SuSIE (Subgoal Synthesis via Image Editing), a method that leverages text-to-image diffusion models (Stable Diffusion, InstructPix2Pix) for zero-shot robot planning.
Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning
Mitsuhiko Nakamoto*, Yuexiang Zhai*, Anikait Singh, Max Sobol Mark, Yi Ma, Chelsea Finn, Aviral Kumar, Sergey Levine.
NeurIPS 2023
arXiv link / project page / video
We propose calibrated Q-learning (Cal-QL), a method for acquiring an offline initialization that facilitates online fine-tuning.
Pre-Training for Robots: Offline RL Enables Learning New Tasks from a Handful of Trials
Aviral Kumar*, Anikait Singh*, Frederik Ebert*, Mitsuhiko Nakamoto, Yanlai Yang, Chelsea Finn, Sergey Levine
Robotics: Science and Systems (RSS) 2023
arXiv link / project page


Journal Papers
Deep Learning Algorithm to Detect Cardiac Sarcoidosis From Echocardiographic Movies
S.Katsushika, S.Kodera, M.Nakamoto, K.Ninomiya, N.Kakuda, H.Shinohara, R.Matsuoka, H.Ieki, M.Uehara, Y.Higashikuni, K.Nakanishi, T.Nakao, N.Takeda, K.Fujiu, M.Daimon, J.Ando, H.Akazawa, H.Morita, I.Komuro.
In Circulation Journal, 2021
paper link
Deep learning model to detect significant aortic regurgitation using electrocardiography
S.Sawano, S.Kodera, S.Katsushika, M.Nakamoto, K.Ninomiya, H.Shinohara, Y.Higashikuni, K.Nakanishi, T.Nakao, T.Seki, N.Takeda, K.Fujiu, M.Daimon, H.Akazawa, H.Morita, I.Komuro.
In Journal of Cardiology, 2021
paper link
Automatic detection of vessel structure by deep learning using intravascular ultrasound images of the coronary arteries
H.Shinohara, S.Kodera, K.Ninomiya, M.Nakamoto, KS.Katsushika, A.Saito, S.Minatsuki, H.Kikuchi, A.Kiyosue, Y.Higashikuni, N.Takeda, K.Fujiu, J.Ando, H.Akazawa, H.Morita, I.Komuro.
In PLOS ONE, 2021
paper link
Side Projects
Teleoperated Robot
  • Technologies: CoppeliaSim, Arduino, Servo Motor, 3D CAD design, 3D Printer, potentiometers
  • [Demo videos]
Wireless Remote Switch Toggler
  • A device for turning on/off the electric switch remotely using a smartphone.
  • Technologies: ESP32, Servo Motor, BLE socket programming, electronic design (EAGLE)
  • [Demo videos]
Misc
Teaching
Services
  • Reviews: NeurIPS (2024, 2025), ICLR (2024, 2025, 2026), ICML (2025, 2026), ICRA (2025, 2026), CoRL (2025, 2026), RA-L
I am from Yokohama, Japan. My name in Japanese is 中本光彦.
During my free time, I enjoy playing and watching sports, especially soccer. I also enjoy playing Japanese chess (aka Shogi).

Template Source