|
Research
My research focuses on multimodal generative models, such as vision language models and visual synthesis models, and applications of deep reinforcement learning algorithms, with the goal of improving reasoning, decision-making, and generalization capabilities of multimodal models.
|
|
|
Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning
Chengshuai Shi*,
Wenzhe Li*,
Xinran Liang*,
Yizhou Lu,
Wenjia Yang,
Ruirong Feng,
Seth Karten,
Ziran Yang,
Zihan Ding,
Gabriel Sarch,
Danqi Chen,
Karthik Narasimhan,
Chi Jin
arxiv preprint, 2026
paper
/
website
We introduce Odysseus, an open training framework for reinforcement learning of vision-language models on 100+ turn decision-making in visually grounded games.
|
|
|
Personalized Generative Models for Contextual Debiasing
Xinran Liang,
Esin Tureci,
Prachi Sinha,
Ye Zhu,
Vikram Ramaswamy,
Olga Russakovsky
CVPR Workshop on Synthetic Data for Computer Vision, 2026
paper
/
code
We introduce a method to decouple contextual patterns in vision datasets: it personalizes text-to-image diffusion models to synthesize training augmentations with uncommon visual contexts while preserving alignment with the original dataset.
|
|
|
ALP: Action-Aware Embodied Learning for Perception
Xinran Liang,
Anthony Han,
Wilson Yan,
Aditi Raghunathan,
Pieter Abbeel
arxiv preprint, 2023
paper
/
website
/
code
An embodied learning framework based on active exploration for visual representations and perception tasks.
We propose to learn representations from action signals implicitly through reinforcement learning and explicitly via inverse dynamics prediction.
|
|
|
Reward Uncertainty for Exploration in Preference-based Reinforcement Learning
Xinran Liang,
Katherine Shu,
Kimin Lee*,
Pieter Abbeel*
International Conference on Learning Representations (ICLR), 2022
paper
/
code
We propose a simple and efficient human-guided exploration method by measuring uncertainty in human instructions as intrinsic rewards.
|
|