About Me

Hello! I am Yi Liu, an algorithm researcher at AgiBot working on embodied intelligence, with a focus on robot foundation models (VLA and WM) and Self-Evolving.

I received my M.S. (2025) and B.Eng. (2022) degrees in Computer Science and Technology from Beihang University (BUAA), where I was advised by Prof. Si Liu and Prof. Lijun Zhang.

My research interest includes Embodied AI, Computer Vision, Image Generation and Representation Learning.

News

  • We released τ₀-VLA, a hierarchical robot foundation model for long-horizon manipulation.
  • Libra-VLA was accepted to ACL 2026, and ACoT-VLA was accepted to CVPR 2026.
  • We introduced GenieReasoner for unified embodied reasoning and robot action.

Selected Publications

Full list on Google Scholar

τ₀-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

Xiaowei Cai, Yunuo Cai, Bingao Chen, Jingxiao Chen, Zhi Chen, Siyuan Feng, Tengyu Hou, Jingshun Huang, Han Jiang, Runkun Ju, Dong Li, Mingxiang Li, Shaowei Li, Xinchen Li, Yifan Li, Yi Liu*, Zhongyuan Liu, Jianlan Luo, Junwen Miao, Ruiqi Ni, Buqing Nie, Mingjie Pan, Xinlin Ren, Jianheng Song, Jiaxu Wang, Peiqi Wang, Sen Wang, Xiaoyan Wang, Dafeng Wei, Dongming Wu, Pengwei Xie, Pu Yang, Hangjian Ye, Xiangyu Yue, Jinyu Zhang*, Qinglin Zhang, Xueyong Zhao, Pengfei Zhou, Yue Zhou

arXiv, 2026 · Authors listed alphabetically · Co-first author · Project lead

A hierarchical robot foundation model that uses execution memory and world-model-guided test-time search for long-horizon manipulation.

Libra-VLA real-world robot manipulation tasks

Libra-VLA: Achieving Learning Equilibrium via Asynchronous Coarse-to-Fine Dual-System

Yifei Wei, Linqing Zhong, Yi Liu, Yuxiang Lu, Xindong He, Maoqing Yao, Guanghui Ren

ACL, 2026

A coarse-to-fine VLA that separates macro-intent planning from continuous action refinement and runs both systems asynchronously.

Architectural overview of ACoT-VLA

ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models

Linqing Zhong, Yi Liu, Yifei Wei, Ziyu Xiong, Maoqing Yao, Si Liu, Guanghui Ren

CVPR, 2026

Introduces action chain-of-thought reasoning through explicit coarse trajectories and implicit action priors, bridging semantic understanding and precise robot control.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

Yi Liu*, Sukai Wang*, Dafeng Wei, Xiaowei Cai, Linqing Zhong, Jiange Yang, Guanghui Ren, Jinyu Zhang, Maoqing Yao, Chuankang Li, Xindong He, Liliang Chen, Jianlan Luo

arXiv, 2025 · Co-first author

GenieReasoner co-trains embodied reasoning and discrete robotic actions with the FACT flow-matching action tokenizer.

AgentIAD multi-round tool-augmented inspection framework

AgentIAD: Tool-Augmented Single-Agent for Industrial Anomaly Detection

Junwen Miao, Penghui Du, Yi Liu, Yu Wang, Yan Wang

arXiv, 2025

A tool-augmented vision-language agent that iteratively zooms into suspicious regions and retrieves normal references for reliable, interpretable industrial anomaly detection.

Conceptual comparison of CoST collaborative perception

CoST: Efficient Collaborative Perception From Unified Spatiotemporal Perspective

Zongheng Tang, Yi Liu, Yifan Sun, Yulu Gao, Jinyu Chen, Runsheng Xu, Si Liu

ICCV, 2025 · Highlight

Unifies spatial multi-agent and temporal fusion while avoiding repeated transmission of static features in collaborative 3D perception.

CycleVAR stylized landscape translation results

CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation

Yi Liu, Shengqian Li, Zuzeng Lin, Feng Wang, Si Liu

ICCV, 2025 · First author

Repurposes pretrained visual autoregressive models for unpaired image translation with multi-scale token prefilling and differentiable quantization.

Visualization of selected and discarded DETR queries in QSKD

Knowledge Distillation via Query Selection for Detection Transformer

Yi Liu, Luting Wang, Zongheng Tang, Yue Liao, Yifan Sun, Lijun Zhang, Si Liu

arXiv, 2024 · First author

Selects informative positive and hard-negative DETR queries for attention-guided feature distillation and locally aligned prediction distillation.

Object-Aware Distillation Pyramid framework

Object-Aware Distillation Pyramid for Open-Vocabulary Object Detection

Luting Wang, Yi Liu, Penghui Du, Zihan Ding, Yue Liao, Qiaosong Qi, Biaolong Chen, Si Liu

CVPR, 2023

Transfers vision-language knowledge to open-vocabulary detectors through object-aware extraction and multi-level distillation.

Experience

Algorithm Researcher

AgiBot

Robot foundation models, hierarchical VLA systems, multimodal co-training, and discrete action representation.

Research Intern

Kling AI, Kuaishou

Controllable virtual try-on with 3D constraints and language-guided, category-general segmentation.

Research Intern

CreateAI — AIGC

Visual autoregressive models for controllable and unsupervised image translation.

Research Intern

Baidu, Visual Technology Department

Multimodal perception and efficient vehicle–infrastructure collaborative perception.

Education

M.S. in Computer Science and Technology

Beihang University · GPA 3.89 / 4.0 · Top 10%

Advisors: Prof. Si Liu and Prof. Lijun Zhang. Research in embodied intelligence, perception, and image generation.

B.Eng. in Computer Science and Technology

Beihang University · GPA 3.85 / 4.0 · Top 5%

Completed a second degree in Mathematics.

Selected Honors

  • National ScholarshipTop 1%
  • Outstanding Graduate StudentTop 3%
  • Beijing Outstanding GraduateTop 5%
  • Top Academic ScholarshipTop 3%