ZIMO WEN

Undergraduate Researcher · Computer Science
Shanghai Jiao Tong University
Zhiyuan Honors Program · MVIG Lab
Zimo Wen's GitHub avatar

Biography

I am an undergraduate researcher in Computer Science at Shanghai Jiao Tong University, in the Zhiyuan Honors Program and MVIG Lab, advised by Professors Cewu Lu and Chuan Wen. My research interests span embodied intelligence, agentic systems, and multimodal models, with a focus on robot self-improvement, unified visual understanding and generation, and efficient learning systems.

News

  • SenseNova-U1.5 technical report is now available.
  • RoboRSI research report and project demos are now available.
  • Argus preprint is available on arXiv.
  • Mage-Flow preprint is available on arXiv.
  • RESOURCE2SKILL preprint is available on arXiv.
  • UniG2U-Bench preprint is available on arXiv.

Selected Projects

Agentic Systems3
Project video
Argus: A General-Purpose Agentic Reasoning Runtime for Long-Horizon Tasks
Boxiu Li*, Zimo Wen*, Yijia Fan*, Chuan Wen, Fan Yang, Hangxi Guo, Jiaao Wu, Jiachen Zhang, Junxiang Lei, Mukai Li, Ruize Tang, Runjing Gu, Shibo Hu, Sihan Chen, Sufeng Guo, Wanbo Zhang, Xian Zhang, Xiaoyu Chen, Xuanhe Zhou, Xuyao Huang, Yifei Gao, Yifei Shen, Yilin Chen, Yuheng Wu, Yuzhe Zhang, Zelong Zhao, Zhijie Deng (* equal contribution)
Agentic systems
Robotics2
Multimodal Learning1
Time Series & Dynamics3
Technical Reports2

Experience

  • Shanghai Jiao Tong University, MVIG LabUndergraduate Research Intern with Professors Cewu Lu and Chuan WenEmbodied intelligence, force-aware interaction, and vision-language-action systems.
    Oct 2024 – Present
  • SenseTimeTop InternSenseNova-U1.5: unified visual understanding and generation.
  • Microsoft Research AsiaTomorrow Star ProgramScaling laws of multimodal generative models, video generation, and multimodal training infrastructure.
    Jun 2025 – Dec 2025
  • Shanghai Jiao Tong University, Collaborative Intelligence Technology LabTime-Series Forecasting ResearchRetrieval-augmented few-shot time-series forecasting, leading to DANet.
    Feb 2024 – Dec 2024

Community Contribution

  • lmms-eval, ContributorA unified evaluation toolkit for multimodal models across text, image, video, and audio.
  • lmms-engine, ContributorA flexible training engine for large multimodal models.
  • Flash Linear Attention, ContributorEfficient implementations and kernels for linear attention and related sequence models.