Pixel-art Yiyan Ji with curly hair and glasses, waving while holding a laptop

A curious mind, working on AI.

Yiyan Ji.

I am a Ph.D. student at the School of Intelligence Science and Technology, Nanjing University. I am also a research intern on Kuaishou's Kling team.

Ph.D. · Nanjing UniversityResearch Intern · Kling, Kuaishou

Understanding the world.
Acting within it.

My research focuses on multimodal large language model reasoning, planning, evaluation, and efficient multimodal systems. I am particularly interested in multimodal agents and embodied intelligence.

Multimodal understandingEfficient modelingEmbodied agents

Selected research

First / co-first author works · * Equal contribution

ICML 2026Co-first author

OmniSIFT

Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models

Keep what matters. Compress video redundancy, then use visual cues to select informative audio tokens.

Yue Ding*, Yiyan Ji*, Jungang Li, Xuyang Liu, Xinlong Chen, Junfei Wu, Bozhou Li, Bohan Zeng, Yang Shi, Yushuo Guan, et al.

ICLR 2026Co-first author

OmniVideoBench

Towards Audio-Visual Understanding Evaluation for Omni MLLMs

Seeing and hearing, together. Evaluating reasoning that requires both visual and audio evidence.

Caorui Li*, Yu Chen*, Yiyan Ji*, Jin Xu, Ziyu Cui, Shiyu Li, Yifei Zhang, Wenrui Wang, Ziyue Song, Dong Zhang, et al.

ACM MM 2025First author

MPCC

A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models

From perception to a workable plan. Testing multimodal planning under complex constraints.

Yiyan Ji, Haoran Chen, Qiguang Chen, Chengyue Wu, Libo Qin, Wanxiang Che

Education

Ph.D., School of Intelligence Science and Technology, Nanjing University

Currently pursuing a Ph.D.

Undergraduate Study, School of Future Technology, Harbin Institute of Technology

Focused on deep learning, large language models, and multimodal large language models.

Experience

Research Intern, Kling Team, Kuaishou Technology

Working on multimodal and omni-modal model research.

Let's exchange ideas.

I am open to discussions and collaboration on multimodal understanding, efficient reasoning, agents, and related applications. Anyone interested in my research is welcome to contact me via email.

jiyiiiyyy@gmail.com