OmniSIFT
Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models
Keep what matters. Compress video redundancy, then use visual cues to select informative audio tokens.
A curious mind, working on AI.
I am a Ph.D. student at the School of Intelligence Science and Technology, Nanjing University. I am also a research intern on Kuaishou's Kling team.
My research focuses on multimodal large language model reasoning, planning, evaluation, and efficient multimodal systems. I am particularly interested in multimodal agents and embodied intelligence.
First / co-first author works · * Equal contribution
Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models
Keep what matters. Compress video redundancy, then use visual cues to select informative audio tokens.
Towards Audio-Visual Understanding Evaluation for Omni MLLMs
Seeing and hearing, together. Evaluating reasoning that requires both visual and audio evidence.
A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models
From perception to a workable plan. Testing multimodal planning under complex constraints.
Currently pursuing a Ph.D.
Focused on deep learning, large language models, and multimodal large language models.
Working on multimodal and omni-modal model research.
I am open to discussions and collaboration on multimodal understanding, efficient reasoning, agents, and related applications. Anyone interested in my research is welcome to contact me via email.
jiyiiiyyy@gmail.com