I am a first-year master's student at Shanghai Jiao Tong University, advised by Prof. Guo Lu. I received my bachelor's degree in Information Engineering from the same university.
My research focuses on omni-modal models, with an emphasis on integrating vision, audio, and language for natural, real-time interaction. I am particularly interested in how these models can perceive, reason, and act throughout an ongoing conversation. My previous work explored efficient video understanding through spatiotemporal token compression.
I am also a research intern at StepFun, working on audio-language model post-training. I contributed to StepAudio 3 Realtime, primarily in full-duplex dialogue and voice agent capabilities. This work aligns with my broader interest in building multimodal assistants that interact naturally and carry out tasks through conversation.
") does not match the recommended repository name for your site ("").
", so that your site can be accessed directly at "http://".
However, if the current repository name is intended, you can ignore this message by removing "{% include widgets/debug_repo_name.html %}" in index.html.
",
which does not match the baseurl ("") configured in _config.yml.
baseurl in _config.yml to "".

Bin Lin, Bo Zhao, Boyang Zhang, …, Jialong Xue, …
arXiv preprint arXiv:2609.14005 2026
StepAudio 3 Realtime integrates audio understanding, full-duplex conversation, parallel reasoning and speech generation, and voice-based task execution. It supports natural turn-taking and interruptions, with asynchronous tool use that keeps conversations flowing. My contributions: primarily full-duplex interaction and Voice Agent capabilities.
Bin Lin, Bo Zhao, Boyang Zhang, …, Jialong Xue, …
arXiv preprint arXiv:2609.14005 2026
StepAudio 3 Realtime integrates audio understanding, full-duplex conversation, parallel reasoning and speech generation, and voice-based task execution. It supports natural turn-taking and interruptions, with asynchronous tool use that keeps conversations flowing. My contributions: primarily full-duplex interaction and Voice Agent capabilities.

Junhao Du*, Jialong Xue*, Anqi Li, Jincheng Dai, Guo Lu
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026
We formulate token compression for Video-LLMs as a unified spatiotemporal allocation problem under a global retention budget. The method preserves strong video understanding performance at ultra-low token retention without retraining.
Junhao Du*, Jialong Xue*, Anqi Li, Jincheng Dai, Guo Lu
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026
We formulate token compression for Video-LLMs as a unified spatiotemporal allocation problem under a global retention budget. The method preserves strong video understanding performance at ultra-low token retention without retraining.