📝 Selected Publications

For a complete list of publications, please visit my Google Scholar profile

Note: * denotes equal contribution

📄 Technical Report 1
Tech Report 2026
LongCat-Next

LongCat-Next: Lexicalizing Modalities as Discrete Tokens
Native Multimodal Any-to-Any Generation Foundation Model
Meituan LongCat Team

[Paper] [Code]

LongCat-Next is a native multimodal model (A3B) that unifies text, vision, and audio under a single autoregressive objective via discrete tokenization, achieving strong performance across multimodal benchmarks.

🤖 Vision-Language Models & VLA 4
ICLR 2026
Video-STAR

Video-STAR: Reinforcing Zero-shot Video Understanding with Tools
Tool-Using Agent Multi-turn RL Zero-shot Video
Yuan Z., Qu X., Qian, C., Chen, R., Tang, J., Sun L., Chu X., Zhang D., Wang Y., Cai Y., Li S.

[Paper] [Code]

Video-STAR proposes a novel framework that reinforces zero-shot video understanding through tool-use agents with multi-turn reasoning.

ICLR 2026
AutoDrive-R²

AutoDrive-R²: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving
Multimodal Reasoning Autonomous Driving Vision-Language-Action
Featured by AutoDrive Heart (自动驾驶之心)
Yuan Z., Tang, J., Luo, J., Chen, R., Qian, C., Sun, L., Cai Y., Zhang D., Li, S.

[Paper] [Code]

AutoDrive-R² introduces a reasoning and self-reflection framework for Vision-Language-Action models in autonomous driving scenarios.

ICML 2026
Reasoning-VLA

Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
Autonomous Driving Fast VLA Real-time Inference
Zhang D.*, Yuan Z.*, Chen Z., Liao C., Chen Y., Shen F., Zhou Q., Chua T.

[Paper]

Reasoning-VLA presents a fast and general VLA reasoning model optimized for real-time autonomous driving applications.

🎨 Generative Foundation Model 2
CVPR 2026
ADE-CoT

ADE-CoT: Adaptive Diffusion Elicits Chain-of-Thought in Image Editing
Diffusion Model Chain-of-Thought Image Editing
Qu X.*, Yuan Z.*, Tang J., Chen R., Tang D., Yu M., Sun L., Bai Y., Chu X., Gou G., Xiong G., Cai Y.

[Paper]

Preprint
Recovering Degradations

Recovering Degradations with Generative Model: A Consistency-aware Distillation Network for Infrared and Visible Image Fusion
Generation Model Image Fusion Infrared-Visible
Yu H.*, Yuan Z.*, Bai Y., Li J., Liu J., Li S., Sun L., Chu X.

📐 3D Vision 6
TCSVT 2025
DVP-MVS++

DVP-MVS++: Synergize Depth-Normal-Edge and Harmonized Visibility Prior for Multi-View Stereo
Multi-View Stereo 3D Reconstruction
Yuan Z., Zhang, D., Li, Z., Qian, C., Chen, J., Chen, Y., Chen K., Mao T., Li Z., Jiang H., Wang, Z.

[Paper] [Code]

DVP-MVS++ advances multi-view stereo through synergistic depth-normal-edge and visibility prior modeling.

TCSVT 2025
SED-MVS

SED-MVS: Segmentation-Driven and Edge-Aligned Deformation Multi-View Stereo with Depth Restoration and Occlusion Constraint
Segmentation-Driven Depth Estimation
Yuan Z., Yang, Z., Cai, Y., Wu, K., Liu, M., Zhang, D., Jiang H, Li Z., Wang, Z.

[Paper] [Code]

SED-MVS introduces segmentation-driven and edge-aligned deformation for robust multi-view stereo with depth restoration.

AAAI 2025
DVP-MVS

DVP-MVS: Synergize Depth-Edge and Visibility Prior for Multi-View Stereo
Visibility Prior 3D Vision
Yuan Z., Luo, J., Shen, F., Li, Z., Liu, C., Mao, T., Wang, Z.

[Paper] [Code]

AAAI 2025
MSP-MVS

MSP-MVS: Multi-granularity segmentation prior guided multi-view stereo
Segmentation Prior Multi-View
Yuan Z., Liu, C., Shen, F., Li, Z., Luo, J., Mao, T., Wang, Z.

[Paper] [Code]

AAAI 2024
SD-MVS

SD-MVS: Segmentation-driven deformation multi-view stereo with spherical refinement and em optimization
Spherical Refinement EM Optimization
Yuan Z., Cao, J., Li, Z., Jiang, H., Wang, Z.

[Paper] [Code]

PR 2024
TSAR-MVS