I am a first-year M.S. student in Computer Science & Technology at Xi'an Jiaotong University. My research centers on large language models and multimodal / omni-modal foundation models, with a focus on the full post-training pipeline.
I have interned at ByteDance Seed, TikTok, and iFLYTEK, working on omni-modal (speech) foundation models, large-scale multimodal content understanding, and medical LLMs.
- ๐ญ Currently working on the speech modality of an omni-modal foundation model @ ByteDance Seed
- ๐ฑ Interested in unifying perception, reasoning, and generation across modalities โ efficiently
- ๐ฌ Happy to chat about multimodal LLMs, omni-modal training, and RL post-training
- ๐ซ Feel free to email me for any form of academic cooperation!
- Multimodal & Omni-modal LLMs โ unifying text, vision, and speech into a single foundation model
- LLM Continued Pre-training ยท Mid-training ยท Post-training โ data recipes, task composition & interference
- Efficient & Unified Multimodal Reasoning โ token pruning, distillation, RL (GRPO) for reasoning
โญ = first author ย ยทย full list on Google Scholar
-
โญ AGTAO: Robust and Stabilized LLM Unlearning via Adversarial Gating Training with Adaptive Orthogonality
ย
-
โญ Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning
ย
-
โญ EMDFNet: Efficient Multi-scale and Diverse Feature Network for Traffic Sign Detection
ย
-
PhysPRM: A Generative Process Reward Model with Fine-grained Diagnosis for Physics Problem Solving
ย
-
AERO: Autonomous Evolutionary Reasoning Optimization via Endogenous Dual-Loop Feedback
ย
-
LogicGraph: Benchmarking Multi-Path Logical Reasoning via Neuro-Symbolic Generation and Verification
ย
| Role | Organization | Period | Focus |
|---|---|---|---|
| Intern | ByteDance Seed | 2026.02 โ Present | Omni-modal foundation models (speech) |
| Intern | TikTok, ByteDance | 2025.02 โ 2025.09 | Multimodal content understanding |
| Algorithm Intern | iFLYTEK | 2024.10 โ 2025.01 | Medical LLM (SFT + DPO) ยท national invention patent |


