PPaperPicks

Xiangyu Zhang

Megvii Inc., Beijing, China

39 papers at tracked venues · 30 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. PRIME: A Process-Outcome Alignment Benchmark for Verifiable Reasoning in Mathematics and Engineering
  2. PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
  3. SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
  4. WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
  5. Auto-Landmark: Acoustic Landmark Dataset and Open-Source Toolkit for Landmark Extraction
    InterSpeech 2025 · Xiangyu Zhang
  6. Beyond Sequences: Two-dimensional Representation and Dependency Encoding for Code Generation
    ACL 2025 · Xiangyu Zhang
  7. DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation
  8. GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
  9. Glad: A Streaming Scene Generator for Autonomous Driving
  10. Holistic Tokenizer for Autoregressive Image Generation
  11. Language Prompt for Autonomous Driving
  12. Multi-matrix Factorization Attention
  13. Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
  14. Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
  15. Perception in Reflection
  16. Perception-R1: Pioneering Perception Policy with Reinforcement Learning
  17. Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMs
  18. Reconstructive Visual Instruction Tuning
  19. Ross3d: Reconstructive Visual Instruction Tuning With 3D-Awareness
  20. SpeechT-RAG: Reliable Depression Detection in LLMs with Retrieval-Augmented Generation Using Speech Timing Information
    ACL 2025 · Xiangyu Zhang
  21. SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control
  22. Taming Teacher Forcing for Masked Autoregressive Video Generation
  23. Unhackable Temporal Reward for Scalable Video MLLMs
  24. Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation
  25. Align-DETR: Enhancing End-to-end Object Detection with Aligned Loss
  26. Binaural Selective Attention Model for Target Speaker Extraction
  27. ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning
  28. Compound Text-Guided Prompt Tuning via Image-Adaptive Cues
  29. DDAE: Towards Deep Dynamic Vision BERT Pretraining
  30. DreamLLM: Synergistic Multimodal Comprehension and Creation
  31. Far3D: Expanding the Horizon for Surround-View 3D Object Detection
  32. Merlin: Empowering Multimodal LLMs with Foresight Minds
  33. OneChart: Purify the Chart Structural Extraction via One Auxiliary Token
  34. Panacea: Panoramic and Controllable Video Generation for Autonomous Driving
  35. Self-Supervised Visual Preference Alignment
  36. Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion Model
    EMNLP 2024 · Xiangyu Zhang
  37. Stream Query Denoising for Vectorized HD-Map Construction
  38. Vary: Scaling up the Vision Vocabulary for Large Vision-Language Model
  39. When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection
    EMNLP 2024 · Xiangyu Zhang