PPaperPicks

Ziyu Guo

25 papers at tracked venues · 21 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking
  2. Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
  3. EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual Insights
  4. Less is More: Improving Motion Diffusion Models with Sparse Keyframes
  5. Let's Verify and Reinforce Image Generation Step by Step
  6. LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding
  7. MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
  8. MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
  9. MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
  10. MMSearch: Unveiling the Potential of Large Models as Multi-modal Search Engines
  11. Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
  12. SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
    ACL 2025 · Ziyu Guo
  13. SmoothCache: A Universal Inference Acceleration Technique for Diffusion Transformers
  14. StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross Fusion
    ICCV 2025 · Ziyu Guo
  15. T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
  16. UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
  17. What We Miss Matters: Learning from the Overlooked in Point Cloud Transformers
  18. LLM-Assisted Multi-Teacher Continual Learning for Visual Question Answering in Robotic Surgery
  19. MATHVERSE: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
  20. No Time to Train: Empowering Non-Parametric Networks for Few-Shot 3D Scene Segmentation
  21. Personalize Segment Anything Model with One Shot
  22. Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation
  23. SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning
  24. Spatio-Temporal Pivotal Graph Neural Networks for Traffic Flow Forecasting
  25. X-former Elucidator: Reviving Efficient Attention for Long Context Language Modeling