PPaperPicks

Yueting Zhuang

49 papers at tracked venues · 45 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. AHEAD: Attention Head Energy-Aware Dynamics for Hallucination Mitigation in MLLMs
  2. Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
  3. Experience-driven Multi-turn Reinforcement Learning for GUI Agents
  4. GUI-G²: Gaussian Reward Modeling for GUI Grounding
  5. MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models
  6. Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
  7. Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
  8. UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization
  9. 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
  10. Align²LLaVA: Cascaded Human and Large Language Model Preference Alignment for Multi-modal Instruction Curation
  11. AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
  12. Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
  13. Chart-HQA: A Benchmark for Hypothetical Question Answering in Charts
  14. Counterfactual Evolution of Multimodal Datasets via Visual Programming
  15. EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
  16. EvolvedGRPO: Unlocking Reasoning in LVLMs via Progressive Instruction Evolution
  17. EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model
  18. HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
  19. Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models
  20. Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
  21. Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
  22. Let LRMs Break Free from Overthinking via Self-Braking Tuning
  23. Logic Distillation: Learning from Code Function by Function for Decision-making Tasks
  24. Mastering Collaborative Multi-Modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness
  25. Meta-Reflection: A Feedback-Free Reflection Learning Framework
  26. Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning
  27. Robust Modality-Incomplete Anomaly Detection: A Modality-Instructive Framework with Benchmark
  28. STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
  29. SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation
  30. TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition
  31. VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
  32. What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
  33. Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization
  34. Auto-Encoding Morph-Tokens for Multimodal LLM
  35. Bridging Local Details and Global Context in Text-Attributed Graphs
  36. DEMON24: ACM MM24 Demonstrative Instruction Following Challenge
  37. Data Shunt: Collaboration of Small and Large Models for Lower Costs and Better Performance
  38. De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
  39. Fact : Teaching MLLMs with Faithful, Concise and Transferable Rationales
  40. Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
  41. HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
  42. Learning Global Controller in Latent Space for Parameter-Efficient Fine-Tuning
  43. Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
  44. Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model
  45. Self-Contrast: Better Reflection Through Inconsistent Solving Perspectives
  46. T2S-GPT: Dynamic Vector Quantization for Autoregressive Sign Language Production from Text
  47. TaskBench: Benchmarking Large Language Models for Task Automation
  48. Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question Answering
  49. WorldGPT: Empowering LLM as Multimodal World Model