PPaperPicks

Siliang Tang

Zhejiang University, College of Computer Science and Technology, Hangzhou, China

49 papers at tracked venues · 45 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. AHEAD: Attention Head Energy-Aware Dynamics for Hallucination Mitigation in MLLMs
  2. CoMoL: Efficient Mixture of LoRA Experts via Dynamic Core Space Merging
  3. Evolving Generalist Virtual Agents with Generative and Associative Memory
  4. MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models
  5. PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models
  6. Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
  7. Align²LLaVA: Cascaded Human and Large Language Model Preference Alignment for Multi-modal Instruction Curation
  8. AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
  9. Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
  10. Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark
  11. Chart-HQA: A Benchmark for Hypothetical Question Answering in Charts
  12. ChatMap: Mining Human Thought Processes for Customer Service Chatbots via Multi-Agent Collaboration
  13. Choice is what matters after Attention
  14. Counterfactual Evolution of Multimodal Datasets via Visual Programming
  15. EvolvedGRPO: Unlocking Reasoning in LVLMs via Progressive Instruction Evolution
  16. Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
  17. GraphCLIP: Enhancing Transferability in Graph Foundation Models for Text-Attributed Graphs
  18. HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
  19. Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
  20. Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
  21. Logic Distillation: Learning from Code Function by Function for Decision-making Tasks
  22. MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query
  23. Mastering Collaborative Multi-Modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness
  24. Meta-Reflection: A Feedback-Free Reflection Learning Framework
  25. On Path to Multimodal Generalist: General-Level and General-Bench
  26. Robust Modality-Incomplete Anomaly Detection: A Modality-Instructive Framework with Benchmark
  27. STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
  28. TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition
  29. The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video Generation
  30. What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
  31. Auto-Encoding Morph-Tokens for Multimodal LLM
  32. Bridging Local Details and Global Context in Text-Attributed Graphs
  33. DEMON24: ACM MM24 Demonstrative Instruction Following Challenge
  34. DIEM: Decomposition-Integration Enhancing Multimodal Insights
  35. Data Shunt: Collaboration of Small and Large Models for Lower Costs and Better Performance
  36. De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
  37. Efficient Tuning and Inference for Large Language Models on Textual Graphs
  38. Fact : Teaching MLLMs with Faithful, Concise and Transferable Rationales
  39. Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
  40. GraphControl: Adding Conditional Control to Universal Graph Pre-trained Models for Graph Domain Transfer Learning
  41. HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
  42. I3: Intent-Introspective Retrieval Conditioned on Instructions
  43. MARIO: Model Agnostic Recipe for Improving OOD Generalization of Graph Contrastive Learning
  44. Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
  45. NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
  46. T2S-GPT: Dynamic Vector Quantization for Autoregressive Sign Language Production from Text
  47. Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
  48. Unified Generative and Discriminative Training for Multi-modal Large Language Models
  49. WorldGPT: Empowering LLM as Multimodal World Model