PPaperPicks

Dahua Lin

Chinese University of Hong Kong, Department of Information Engineering, CUHK - SenseTime Joint Lab, Hong Kong

120 papers at tracked venues · 99 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Half-S: Halving the Scale for Near-Lossless 4-Bit LLM Training
  2. MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy
  3. MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
  4. Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic
  5. Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
  6. Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs
  7. 3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusion
  8. 3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation
  9. Bootstrap3D: Improving Multi-View Diffusion Model with Synthetic Data
  10. ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way
  11. Conical Visual Concentration for Efficient Large Vision-Language Models
  12. Consultant Decoding: Yet Another Synergistic Mechanism
  13. Creation-Mmbench: Assessing Context-Aware Creative Intelligence in Mllms
  14. Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
  15. GRAIT: Gradient-Driven Refusal-Aware Instruction Tuning for Effective Hallucination Mitigation
  16. GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography
  17. Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity
  18. HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance
  19. Hierachical Balance Packing: Towards Efficient Supervised Fine-tuning for Long-Context LLM
  20. Horizon-GS: Unified 3D Gaussian Splatting for Large-Scale Aerial-to-Ground Scenes
  21. IDArb: Intrinsic Decomposition for Arbitrary Number of Input Views and Illuminations
  22. Imagine360: Immersive 360 Video Generation from Perspective Anchor
  23. InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
  24. Keyframe-Guided Creative Video Inpainting
  25. LEGION: Learning to Ground and Explain for Synthetic Image Detection
  26. LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models
  27. Long Context Tuning for Video Generation
  28. MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models
  29. MM-IFEngine: Towards Multimodal Instruction Following
  30. Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
  31. Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of Go
  32. More Data or Better Data? A Critical Analysis of Data Selection and Synthesis for Mathematical Reasoning
  33. Multi-Identity Human Image Animation with Structural Video Diffusion
  34. MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design
  35. OmniBal: Towards Fast Instruction-Tuning for Vision-Language Models via Omniverse Computation Balance
  36. OpenHuEval: Evaluating Large Language Model on Hungarian Specifics
  37. Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation
  38. Proc-GS: Procedural Building Generation for City Assembly with 3D Gaussians
  39. ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuning
  40. SAM2LONG: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree
  41. SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
  42. Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning
  43. SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
  44. SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation
  45. Training Language Models to Critique With Multi-agent Feedback
  46. UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
  47. Utilize the Flow Before Stepping into the Same River Twice: Certainty Represented Knowledge Flow for Refusal-Aware Instruction Tuning
  48. VFLowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
  49. Video World Models with Long-term Spatial Memory
  50. VideoRoPE: What Makes for Good Video Rotary Position Embedding?
  51. Visual-RFT: Visual Reinforcement Fine-Tuning
  52. What are the Essential Factors in Crafting Effective Long Context Multi-Hop Instruction Datasets? Insights and Best Practices
  53. X-Prompt: Generalizable Auto-Regressive Visual Learning with In-Context Prompting
  54. daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
  55. ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models
  56. ANAH: Analytical Annotation of Hallucinations in Large Language Models
  57. Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
  58. Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models
  59. AlchemistCoder: Harmonizing and Eliciting Code Capability by Hindsight Tuning on Multi-source Data
  60. Alpha-CLIP: A CLIP Model Focusing on Wherever you Want
  61. AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
  62. Are We on the Right Way for Evaluating Large Vision-Language Models?
  63. Balanced Data Sampling for Language Model Training with Clustering
  64. Betrayed by Attention: A Simple yet Effective Approach for Self-supervised Video Object Segmentation
  65. BotChat: Evaluating LLMs' Capabilities of Having Multi-Turn Dialogues
  66. Cinematic Behavior Transfer via NeRF-based Differentiable Filming
  67. Code Needs Comments: Enhancing Code LLMs with Comment Augmentation
  68. CriticEval: Evaluating Large-scale Language Model as Critic
  69. EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI
  70. F-Eval: Asssessing Fundamental Abilities with Refined Evaluation Methods
  71. FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models
  72. Flames: Benchmarking Value Alignment of LLMs in Chinese
  73. From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
  74. GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
  75. GPT4Point: A Unified Framework for Point-Language Understanding and Generation
  76. GeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single Image
  77. HumanGaussian: Text-Driven 3D Human Generation with Gaussian Splatting
  78. HumanVid: Demystifying Training Data for Camera-controllable Human Image Animation
  79. HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion
  80. Identifying Semantic Induction Heads to Understand In-Context Learning
  81. InterControl: Zero-shot Human Interaction Generation by Controlling Every Joint
  82. InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
  83. Lean Workbook: A large-scale Lean problem set formalized from natural language math problems
  84. Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and Feedback
  85. LongWanjuan: Towards Systematic Measurement for Long Text Quality
  86. MGF: Mixed Gaussian Flow for Diverse Trajectory Prediction
  87. MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
  88. MMBench: Is Your Multi-modal Model an All-Around Player?
  89. MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
  90. MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
  91. Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials
  92. MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
  93. MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
  94. Navigating the OverKill in Large Language Models
  95. OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
  96. OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
  97. Omni6D: Large-Vocabulary 3D Object Dataset for Category-Level 6D Object Pose Estimation
  98. OneLLM: One Framework to Align All Modalities with Language
  99. PointLLM: Empowering Large Language Models to Understand Point Clouds
  100. Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
  101. ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
  102. Rethinking Image-to-Video Adaptation: An Object-Centric Perspective
  103. SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
  104. SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction
  105. Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering
  106. Scaling Behavior for Large Language Models regarding Numeral Systems: An Example using Pythia
  107. Scaling Laws of RoPE-based Extrapolation
  108. ShareGPT4V: Improving Large Multi-modal Models with Better Captions
  109. ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
  110. SparseCtrl: Adding Sparse Controls to Text-to-Video Diffusion Models
  111. Streaming Long Video Understanding with Large Language Models
  112. T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
  113. Towards Text-guided 3D Scene Composition
  114. Turn Waste into Worth: Rectifying Top-k Router of MoE
  115. Uncertainty Aware Learning for Language Model Alignment
  116. Unified Human-Scene Interaction via Prompted Chain-of-Contacts
  117. VBench: Comprehensive Benchmark Suite for Video Generative Models
  118. VLMEvalKit: An Open-Source ToolKit for Evaluating Large Multi-Modality Models
  119. VideoBooth: Diffusion-based Video Generation with Image Prompts
  120. X-neuron: Interpreting, Locating and Editing of Neurons in Reinforcement Learning Policy