PPaperPicks

Hongsheng Li

Chinese University of Hong Kong, Department of Electrical Engineering, CUHK-SenseTime Joint Laboratory, Hong Kong

108 papers at tracked venues · 87 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. A3: Android Agent Arena for Mobile GUI Agents with Essential-State Procedural Evaluation
  2. From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-Bench
  3. MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
  4. Self-NPO: Data-Free Diffusion Model Enhancement via Truncated Diffusion Fine-Tuning
  5. TIDE: Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
  6. Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format Reasoning
  7. UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning
  8. AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents
  9. Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding
  10. Alignment with Fill-In-the-Middle for Enhancing Code Generation
  11. BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
  12. BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices
  13. CameraCtrl II: Dynamic Scene Exploration via Camera-Controlled Video Diffusion Models
  14. CameraCtrl: Enabling Camera Control for Video Diffusion Models
  15. CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection
  16. Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
  17. Diffusion-NPO: Negative Preference Optimization for Better Preference Aligned Generation of Diffusion Models
  18. Docopilot: Improving Multimodal Models for Document-Level Understanding
  19. Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
  20. EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM
  21. EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation
  22. FlexDrive: Toward Trajectory Flexibility in Driving Scene Gaussian Splatting Reconstruction and Rendering
  23. FreeSim: Toward Free-viewpoint Camera Simulation in Driving Scenes
  24. From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
  25. GS-DiT: Advancing Video Generation with Dynamic 3D Gaussian Fields through Efficient Dense 3D Point Tracking
  26. GaussianPainter: Painting Point Cloud into 3D Gaussians with Normal Guidance
  27. GenieBlue: Integrating Both Linguistic and Multimodal Capabilities for Large Language Models on Mobile Devices
  28. GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and Editing
  29. LLaVA-MoD: Making LLaVA Tiny via MoE-Knowledge Distillation
  30. LM-Searcher: Cross-domain Neural Architecture Search with LLMs via Unified Numerical Encoding
  31. Let's Verify and Reinforce Image Generation Step by Step
  32. LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding
  33. Lumina-Image 2.0: a Unified and Efficient Image Generative Framework
  34. Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution Generation
  35. M3Net: Multimodal Multi-task Learning for 3D Detection, Segmentation, and Occupancy Prediction in Autonomous Driving
  36. MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
  37. MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
  38. MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
  39. MMSearch: Unveiling the Potential of Large Models as Multi-modal Search Engines
  40. MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning
  41. MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code
  42. Mixture Compressor for Mixture-of-Experts LLMs Gains More
  43. NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
  44. NopeRoomGS: Indoor 3D Gaussian Splatting Optimization without Camera Pose Input
  45. One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual Representation
  46. PUMA: Empowering Unified MLLM with Multi-Granular Visual Generation
  47. Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
  48. PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
  49. Point Cluster: A Compact Message Unit for Communication-Efficient Collaborative Perception
  50. Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
  51. Rectified Diffusion: Straightness Is Not Your Need in Rectified Flow
  52. ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation
  53. SKT: Integrating State-Aware Keypoint Trajectories with Vision-Language Models for Robotic Garment Manipulation
  54. SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
  55. SmartBench: Is Your LLM Truly a Good Chinese Smartphone Assistant?
  56. SmartPretrain: Model-Agnostic and Dataset-Agnostic Representation Learning for Motion Prediction
  57. SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
  58. T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
  59. Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and Methodology
  60. UAV-Flow Colosseo: A Real-World Benchmark for Flying-on-a-Word UAV Imitation Learning
  61. UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
  62. UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models
  63. VBCD: A Voxel-Based Framework for Personalized Dental Crown Design
  64. Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
  65. VividFace: A Robost and High-Fidelity Video Face Swapping Framework
  66. WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
  67. A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding
  68. ADDP: Learning General Representations for Image Recognition and Generation with Alternating Denoising Diffusion Process
  69. Any2Point: Empowering Any-Modality Large Models for Efficient 3D Understanding
  70. Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft
  71. Be-Your-Outpainter: Mastering Video Outpainting Through Input-Specific Adaptation
  72. BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation Using RGB Frames and Events
  73. CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
  74. Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
  75. Deep Reward Supervisions for Tuning Text-to-Image Diffusion Models
  76. Delving Deep into Engagement Prediction of Short Videos
  77. DiffInDScene: Diffusion-Based High-Quality 3D Indoor Scene Generation
  78. Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
  79. Empowering Character-level Text Infilling by Eliminating Sub-Tokens
  80. Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models
  81. FouriScale: A Frequency Perspective on Training-Free High-Resolution Image Synthesis
  82. GLID: Pre-training a Generalist Encoder-Decoder Vision Model
  83. GiT: Towards Generalist Vision Transformer Through Universal Language Interface
  84. LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero-initialized Attention
  85. LMDrive: Closed-Loop End-to-End Driving with Large Language Models
  86. Learning 1D Causal Visual Representation with De-focus Attention Networks
  87. Lumina-Next : Making Lumina-T2X Stronger and Faster with Next-DiT
  88. MATHVERSE: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
  89. ManipVQA: Injecting Robotic Affordance and Physically Grounded Information into Multi-Modal Large Language Models
  90. MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning
  91. MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs
  92. Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
  93. MoVA: Adapting Mixture of Vision Experts to Multimodal Context
  94. Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
  95. Personalize Segment Anything Model with One Shot
  96. Phased Consistency Models
  97. Ponymation: Learning Articulated 3D Animal Motions from Unlabeled Online Videos
  98. SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
  99. SPHINX: A Mixer of Weights, Visual Embeddings and Image Scales for Multi-modal Large Language Models
  100. SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models
  101. SmartRefine: A Scenario-Adaptive Refinement Framework for Efficient Motion Prediction
  102. Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification
  103. Three Things We Need to Know About Transferring Stable Diffusion to Visual Dense Prediction Tasks
  104. VeloVox: A Low-Cost and Accurate 4D Object Detector with Single-Frame Point Cloud of Livox LiDAR
  105. Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
  106. ZOPP: A Framework of Zero-shot Offboard Panoptic Perception for Autonomous Driving
  107. ZoLA: Zero-Shot Creative Long Animation Generation with Short Video Model
  108. nuCraft: Crafting High Resolution 3D Semantic Occupancy for Unified 3D Scene Understanding