PPaperPicks

Ping Luo

University of Hong Kong, Shanghai AI Laboratory, Department of Computer Science, Hong Kong

85 papers at tracked venues · 73 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Beyond Query Memorization: Large Language Model Routing with Query Decomposition and Historical Matching
  2. FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
  3. Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers
  4. R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual Scenarios
  5. TVWorld: Foundations for Remote-Control TV Agents
  6. AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation
  7. AnalogCoder: Analog Circuit Design via Training-Free Code Generation
  8. AutoMMLab: Automatically Generating Deployable Models from Language Instructions for Computer Vision Tasks
  9. BOOD: Boundary-based Out-Of-Distribution Data Generation
  10. CompGS: Unleashing 2D Compositionality for Compositional Text-to-3D via Dynamically Optimizing 3D Gaussians
  11. DexHandDiff: Interaction-aware Diffusion Planning for Adaptive Dexterous Manipulation
  12. DiffusionMat: Alpha Matting as Deterministic Sequential Refinement Learning
  13. Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language Bootstrapping
  14. Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-Time Open-Vocabulary Object Detection
  15. EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
  16. End-to-End Autonomous Driving Through V2X Cooperation
  17. FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
  18. G3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object Manipulation
  19. GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices
  20. Goku: Flow Based Video Generative Foundation Models
  21. HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language Model
  22. IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model
  23. Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
  24. JiSAM: Alleviate Labeling Burden and Corner Case Problems in Autonomous Driving via Minimal Real-World Data
  25. Learning Humanoid Locomotion with Perceptive Internal Model
  26. LiT: Delving into a Simple Linear Diffusion Transformer for Image Generation
  27. MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
  28. MangaNinja: Line Art Colorization with Precise Reference Following
  29. NADER: Neural Architecture Design via Multi-Agent Collaboration
  30. OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
  31. OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis
  32. Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots
  33. Prompt-A-Video: Prompt your Video Diffusion Model via Preference-Aligned LLM
  34. Research Challenges and Progress in the End-to-End V2X Cooperative Autonomous Driving Competition
  35. RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins
  36. SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement
  37. SpecEM: Training-Free LLM Ensembling via Iterative Drafting, Verification, and Online Feedback
  38. TREND: Unsupervised 3D Representation Learning via Temporal Forecasting for LiDAR Perception
  39. Text2World: Benchmarking Large Language Models for Symbolic World Model Generation
  40. Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
  41. Unsupervised Continual Domain Shift Learning with Multi-Prototype Modeling
  42. Whether LLMs Know If They Know: Identifying Knowledge Boundaries via Debiased Historical In-Context Learning
  43. WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
  44. BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity Allocation
  45. Cached Transformers: Improving Transformers with Differentiable Memory Cachde
  46. ChartAssistant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
  47. Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation
  48. ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models
  49. DeepAccident: A Motion and Accident Prediction Benchmark for V2X Autonomous Driving
  50. DiffAgent: Fast and Accurate Text-to-Image API Selection with Large Language Model
  51. DriveLM: Driving with Graph Visual Question Answering
  52. GKGNet: Group K-Nearest Neighbor Based Graph Convolutional Network for Multi-label Image Recognition
  53. GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
  54. GenTron: Diffusion Transformers for Image and Video Generation
  55. Generalized Predictive Model for Autonomous Driving
  56. Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
  57. InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
  58. LLaMA Pro: Progressive LLaMA with Block Expansion
  59. Large Language Models as Automated Aligners for benchmarking Vision-Language Models
  60. Learning Manipulation by Predicting Interaction
  61. MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
  62. Mind the Boundary: Coreset Selection via Reconstructing the Decision Boundary
  63. MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank Experts
  64. Needle In A Multimodal Haystack
  65. OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM
  66. OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
  67. PIXART-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
  68. PROGRAM: PROtotype GRAph Model based Pseudo-Label Learning for Test-Time Adaptation
  69. PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
  70. Position: Towards Implicit Prompt For Text-To-Image Models
  71. RegionGPT: Towards Region Understanding Vision Language Model
  72. Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability, Reproducibility, and Practicality
  73. RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis
  74. RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (Early Version)
  75. Scalable and Effective Arithmetic Tree Generation for Adder and Multiplier Designs
  76. Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies
  77. SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge
  78. Segment, Lift and Fit: Automatic 3D Shape Labeling from 2D Prompts
  79. SkillDiffuser: Interpretable Hierarchical Planning via Skill Abstractions in Diffusion-Based Task Execution
  80. Tree-Planner: Efficient Close-loop Task Planning with Large Language Models
  81. UniFS: Universal Few-Shot Instance Perception with Point Representations
  82. VDT: General-purpose Video Diffusion Transformers via Mask Modeling
  83. VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
  84. When Pedestrian Detection Meets Multi-modal Learning: Generalist Model and Benchmark Dataset
  85. You Only Learn One Query: Learning Unified Human Query for Single-Stage Multi-person Multi-task Human-Centric Perception