PPaperPicks

Wanli Ouyang

Chinese University of Hong Kong, Electronic Engineering, Hong Kong

98 papers at tracked venues · 84 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement
  2. ARCHE: A Novel Task to Evaluate LLMs on Latent Reasoning Chain Extraction
  3. Mitigating Low-Quality Reasoning in MLLMs: Self-Driven Refined Multimodal CoT with Selective Thinking and Step-wise Visual Enhancement
  4. Nature-Inspired Population-Based Evolution of Large Language Models
  5. ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition
  6. VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models
  7. A CLIP-Powered Framework for Robust and Generalizable Data Selection
  8. Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
  9. Biology-Instructions: A Dataset and Benchmark for Multi-Omics Sequence Understanding Capability of Large Language Models
  10. Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta Compression
  11. CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation
  12. CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming
  13. CSBrain: A Cross-scale Spatiotemporal Brain Foundation Model for EEG Decoding
  14. ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area
  15. ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems
  16. Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
  17. Depth Any Video with Scalable Synthetic Data
  18. Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback
  19. ECLR'25: 2nd Workshop on Efficient Computing Under Limited Resources: Visual Computing
  20. EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds
  21. Flow-GRPO: Training Flow Matching Models via Online RL
  22. FuncGenFoil: Airfoil Generation and Editing Model in Function Space
  23. GigaGS: 3D Gaussian Based Planar Representation for Large-Scene Surface Reconstruction
  24. HiSplat: Hierarchical 3D Gaussian Splatting for Generalizable Sparse-View Reconstruction
  25. Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
  26. Improving Video Generation with Human Feedback
  27. LLaMA-Berry: Pairwise Optimization for Olympiad-level Mathematical Reasoning via O1-like Monte Carlo Tree Search
  28. LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents
  29. Learning to Predict the Future from Monocular Vision for Efficient Human-Aware Navigation
  30. MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search
  31. MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses
  32. Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System
  33. MindAligner: Explicit Brain Functional Alignment for Cross-Subject Visual Decoding from Limited fMRI Data
  34. Multi-Modal Latent Variables for Cross-Individual Primary Visual Cortex Modeling and Analysis
  35. Native-Resolution Image Synthesis
  36. Neural Representational Consistency Emerges from Probabilistic Neural-Behavioral Representation Alignment
  37. Neuro-3D: Towards 3D Visual Decoding from EEG Signals
  38. PRING: Rethinking Protein-Protein Interaction Prediction from Pairs to Graphs
  39. PostCast: Generalizable Postprocessing for Precipitation Nowcasting via Unsupervised Blurriness Modeling
  40. RH20T-P: A Primitive-Level Robotic Manipulation Dataset towards Composable Generalization Agents in Real-world Scenarios
  41. ROGRAG: A Robustly Optimized GraphRAG Framework
  42. Retrieval is Not Enough: Enhancing RAG through Test-Time Critique and Optimization
  43. STAR: A Benchmark for Astronomical Star Fields Super-Resolution
  44. Satellite Observations Guided Diffusion Model for Accurate Meteorological States at Arbitrary Resolution
  45. ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
  46. SparseFlex: High-Resolution and Arbitrary-Topology 3D Shape Modeling
  47. SynBrain: Enhancing Visual-to-fMRI Synthesis via Probabilistic Representation Learning
  48. TAR3D: Creating High-Quality 3D Assets Via Next-Part Prediction
  49. Towards Efficient and Intelligent Laser Weeding: Method and Dataset for Weed Stem Detection
  50. Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
  51. UniSTD: Towards Unified Spatio-Temporal Learning across Diverse Disciplines
  52. Venus-MAXWELL: Efficient Learning of Protein-Mutation Stability Landscapes using Protein Language Models
  53. WeatherGFM: Learning a Weather Generalist Foundation Model via In-context Learning
  54. Where Am I and What Will I See: An Auto-Regressive Model for Spatial Localization and View Prediction
  55. scMRDR: A scalable and flexible framework for unpaired single-cell multi-omics data integration
  56. A Perspective of Q-value Estimation on Offline-to-Online Reinforcement Learning
  57. AFBench: A Large-scale Benchmark for Airfoil Design
  58. Agent3D-Zero: An Agent for Zero-Shot 3D Understanding
  59. An Embarrassingly Simple Approach to Enhance Transformer Performance in Genomic Selection for Crop Breeding
  60. BEACON: Benchmark for Comprehensive RNA Tasks and Language Models
  61. Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
  62. Boosting Residual Networks with Group Knowledge
  63. CasCast: Skillful High-resolution Precipitation Nowcasting via Cascaded Modelling
  64. ConceptMath: A Bilingual Concept-wise Benchmark for Measuring Mathematical Reasoning of Large Language Models
  65. ContraNovo: A Contrastive Learning Approach to Enhance De Novo Peptide Sequencing
  66. Dense Connector for MLLMs
  67. DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM
  68. DiffBIR: Toward Blind Image Restoration with Generative Diffusion Prior
  69. DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware Diffusion
  70. EMR-Merging: Tuning-Free High-Performance Model Merging
  71. Empowering and Assessing the Utility of Large Language Models in Crop Science
  72. Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!
  73. FNP: Fourier Neural Processes for Arbitrary-Resolution Data Assimilation
  74. FiT: Flexible Vision Transformer for Diffusion Model
  75. Frozen CLIP Transformer Is an Efficient Point Cloud Encoder
  76. GVGEN: Text-to-3D Generation with Volumetric Representation
  77. Generalizing Weather Forecast to Fine-grained Temporal Scales via Physics-AI Hybrid Modeling
  78. GraphReader: Building Graph-based Agent to Enhance Long-Context Abilities of Large Language Models
  79. Instruct-ReID: A Multi-Purpose Person Re-Identification Task with Instructions
  80. LOCR: Location-Guided Transformer for Optical Character Recognition
  81. Lumina-Next : Making Lumina-T2X Stronger and Faster with Next-DiT
  82. MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
  83. Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNA
  84. MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators
  85. NeuRodin: A Two-stage Framework for High-Fidelity Neural Surface Reconstruction
  86. Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
  87. Point Cloud Matters: Rethinking the Impact of Different Observation Spaces on Robot Learning
  88. Point Cloud Pre-Training with Diffusion Models
  89. Point Transformer V3: Simpler, Faster, Stronger
  90. PredBench: Benchmarking Spatio-Temporal Prediction Across Diverse Disciplines
  91. ProSST: Protein Language Modeling with Quantized Structure and Disentangled Attention
  92. RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models
  93. Semi-supervised 3D Object Detection with PatchTeacher and PillarMix
  94. TASeg: Temporal Aggregation Network for LiDAR Semantic Segmentation
  95. Taming Stable Diffusion for Text to 360° Panorama Image Generation
  96. Towards a Self-contained Data-driven Global Weather Forecasting Framework
  97. UniDream: Unifying Diffusion Priors for Relightable Text-to-3D Generation
  98. UniPAD: A Universal Pre-Training Paradigm for Autonomous Driving