PPaperPicks

Xiawu Zheng

46 papers at tracked venues · 43 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. ALGOGEN: Tool-Generated Verifiable Traces for Reliable Algorithm Visualization
  2. Connecting the Dots: Training-Free Visual Grounding via Agentic Reasoning
  3. QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension
  4. Relaxing the Constraints: A Dual-Importance Projection Mechanism for Lifelong Model Editing
  5. ALLGCD: Leveraging All Unlabeled Data for Generalized Category Discovery
  6. Aligning Instance Brownian Bridge with Texts for Open-Vocabulary Video Instance Segmentation
  7. Automated Fine-Grained Mixture-of-Experts Quantization
  8. BAME: Block-Aware Mask Evolution for Efficient N: M Sparse Training
  9. Benchmarking Abstract and Reasoning Abilities Through A Theoretical Perspective
  10. Data Interpreter: An LLM Agent for Data Science
  11. Determining Layer-wise Sparsity for Large Language Models Through a Theoretical Perspective
  12. Discovering Important Experts for Mixture-of-Experts Models Pruning Through a Theoretical Perspective
  13. Distilling Spatially-Heterogeneous Distortion Perception for Blind Image Quality Assessment
  14. Dynamic Clustering Convolutional Neural Network
  15. Dynamic Low-Rank Sparse Adaptation for Large Language Models
  16. Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
  17. Feature Denoising Diffusion Model for Blind Image Quality Assessment
  18. Few-Shot Image Quality Assessment via Adaptation of Vision-Language Models
  19. From Objects to Events: Unlocking Complex Visual Understanding in Object Detectors Via LLM-guided Symbolic Reasoning
  20. Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
  21. MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
  22. Multimodal Quantitative Language for Generative Recommendation
  23. VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
  24. Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
  25. Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
  26. Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
  27. polybasic Speculative Decoding Through a Theoretical Perspective
  28. Adaptive Feature Selection for No-Reference Image Quality Assessment by Mitigating Semantic Noise Sensitivity
  29. AffineQuant: Affine Transformation Quantization for Large Language Models
  30. Bilateral Event Mining and Complementary for Event Stream Super-Resolution
  31. Binding-Adaptive Diffusion Models for Structure-Based Drug Design
  32. Cantor: Inspiring Multimodal Chain-of-Thought of MLLM
  33. Efficient Event Stream Super-Resolution with Recursive Multi-Branch Fusion
  34. GraCo: Granularity-Controllable Interactive Segmentation
  35. Integrating Global Context Contrast and Local Sensitivity for Blind Image Quality Assessment
  36. Interaction-based Retrieval-augmented Diffusion Models for Protein-specific 3D Molecule Generation
  37. MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
  38. Motion-aware Latent Diffusion Models for Video Frame Interpolation
  39. Multi-branch Collaborative Learning Network for 3D Visual Grounding
  40. Multimodal Inplace Prompt Tuning for Open-set Object Detection
  41. Outlier-aware Slicing for Post-Training Quantization in Vision Transformer
  42. Protein-Ligand Interaction Prior for Binding-aware 3D Molecule Diffusion Models
  43. RepAn: Enhanced Annealing through Re-parameterization
  44. Semi-Supervised Blind Image Quality Assessment through Knowledge Distillation and Incremental Learning
  45. Solving the Catastrophic Forgetting Problem in Generalized Category Discovery
  46. Textual Grounding for Open-Vocabulary Visual Information Extraction in Layout-Diversified Documents