PPaperPicks

Yunhong Wang

Beihang University, School of Computer Science and Engineering, Beijing, China

33 papers at tracked venues · 26 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Live-Aid: A Large-Scale Dialogue Dataset and Benchmark for Interleaved Multi-party Interactions in Live Streaming
  2. Mem²Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation
  3. APHQ-ViT: Post-Training Quantization with Average Perturbation Hessian Based Reconstruction for Vision Transformers
  4. AttriPrompt: Dynamic Prompt Composition Learning for CLIP
  5. FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix Approximation
  6. GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art
  7. GeoBEV: Learning Geometric BEV Representation for Multi-view 3D Object Detection
  8. KwaiChat: A Large-Scale Video-Driven Multilingual Mixed-Type Dialogue Corpus
  9. Multi-modal Deepfake Detection via Multi-task Audio-Visual Prompt Learning
  10. OpenRSD: Towards Open-Prompts for Object Detection in Remote Sensing Images
  11. RETAIL: Towards Real-world Travel Planning for Large Language Models
  12. RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models
  13. SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual Tracking
  14. SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
  15. TCAQ-DM: Timestep-Channel Adaptive Quantization for Diffusion Models
  16. ToolSpectrum: Towards Personalized Tool Utilization for Large Language Models
  17. Towards Robust and Controllable Text-to-Motion via Masked Autoregressive Diffusion
  18. TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments
  19. Unified Knowledge Maintenance Pruning and Progressive Recovery with Weight Recalling for Large Vision-Language Models
  20. Weak2Wise: An Automated, Lightweight Framework for Weak-LLM-Friendly Reasoning Synthesis
  21. 4Diffusion: Multi-view Video Diffusion Model for 4D Generation
  22. AGS: Affordable and Generalizable Substitute Training for Transferable Adversarial Attack
  23. ActiveDC: Distribution Calibration for Active Finetuning
  24. AdaLog: Post-training Quantization for Vision Transformers with Adaptive Logarithm Quantizer
  25. DSD-DA: Distillation-based Source Debiasing for Domain Adaptive Object Detection
  26. FSD-BEV: Foreground Self-distillation for Multi-view 3D Object Detection
  27. GLGait: A Global-Local Temporal Receptive Field Network for Gait Recognition in the Wild
  28. HIPTrack: Visual Tracking with Historical Prompts
  29. Leveraging Predicate and Triplet Learning for Scene Graph Generation
  30. MutDet: Mutually Optimizing Pre-training for Remote Sensing Object Detection
  31. Rotation Has Two Sides: Evaluating Data Augmentation for Deep One-class Classification
  32. Transforming Vision Transformer: Towards Efficient Multi-Task Asynchronous Learner
  33. Understanding Heterophily for Graph Neural Networks