PPaperPicks

Wenhai Wang

40 papers at tracked venues · 34 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models
  2. LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment
  3. LLMQuA: Practical Backdoor Injection on Large Language Model Quantization
  4. SDAR: A Synergistic Diffusion-AutoRegression Paradigm for Scalable Sequence Generation
  5. Selective Knowledge Distillation: Fusing LLM Semantic Strengths with DNN Efficiency for Binary Code Similarity Detection
  6. Watch Out Your Industrial Copilots: Stealthy Backdoor Attack Against LLM-Based PLC Code Generation
  7. ArchCAD-400K: A Large-Scale CAD drawings Dataset and New Baseline for Panoptic Symbol Spotting
  8. ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area
  9. CoMemo: LVLMs Need Image Context with Image Memory
  10. Diffuse&Refine: Intrinsic Knowledge Generation and Aggregation for Incremental Object Detection
  11. Docopilot: Improving Multimodal Models for Document-Level Understanding
  12. HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
  13. Lumina-Image 2.0: a Unified and Efficient Image Generative Framework
  14. MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible Cost
  15. NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
  16. OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance Information
  17. OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis
  18. OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
  19. OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
  20. PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
  21. Point or Line? Using Line-based Representation for Panoptic Symbol Spotting in CAD Drawings
  22. Sticking to the Mean: Detecting Sticky Tokens in Text Embedding Models
  23. UltraModel: A Modeling Paradigm for Industrial Objects
  24. Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction
  25. Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting
  26. Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
  27. AVSegFormer: Audio-Visual Segmentation with Transformer
  28. Bounding Box Stability against Feature Dropout Reflects Detector Generalization across Environments
  29. ControlLLM: Augment Language Models with Tools by Searching on Graphs
  30. Distilling Knowledge from Large-Scale Image Models for Object Detection
  31. Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
  32. Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
  33. InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
  34. Needle In A Multimodal Haystack
  35. RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis
  36. The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
  37. The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World
  38. Tram: A Token-level Retrieval-augmented Mechanism for Source Code Summarization
  39. Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning
  40. VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks