PPaperPicks

Ran He

Center for Research on Intelligent Perception and Computing, Beijing, China

40 papers at tracked venues · 38 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. CoGrad3D: Spatially-Coupled Timestep Optimization with Orthogonal Gradient Fusion for 3D Generation
  2. Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
  3. What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time
  4. Breaking Mental Set to Improve Reasoning through Diverse Multi-Agent Debate
  5. Breaking the Low-Rank Dilemma of Linear Attention
  6. Cooperative Pseudo Labeling for Unsupervised Federated Classification
  7. DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling
  8. Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
  9. Exploring Vacant Classes in Label-Skewed Federated Learning
  10. InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning
  11. LoRA-Pro: Are Low-Rank Adapters Properly Optimized?
  12. MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
  13. Protecting Model Adaptation from Trojans in the Unlabeled Data
  14. R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning
  15. Rectifying Magnitude Neglect in Linear Attention
  16. Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability Theory
  17. Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens
  18. The Illusion of Progress? A Critical Look at Test-Time Adaptation for Vision-Language Models
  19. Towards Robust Defense Against Customization via Protective Perturbation Resistant to Diffusion-based Purification
  20. VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
  21. Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
  22. Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
  23. ZeroPatcher: Training-free Sampler for Video Inpainting and Editing
  24. A Hard-to-Beat Baseline for Training-free CLIP-based Adaptation
  25. Backdoor Defense via Test-Time Detecting and Repairing
  26. Connecting the Dots: Collaborative Fine-tuning for Black-Box Vision-Language Models
  27. DeVAn: Dense Video Annotation for Video-Language Models
  28. Hallo3D: Multi-Modal Hallucination Detection and Mitigation for Consistent 3D Content Generation
  29. Heterogeneous Test-Time Training for Multi-Modal Person Re-identification
  30. InfiMM: Advancing Multimodal Understanding with an Open-Sourced Visual Language Model
  31. Multimodal Prompt Perceiver: Empower Adaptiveness, Generalizability and Fidelity for All-in-One Image Restoration
  32. RMT: Retentive Networks Meet Vision Transformers
  33. Realistic Unsupervised CLIP Fine-tuning with Universal Entropy Optimization
  34. STAMP: Outlier-Aware Test-Time Adaptation with Stable Memory Replay
  35. ScaleCrafter: Tuning-free Higher-Resolution Visual Generation with Diffusion Models
  36. Thought Propagation: an Analogical Approach to Complex Reasoning with Large Language Models
  37. Towards Eliminating Hard Label Constraints in Gradient Inversion Attacks
  38. Uncertainty-Aware Source-Free Adaptive Image Super-Resolution with Wavelet Augmentation Transformer
  39. Visual Anchors Are Strong Information Aggregators For Multimodal Large Language Model
  40. ZePo: Zero-Shot Portrait Stylization with Faster Sampling