PPaperPicks

Hehe Fan

31 papers at tracked venues · 26 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. DLVINet: Advancing Dual-Lens Video Inpainting Beyond Parallax Constraints
  2. GraphTARIF: Linear Graph Transformer with Augmented Rank and Improved Focus
  3. One Refiner to Unlock Them All: Inference-Time Reasoning Elicitation via Reinforcement Query Refinement
  4. RegionSLM: Region-aware Question Answering on Document Screenshots
  5. Adapting Text-to-Image Generation with Feature Difference Instruction for Generic Image Restoration
  6. BVINet: Unlocking Blind Video Inpainting With Zero Annotations
  7. Drafting and Revision: Advancing High-Fidelity Video Inpainting
  8. DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization
  9. Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
  10. EnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent Space
  11. InfiniDreamer: Arbitrarily Long Human Motion Generation Via Segment Score Distillation
  12. MMAD: Multi-Label Micro-Action Detection in Videos
  13. Motion Matters: Motion-guided Modulation Network for Skeleton-based Micro-Action Recognition
  14. OSDA Agent: Leveraging Large Language Models for De Novo Design of Organic Structure Directing Agents
  15. Prompt-Aware Controllable Shadow Removal
  16. ProtChatGPT: Towards Understanding Proteins with Hybrid Representation and Large Language Models
  17. Prototypical Calibrating Ambiguous Samples for Micro-Action Recognition
  18. Reaction Graph: Towards Reaction-Level Modeling for Chemical Reactions with 3D Structures
  19. TSGS: Improving Gaussian Splatting for Transparent Surface Reconstruction via Normal and De-lighting Priors
  20. VideoGrain: Modulating Space-Time Attention for Multi-Grained Video Editing
  21. Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
  22. ZeroMamba: Exploring Visual State Space Model for Zero-Shot Learning
  23. Clustering for Protein Representation Learning
  24. DocMSU: A Comprehensive Benchmark for Document-Level Multimodal Sarcasm Understanding
  25. Hand-Centric Motion Refinement for 3D Hand-Object Interaction via Hierarchical Spatial-Temporal Modeling
  26. HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting
  27. Improving Context Understanding in Multimodal Large Language Models via Multimodal Composition Learning
  28. Progressive Point Cloud Denoising with Cross-Stage Cross-Coder Adaptive Edge Graph Convolution Network
  29. TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
  30. Uncovering what, why and How: A Comprehensive Benchmark for Causation Understanding of Video Anomaly
  31. VividDreamer: Invariant Score Distillation for Hyper-Realistic Text-to-3D Generation