PPaperPicks

Chenliang Xu

27 papers at tracked venues · 22 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
  2. DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning
  3. Tuning the Face: Modulating Facial Expressions for Realistic Self-Avatars in Virtual Reality
  4. $\pi$-AVAS: Can Physics-Integrated Audio-Visual Modeling Boost Neural Acoustic Synthesis?
  5. BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching Models
  6. CaRDiff: Video Salient Object Ranking Chain of Thought Reasoning for Saliency Prediction with Diffusion
  7. Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
  8. Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
  9. Generative AI for Cel-Animation: A Survey
  10. GestureLSM: Latent Shortcut Based Co-Speech Gesture Generation with Spatial-Temporal Modeling
  11. Harnessing the Computation Redundancy in ViTs to Boost Adversarial Transferability
  12. Learning to Highlight Audio by Watching Movies
  13. MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
  14. Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives
  15. Targeted Forgetting of Image Subgroups in CLIP Models
  16. Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach
  17. V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning
  18. VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
  19. ZeroSep: Separate Anything in Audio with Zero Training
  20. Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
  21. Discover and Mitigate Multiple Biased Subgroups in Image Classifiers
  22. EAGLE: Egocentric AGgregated Language-video Engine
  23. Learning to Transform Dynamically for Better Adversarial Transferability
  24. Modeling and Driving Human Body Soundfields Through Acoustic Primitives
  25. OSCaR: Object State Captioning and State Change Representation
  26. One Forward is Enough for Neural Network Training via Likelihood Ratio Method
  27. Tri2-plane: Thinking Head Avatar via Feature Pyramid