PPaperPicks

Xu Sun

Peking University, School of EECS, MOE Key Lab of Computational Linguistics, China

19 papers at tracked venues · 14 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
  2. TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
  3. ATLANTIS: Weak-to-Strong Learning via Importance Sampling
  4. Generative Frame Sampler for Long Video Understanding
  5. InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation
  6. PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
  7. RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
  8. Temporal Reasoning Transfer from Text to Video
  9. TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
  10. VidTwin: Video VAE with Decoupled Structure and Dynamics
  11. A Survey on In-context Learning
  12. Edit As You Wish: Video Caption Editing with Multi-grained User Control
  13. Enhancing Byzantine-Resistant Aggregations with Client Embedding
  14. LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?
  15. TempCompass: Do Video LLMs Really Understand Videos?
  16. TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
  17. Towards Codable Watermarking for Injecting Multi-Bits Information to LLMs
  18. VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
  19. Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents