PPaperPicks

Xiaoye Qu

39 papers at tracked venues · 32 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models
  2. Rethinking Video-Language Model from the Language Input Perspective
  3. Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
  4. Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment Retrieval
  5. CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
  6. Cooperative or Competitive? Understanding the Interaction between Attention Heads From A Game Theory Perspective
    ACL 2025 · Xiaoye Qu
  7. Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement Learning
  8. Dynamic Data Mixing Maximizes Instruction Tuning for Mixture-of-Experts
  9. Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think
  10. Fit the Distribution: Cross-Image/Prompt Adversarial Attacks on Multimodal Large Language Models
  11. From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration
  12. LLM-Assisted Entropy-Based Adaptive Distillation for Unsupervised Fine-Grained Visual Representation Learning
  13. Learning to Reason under Off-Policy Guidance
  14. Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment
  15. Multi-level Association Refinement Network for Dialogue Aspect-based Sentiment Quadruple Analysis
  16. Open-World Fine-Grained Fashion Retrieval with LLM-based Commonsense Knowledge Infusion
  17. PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models
  18. SEE: Continual Fine-tuning with Sequential Ensemble of Experts
  19. Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback
  20. Towards Building Model/Prompt-Transferable Attackers against Large Vision-Language Models
  21. Towards Stabilized and Efficient Diffusion Transformers Through Long-Skip-Connections With Spectral Constraints
  22. Confidence is not Timeless: Modeling Temporal Validity for Rule-based Temporal Knowledge Graph Forecasting
  23. ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLMs
  24. Enhancing Low-Resource Relation Representations through Multi-View Decoupling
  25. Frequency-Aware GAN for Imperceptible Transfer Attack on 3D Point Clouds
  26. GIST: Improving Parameter Efficient Fine-Tuning via Knowledge Interaction
  27. Improving Pseudo Labels with Global-Local Denoising Framework for Cross-lingual Named Entity Recognition
  28. LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-Training
  29. Learning the Unlearned: Mitigating Feature Suppression in Contrastive Learning
  30. Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?
  31. Mitigating Boundary Ambiguity and Inherent Bias for Text Classification in the Era of Large Language Models
  32. Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval using Language
  33. On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits Fusion
  34. Pandora's Box: Towards Building Universal Attackers against Real-World Large Vision-Language Models
  35. Rethinking Weakly-Supervised Video Temporal Grounding From a Game Perspective
  36. SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information
  37. Temporal Sentence Grounding with Relevance Feedback in Videos
  38. Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging
  39. Unsupervised Domain Adaptative Temporal Sentence Localization with Mutual Information Maximization