PPaperPicks

Linchao Zhu

29 papers at tracked venues · 28 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Attention as Selector: Unlocking VLM Attention for Long Document Page Retrieval
  2. How to Improve LLMs' Performance on Specific Languages: A Perspective on LLM-Derived Language Similarity
  3. N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization
  4. Verifier-Free RL for LLMs via Intrinsic Gradient-Norm Reward
  5. 3DID: Direct 3D Inverse Design for Aerodynamics with Physics-Aware Optimization
  6. Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks
  7. DeltaPhi: Physical States Residual Learning for Neural Operators in Data-Limited PDE Solving
  8. Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
  9. FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
  10. From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-Reward Alignment
  11. H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
  12. HUST: High-Fidelity Unbiased Skin Tone Estimation via Texture Quantization
  13. Holistic Physics Solver: Learning PDEs in a Unified Spectral-Physical Space
  14. Long-horizon Visual Instruction Generation with Logic and Attribute Self-reflection
  15. MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
  16. MuTIS: Enhancing Reasoning Efficiency through Multi Turn Intervention Sampling in Reinforcement Learning
  17. Scalable Vision-Language Understanding and Generation
    AAAI 2025 · Linchao Zhu
  18. VideoGrain: Modulating Space-Time Attention for Multi-Grained Video Editing
  19. CapHuman: Capture Your Moments in Parallel Universes
  20. DGL: Dynamic Global-Local Prompt Tuning for Text-Video Retrieval
  21. FragRel: Exploiting Fragment-level Relations in the External Memory of Large Language Models
  22. FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention
  23. GG-Editor: Locally Editing 3D Avatars with Multimodal Large Language Model Guidance
  24. Knowledge-Enhanced Dual-Stream Zero-Shot Composed Image Retrieval
  25. MoS2: Mixture of Scale and Shift Experts for Text-Only Video Captioning
  26. Neural Interaction Energy for Multi-Agent Trajectory Prediction
  27. Stitching Segments and Sentences towards Generalization in Video-Text Pre-training
  28. Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
  29. VillagerAgent: A Graph-Based Multi-Agent Framework for Coordinating Complex Task Dependencies in Minecraft