PPaperPicks

Jiaqi Wang

Shanghai Artificial Intelligence Laboratory, China

54 papers at tracked venues · 49 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Game Ground Bench: Probing the Limits of LVLMs in Complex Semantic Grounding Across Game Universes
  2. MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
  3. SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition
    AAAI 2026 · Jiaqi Wang
  4. Spikingformer: A Key Foundation Model for Spiking Neural Networks
  5. WebSynthesis: World Model-Guided Monte Carlo Tree Search for Efficient WebAgent Trajectory Synthesis
  6. Bootstrap3D: Improving Multi-View Diffusion Model with Synthetic Data
  7. BrainECHO: Semantic Brain Signal Decoding through Vector-Quantized Spectrogram Reconstruction for Whisper-Enhanced Text Generation
  8. ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way
  9. Conical Visual Concentration for Efficient Large Vision-Language Models
  10. Deciphering Cross-Modal Alignment in Large Vision-Language Models Via Modality Integration Rate
  11. Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
  12. HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance
  13. IDArb: Intrinsic Decomposition for Arbitrary Number of Input Views and Illuminations
  14. InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
  15. Light-a-Video: Training-Free Video Relighting via Progressive Light Fusion
  16. MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models
  17. MM-IFEngine: Towards Multimodal Instruction Following
  18. MotionClone: Training-Free Motion Cloning for Controllable Video Generation
  19. OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
  20. OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
  21. S$2$M-Former: Spiking Symmetric Mixing Branchformer for Brain Auditory Attention Detection
    NeurIPS 2025 · Jiaqi Wang
  22. SAM2LONG: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree
  23. SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
  24. SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation
  25. Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
    NeurIPS 2025 · Jiaqi Wang
  26. Thread the Needle: Genomics-Guided Prompt-Bridged Attention Model for Survival Prediction of Glioma Based on MRI Images
  27. Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings
  28. Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
  29. Utilize the Flow Before Stepping into the Same River Twice: Certainty Represented Knowledge Flow for Refusal-Aware Instruction Tuning
  30. VideoRoPE: What Makes for Good Video Rotary Position Embedding?
  31. Visual-RFT: Visual Reinforcement Fine-Tuning
  32. X-Prompt: Generalizable Auto-Regressive Visual Learning with In-Context Prompting
  33. AIGCs Confuse AI Too: Investigating and Explaining Synthetic Image-induced Hallucinations in Large Vision-Language Models
  34. Adversarial Prompt Tuning for Vision-Language Models
  35. Alpha-CLIP: A CLIP Model Focusing on Wherever you Want
  36. Are We on the Right Way for Evaluating Large Vision-Language Models?
  37. CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers
  38. Enhancing EEG-to-Text Decoding through Transferable Representations from Pre-trained Contrastive EEG-Text Masked Autoencoder
    ACL 2024 · Jiaqi Wang
  39. FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models
  40. GPT4Point: A Unified Framework for Point-Language Understanding and Generation
  41. InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
  42. Long-CLIP: Unlocking the Long-Text Capability of CLIP
  43. MMBench: Is Your Multi-modal Model an All-Around Player?
  44. MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
  45. MMLONGBENCH-DOC: Benchmarking Long-context Document Understanding with Visualizations
  46. Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials
  47. OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
  48. OneLLM: One Framework to Align All Modalities with Language
  49. Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
  50. ShareGPT4V: Improving Large Multi-modal Models with Better Captions
  51. ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
  52. Streaming Long Video Understanding with Large Language Models
  53. VIGC: Visual Instruction Generation and Correction
  54. VLMEvalKit: An Open-Source ToolKit for Evaluating Large Multi-Modality Models