PPaperPicks

Shengqiong Wu

19 papers at tracked venues · 19 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Orthogonal Spatial-temporal Distributional Transfer for 4D Generation
  2. Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
    AAAI 2025 · Shengqiong Wu
  3. JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
  4. Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual Scene
    CVPR 2025 · Shengqiong Wu
  5. On Path to Multimodal Generalist: General-Level and General-Bench
  6. The ACM Multimedia 2025 Grand Challenge of Multimodal Conversational Aspect-based Sentiment Analysis
  7. Towards Semantic Equivalence of Tokenization in Multimodal LLM
    ICLR 2025 · Shengqiong Wu
  8. Universal Scene Graph Generation
    CVPR 2025 · Shengqiong Wu
  9. VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
  10. Dysen-VDM: Empowering Dynamics-Aware Text-to-Video Diffusion with LLMs
  11. MMLSCU: A Dataset for Multi-modal Multi-domain Live Streaming Comment Understanding
  12. NExT-GPT: Any-to-Any Multimodal LLM
    ICML 2024 · Shengqiong Wu
  13. OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
  14. PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis
  15. Recognizing Everything from All Modalities at Once: Grounded Multimodal Universal Information Extraction
  16. Self-Adaptive Fine-grained Multi-modal Data Augmentation for Semi-supervised Muti-modal Coreference Resolution
  17. SpeechEE: A Novel Benchmark for Speech Event Extraction
  18. Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
  19. Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing