PPaperPicks

Shih-Fu Chang

14 papers at tracked venues · 9 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. M²-TabFact: Multi-Document Multi-Modal Fact Verification with Visual and Textual Representations of Tabular Data
  2. PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction
  3. Beyond Grounding: Extracting Fine-Grained Event Hierarchies across Modalities
  4. Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
  5. Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning
  6. Ferret: Refer and Ground Anything Anywhere at Any Granularity
  7. JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images
  8. MoDE: CLIP Data Experts via Clustering
  9. Personalized Video Comment Generation
  10. RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos
  11. SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos
  12. Training-free Deep Concept Injection Enables Language Models for Video Question Answering
  13. VIEWS: Entity-Aware News Video Captioning
  14. What, When, and Where? Self-Supervised Spatio- Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions