PPaperPicks

R. Manmatha

9 papers at tracked venues · 5 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
  2. Scaling up Image Segmentation across Data and Tasks
  3. DEED: Dynamic Early Exit on Decoder for Accelerating Encoder-Decoder Transformer Models
  4. DocFormerv2: Local Features for Document Understanding
  5. DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
  6. Multiple-Question Multiple-Answer Text-VQA
  7. No Head Left Behind - Multi-Head Alignment Distillation for Transformers
  8. On the Scalability of Diffusion-based Text-to-Image Generation
  9. VisFocus: Prompt-Guided Vision Encoders for OCR-Free Dense Document Understanding