PPaperPicks

Yong Man Ro

Korea Advanced Institute of Science and Technology, School of Electrical Engineering, Image and Video Systems Lab, Daejeon, South Korea

21 papers at tracked venues · 17 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Emotion-Coherent Reasoning for Multimodal LLMs via Emotional Rationale Verifier
  2. Focus Where It Matters: LLM-Guided Regional Identification for Instruction-based Image Editing
  3. Long-Form Speech Generation with Spoken Language Models
  4. MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
  5. Personalized Lip Reading: Adapting to Your Unique Lip Movements with Vision and Language
  6. SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
  7. Unified Reinforcement and Imitation Learning for Vision-Language Models
  8. VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models
  9. Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
  10. AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation
  11. CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models
  12. Causal Mode Multiplexer: A Novel Framework for Unbiased Multispectral Pedestrian Detection
  13. CoLLaVO: Crayon Large Language and Vision mOdel
  14. Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation
  15. Improving Open Set Recognition via Visual Prompts Distilled from Common-Sense Knowledge
  16. Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation
  17. Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models
  18. MoAI: Mixture of All Intelligence for Large Language and Vision Models
  19. TroL: Traversal of Layers for Large Language and Vision Models
  20. What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models
  21. Where Visual Speech Meets Language: VSP-LLM Framework for Efficient and Context-Aware Visual Speech Processing