PPaperPicks

Youngjae Yu

43 papers at tracked venues · 26 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. DUSK: Do Not Unlearn Shared Knowledge
  2. Do Language Models Associate Sound with Meaning? A Multimodal Study of Sound Symbolism
  3. Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design Understanding
  4. Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation
  5. GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance
  6. Investigating Counterfactual Unfairness in LLMs towards Identities through Humor
  7. Right at My Level: A Unified Multilingual Framework for Proficiency-Aware Text Simplification
  8. Tracing Mathematical Proficiency Through Problem-Solving Processes
  9. Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?
  10. CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot Interaction
  11. C²: Scalable Auto-Feedback for LLM-based Chart Generation
  12. DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face Animation
  13. Diffusion-Driven Two-Stage Active Learning for Low-Budget Semantic Segmentation
  14. DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding
  15. Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics
  16. EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
  17. ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO
  18. KL Penalty Control via Perturbation for Direct Preference Optimization
  19. MASS: Overcoming Language Bias in Image-Text Matching
  20. MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translation
  21. Multimodal UNcommonsense: From Odd to Ordinary and Ordinary to Odd
  22. Persona Dynamics: Unveiling the Impact of Persona Traits on Agents in Text-Based Games
  23. Representation Bending for Large Language Model Safety
  24. Revisiting Residual Connections: Orthogonal Updates for Stable and Efficient Deep Networks
  25. Scalp Diagnostic System with Label-Free Segmentation and Training-Free Image Translation
  26. Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
  27. Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision Making
  28. V.I.P.: Iterative Online Preference Distillation for Efficient Video Diffusion Models
  29. VAGUE: Visual Contexts Clarify Ambiguous Expressions
  30. VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape Rooms
  31. Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation
  32. ActionSwitch: Class-Agnostic Detection of Simultaneous Actions in Streaming Videos
  33. Aligning Large Language Models by On-Policy Self-Judgment
  34. Cactus: Towards Psychological Counseling Conversations using Cognitive Behavioral Theory
  35. Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation
  36. Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!
  37. How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models
  38. Language Models as Compilers: Simulating Pseudocode Execution Improves Algorithmic Reasoning in Language Models
  39. Pearl: A Review-driven Persona-Knowledge Grounded Conversational Recommendation Dataset
  40. SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models
  41. Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
  42. Towards Visual Text Design Transfer Across Languages
  43. Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback