PPaperPicks

Mohit Bansal

89 papers at tracked venues · 60 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
  2. DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
  3. GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
  4. Instruction Tuning with and without Context: Behavioral Shifts and Downstream Impact
  5. PRInTS: Reward Modeling for Long-Horizon Information Seeking
  6. PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise
  7. RotBench: Evaluating Multi-modal Large Language Models on Identifying Image Rotation
  8. Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection
  9. Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
  10. Stabilizing Efficient Reasoning with Step-Level Advantage Selection
  11. TimeRefine: Temporal Grounding with Time Refining Video LLM
  12. 4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time
  13. AdaCAD: Adaptively Decoding to Balance Conflicts between Contextual and Parametric Knowledge
  14. Adapt-∞: Scalable Continual Multimodal Instruction Tuning via Dynamic Data Selection
  15. Anyprefer: An Agentic Framework for Preference Data Synthesis
  16. Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
  17. Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
  18. CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
  19. CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
  20. Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
  21. Dam: Dynamic Adapter Merging for Continual Video QA Learning
  22. DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback
  23. FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipeline
  24. Glider: Global and Local Instruction-Driven Expert Router
  25. Improving Faithfulness of Text-to-Image Diffusion Models through Inference Intervention
  26. LAQuer: Localized Attribution Queries in Content-grounded Generation
  27. LASeR: Learning to Adaptively Select Reward Models with Multi-Arm Bandits
  28. Language Models Identify Ambiguities and Exploit Loopholes
  29. M3DocVQA: Multi-Modal Multi-Page Multi-Document Understanding
  30. MAMM-Refine: A Recipe for Improving Faithfulness in Generation with Multi-Agent Collaboration
  31. MAgICoRe: Multi-Agent, Iterative, Coarse-to-Fine Refinement for Reasoning
  32. MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
  33. Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
  34. Multi-Attribute Steering of Language Models via Targeted Intervention
  35. On Positional Bias of Faithfulness for Long-form Summarization
  36. RACCooN: Versatile Instructional Video Editing with Auto-Generated Narratives
  37. ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
  38. Reverse Thinking Makes LLMs Stronger Reasoners
  39. SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
  40. SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
  41. See It from My Perspective: How Language Affects Cultural Bias in Image Understanding
  42. Self-Consistency Preference Optimization
  43. System 1.x: Learning to Balance Fast and Slow Planning with Language Models
  44. Teaching Models to Balance Resisting and Accepting Persuasion
  45. Unbounded: A Generative Infinite Game of Character Life Simulation
  46. VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning
  47. VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
  48. Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
  49. Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
  50. VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
  51. A Simple LLM Framework for Long-Range Video Question-Answering
  52. ACUEval: Fine-grained Hallucination Evaluation and Correction for Abstractive Summarization
  53. ADaPT: As-Needed Decomposition and Planning with Language Models
  54. Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
  55. Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
  56. Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks
  57. CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation
  58. Contrastive Region Guidance: Improving Grounding in Vision-Language Models Without Training
  59. D2 Pruning: Message Passing for Balancing Diversity & Difficulty in Data Pruning
  60. Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
  61. Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
  62. ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
  63. Evaluating Very Long-Term Conversational Memory of LLM Agents
  64. Explaining and Improving Contrastive Decoding by Extrapolating the Probabilities of a Huge and Hypothetical LM
  65. Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024
  66. GTBench: Uncovering the Strategic Reasoning Capabilities of LLMs via Game-Theoretic Evaluations
  67. Inducing Systematicity in Transformers by Attending to Structurally Quantized Embeddings
  68. Knowledge-Aware Reasoning over Multimodal Semi-structured Tables
  69. LACIE: Listener-Aware Finetuning for Calibration in Large Language Models
  70. LLM Self-Correction with DeCRIM: Decompose, Critique, and Refine for Enhanced Following of Instructions with Multiple Constraints
  71. MAGDi: Structured Distillation of Multi-Agent Interaction Graphs Improves Reasoning in Smaller Language Models
  72. MIRACLE: An Online, Explainable Multimodal Interactive Concept Learning System
  73. Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
  74. Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
  75. Multimodal Representation Learning by Alternating Unimodal Adaptation
  76. Position: TrustLLM: Trustworthiness in Large Language Models
  77. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024
  78. Prompting Vision-Language Models For Aspect-Controlled Generation of Referring Expressions
  79. REFINESUMM: Self-Refining MLLM for Generating a Multimodal Summarization Dataset
  80. ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
  81. ReGAL: Refactoring Programs to Discover Generalizable Abstractions
  82. Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models
  83. Rethinking Interactive Image Segmentation with Low Latency, High Quality, and Diverse Prompts
  84. SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
  85. Soft Self-Consistency Improves Language Models Agents
  86. The Power of Summary-Source Alignments
  87. The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
  88. VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation
  89. Zero-Shot Controllable Image-to-Video Animation via Motion Decomposition