PPaperPicks

Dinesh Manocha

University of Maryland at College Park, MD, USA

101 papers at tracked venues · 48 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Bi-VLM: Binary Post-Training Quantization for Vision-Language Models
  2. DIAGRAMS : A Review Framework for Reasoning-Level Attribution in Diagram QA
    ACL 2026 ·
    Anirudh Iyengar Kaniyar Narayana Iyengar
  3. Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models
  4. FIGMA: Towards FIne-Grained Music retrievAl
  5. MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
  6. MoRe: Monocular Geometry Refinement via Graph Optimization for Cross-View Consistency
  7. PolyAudio: Advancing Multi-Audio Reasoning in Large Audio Language Models with Interleaved Multi-Audio Contexts
  8. Structured Uncertainty guided Clarification for LLM Agents
  9. UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery Using Gaussian Splatting
  10. A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality
  11. AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
  12. Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
  13. Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
  14. Aurelia: Test-Time Reasoning Distillation in Audio-Visual LLMs
  15. AutoSpatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning
  16. Behav: Behavioral Rule Guided Autonomy Using VLMs for Robot Navigation in Outdoor Scenes
  17. Better Features, Better Calibration: A Simple Fix for Overconfident Networks
  18. Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
    ICML 2025 ·
    Mohamad Fares El Hajj Chehade
  19. CROSS-GAiT: Cross-Attention-Based Multimodal Representation Fusion for Parametric Gait Adaptation in Complex Terrains
  20. ChartLens: Fine-grained Visual Attribution in Charts
  21. Collab: Controlled Decoding using Mixture of Agents for LLM Alignment
  22. Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation
  23. Do Audio-Language Models Understand Linguistic Variations?
  24. Does Thinking More Always Help? Mirage of Test-Time Scaling in Reasoning Models
  25. EDM: Equirectangular Projection-Oriented Dense Kernelized Feature Matching
  26. EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
  27. ET-Former: Efficient Triplane Deformable Attention for 3D Semantic Scene Completion From Monocular Camera
  28. EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception
  29. Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation
  30. Financial Models meets Generative Art: Black-Scholes-Inspired Concept Blending in Text-to-Image Diffusion
  31. Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents
  32. Gnd: Global Navigation Dataset With Multi-Modal Perception and Multi-Category Traversability in Outdoor Campus Environments
  33. How Learnable Grids Recover Fine Detail in Low Dimensions: A Neural Tangent Kernel Analysis of Multigrid Parametric Encodings
  34. IM360: Large-Scale Indoor Mapping with 360 Cameras
  35. Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
  36. Improving Zero-Shot ObjectNav with Generative Communication
  37. Is the House Ready For Sleeptime? Generating and Evaluating Situational Queries for Embodied Question Answering
  38. LBAP: Improved Uncertainty Alignment of LLM Planners using Bayesian Inference
  39. MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
  40. MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
  41. MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
  42. On the Vulnerability of LLM/VLM-Controlled Robotics
  43. PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
  44. ProSE: Diffusion Priors for Speech Enhancement
  45. PromptRefine: Enhancing Few-Shot Performance on Low-Resource Indic Languages with Example Selection from related Example Banks
  46. RELIC: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples
  47. RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph Optimization
  48. Sensible Agent: A Framework for Unobtrusive Interaction with Proactive AR Agents
  49. Since U Been Gone: Augmenting Context-Aware Transcriptions for Re-Engaging in Immersive VR Meetings
  50. Social-LLaVA: Enhancing Social Robot Navigation through Human-Language Reasoning
  51. Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic Data
  52. TK-Planes: Tiered K-Planes with High Dimensional Feature Vectors for Dynamic UAV-based Scenes
  53. Towards Optimal Multi-draft Speculative Decoding
  54. VLM-GroNav: Robot Navigation Using Physically Grounded Vision-Language Models in Outdoor Environments
  55. VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
  56. VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation
  57. Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
  58. ZSORN: Language-Driven Object-Centric Zero-Shot Object Retrieval and Navigation
  59. A Closer Look at the Limitations of Instruction Tuning
  60. ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract Descriptions
  61. AG-Cvg: Coverage Planning with a Mobile Recharging UGV and an Energy-Constrained UAV
  62. AGL-Net: Aerial-Ground Cross-Modal Global Localization with Varying Scales
  63. AMCO: Adaptive Multimodal Coupling of Vision and Proprioception for Quadruped Robot Navigation in Outdoor Environments
  64. ASPIRE: Language-Guided Data Augmentation for Improving Robustness Against Spurious Correlations
  65. AV-RIR: Audio-Visual Room Impulse Response Estimation
  66. AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
  67. Can LLM's Generate Human-Like Wayfinding Instructions? Towards Platform-Agnostic Embodied Instruction Synthesis
  68. CoDa: Constrained Generation based Data Augmentation for Low-Resource NLP
    NAACL 2024 ·
    Chandra Kiran Reddy Evuru
  69. CoNVOI: Context-aware Navigation using Vision Language Models in Outdoor and Indoor Environments
    IROS 2024 ·
    Adarsh Jagan Sathyamoorthy
  70. CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
  71. DTG : Diffusion-based Trajectory Generation for Mapless Global Navigation
  72. Do Vision-Language Models Understand Compound Nouns?
  73. Doc2Command: Furthering Language Guided Document Editing
  74. DocEdit-v2: Document Structure Editing Via Multimodal LLM Grounding
  75. EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
  76. GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
  77. Hallusionbench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
  78. IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning
  79. LANCAR: Leveraging Language for Context-Aware Robot Locomotion in Unstructured Environments
  80. LTM: Lightweight Textured Mesh Extraction and Refinement of Large Unbounded Scenes for Efficient Storage and Real-Time Rendering
  81. LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
  82. MEERKAT: Audio-Visual Large Language Model for Grounding in Space and Time
  83. MELFuSION: Synthesizing Music from Image and Language Cues Using Diffusion Models
  84. MIM: Indoor and Outdoor Navigation in Complex Environments Using Multi-Layer Intensity Maps
    ICRA 2024 ·
    Adarsh Jagan Sathyamoorthy
  85. MTG: Mapless Trajectory Generator with Traversability Coverage for Outdoor Navigation
  86. MaxMin-RLHF: Alignment with Diverse Human Preferences
  87. PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human Feedback
  88. PoCo: Point Context Cluster for RGBD Indoor Place Recognition
  89. Position: On the Possibilities of AI-Generated Text Detection
  90. SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition
  91. Sim-to-Real Robotic Sketching using Behavior Cloning and Reinforcement Learning
  92. Speech2UnifiedExpressions: Synchronous Synthesis of Co-Speech Affective Face and Body Expressions from Affordable Inputs
  93. TAME-RD: Text Assisted Replication of Image Multi-Adjustments for Reverse Designing
  94. Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time Oracles
  95. Transfer Q-star : Principled Decoding for LLM Alignment
  96. UAV-Sim: NeRF-based Synthetic Data Generation for UAV-based Perception
  97. Unconstrained Model Predictive Control for Robot Navigation under Uncertainty
  98. V-Trans4Style: Visual Transition Recommendation for Video Production Style Adaptation
  99. VAPOR: Legged Robot Navigation in Unstructured Outdoor Environments using Offline Reinforcement Learning
  100. VLPG-Nav: Object Navigation Using Visual Language Pose Graph and Object Localization Probability Maps
  101. When, What, and with Whom to Communicate: Enhancing RL-based Multi-Robot Navigation through Selective Communication