PPaperPicks

Chitta Baral

Arizona State University, Tempe, Arizona, USA

40 papers at tracked venues · 22 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. FAMA: Failure-Aware Meta-Agentic Framework for Open-Source LLMs in Interactive Tool Use Environments
  2. From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents
  3. MMTabReal: Real-World Benchmark for Multimodal Table Understanding
    ACL 2026 ·
    Prasham Yatinkumar Titiya
  4. The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
  5. AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
  6. ActionReasoningBench: Reasoning about Actions with and without Ramification Constraints
  7. EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven Alignment
  8. GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
  9. How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on tau-bench
  10. Hypothesis Generation for Materials Discovery and Design Using Goal-Driven and Constraint-Guided LLM Agents
  11. Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks
  12. Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning
  13. Map&Make: Schema Guided Text to Table Generation
  14. PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem Solving
  15. PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem Solving
  16. QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
  17. RefEdit: A Benchmark and Method for Improving Instruction-Based Image Editing Model on Referring Expressions
  18. Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
  19. TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
  20. ThinkTuning: Instilling Cognitive Reflections without Distillation
  21. ToW: Thoughts of Words Improve Reasoning in Large Language Models
  22. UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization
  23. VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
  24. Chaos with Keywords: Exposing Large Language Models Sycophancy to Misleading Keywords and Evaluating Defense Strategies
  25. ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
  26. ECLIPSE: A Resource-Efficient Text-to-Image Prior for Image Generations
  27. Evaluating Multimodal Large Language Models across Distribution Shifts and Augmentations
  28. Getting it Right: Improving Spatial Consistency in Text-to-Image Models
  29. Grounding Stylistic Domain Generalization with Quantitative Domain Shift Measures and Synthetic Scene Images
  30. InstructABSA: Instruction Learning for Aspect Based Sentiment Analysis
  31. Investigating Acceleration of LLaMA Inference by Enabling Intermediate Layer Decoding via Instruction Tuning with 'LITE'
  32. Learning Temporally Composable Task Segmentations with Language
  33. LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
  34. Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts
  35. Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
  36. On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
  37. REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
  38. Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
  39. The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness
  40. TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives