PPaperPicks

Benjamin Van Durme

Johns Hopkins University, USA

54 papers at tracked venues · 31 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Bonsai: Interpretable Tree-Adaptive Grounded Reasoning
  2. Can LLMs Identify Tax Abuse?
  3. CoverageBench: Evaluating Information Coverage across Tasks and Domains
  4. Does Reasoning Make Search More Fair? Comparing Fairness in Reasoning and Non-reasoning Rerankers
  5. How Grounded is Wikipedia? A Study on Structured Evidential Support and Retrieval
  6. Language Models and Logic Programs for Trustworthy Tax Reasoning
  7. Multi-Vector Index Compression in Any Modality
  8. WikiVideo: Article Generation from Multiple Videos
  9. arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation
  10. CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?
  11. CLERC: A Dataset for U. S. Legal Case Retrieval and Retrieval-Augmented Analysis Generation
  12. Certified Mitigation of Worst-Case LLM Copyright Infringement
  13. Conformal Linguistic Calibration: Trading-off between Factuality and Specificity
  14. Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
  15. Core: Robust Factual Precision with Informative Sub-Claim Identification
  16. DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation
  17. FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions
  18. From Models to Microtheories: Distilling a Model's Topical Knowledge for Grounded Question-Answering
  19. Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass
  20. Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering
  21. Jailbreak Distillation: Renewable Safety Benchmarking
  22. LLM Agents for Coordinating Multi-User Information Gathering
  23. MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
  24. Multi-Field Adaptive Retrieval
  25. MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval
  26. Promptriever: Instruction-Trained Retrievers Can Be Prompted Like Language Models
  27. RATIONALYST: Pre-training Process-Supervision for Improving Reasoning
  28. RE-AdaptIR: Improving Information Retrieval through Reverse Engineered Adaptation
  29. SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
  30. TurkingBench: A Challenge Benchmark for Web Agents
  31. Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
  32. Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval
  33. WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
  34. mFollowIR: A Multilingual Benchmark for Instruction Following in Retrieval
  35. A Survey of Video Datasets for Grounded Event Understanding
  36. Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
  37. Do Androids Know They're Only Dreaming of Electric Sheep?
  38. Dodo: Dynamic Contextual Compression for Decoder-only LMs
  39. Enhancing Systematic Decompositional Natural Language Inference Using Informal Logic
  40. FAMuS: Frames Across Multiple Sources
  41. Grounding Partially-Defined Events in Multimodal Data
  42. Interpreting User Requests in the Context of Natural Language Standing Instructions
  43. LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
  44. LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
  45. Language-to-Code Translation with a Single Labeled Example
  46. Learning to Retrieve Iteratively for In-Context Learning
  47. NELLIE: A Neuro-Symbolic Inference Engine for Grounded, Compositional, and Explainable Reasoning
  48. Narrowing the Gap between Zero- and Few-shot Machine Translation by Matching Styles
  49. Natural Language Decomposition and Interpretation of Complex Utterances
  50. Ontologically Faithful Generation of Non-Player Character Dialogues
  51. RORA: Robust Free-Text Rationale Evaluation
  52. SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation
  53. TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
  54. Zero and Few-shot Semantic Parsing with Ambiguous Inputs