P
PaperPicks
Conferences
Arman Cohan
82 papers at tracked venues · 42 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID search ↗
Venues
ACL
×31
EMNLP
×25
NAACL
×11
NeurIPS
×4
EACL
×2
ECIR
×2
ICLR
×2
ICML
×2
AAAI
×1
CVPR
×1
SIGIR
×1
Frequent coauthors
Yilun Zhao
DBLP profile ↗
ORCID search ↗
×10
Yixin Liu
DBLP profile ↗
ORCID search ↗
×5
Xiangru Tang
DBLP profile ↗
ORCID search ↗
×4
Gabrielle Kaili-May Liu
DBLP profile ↗
ORCID search ↗
×3
Simeng Han
DBLP profile ↗
ORCID search ↗
×3
Benlu Wang
DBLP profile ↗
ORCID search ↗
×2
Sihong Wu
DBLP profile ↗
ORCID search ↗
×2
Orion Weller
DBLP profile ↗
ORCID search ↗
×2
Zhaojian Yu
DBLP profile ↗
ORCID search ↗
×2
Aniketh Garikaparthi
DBLP profile ↗
ORCID search ↗
×2
Chunyuan Deng
DBLP profile ↗
ORCID search ↗
×2
Linyong Nan
DBLP profile ↗
ORCID search ↗
×2
Papers
A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning
ACL 2026
·
Tianyu Yang
DBLP profile ↗
ORCID search ↗
A Survey on Evaluation of LLM-based Agents
ACL 2026
·
Asaf Yehudai
DBLP profile ↗
ORCID search ↗
Anchor: Branch-Point Data Generation for GUI Agents
ACL 2026
·
Jinbiao Wei
DBLP profile ↗
ORCID search ↗
CPTCoder: A Reliable LLM System for Medical Procedure Code Prediction
ACL 2026
·
Benlu Wang
DBLP profile ↗
ORCID search ↗
Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future
ACL 2026
·
Sihong Wu
DBLP profile ↗
ORCID search ↗
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics
ACL 2026
·
Jinu Lee
DBLP profile ↗
ORCID search ↗
Investigating Retrieval-Augmented Generation Systems on Unanswerable, Uncheatable, Realistic, Multi-hop Queries
ECIR 2026
·
Gabrielle Kaili-May Liu
DBLP profile ↗
ORCID search ↗
MMSciCode: Real-world Evaluation of Multilingual Multi-Discipline Scientific Research Coding
ACL 2026
·
Xue Xia
DBLP profile ↗
ORCID search ↗
MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application
ACL 2026
·
Xueqing Peng
DBLP profile ↗
ORCID search ↗
Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL
EACL 2026
·
Yifei Shen
DBLP profile ↗
ORCID search ↗
RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation
ACL 2026
·
Sihong Wu
DBLP profile ↗
ORCID search ↗
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
ACL 2026
·
Yilun Zhao
DBLP profile ↗
ORCID search ↗
SciMDR: Advancing Scientific Multimodal Document Reasoning
ACL 2026
·
Ziyu Chen
DBLP profile ↗
ORCID search ↗
SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature
EACL 2026
·
Hang Ding
DBLP profile ↗
ORCID search ↗
A Large-Scale Study of Reranker Relevance Feedback at Inference
SIGIR 2025
·
Revanth Gangi Reddy
DBLP profile ↗
ORCID search ↗
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research
ACL 2025
·
Yilun Zhao
DBLP profile ↗
ORCID search ↗
Are Multimodal LLMs Robust Against Adversarial Perturbations? RoMMath: A Systematic Evaluation on Multimodal Math Reasoning
NAACL 2025
·
Yilun Zhao
DBLP profile ↗
ORCID search ↗
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
ACL 2025
·
Zhijian Xu
DBLP profile ↗
ORCID search ↗
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
ACL 2025
·
Yilun Zhao
DBLP profile ↗
ORCID search ↗
ChemAgent: Self-updating Memories in Large Language Models Improves Chemical Reasoning
ICLR 2025
·
Xiangru Tang
DBLP profile ↗
ORCID search ↗
CourtReasoner: Can LLM Agents Reason Like Judges?
EMNLP 2025
·
Sophia Simeng Han
DBLP profile ↗
ORCID search ↗
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective
EMNLP 2025
·
Siyue Zhang
DBLP profile ↗
ORCID search ↗
DyFlow: Dynamic Workflow Framework for Agentic Reasoning
NeurIPS 2025
·
Yanbo Wang
DBLP profile ↗
ORCID search ↗
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
EMNLP 2025
·
Yitao Long
DBLP profile ↗
ORCID search ↗
FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
EMNLP 2025
·
Tiansheng Hu
DBLP profile ↗
ORCID search ↗
FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions
NAACL 2025
·
Orion Weller
DBLP profile ↗
ORCID search ↗
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations
EMNLP 2025
·
Benlu Wang
DBLP profile ↗
ORCID search ↗
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Task
ACL 2025
·
Zhaojian Yu
DBLP profile ↗
ORCID search ↗
IRIS: Interactive Research Ideation System for Accelerating Scientific Discovery
ACL 2025
·
Aniketh Garikaparthi
DBLP profile ↗
ORCID search ↗
Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplification and Resistance in Multi-Agent Based LLM-as-Judge
EMNLP 2025
·
Chiyu Ma
DBLP profile ↗
ORCID search ↗
LimRank: Less is More for Reasoning-Intensive Information Reranking
EMNLP 2025
·
Tingyu Song
DBLP profile ↗
ORCID search ↗
LocAgent: Graph-Guided LLM Agents for Code Localization
ACL 2025
·
Zhaoling Chen
DBLP profile ↗
ORCID search ↗
MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search
EMNLP 2025
·
Yunhai Hu
DBLP profile ↗
ORCID search ↗
MDCure: A Scalable Pipeline for Multi-Document Instruction-Following
ACL 2025
·
Gabrielle Kaili-May Liu
DBLP profile ↗
ORCID search ↗
MIR: Methodology Inspiration Retrieval for Scientific Research Problems
ACL 2025
·
Aniketh Garikaparthi
DBLP profile ↗
ORCID search ↗
MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
CVPR 2025
·
Yilun Zhao
DBLP profile ↗
ORCID search ↗
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
NeurIPS 2025
·
Andrew M. Bean
DBLP profile ↗
ORCID search ↗
MedTutor: A Retrieval-Augmented LLM System for Case-Based Medical Education
EMNLP 2025
·
Dongsuk Jang
DBLP profile ↗
ORCID search ↗
MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
EMNLP 2025
·
Gabrielle Kaili-May Liu
DBLP profile ↗
ORCID search ↗
On Evaluating LLM Alignment by Evaluating LLMs as Judges
NeurIPS 2025
·
Yixin Liu
DBLP profile ↗
ORCID search ↗
Physics: Benchmarking Foundation Models on University-Level Physics Problem Solving
ACL 2025
·
Kaiyue Feng
DBLP profile ↗
ORCID search ↗
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
NAACL 2025
·
Mingqi Gao
DBLP profile ↗
ORCID search ↗
ReIFE: Re-evaluating Instruction-Following Evaluation
NAACL 2025
·
Yixin Liu
DBLP profile ↗
ORCID search ↗
Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language Models
ACL 2025
·
Junjie Wu
DBLP profile ↗
ORCID search ↗
RouterRetriever: Routing over a Mixture of Expert Embedding Models
AAAI 2025
·
Hyunji Lee
DBLP profile ↗
ORCID search ↗
SCIURus: Shared Circuits for Interpretable Uncertainty Representations in Language Models
NAACL 2025
·
Carter Teplica
DBLP profile ↗
ORCID search ↗
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
NeurIPS 2025
·
Yilun Zhao
DBLP profile ↗
ORCID search ↗
SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific Literature
EMNLP 2025
·
David Wadden
DBLP profile ↗
ORCID search ↗
SciSketch: An Open-source Framework for Automated Schematic Diagram Generation in Scientific Papers
EMNLP 2025
·
Zihang Wang
DBLP profile ↗
ORCID search ↗
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
ACL 2025
·
Chengye Wang
DBLP profile ↗
ORCID search ↗
TESS 2: A Large-Scale Generalist Diffusion Language Model
ACL 2025
·
Jaesung Tae
DBLP profile ↗
ORCID search ↗
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
ICLR 2025
·
Ziyao Shangguan
DBLP profile ↗
ORCID search ↗
Table-R1: Inference-Time Scaling for Table Reasoning Tasks
EMNLP 2025
·
Zheyuan Yang
DBLP profile ↗
ORCID search ↗
Understanding Reference Policies in Direct Preference Optimization
NAACL 2025
·
Yixin Liu
DBLP profile ↗
ORCID search ↗
Z1: Efficient Test-time Scaling with Code
EMNLP 2025
·
Zhaojian Yu
DBLP profile ↗
ORCID search ↗
mFollowIR: A Multilingual Benchmark for Instruction Following in Retrieval
ECIR 2025
·
Orion Weller
DBLP profile ↗
ORCID search ↗
Bayesian Calibration of Win Rate Estimation with LLM Evaluators
EMNLP 2024
·
Yicheng Gao
DBLP profile ↗
ORCID search ↗
Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
NAACL 2024
·
Yixin Liu
DBLP profile ↗
ORCID search ↗
Calibrating Long-form Generations From Large Language Models
EMNLP 2024
·
Yukun Huang
DBLP profile ↗
ORCID search ↗
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Financial Documents
ACL 2024
·
Yilun Zhao
DBLP profile ↗
ORCID search ↗
FOLIO: Natural Language Reasoning with First-Order Logic
EMNLP 2024
·
Simeng Han
DBLP profile ↗
ORCID search ↗
FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents
EMNLP 2024
·
Yilun Zhao
DBLP profile ↗
ORCID search ↗
Investigating Data Contamination in Modern Benchmarks for Large Language Models
NAACL 2024
·
Chunyuan Deng
DBLP profile ↗
ORCID search ↗
KnowledgeFMath: A Knowledge-Intensive Math Reasoning Dataset in Finance Domains
ACL 2024
·
Yilun Zhao
DBLP profile ↗
ORCID search ↗
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
EMNLP 2024
·
Chuhan Li
DBLP profile ↗
ORCID search ↗
MIMIR: A Customizable Agent Tuning Platform for Enhanced Scientific Applications
EMNLP 2024
·
Xiangru Tang
DBLP profile ↗
ORCID search ↗
MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning
ACL 2024
·
Xiangru Tang
DBLP profile ↗
ORCID search ↗
NExT: Teaching Large Language Models to Reason about Code Execution
ICML 2024
·
Ansong Ni
DBLP profile ↗
ORCID search ↗
OLMo: Accelerating the Science of Language Models
ACL 2024
·
Dirk Groeneveld
DBLP profile ↗
ORCID search ↗
OMG-QA: Building Open-Domain Multi-Modal Generative Question Answering Systems
EMNLP 2024
·
Linyong Nan
DBLP profile ↗
ORCID search ↗
Observable Propagation: Uncovering Feature Vectors in Transformers
ICML 2024
·
Jacob Dunefsky
DBLP profile ↗
ORCID search ↗
On Evaluating the Integration of Reasoning and Action in LLM Agents with Database Question Answering
NAACL 2024
·
Linyong Nan
DBLP profile ↗
ORCID search ↗
On Learning to Summarize with Large Language Models as References
NAACL 2024
·
Yixin Liu
DBLP profile ↗
ORCID search ↗
OpenT2T: An Open-Source Toolkit for Table-to-Text Generation
EMNLP 2024
·
Haowei Zhang
DBLP profile ↗
ORCID search ↗
P-FOLIO: Evaluating and Improving Logical Reasoning with Abundant Human-Written Reasoning Chains
EMNLP 2024
·
Simeng Han
DBLP profile ↗
ORCID search ↗
Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
ACL 2024
·
Martin Riddell
DBLP profile ↗
ORCID search ↗
Rethinking Efficient Multilingual Text Summarization Meta-Evaluation
ACL 2024
·
Rilyn Han
DBLP profile ↗
ORCID search ↗
SciDQA: A Deep Reading Comprehension Dataset over Scientific Papers
EMNLP 2024
·
Shruti Singh
DBLP profile ↗
ORCID search ↗
Struc-Bench: Are Large Language Models Good at Generating Complex Structured Tabular Data?
NAACL 2024
·
Xiangru Tang
DBLP profile ↗
ORCID search ↗
TAIL: A Toolkit for Automatic and Realistic Long-Context Large Language Model Evaluation
EMNLP 2024
·
Gefei Gu
DBLP profile ↗
ORCID search ↗
TaPERA: Enhancing Faithfulness and Interpretability in Long-Form Table QA by Content Planning and Execution-based Reasoning
ACL 2024
·
Yilun Zhao
DBLP profile ↗
ORCID search ↗
Unveiling the Spectrum of Data Contamination in Language Model: A Survey from Detection to Remediation
ACL 2024
·
Chunyuan Deng
DBLP profile ↗
ORCID search ↗