P
PaperPicks
Conferences
Caiming Xiong
74 papers at tracked venues · 50 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID search ↗
Venues
ICLR
×14
EMNLP
×12
ACL
×11
NeurIPS
×10
NAACL
×8
ICML
×6
CVPR
×4
ECCV
×3
ICCV
×3
AAAI
×1
ICRA
×1
UIST
×1
Frequent coauthors
Yiheng Xu
DBLP profile ↗
ORCID search ↗
×3
Kung-Hsiang Huang
DBLP profile ↗
ORCID search ↗
×3
Artemis Panagopoulou
DBLP profile ↗
ORCID search ↗
×3
Jianguo Zhang
DBLP profile ↗
ORCID search ↗
×2
Le Xue
DBLP profile ↗
ORCID search ↗
×2
Prafulla Kumar Choubey
DBLP profile ↗
ORCID search ↗
×2
Zhiwei Liu
DBLP profile ↗
ORCID search ↗
×2
Xiangyu Peng
DBLP profile ↗
ORCID search ↗
×2
Tianbao Xie
DBLP profile ↗
ORCID search ↗
×2
Philippe Laban
DBLP profile ↗
ORCID search ↗
×2
Simeng Han
DBLP profile ↗
ORCID search ↗
×2
Zhiyuan Hu
DBLP profile ↗
ORCID search ↗
×1
Papers
Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models
ACL 2026
·
Zhiyuan Hu
DBLP profile ↗
ORCID search ↗
From Passive Metric to Active Signal: The Evolving Role of Uncertainty Quantification in Large Language Models
ACL 2026
·
Jiaxin Zhang
DBLP profile ↗
ORCID search ↗
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
ACL 2026
·
Shrey Pandit
DBLP profile ↗
ORCID search ↗
J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization
ACL 2026
·
Austin Xu
DBLP profile ↗
ORCID search ↗
APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
NeurIPS 2025
·
Akshara Prabhakar
DBLP profile ↗
ORCID search ↗
ActionStudio: A Lightweight Framework for Data and Training of Large Action Models
EMNLP 2025
·
Jianguo Zhang
DBLP profile ↗
ORCID search ↗
AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials
ICLR 2025
·
Yiheng Xu
DBLP profile ↗
ORCID search ↗
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
ICML 2025
·
Yiheng Xu
DBLP profile ↗
ORCID search ↗
Automatic Curriculum Expert Iteration for Reliable LLM Reasoning
ICLR 2025
·
Zirui Zhao
DBLP profile ↗
ORCID search ↗
BLIP-3: A Family of Open Large Multimodal Models
ICCV 2025
·
Le Xue
DBLP profile ↗
ORCID search ↗
Benchmarking Deep Search over Heterogeneous Enterprise Data
EMNLP 2025
·
Prafulla Kumar Choubey
DBLP profile ↗
ORCID search ↗
Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning
NeurIPS 2025
·
Jiayu Wang
DBLP profile ↗
ORCID search ↗
BingoGuard: LLM Content Moderation Tools with Risk Levels
ICLR 2025
·
Fan Yin
DBLP profile ↗
ORCID search ↗
CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments
NAACL 2025
·
Kung-Hsiang Huang
DBLP profile ↗
ORCID search ↗
CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
NAACL 2025
·
Jierui Li
DBLP profile ↗
ORCID search ↗
Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D
EMNLP 2025
·
Artemis Panagopoulou
DBLP profile ↗
ORCID search ↗
Demystifying Domain-adaptive Post-training for Financial LLMs
EMNLP 2025
·
Zixuan Ke
DBLP profile ↗
ORCID search ↗
Direct Judgement Preference Optimization
EMNLP 2025
·
Peifeng Wang
DBLP profile ↗
ORCID search ↗
Diversity Empowers Intelligence: Integrating Expertise of Software Engineering Agents
ICLR 2025
·
Kexun Zhang
DBLP profile ↗
ORCID search ↗
Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage
NAACL 2025
·
Kaige Xie
DBLP profile ↗
ORCID search ↗
DyMU: Dynamic Merging and Virtual Unmerging for Efficient Variable-Length VLMs
NeurIPS 2025
·
Zhenhailong Wang
DBLP profile ↗
ORCID search ↗
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
ICML 2025
·
Yilun Zhou
DBLP profile ↗
ORCID search ↗
FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
ICLR 2025
·
Yifei Ming
DBLP profile ↗
ORCID search ↗
GReaTer: Gradients Over Reasoning Makes Smaller Language Models Strong Prompt Optimizers
ICLR 2025
·
Sarkar Snigdha Sarathi Das
DBLP profile ↗
ORCID search ↗
Generative Frame Sampler for Long Video Understanding
ACL 2025
·
Linli Yao
DBLP profile ↗
ORCID search ↗
LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback
ACL 2025
·
Thai Quoc Hoang
DBLP profile ↗
ORCID search ↗
LATTE: Learning to Think with Vision Specialists
EMNLP 2025
·
Zixian Ma
DBLP profile ↗
ORCID search ↗
MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models
EMNLP 2025
·
Zhiwei Liu
DBLP profile ↗
ORCID search ↗
Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts
ICML 2025
·
Xu Liu
DBLP profile ↗
ORCID search ↗
PersonaBench: Evaluating AI Models on Understanding Personal Information through Accessing (Synthetic) Private User Data
ACL 2025
·
Juntao Tan
DBLP profile ↗
ORCID search ↗
ReGenesis: LLMs can Grow into Reasoning Generalists via Self-Improvement
ICLR 2025
·
Xiangyu Peng
DBLP profile ↗
ORCID search ↗
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
ICML 2025
·
Baohao Liao
DBLP profile ↗
ORCID search ↗
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
NeurIPS 2025
·
Tianbao Xie
DBLP profile ↗
ORCID search ↗
SiReRAG: Indexing Similar and Related Information for Multihop Reasoning
ICLR 2025
·
Nan Zhang
DBLP profile ↗
ORCID search ↗
SlackAgents: Scalable Collaboration of AI Agents in Workspaces
EMNLP 2025
·
Zhiwei Liu
DBLP profile ↗
ORCID search ↗
Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
ICLR 2025
·
Fangyu Lei
DBLP profile ↗
ORCID search ↗
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
ICCV 2025
·
Honglu Zhou
DBLP profile ↗
ORCID search ↗
Text2Data: Low-Resource Data Generation with Textual Control
AAAI 2025
·
Shiyu Wang
DBLP profile ↗
ORCID search ↗
ThinK: Thinner Key Cache by Query-Driven Pruning
ICLR 2025
·
Yuhui Xu
DBLP profile ↗
ORCID search ↗
Trust but Verify: Programmatic VLM Evaluation in the Wild
ICCV 2025
·
Viraj Prabhu
DBLP profile ↗
ORCID search ↗
Turning Conversations into Workflows: A Framework to Extract and Evaluate Dialog Workflows for Service AI Agents
ACL 2025
·
Prafulla Kumar Choubey
DBLP profile ↗
ORCID search ↗
Unanswerability Evaluation for Retrieval Augmented Generation
ACL 2025
·
Xiangyu Peng
DBLP profile ↗
ORCID search ↗
ViUniT: Visual Unit Tests for More Robust Visual Programming
CVPR 2025
·
Artemis Panagopoulou
DBLP profile ↗
ORCID search ↗
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
ACL 2025
·
Kung-Hsiang Huang
DBLP profile ↗
ORCID search ↗
xLAM: A Family of Large Action Models to Empower AI Agent Systems
NAACL 2025
·
Jianguo Zhang
DBLP profile ↗
ORCID search ↗
APIGen: Automated PIpeline for Generating Verifiable and Diverse Function-Calling Datasets
NeurIPS 2024
·
Zuxin Liu
DBLP profile ↗
ORCID search ↗
ARM: Alignment with Residual Energy-Based Model
NAACL 2024
·
Bo Pang
DBLP profile ↗
ORCID search ↗
Beyond the Chat: Executable and Verifiable Text-Editing with LLMs
UIST 2024
·
Philippe Laban
DBLP profile ↗
ORCID search ↗
Consent in Crisis: The Rapid Decline of the AI Data Commons
NeurIPS 2024
·
Shayne Longpre
DBLP profile ↗
ORCID search ↗
Diffusion Model Alignment Using Direct Preference Optimization
CVPR 2024
·
Bram Wallace
DBLP profile ↗
ORCID search ↗
Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles
NAACL 2024
·
Kung-Hsiang Huang
DBLP profile ↗
ORCID search ↗
FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability
ACL 2024
·
Congying Xia
DBLP profile ↗
ORCID search ↗
FOLIO: Natural Language Reasoning with First-Order Logic
EMNLP 2024
·
Simeng Han
DBLP profile ↗
ORCID search ↗
Fair Abstractive Summarization of Diverse Perspectives
NAACL 2024
·
Yusen Zhang
DBLP profile ↗
ORCID search ↗
HIVE: Harnessing Human Feedback for Instructional Visual Editing
CVPR 2024
·
Shu Zhang
DBLP profile ↗
ORCID search ↗
Hierarchical Point Attention for Indoor 3D Object Detection
ICRA 2024
·
Manli Shu
DBLP profile ↗
ORCID search ↗
How Do Transformers Learn In-Context Beyond Simple Functions? A Case Study on Learning with Representations
ICLR 2024
·
Tianyu Guo
DBLP profile ↗
ORCID search ↗
INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness
NeurIPS 2024
·
Hung Le
DBLP profile ↗
ORCID search ↗
LayoutDETR: Detection Transformer Is a Good Multimodal Layout Designer
ECCV 2024
·
Ning Yu
DBLP profile ↗
ORCID search ↗
Lemur: Harmonizing Natural Language and Code for Language Agents
ICLR 2024
·
Yiheng Xu
DBLP profile ↗
ORCID search ↗
MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
NeurIPS 2024
·
Anas Awadalla
DBLP profile ↗
ORCID search ↗
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
NeurIPS 2024
·
Tianbao Xie
DBLP profile ↗
ORCID search ↗
P-FOLIO: Evaluating and Improving Logical Reasoning with Abundant Human-Written Reasoning Chains
EMNLP 2024
·
Simeng Han
DBLP profile ↗
ORCID search ↗
Position: TrustLLM: Trustworthiness in Large Language Models
ICML 2024
·
Yue Huang
DBLP profile ↗
ORCID search ↗
Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization
ICLR 2024
·
Weiran Yao
DBLP profile ↗
ORCID search ↗
Sample-Efficient Learning of POMDPs with Multiple Observations In Hindsight
ICLR 2024
·
Jiacheng Guo
DBLP profile ↗
ORCID search ↗
Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?
NeurIPS 2024
·
Ruisheng Cao
DBLP profile ↗
ORCID search ↗
Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems
EMNLP 2024
·
Philippe Laban
DBLP profile ↗
ORCID search ↗
ULIP-2: Towards Scalable Multimodal Pre-Training for 3D Understanding
CVPR 2024
·
Le Xue
DBLP profile ↗
ORCID search ↗
Unified Training of Universal Time Series Forecasting Transformers
ICML 2024
·
Gerald Woo
DBLP profile ↗
ORCID search ↗
Unlocking Anticipatory Text Generation: A Constrained Approach for Large Language Models Decoding
EMNLP 2024
·
Lifu Tu
DBLP profile ↗
ORCID search ↗
What Are We Measuring When We Evaluate Large Vision-Language Models? An Analysis of Latent Factors and Biases
NAACL 2024
·
Anthony Meng Huat Tiong
DBLP profile ↗
ORCID search ↗
X-InstructBLIP: A Framework for Aligning Image, 3D, Audio, Video to LLMs and its Emergent Cross-Modal Reasoning
ECCV 2024
·
Artemis Panagopoulou
DBLP profile ↗
ORCID search ↗
xGen-VideoSyn-1: High-Fidelity Text-to-Video Synthesis with Compressed Representations
ECCV 2024
·
Can Qin
DBLP profile ↗
ORCID search ↗