PPaperPicks

Nigel Collier

University of Cambridge, UK

30 papers at tracked venues · 19 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Confidence Estimation for LLMs in Multi-turn Interactions
  2. Demystifying Multi-Agent Debate: The Role of Confidence and Diversity
  3. Failure Modes in Multi-Hop QA: The Weakest Link Effect and the Recognition Bottleneck
  4. LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generation
  5. Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
  6. Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning
  7. Value of Information: A Framework for Human-Agent Communication
  8. 500xCompressor: Generalized Prompt Compression for Large Language Models
  9. Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language Models
  10. All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoning
  11. Conformity in Large Language Models
  12. LoGU: Long-form Generation with Uncertainty Expressions
  13. PT-MoE: An Efficient Finetuning Framework for Integrating Mixture-of-Experts into Prompt Tuning
  14. Prompt Compression for Large Language Models: A Survey
  15. ReasonGraph: Visualization of Reasoning Methods and Extended Inference Paths
  16. Time to Revisit Exact Match
  17. UNCLE: Benchmarking Uncertainty Expressions in Long-Form Generation
  18. When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning
  19. iNews: A Multimodal Dataset for Modeling Personalized Affective Responses to News
  20. An Individualized News Affective Response Dataset
  21. BAND: Biomedical Alert News Dataset
  22. Can LLM be a Personalized Judge?
  23. Can We Instruct LLMs to Compensate for Position Bias?
  24. Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments
  25. LUQ: Long-text Uncertainty Quantification for LLMs
  26. PiVe: Prompting with Iterative Verification Improving Graph-based Generative Capability of LLMs
  27. Quantifying the Persona Effect in LLM Simulations
  28. TOAD: Task-Oriented Automatic Dialogs with Diverse Response Styles
  29. TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
  30. Unlocking Structure Measuring: Introducing PDD, an Automatic Metric for Positional Discourse Coherence