PPaperPicks

Jimmy Lin

University of Waterloo, David R. Cheriton School of Computer Science

65 papers at tracked venues · 38 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Assembling Your Personal AI Council in Yupp to Provide Multiple Perspectives
    WSDM 2026 · Jimmy Lin
  2. Automating Generation of Long-Form Queries
  3. BrowseComp-Plus: A Fair and Disentangled Evaluation Benchmark for Deep Search Agents
  4. Contrastive Learning Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data
  5. Do We Still Need Text Features for Video Retrieval in the Era of Vision-Language Models?
  6. LACONIC: Dense-Level Effectiveness for Scalable Sparse Retrieval via a Two-Phase Training Curriculum
  7. Lighting the Way for BRIGHT: Reproducible Baselines with Anserini, Pyserini, and RankLLM
  8. MCP Servers for Pyserini and RankLLM: Enabling Agentic Retrieval-Augmented Generation
  9. NanoKnow: How to Know What Your Language Model Knows
  10. Rank-R1: Enhancing Reasoning in LLM-based Document Rerankers via Reinforcement Learning
  11. Rerank Before You Reason: Analyzing Reranking Tradeoffs through Effective Token Cost in Deep Search Agents
  12. Revisiting BM25 Feedback Models using HyDE
  13. Understanding Multi-Structured Documents via LLMs'
  14. rosaOS: Agentic Operating System for Embodied LLMs
  15. Accelerating Listwise Reranking: Reproducing and Enhancing FIRST
  16. AfroBench: How Good are Large Language Models on African Languages?
  17. Assessing Support for the TREC 2024 RAG Track: A Large-Scale Comparative Study of LLM and Human Evaluations
  18. Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
  19. Beyond Quacking: Deep Integration of Language Models and RAG into DuckDB
  20. CURE: A dataset for Clinical Understanding & Retrieval Evaluation
  21. Can't Hide Behind the API: Stealing Black-Box Commercial Embedding Models
  22. DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense Retrievers
  23. Document Screenshot Retrievers are Vulnerable to Pixel Poisoning Attacks
  24. FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
  25. Gosling Grows Up: Retrieval with Learned Dense and Sparse Representations Using Anserini
    SIGIR 2025 · Jimmy Lin
  26. Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
  27. LiT and Lean: Distilling Listwise Rerankers Into Encoder-Decoder Models
  28. MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems
  29. Mm-Embed: Universal Multimodal Retrieval with Multimodal LLMS
  30. Operational Advice for Dense and Sparse Retrievers: HNSW, Flat, or Inverted Indexes?
    ACL 2025 · Jimmy Lin
  31. Patience in Proximity: A Simple Early Termination Strategy for HNSW Graph Traversal in Approximate k-Nearest Neighbor Search
  32. QuackIR: Retrieval in DuckDB and Other Relational Database Management Systems
  33. Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track
  34. Rank-Without-GPT: Building GPT-Independent Listwise Rerankers on Open-Source Large Language Models
  35. RankLLM: A Python Package for Reranking with LLMs
  36. Study on LLMs for Promptagator-Style Dense Retriever Training
  37. Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality
  38. The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models
  39. The Impact of Incidental Multilingual Text on Cross-Lingual Transfer in Monolingual Retrieval
  40. Tomato, Tomahto, Tomate: Do Multilingual Language Models Understand Based on Subword-Level Semantic Concepts?
  41. UniRAG: Universal Retrieval Augmentation for Large Vision Language Models
  42. VISA: Retrieval Augmented Generation with Visual Source Attribution
  43. Zero-Shot ATC Coding with Large Language Models for Clinical Assessments
  44. "Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation
  45. CELI: Simple yet Effective Approach to Enhance Out-of-Domain Generalization of Cross-Encoders
  46. CIRAL: A Test Collection for CLIR Evaluations in African Languages
  47. Can Query Expansion Improve Generalization of Strong Cross-Encoder Rankers?
  48. ConvKGYarn: Spinning Configurable and Scalable Conversational Knowledge Graph QA Datasets with Large Language Models
  49. EWEK-QA : Enhanced Web and Efficient Knowledge Graph Retrieval for Citation-based Question Answering Systems
  50. FLAME : Factuality-Aware Alignment for Large Language Models
  51. Fine-Tuning LLaMA for Multi-Stage Text Retrieval
  52. Found in the Middle: Permutation Self-Consistency Improves Listwise Ranking in Large Language Models
  53. Jointly Modeling Spatio-Temporal Features of Tactile Signals for Action Classification
    AAAI 2024 · Jimmy Lin
  54. Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval
  55. Nearest Neighbor Speculative Decoding for LLM Generation and Attribution
  56. On Backbones and Training Regimes for Dense Retrieval in African Languages
  57. PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval
  58. Reflections on the Coding Ability of LLMs for Analyzing Market Research Surveys
  59. Resources for Brewing BEIR: Reproducible Reference Models and Statistical Analyses
  60. Systematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR
  61. Towards Robust QA Evaluation via Open LLMs
  62. Unifying Multimodal Retrieval via Document Screenshot Embedding
  63. Vector Search with OpenAI Embeddings: Lucene Is All You Need
  64. Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation
  65. Zero-Shot Cross-Lingual Reranking with Large Language Models for Low-Resource Languages