PPaperPicks

Luca Soldaini

Amazon Alexa, CA, USA

24 papers at tracked venues · 18 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. The olmOCR Project: Building Fully Open OCR using VLMs
  2. WSDM CUP 2026: Multilingual Retrieval
  3. DataDecide: How to Predict Best Pretraining Data with Small Experiments
  4. DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images
  5. FlexOLMo: Open Language Models for Flexible Data Use
  6. FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions
  7. Language models scale reliably with over-training and on downstream tasks
  8. Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
  9. OLMoE: Open Mixture-of-Experts Language Models
  10. OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens
  11. Organize the Web: Constructing Domains Enhances Pre-Training Data Curation
  12. RouterRetriever: Routing over a Mixture of Expert Embedding Models
  13. SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific Literature
  14. The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
  15. mFollowIR: A Multilingual Benchmark for Instruction Following in Retrieval
  16. AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters
  17. DataComp-LM: In search of the next generation of training sets for language models
  18. Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
    ACL 2024 · Luca Soldaini
  19. KIWI: A Dataset of Knowledge-Intensive Writing Instructions for Answering Research Questions
  20. MathFish: Evaluating Language Model Math Reasoning via Grounding in Educational Curricula
  21. OLMo: Accelerating the Science of Language Models
  22. On the Evaluation of Machine-Generated Reports
  23. Paloma: A Benchmark for Evaluating Language Model Fit
  24. What's In My Big Data?