PPaperPicks

Percy Liang

Stanford University, Computer Science Department

35 papers at tracked venues · 32 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Beyond a Single Extractor: Re-thinking HTML-to-Text Extraction for LLM Pre-training
  2. AIR-BENCH 2024: A Safety Benchmark based on Regulation and Policies Specified Risk Categories
  3. Auditing Prompt Caching in Language Model APIs
  4. Audits Under Resource, Data, and Access Constraints: Scaling Laws For Less Discriminatory Alternatives
  5. AutoBencher: Towards Declarative Benchmark Construction
  6. BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments
  7. Blackbox Model Provenance via Palimpsestic Membership Inference
  8. BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
  9. Eliciting Language Model Behaviors with Investigator Agents
  10. Establishing Best Practices in Building Rigorous Agentic Benchmarks
  11. Independence Tests for Language Models
  12. Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
  13. LawInstruct: A Resource for Studying Language Model Adaptation to the Legal Domain
  14. MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering
  15. Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
  16. Model Equality Testing: Which Model is this API Serving?
  17. On the Entropy Calibration of Language Models
  18. Position: In-House Evaluation Is Not Enough. Towards Robust Third-Party Evaluation and Flaw Disclosure for General-Purpose AI
  19. Position: Language model developers should report train-test overlap
  20. Reliable and Efficient Amortized Model-based Evaluation
  21. Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape View
  22. s1: Simple test-time scaling
  23. Benchmarking and Improving Generator-Validator Consistency of Language Models
  24. Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
  25. Image2Struct: Benchmarking Structure Extraction for Vision-Language Models
    NeurIPS 2024 ·
    Josselin Somerville Roberts
  26. Large Language Models as Analogical Reasoners
  27. MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
  28. MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical Records
  29. On the Learnability of Watermarks for Language Models
  30. Position: A Safe Harbor for AI Evaluation and Red Teaming
  31. Position: On the Societal Impact of Open Foundation Models
  32. Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
  33. RedPajama: an Open Dataset for Training Large Language Models
  34. Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
  35. VHELM: A Holistic Evaluation of Vision Language Models