PPaperPicks

Bill Yuchen Lin

University of Washington, Seattle, WA, USA

27 papers at tracked venues · 18 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Temporal Sampling for Forgotten Reasoning in LLMs
  2. ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
  3. CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-Teaming
  4. Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
  5. L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects
  6. Latent Action Pretraining from Videos
  7. Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
  8. RewardBench: Evaluating Reward Models for Language Modeling
  9. SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
  10. SimulBench: Evaluating Language Models with Creative Simulation Tasks
  11. Small Models Struggle to Learn from Strong Reasoners
  12. Stronger Models are Not Always Stronger Teachers for Instruction Tuning
  13. The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
  14. The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
  15. VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
  16. WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
    ICLR 2025 · Bill Yuchen Lin
  17. ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
    ICML 2025 · Bill Yuchen Lin
  18. Agent Lumos: Unified and Modular Training for Open-Source Language Agents
  19. OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
  20. Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
  21. SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
  22. Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
  23. The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning
    ICLR 2024 · Bill Yuchen Lin
  24. Trial and Error: Exploration-Based Trajectory Optimization of LLM Agents
  25. VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation
  26. WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
  27. WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences