PPaperPicks

Daniel Khashabi

32 papers at tracked venues · 17 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Principled Context Engineering for RAG: Statistical Guarantees via Conformal Prediction
  2. Query Decomposition for RAG: Balancing Exploration-Exploitation
  3. arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation
  4. Benchmarking Language Model Creativity: A Case Study on Code Generation
  5. CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?
  6. Certified Mitigation of Worst-Case LLM Copyright Infringement
  7. Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
  8. Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
  9. Core: Robust Factual Precision with Informative Sub-Claim Identification
  10. Evaluating the Evaluators: Are readability metrics good measures of readability?
  11. FEEDBACK FRICTION: LLMs Struggle to Fully Incorporate External Feedback
  12. GenEx: Generating an Explorable World
  13. ICL CIPHERS: Quantifying "Learning" in In-Context Learning via Substitution Ciphers
  14. Jailbreak Distillation: Renewable Safety Benchmarking
  15. RATIONALYST: Pre-training Process-Supervision for Improving Reasoning
  16. SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
  17. SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning
  18. TurkingBench: A Challenge Benchmark for Web Agents
  19. Upsample or Upweight? Balanced Training on Heavily Imbalanced Datasets
  20. Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
  21. WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
  22. AnaloBench: Benchmarking the Identification of Abstract and Long-context Analogies
  23. DiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech Translation
  24. Efficient Large Multi-modal Models via Visual Context Compression
  25. Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation Models
  26. Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell
  27. Position: Do pretrained Transformers Learn In-Context by Gradient Descent?
  28. RORA: Robust Free-Text Rationale Evaluation
  29. SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation
  30. The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts
  31. The Trickle-down Impact of Reward Inconsistency on RLHF
  32. k-SemStamp: A Clustering-Based Semantic Watermark for Detection of Machine-Generated Text