PPaperPicks

Dawn Song

University of California, Berkeley, Computer Science Division

49 papers at tracked venues · 45 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Can Editing LLMs Inject Harm?
  2. Fico: Evaluating Vision-Language Models under Visual Fidelity and Compression at Scale
  3. Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption
  4. Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities
  5. dLLM: Simple Diffusion Language Modeling
  6. A Sustainable AI Economy Needs Data Deals That Work for Generators
  7. AGENTVIGIL: Automatic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
  8. AIR-BENCH 2024: A Safety Benchmark based on Regulation and Policies Specified Risk Categories
  9. An Undetectable Watermark for Generative Image Models
  10. BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
  11. COSMIC: Generalized Refusal Direction Identification in LLM Activations
  12. Capturing the Temporal Dependence of Training Data Influence
  13. CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
  14. Data Shapley in One Training Run
  15. GuardAgent: Safeguard LLM Agents via Knowledge-Enabled Reasoning
  16. HADES: Range-Filtered Private Aggregation on Public Data
  17. Improving LLM Safety Alignment with Dual-Objective Optimization
  18. Multimodal Situational Safety
  19. OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization
  20. OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models
  21. Position: Formal Mathematical Reasoning - A New Frontier in AI
  22. Position: In-House Evaluation Is Not Enough. Towards Robust Third-Party Evaluation and Flaw Disclosure for General-Purpose AI
  23. Position: Political Neutrality in AI Is Impossible - But Here Is How to Approximate It
  24. SECODEPLT: A Unified Benchmark for Evaluating the Security Risks and Capabilities of Code GenAI
  25. SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning
  26. Scalable Best-of-N Selection for Large Language Models via Self-Certainty
  27. Tamper-Resistant Safeguards for Open-Weight LLMs
  28. VMDT: Decoding the Trustworthiness of Video Foundation Models
  29. Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
  30. "I Can't Believe It's Not Custodial!": Usable Trustless Decentralized Key Management
  31. Agent Instructs Large Language Models to be General Zero-Shot Reasoners
  32. AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
  33. BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models
  34. Boosting Alignment for Post-Unlearning Text-to-Image Generative Models
  35. C-RAG: Certified Generation Risks for Retrieval-Augmented Language Models
  36. Data Free Backdoor Attacks
  37. Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
  38. Effective and Efficient Federated Tree Learning on Hybrid Data
  39. GRATH: Gradual Self-Truthifying for Large Language Models
  40. GREATS: Online Selection of High-Quality Data for LLM Training in Every Iteration
  41. Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters
  42. LLM-PBE: Assessing Data Privacy in Large Language Models
  43. Position: Evolving AI Collectives Enhance Human Diversity and Enable Self-Regulation
  44. Position: On the Societal Impact of Open Foundation Models
  45. Re-Tuning: Overcoming the Compositionality Limits of Large Language Models with Recursive Tuning
  46. RedCode: Risky Code Execution and Generation Benchmark for Code Agents
  47. RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content
  48. SHINE: Shielding Backdoors in Deep Reinforcement Learning
  49. The False Promise of Imitating Proprietary Language Models