PPaperPicks

Xuandong Zhao

27 papers at tracked venues · 22 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption
  2. A Practical Examination of AI-Generated Text Detectors for Large Language Models
  3. A Technical Report on "Erasing the Invisible": The 2024 NeurIPS Competition on Stress Testing Image Watermarks
  4. AGENTVIGIL: Automatic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
  5. An Undetectable Watermark for Generative Image Models
  6. CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
  7. DIS-CO: Discovering Copyrighted Content in VLMs Training Data
  8. Efficiently Identifying Watermarked Segments in Mixed-Source Texts
    ACL 2025 · Xuandong Zhao
  9. Improving LLM Safety Alignment with Dual-Objective Optimization
    ICML 2025 · Xuandong Zhao
  10. MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
  11. Multimodal Situational Safety
  12. OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models
  13. Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
    ICLR 2025 · Xuandong Zhao
  14. SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning
  15. Scalable Best-of-N Selection for Large Language Models via Self-Certainty
  16. Weak-to-Strong Jailbreaking on Large Language Models
    ICML 2025 · Xuandong Zhao
  17. A Survey on Detection of LLMs-Generated Content
  18. Bileve: Securing Text Provenance in Large Language Models Against Spoofing with Bi-level Signature
  19. Chatbot and Fatigued Driver: Exploring the Use of LLM-Based Voice Assistants for Driving Fatigue
  20. DE-COP: Detecting Copyrighted Content in Language Models Training Data
  21. GumbelSoft: Diversified Language Model Watermarking via the GumbelMax-trick
  22. Invisible Image Watermarks Are Provably Removable Using Generative AI
    NeurIPS 2024 · Xuandong Zhao
  23. MarkLLM: An Open-Source Toolkit for LLM Watermarking
  24. Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
  25. Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
  26. Provable Robust Watermarking for AI-Generated Text
    ICLR 2024 · Xuandong Zhao
  27. Watermarking for Large Language Models
    ACL 2024 · Xuandong Zhao