PPaperPicks

Han Qiu

Tsinghua University, Beijing, China

25 papers at tracked venues · 20 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
  2. LeakDojo: Decoding the Leakage Threats of RAG Systems
  3. New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs
  4. Revisiting the Reliability of Language Models in Instruction-Following
  5. TOXIFRENCH: Benchmarking and Enhancing Language Models via CoT Fine-Tuning for French Toxicity Detection
  6. The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning
  7. When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity
  8. "I've Decided to Leak": Probing Internals Behind Prompt Leakage Intents
  9. A Benchmark for Semantic Sensitive Information in LLMs Outputs
  10. An Engorgio Prompt Makes Large Language Model Babble on
  11. Cowpox: Towards the Immunity of VLM-based Multi-Agent Systems
  12. Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
  13. Mask Image Watermarking
  14. ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs: ShieldVLM
  15. Speculating LLMs' Chinese Training Data Pollution from Their Tokens
  16. Understanding the Dark Side of LLMs' Intrinsic Self-Correction
  17. VISO: Accelerating In-Orbit Object Detection with Language-Guided Mask Learning and Sparse Inference
  18. VideoShield: Regulating Diffusion-based Video Generation Models via Watermarking
  19. When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
  20. COSMIC: Compress Satellite Image Efficiently via Diffusion Compensation
  21. Course-Correction: Safety Alignment Using Synthetic Preferences
  22. Purifying Quantization-conditioned Backdoors via Layer-wise Activation Correction with Distribution Approximation
  23. The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
  24. Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias
  25. You Only Query Once: An Efficient Label-Only Membership Inference Attack