PPaperPicks

Zhexin Zhang

13 papers at tracked venues · 12 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
    ACL 2026 · Zhexin Zhang
  2. LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
  3. New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs
  4. When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity
  5. Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
  6. JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering
  7. Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
  8. LongSafety: Evaluating Long-Context Safety of Large Language Models
  9. ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs: ShieldVLM
  10. A Design of Interface for Visual-Impaired People to Access Visual Information from Images Featuring Large Language Models and Visual Language Models
    CHI 2024 · Zhexin Zhang
  11. Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
    ACL 2024 · Zhexin Zhang
  12. SafetyBench: Evaluating the Safety of Large Language Models
    ACL 2024 · Zhexin Zhang
  13. ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors
    EMNLP 2024 · Zhexin Zhang