PPaperPicks

Jun Sun

Singapore Management University, Singapore

22 papers at tracked venues · 16 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Be Responsible in Your Answers! Monitoring Out-of-Domain Behaviors in Domain-Specific LLMs
  2. LLMQuA: Practical Backdoor Injection on Large Language Model Quantization
  3. SafetyReminder: Reviving Delayed Safety Awareness of Vision-Language Models to Defend Against Jailbreak Attacks
  4. Towards Provably Unlearnable Examples via Bayes Error Optimization
  5. Train in Vain: Functionality-Preserving Poisoning to Prevent Unauthorized Use of Code Datasets
  6. Unleashing the Unseen: Harnessing Benign Datasets for Jailbreaking Large Language Models
  7. BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
  8. CROW: Eliminating Backdoors from Large Language Models via Internal Consistency Regularization
  9. Causal Contrastive Learning with Data Augmentations for Imitation-Based Planning
  10. Democratic Training Against Universal Adversarial Perturbations
  11. Do Influence Functions Work on Large Language Models?
  12. Evaluating and Mitigating Linguistic Discrimination in Large Language Models: Perspectives on Safety Equity and Knowledge Equity
  13. LLMScan: Causal Scan for LLM Misbehavior Detection
  14. Position: Trustworthy AI Agents Require the Integration of Large Language Models and Formal Methods
  15. Quantitative Runtime Monitoring of Ethereum Transaction Attacks
  16. Training Verification-Friendly Neural Networks via Neuron Behavior Consistency
  17. Unleashing the Power of Visual Foundation Models for Generalizable Semantic Segmentation
  18. Zero-Shot Defense Against Toxic Images via Inherent Multimodal Alignment in LVLMs
  19. ALI-Agent: Assessing LLMs' Alignment with Human Values via Agent-based Evaluation
  20. Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models
  21. Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing
  22. Semantic Conformance Testing of Relational DBMS