PPaperPicks

Tsung-Yi Ho

Chinese University of Hong Kong, Hong Kong, SAR, China

20 papers at tracked venues · 19 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. GRE Score: Generative Risk Evaluation for Large Language Models
  2. Hey, That's My Data! Token-Only Dataset Inference in Large Language Models
  3. KCLNet: Electrically Equivalence-Oriented Graph Representation Learning for Analog Circuits
  4. Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
  5. CARE: Decoding-Time Safety Alignment via Rollback and Introspection Intervention
  6. CoP: Agentic Red-teaming for Large Language Models using Composition of Principles
  7. Defensive Prompt Patch: A Robust and Generalizable Defense of Large Language Models against Jailbreak Attacks
  8. PermLLM: Learnable Channel Permutation for N: M Sparse Large Language Models
  9. Retention Score: Quantifying Jailbreak Risks for Vision Language Models
  10. Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models
  11. Achieving Fairness Through Channel Pruning for Dermatological Disease Diagnosis
  12. AutoVP: An Automated Visual Prompting Framework and Benchmark
  13. Be Your Own Neighborhood: Detecting Adversarial Examples by the Neighborhood Relations Built on Self-Supervised Learning
  14. Elijah: Eliminating Backdoors Injected in Diffusion Models via Distribution Shift
  15. GREAT Score: Global Robustness Evaluation of Adversarial Perturbation using Generative Models
  16. Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
  17. MMA-Diffusion: MultiModal Attack on Diffusion Models
  18. NeuralFuse: Learning to Recover the Accuracy of Access-Limited Neural Network Inference in Low-Voltage Regimes
  19. Rethinking Backdoor Attacks on Dataset Distillation: A Kernel Method Perspective
  20. The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Language Models