PPaperPicks

Oleg Rogov

8 papers at tracked venues · 5 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs
  2. Feature Drift: How Fine-Tuning Repurposes Representations in LLMs
  3. I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders
  4. POLAR: A Benchmark for Multilingual, Multicultural, and Multi-Event Online Polarization
  5. CLEAR: Character Unlearning in Textual and Visual Modalities
  6. Certification of Speaker Recognition Models to Additive Perturbations
  7. Novel Loss-Enhanced Universal Adversarial Patches for Sustainable Speaker Privacy
  8. Probabilistically Robust Watermarking of Neural Networks