PPaperPicks

Yonatan Belinkov

33 papers at tracked venues · 23 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. CRISP: Persistent Concept Unlearning via Sparse Autoencoders
  2. Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models
  3. Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
  4. Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
  5. Will it Merge? On The Causes of Model Mergeability
  6. Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
  7. Arithmetic Without Algorithms: Language Models Solve Math with a Bag of Heuristics
  8. Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models
  9. CtD: Composition through Decomposition in Emergent Communication
  10. LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
  11. MIB: A Mechanistic Interpretability Benchmark
  12. Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
  13. Padding Tone: A Mechanistic Analysis of Padding Tokens in T2I Models
  14. Position-aware Automatic Circuit Discovery
  15. REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
  16. Reverse-Engineering the Retrieval Process in GenIR Models
  17. SAEs Are Good for Steering - If You Select the Right Features
  18. Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
  19. Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
  20. Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
  21. Unsupervised Translation of Emergent Communication
  22. Accelerating the Global Aggregation of Local Explanations
  23. Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
  24. Concept-Best-Matching: Evaluating Compositionality In Emergent Communication
  25. Confidence Regulation Neurons in Language Models
  26. ContraSim - Analyzing Neural Representations Based on Contrastive Learning
  27. Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
  28. Fast Forwarding Low-Rank Training
  29. Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking
  30. Leveraging Prototypical Representations for Mitigating Social Bias without Demographic Information
  31. Linearity of Relation Decoding in Transformer Language Models
  32. ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
  33. Semantics and Spatiality of Emergent Communication