PPaperPicks

Stuart Russell

University of California, Berkeley, Department of Electrical Engineering and Computer Sciences, CA, USA

23 papers at tracked venues · 20 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. AssistanceZero: Scalably Solving Assistance Games
  2. Avoiding Catastrophe in Online Learning by Asking for Help
  3. BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping
  4. Diffusion On Syntax Trees For Program Synthesis
  5. Extractive Structures Learned in Pretraining Enable Generalization on Finetuned Facts
  6. Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
  7. Monitoring Latent World States in Language Models with Propositional Probes
  8. Observation Interference in Partially Observable Assistance Games
  9. RL, but don't do anything I wouldn't do
  10. Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
  11. Robust and Diverse Multi-Agent Learning via Rational Policy Gradient
  12. The Partially Observable Off-Switch Game
  13. A Generalized Acquisition Function for Preference-based Reward Learning
  14. AI Alignment with Changing and Influenceable Reward Functions
  15. Ethically Compliant Autonomous Systems under Partial Observability
  16. Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
  17. Image Hijacks: Adversarial Images can Control Generative Models at Runtime
  18. On Representation Complexity of Model-based and Model-free Reinforcement Learning
  19. Position: Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback
  20. Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game
  21. The Effective Horizon Explains Deep RL Performance in Stochastic Environments
  22. Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics
  23. When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback