PPaperPicks

Song-Chun Zhu

University of California, Department of Statistics, Los Angeles, CA, USA

34 papers at tracked venues · 30 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. JurisBench: A Deep Benchmark for Assessing Large Language Models in Professional Legal Practice
  2. SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning
  3. TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
  4. Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
  5. v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound
  6. Building Interactable Replicas of Complex Articulated Objects via Gaussian Splatting
  7. Decompositional Neural Scene Reconstruction with Generative Diffusion Prior
  8. Differentiable Information Enhanced Model-Based Reinforcement Learning
  9. Enhancing LLM-Based Social Bot via an Adversarial Learning Framework
  10. Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
  11. METASCENES: Towards Automated Replica Creation for Real-world 3D Scans
  12. Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
  13. ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection
  14. Social World Model-Augmented Mechanism Design Policy Learning
  15. Unveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-Analysis
  16. World Models Should Prioritize the Unification of Physical and Social Dynamics
  17. AdaSociety: An Adaptive Environment with Social Structures for Multi-Agent Decision-Making
  18. Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations
  19. An Embodied Generalist Agent in 3D World
  20. Bongard-OpenWorld: Few-Shot Reasoning for Free-form Visual Concepts in the Real World
  21. CLOVA: A Closed-LOop Visual Assistant with Tool Usage and Update
  22. CivRealm: A Learning and Reasoning Odyssey in Civilization for Decision-Making Agents
  23. Efficient Adaptation in Mixed-Motive Environments via Hierarchical Opponent Modeling and Planning
  24. FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal Models
  25. Fast Peer Adaptation with Context-aware Exploration
  26. INTERPRET: Interactive Predicate Learning from Language Feedback for Generalizable Task Planning
  27. LLM3: Large Language Model-based Task and Motion Planning with Motion Failure Reasoning
  28. LangSuit·E: Planning, Controlling and Interacting with Large Language Models in Embodied Text Environments
  29. Learning to Balance Altruism and Self-interest Based on Empathy in Mixed-Motive Games
  30. Mars: Situated Inductive Reasoning in an Open-World Environment
  31. Neural-Symbolic Recursive Machine for Systematic Generalization
  32. PhyRecon: Physically Plausible Neural Scene Reconstruction
  33. ProAgent: Building Proactive Cooperative Agents with Large Language Models
  34. RulE: Knowledge Graph Reasoning with Rule Embedding