PPaperPicks

Silvio Savarese

Stanford University, Department of Computer Science, Stanford, CA, USA

31 papers at tracked venues · 20 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
  2. ActionStudio: A Lightweight Framework for Data and Training of Large Action Models
  3. BLIP-3: A Family of Open Large Multimodal Models
  4. CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments
  5. CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
  6. Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D
  7. Diversity Empowers Intelligence: Integrating Expertise of Software Engineering Agents
  8. DyMU: Dynamic Merging and Virtual Unmerging for Efficient Variable-Length VLMs
  9. LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback
  10. LATTE: Learning to Think with Vision Specialists
  11. MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models
  12. Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts
  13. PersonaBench: Evaluating AI Models on Understanding Personal Information through Accessing (Synthetic) Private User Data
  14. Reward-Guided Speculative Decoding for Efficient LLM Reasoning
  15. SlackAgents: Scalable Collaboration of AI Agents in Workspaces
  16. Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
  17. Text2Data: Low-Resource Data Generation with Textual Control
  18. ViUniT: Visual Unit Tests for More Robust Visual Programming
  19. xLAM: A Family of Large Action Models to Empower AI Agent Systems
  20. APIGen: Automated PIpeline for Generating Verifiable and Diverse Function-Calling Datasets
  21. HIVE: Harnessing Human Feedback for Instructional Visual Editing
  22. How Do Transformers Learn In-Context Beyond Simple Functions? A Case Study on Learning with Representations
  23. INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness
  24. MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
  25. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
  26. Online Distribution Shift Detection via Recency Prediction
  27. Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization
  28. ULIP-2: Towards Scalable Multimodal Pre-Training for 3D Understanding
  29. Unified Training of Universal Time Series Forecasting Transformers
  30. X-InstructBLIP: A Framework for Aligning Image, 3D, Audio, Video to LLMs and its Emergent Cross-Modal Reasoning
  31. xGen-VideoSyn-1: High-Fidelity Text-to-Video Synthesis with Compressed Representations