PPaperPicks

Chuang Gan

University of Massachusetts, Amherst, MA, USA

56 papers at tracked venues · 49 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Steering LLM Thinking with Budget Guidance
  2. Tailored Primitive Initialization is the Secret Key to Reinforcement Learning
  3. 3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
  4. ABNet: Adaptive explicit-Barrier Net for Safe and Scalable Robot Learning
  5. AdaWorld: Learning Adaptable World Models with Latent Actions
  6. COMBO: Compositional World Models for Embodied Multi-Agent Cooperation
  7. CommVQ: Commutative Vector Quantization for KV Cache Compression
  8. Delta: Dense Efficient Long-Range 3D tracking for any video
  9. LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences
  10. LangSplatV2: High-dimensional 3D Language Gaussian Splatting with 450+ FPS
  11. Learning 3D Persistent Embodied World Models
  12. Learning 4D Embodied World Models
  13. MatchMaker: Automated Asset Generation for Robotic Assembly
  14. MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
  15. QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
  16. RapVerse: Coherent Vocals and Whole-Body Motion Generation from Text
  17. RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills
  18. SafeDiffuser: Safe Planning with Diffusion Probabilistic Models
  19. Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
  20. Scaling Autonomous Agents via Automatic Reward Modeling And Planning
  21. TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation
  22. TopoGaussian: Inferring Internal Topology Structures from Visual Clues
  23. Towards Understanding Camera Motions in Any Video
  24. UniMuMo: Unified Text, Music, and Motion Generation
  25. VCA: Video Curious Agent for Long Video Understanding
  26. Your Language Model May Think Too Rigidly: Achieving Reasoning Consistency with Symmetry-Enhanced Training
  27. 3D-VLA: A 3D Vision-Language-Action Generative World Model
  28. AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
  29. Aligning Large Multimodal Models with Factually Augmented RLHF
  30. Architect: Generating Vivid and Interactive 3D Scenes with Hierarchical 2D Inpainting
  31. Building Cooperative Embodied Agents Modularly with Large Language Models
  32. CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding
  33. ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
  34. ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning
  35. Constrained Human-AI Cooperation: An Inclusive Embodied Social Intelligence Challenge
  36. ContPhy: Continuum Physical Concept Learning and Reasoning from Videos
  37. DIFFTACTILE: A Physics-based Differentiable Tactile Simulator for Contact-rich Robotic Manipulation
  38. Disentangled Acoustic Fields For Multimodal Physical Scene Understanding
  39. Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
  40. FlexAttention for Efficient High-Resolution Vision-Language Models
  41. GENOME: Generative Neuro-Symbolic Visual Reasoning by Growing and Reusing Modules
  42. HAZARD Challenge: Embodied Decision Making in Dynamically Changing Environments
  43. LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery
  44. Multi-Agent Alternate Q-Learning
  45. MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World
  46. Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance
  47. Physically Compatible 3D Object Modeling from a Single Image
  48. RILA: Reflective and Imaginative Language Agent for Zero-Shot Semantic Audio-Visual Navigation
  49. RoboDreamer: Learning Compositional World Models for Robot Imagination
  50. RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation
  51. SALMON: Self-Alignment with Instructable Reward Models
  52. SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World Knowledge
  53. SocialGPT: Prompting LLMs for Social Relation Reasoning via Greedy Segment Optimization
  54. Speech Self-Supervised Learning Using Diffusion Model Synthetic Data
  55. Thin-Shell Object Manipulations With Differentiable Physics Simulations
  56. Visual Chain-of-Thought Prompting for Knowledge-Based Visual Reasoning