PPaperPicks

Shuicheng Yan

National University of Singapore, Department of Electrical and Computer Engineering

49 papers at tracked venues · 43 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation
  2. EvoRoute: Experience-Driven Self-Routing LLM Agent Systems
  3. PointDGRWKV: Generalizing RWKV-like Architecture to Unseen Domains for Point Cloud Classification
  4. AR 2O Painter: An Artistic Oriented Realtime Realistic Oil Painting Agent Powered by Efficient Fluid Simulation
  5. AgentStudio: A Toolkit for Building General Virtual Agents
  6. Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
  7. Beyond Isolated Words: Diffusion Brush for Handwritten Text-Line Generation
  8. Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
  9. Cradle: Empowering Foundation Agents towards General Computer Control
  10. EasyInv: Toward Fast and Better DDIM Inversion
  11. EditWorld: Simulating World Dynamics for Instruction-Following Image Editing
  12. Explore In-Context Segmentation via Latent Diffusion Models
  13. From Outline to Detail: An Hierarchical End-to-end Framework for Coherent and Consistent Visual Novel Generation and Assembly
  14. G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems
  15. IPDreamer: Appearance-Controllable 3D Object Generation with Complex Image Prompts
  16. JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
  17. JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
  18. MEGA: Memory-Efficient 4D Gaussian Splatting for Dynamic Scenes
  19. Masks Can be Learned as an Alternative to Experts
  20. Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
  21. MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
  22. MoH: Multi-Head Attention as Mixture-of-Head Attention
  23. On Path to Multimodal Generalist: General-Level and General-Bench
  24. Pask: Providing Answer before AsKing toward Proactive AI agent
  25. Point Cloud Mamba: Point Cloud Learning via State Space Model
  26. PointDGMamba: Domain Generalization of Point Cloud Classification via Generalized State Space Model
  27. Poison-splat: Computation Cost Attack on 3D Gaussian Splatting
  28. Policy Optimization under Imperfect Human Interactions with Agent-Gated Shared Autonomy
  29. Policy Regularization on Globally Accessible States in Cross-Dynamics Reinforcement Learning
  30. Removing Prompt-template Bias in Reinforcement Learning from Human Feedback
  31. RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
  32. SuperCorrect: Advancing Small LLM Reasoning with Thought Template Distillation and Self-Correction
  33. Towards Semantic Equivalence of Tokenization in Multimodal LLM
  34. Action Imitation in Common Action Space for Customized Action Image Synthesis
  35. Auto-Encoding Morph-Tokens for Multimodal LLM
  36. Automating Dataset Updates Towards Reliable and Timely Evaluation of Large Language Models
  37. BAFFLE: A Baseline of Backpropagation-Free Federated Learning
  38. DGMamba: Domain Generalization via Generalized State Space Model
  39. From Multimodal LLM to Human-level AI: Modality, Instruction, Reasoning and Beyond
  40. Generative AI in Multimedia: Challenges and Opportunities for Academic and Industrial Impact
  41. Improving Video Segmentation via Dynamic Anchor Queries
  42. InceptionNeXt: When Inception Meets ConvNeXt
  43. LLMs-as-Instructors: Learning from Errors Toward Automating Model Improvement
  44. MVGamba: Unify 3D Content Generation as State Space Sequence Modeling
  45. Non-confusing Generation of Customized Concepts in Diffusion Models
  46. OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
  47. Region-Native Visual Tokenization
  48. Reinforcement Learning from Diverse Human Preferences
  49. Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing