PPaperPicks

Bohan Zhuang

29 papers at tracked venues · 22 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. CoV: Chain-of-View Prompting for Spatial Reasoning
  2. OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
  3. Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning
  4. Are Large Vision Language Models Good Game Players?
  5. Channel Merging: Preserving Specialization for Merged Experts
  6. FPN-in-FPN: A Nested Multi-scale Aggregation Network for Polyp Segmentation
  7. FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
  8. Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image Synthesis
  9. McCaD: Multi-Contrast MRI Conditioned, Adaptive Adversarial Diffusion Model for High-Fidelity MRI Synthesis
  10. Neighboring Autoregressive Modeling for Efficient Visual Generation
  11. T-Stitch: Accelerating Sampling in Pre-Trained Diffusion Models with Trajectory Stitching
  12. ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS
  13. ZipAR: Parallel Autoregressive Image Generation through Spatial Locality
  14. ZipVL: Accelerating Vision-Language Models Through Dynamic Token Sparsity
  15. Efficient Stitchable Task Adaptation
  16. EfficientDM: Efficient Quantization-Aware Fine-Tuning of Low-Bit Diffusion Models
  17. GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI
  18. LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning
  19. LongVLM: Efficient Long Video Understanding via Large Language Models
  20. MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse Views
  21. MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-view Images
  22. MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
  23. ModaVerse: Efficiently Transforming Modalities with LLMs
  24. Motion Mamba: Efficient and Long Sequence Motion Generation
  25. Object-Aware Inversion and Reassembly for Image Editing
  26. QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models
  27. SAM-Med3D-MoE: Towards a Non-Forgetting Segment Anything Model via Mixture of Experts for 3D Medical Image Segmentation
  28. Stitched ViTs are Flexible Vision Backbones
  29. ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification