PPaperPicks

Tao Chen

Fudan University, School of Information Science and Technology, Shanghai, China

40 papers at tracked venues · 32 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement
  2. Mitigating Low-Quality Reasoning in MLLMs: Self-Driven Refined Multimodal CoT with Selective Thinking and Step-wise Visual Enhancement
  3. Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers
  4. All-in-One: Transferring Vision Foundation Models into Stereo Matching
  5. Biology-Instructions: A Dataset and Benchmark for Multi-Omics Sequence Understanding Capability of Large Language Models
  6. Boost Embodied AI Models with Robust Compression Boundary
  7. Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta Compression
  8. Causal Motion Tokenizer for Streaming Motion Generation
  9. Chimera: Improving Generalist Model with Domain-Specific Experts
  10. Consistency-aware Self-Training for Iterative-based Stereo Matching
  11. Cross-Modal Graph Learning for Perivascular Spaces Segmentation
    MICCAI 2025 · Tao Chen
  12. DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models
  13. Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback
  14. DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes
  15. FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding
  16. GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
  17. HiSplat: Hierarchical 3D Gaussian Splatting for Generalizable Sparse-View Reconstruction
  18. Once-Tuning-Multiple-Variants: Tuning Once and Expanded as Multiple Vision-Language Model Variants
  19. PRING: Rethinking Protein-Protein Interaction Prediction from Pairs to Graphs
  20. PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding
  21. Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning
  22. SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
  23. 3DET-Mamba: Causal Sequence Modelling for End-to-End 3D Object Detection
  24. Bifröst: 3D-Aware Image Compositing with Language Instructions
  25. Boosting Residual Networks with Group Knowledge
  26. EMR-Merging: Tuning-Free High-Performance Model Merging
  27. Enhanced Sparsification via Stimulative Training
  28. FNP: Fourier Neural Processes for Arbitrary-Resolution Data Assimilation
  29. LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning
  30. M3DBench: Towards Omni 3D Assistant with Interleaved Multi-modal Instructions
  31. MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer
  32. MeshXL: Neural Coordinate Field for Generative 3D Foundation Models
  33. Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer Compression
  34. PM-INR: Prior-Rich Multi-Modal Implicit Large-Scale Scene Neural Representation
  35. ReSimAD: Zero-Shot 3D Domain Transfer for Autonomous Driving with Source Reconstruction and Target Simulation
  36. Reg-TTA3D: Better Regression Makes Better Test-Time Adaptive 3D Object Detection
  37. S2HPruner: Soft-to-Hard Distillation Bridges the Discretization Gap in Pruning
  38. Spear: Evaluate the Adversarial Robustness of Compressed Neural Models
  39. Through the Real World Haze Scenes: Navigating the Synthetic-to-Real Gap in Challenging Image Dehazing
  40. Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy