PPaperPicks

Yixiao Ge

29 papers at tracked venues · 24 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
  2. ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
  3. AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction
  4. Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
  5. Equivariant Filter Design for Range-Only SLAM
    ICRA 2025 · Yixiao Ge
  6. GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers
  7. HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding
  8. LoRA-Gen: Specializing Large Language Model via Online LoRA Generation
  9. Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
  10. Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots
  11. SEED-Story: Multimodal Long Story Generation with Large Language Model
  12. Scalable Image Tokenization with Index Backpropagation Quantization
  13. VoCo-LLaMA: Towards Vision Compression with Large Language Models
  14. An Equivariant Approach to Robust State Estimation for the ArduPilot Autopilot System
  15. BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
  16. Cached Transformers: Improving Transformers with Differentiable Memory Cachde
  17. DreamDiffusion: High-Quality EEG-to-Image Generation with Temporal Masked Signal Modeling and CLIP Alignment
  18. LLaMA Pro: Progressive LLaMA with Block Expansion
  19. Low-Rank Approximation for Sparse Attention in Multi-Modal LLMs
  20. Making LLaMA SEE and Draw with SEED Tokenizer
  21. MambaTree: Tree Topology is All You Need in State Space Model
  22. Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
  23. Rethinking the Objectives of Vector-Quantized Tokenizers for Image Synthesis
  24. SEED-Bench: Benchmarking Multimodal Large Language Models
  25. ST-LLM: Large Language Models Are Effective Temporal Learners
  26. SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language Models
  27. UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
  28. VIT-LENS: Towards Omni-modal Representations
  29. YOLO-World: Real-Time Open-Vocabulary Object Detection