PPaperPicks

Jiuxiang Gu

31 papers at tracked venues · 23 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. A Survey on LLM-based Conversational User Simulation
  2. MENTOR: Efficient Autoregressive Image Generation with Balanced Multimodal Control
  3. OIDA-QA: A Multimodal Benchmark for Analyzing the Opioid Industry Documents Archive
  4. Unveiling Inherent Visual Grounding in Multimodal LLMs for Text-Rich Images
  5. VipAct: Visual-Perception Enhancement via Specialized VLM Agent Collaboration and Tool-use
  6. ARTIST: Improving the Generation of Text-Rich Images with Disentangled Diffusion Models and Large Language Models
  7. CoMMIT: Coordinated Multimodal Instruction Tuning
  8. DiffIP: Representation Fingerprints for Robust IP Protection of Diffusion Models
  9. Differential Privacy Mechanisms in Neural Tangent Kernel Regression
    WACV 2025 · Jiuxiang Gu
  10. From Selection to Generation: A Survey of LLM-based Active Learning
  11. ImageFolder: Autoregressive Image Generation with Folded Tokens
  12. LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers
  13. METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
  14. MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data
  15. Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
  16. Numerical Pruning for Efficient Autoregressive Models
  17. QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the Edge
  18. R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
  19. Refer to Any Segmentation Mask Group with Vision-Language Prompts
  20. SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding
  21. Self-Debiasing Large Language Models: Zero-Shot Recognition and Reduction of Stereotypes
  22. ADOPD: A Large-Scale Document Page Decomposition Dataset
    ICLR 2024 · Jiuxiang Gu
  23. Advancing Vision-Language Models with Adapter Ensemble Strategies
  24. Category-Aware Active Domain Adaptation
  25. Customization Assistant for Text-to-image Generation
  26. LRM: Large Reconstruction Model for Single Image to 3D
  27. SOHES: Self-supervised Open-world Hierarchical Entity Segmentation
  28. Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
  29. Self-Cleaning: Improving a Named Entity Recognizer Trained on Noisy Data with a Few Clean Instances
  30. TRINS: Towards Multimodal Language Models that Can Read
  31. TextLap: Customizing Language Models for Text-to-Layout Planning