PPaperPicks

Ji Zhang

Alibaba Group, DAMO Academy, Hangzhou, China

36 papers at tracked venues · 27 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
  2. ProFuser: Progressive Fusion of Large Language Models
  3. A Simple yet Effective Layout Token in Large Language Models for Document Understanding
  4. AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization
  5. DISC: Plug-and-Play Decoding Intervention with Similarity of Characters for Chinese Spelling Check
  6. Exploiting Presentative Feature Distributions for Parameter-Efficient Continual Learning of Large Language Models
  7. SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization
  8. VLM-R³: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
  9. WritingBench: A Comprehensive Benchmark for Generative Writing
  10. mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
  11. mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
  12. A Simple yet Effective Training-free Prompt-free Approach to Chinese Spelling Correction Based on Large Language Models
  13. Breaking Barriers of System Heterogeneity: Straggler-Tolerant Multimodal Federated Learning via Knowledge Distillation
  14. Bridging the Space Gap: Unifying Geometry Knowledge Graph Embedding with Optimal Transport
  15. Browse and Concentrate: Comprehending Multimodal Content via Prior-LLM Context Fusion
  16. Budget-Constrained Tool Learning with Planning
  17. CycleAlign: Iterative Distillation from Black-box LLM to White-box Models for Better Human Alignment
  18. DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
  19. Enhancing Zero-shot Audio Classification using Sound Attribute Knowledge from Large Language Models
  20. From Skepticism to Acceptance: Simulating the Attitude Dynamics Toward Fake News
  21. Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
  22. Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
  23. MIBench: Evaluating Multimodal Large Language Models over Multiple Images
  24. MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
  25. Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
  26. Model Composition for Multimodal Large Language Models
  27. PANDA: Preference Adaptation for Enhancing Domain-Specific Abilities of LLMs
  28. Revisiting Unsupervised Temporal Action Localization: The Primacy of High-Quality Actionness and Pseudolabels
  29. Small LLMs Are Weak Tool Learners: A Multi-LLM Agent
  30. SocialBench: Sociality Evaluation of Role-Playing Conversational Agents
  31. TiMix: Text-Aware Image Mixing for Effective Vision-Language Pre-training
  32. TinyChart: Efficient Chart Understanding with Program-of-Thoughts Learning and Visual Token Merging
  33. Towards Better Utilization of Multi-Reference Training Data for Chinese Grammatical Error Correction
  34. mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
  35. mPLUG-OwI2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
  36. mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model