PPaperPicks

Min Zhang

Harbin Institute of Technology, School of Computer Science and Technology, Institute of Computing and Intelligence, Shenzhen, China

200 papers at tracked venues · 161 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. A Graph-Enhanced MLLM for Hierarchical Multimodal Emotion Understanding and Support in Conversations
  2. BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs
  3. Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering
  4. Beyond Unimodal Shortcuts: MLLMs as Cross-Modal Reasoners for Grounded Named Entity Recognition
  5. ComfyFlow: Benchmarking LLMs for AIGC Workflow Generation
  6. ComfyUI-R1: Exploring Reasoning Models for Workflow Generation
  7. Crossing the Reward Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse Domains
  8. D-QRELO: Training- and Data-Free Delta Compression for Large Language Models via Quantization and Residual Low-Rank Approximation
  9. DUAL RM: Beyond Rule-based Preference Reward Modeling via Meta-Reward
  10. DeReA: Improving Idiom Translation with Detect-Retrieve-Arbitrate Reasoning
  11. Diagnosing and Remedying Representation Deficiencies for Deterministic Reasoning in KGQA
  12. Dynamic Long Context Reasoning over Compressed Memory via End-to-End Reinforcement Learning
  13. Dynamic Sampling that Adapts: Self-Aware Iterative Data Persistent Optimization for Mathematical Reasoning
  14. Efficient Reasoning for LLMs through Speculative Chain-of-Thought
  15. Escaping the Echo Trap: On Credit Assignment Failure in Multi-turn LLM Self-Reflection
  16. E³-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning
  17. From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
  18. Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs
  19. IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural Thinking
  20. Improving Value-based Process Verifier via Low-Cost Variance Reduction
  21. Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning
  22. Learning to Extract Rational Evidence via Reinforcement Learning for Retrieval-Augmented Generation
  23. Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework
  24. Listing Minimal Cores in Large Real-World Graphs
  25. LycheeCluster: Efficient Long-Context Inference with Structure-Aware Chunking and Hierarchical KV Indexing
  26. MIND Your Reasoning: A Meta-Cognitive Intuitive-Reflective Network for Dual-Reasoning in Multimodal Stance Detection
  27. MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation
  28. MathAgent: Adversarial Evolution of Constraint Graphs for Mathematical Reasoning Data Synthesis
  29. Negotiating the Punchline: Contextual Meme Understanding via Discrete Semantic Energy Minimization
  30. SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive Thinking
  31. SocialDropout: Dynamic Agent Dropout for Social Simulation
  32. Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment
  33. ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution
  34. Towards Closed-Loop Embodied Empathy Evolution: Probing LLM-Centric Lifelong Empathic Motion Generation in Unseen Scenarios
  35. UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity Mixture-of-Experts
  36. When Is Thinking Enough? Early Exit via Sufficiency Assessment for Efficient Reasoning
  37. When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning
  38. WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments
  39. A Survey on the Feedback Mechanism of LLM-based AI Agents
  40. A Training-free LLM-based Approach to General Chinese Character Error Correction
  41. A Unified Agentic Framework for Evaluating Conditional Image Generation
  42. ALW: Adaptive Layer-Wise contrastive decoding enhancing reasoning ability in Large Language Models
  43. APT: Improving Specialist LLM Performance with Weakness Case Acquisition and Iterative Preference Training
  44. AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMs
  45. Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation
  46. Accurate KV Cache Quantization with Outlier Tokens Tracing
  47. Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing
  48. AgentDropout: Dynamic Agent Elimination for Token-Efficient and High-Performance LLM-Based Multi-Agent Collaboration
  49. AgentInit: Initializing LLM-based Multi-Agent Systems via Diversity and Expertise Orchestration for Effective and Efficient Collaboration
  50. Alignment-Augmented Speculative Decoding with Alignment Sampling and Conditional Verification
  51. An Empirical Study of Iterative Refinements for Non-autoregressive Translation
  52. An Evaluation Resource for Grounding Translation Errors
  53. Basic Reading Distillation
  54. Benchmarking LLMs for Translating Classical Chinese Poetry: Evaluating Adequacy, Fluency, and Elegance
  55. Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
  56. Beware of Calibration Data for Pruning Large Language Models
  57. BrainECHO: Semantic Brain Signal Decoding through Vector-Quantized Spectrogram Reconstruction for Whisper-Enhanced Text Generation
  58. Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language Models
  59. CDT: A Comprehensive Capability Framework for Large Language Models Across Cognition, Domain, and Task
  60. CMT: A Memory Compression Method for Continual Knowledge Learning of Large Language Models
  61. CogMAEC'25: The 1st Workshop on Cognition-oriented Multimodal Affective and Empathetic Computing
  62. ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development
  63. Contrastive Learning on LLM Back Generation Treebank for Cross-domain Constituency Parsing
  64. DISC: Plug-and-Play Decoding Intervention with Similarity of Characters for Chinese Spelling Check
  65. DRPruning: Efficient Large Language Model Pruning through Distributionally Robust Optimization
  66. Decoder-Only LLMs can be Masked Auto-Encoders
  67. DelTA: An Online Document-Level Translation Agent Based on Multi-Level Memory
  68. DoCIA: An Online Document-Level Context Incorporation Agent for Speech Translation
  69. DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
  70. Efficient Safety Alignment of Large Language Models via Preference Re-ranking and Representation-based Reward Modeling
  71. Efficient Speech Language Modeling via Energy Distance in Continuous Latent Space
  72. Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
  73. Exploring the Translation Mechanism of Large Language Models
  74. FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing
  75. FlexRAG: A Flexible and Comprehensive Framework for Retrieval-Augmented Generation
  76. From Awareness to Adaptability: Enhancing Tool Utilization for Scientific Reasoning
  77. Function-to-Style Guidance of LLMs for Code Translation
  78. FunnelRAG: A Coarse-to-Fine Progressive Retrieval Paradigm for RAG
  79. GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent
  80. Generative Reward Modeling via Synthetic Criteria Preference Learning
  81. Generator-Assistant Stepwise Rollback Framework for Large Language Model Agent
  82. Improving Rationality in the Reasoning Process of Language Models through Self-playing Game
  83. InImageTrans: Multimodal LLM-based Text Image Machine Translation
  84. Knowledge Editing with Dynamic Knowledge Graphs for Multi-Hop Question Answering
  85. L-CiteEval: A Suite for Evaluating Fidelity of Long-context Models
  86. LLM-based Translation Inference with Iterative Bilingual Understanding
  87. LLMs Can Also Do Well! Breaking Barriers in Semantic Role Labeling via Large Language Models
  88. LOGO - Long cOntext aliGnment via efficient preference Optimization
  89. Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization
  90. Locate-and-Focus: Enhancing Terminology Translation in Speech Language Models
  91. Look Before You Leap: Enhance Attention and Vigilance Regarding Harmful Content with GuidelineLLM
  92. MASTER: Enhancing Large Language Model via Multi-Agent Simulated Teaching
  93. MMA: Cross-Domain Knowledge Integration via Mixture of Multi-Domain Agents
  94. MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
  95. Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation
  96. MeKB-Sim: Personal Knowledge Base-Powered Multi-Agent Simulation
  97. Memory-augmented Query Reconstruction for LLM-based Knowledge Graph Reasoning
  98. MoDification: Mixture of Depths Made Easy
  99. Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling
  100. Neural Parameter Search for Slimmer Fine-Tuned Models and Better Transfer
  101. ORPP: Self-Optimizing Role-playing Prompts to Enhance Language Model Capabilities
  102. Overcoming Non-monotonicity in Transducer-based Streaming Generation
  103. PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models
  104. Reflection on Knowledge Graph for Large Language Models Reasoning
  105. Revealing and Mitigating Over-Attention in Knowledge Editing
  106. Revealing and Mitigating the Local Pattern Shortcuts of Mamba
  107. SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
  108. SSRB: Direct Natural Language Querying to Massive Heterogeneous Semi-Structured Data
  109. Safety Alignment via Constrained Knowledge Unlearning
  110. SeaPO: Strategic Error Amplification for Robust Preference Optimization of Large Language Models
  111. Skynet-V1: Towards Early Warning of Video Abnormal Events via A Spatial-temporal Causal-enhanced MoE Framework
  112. Speed Up Your Code: Progressive Code Acceleration Through Bidirectional Tree Editing
  113. The ACM Multimedia 2025 Grand Challenge of Avatar-based Multimodal Empathetic Conversation
  114. The Rise of Darkness: Safety-Utility Trade-Offs in Role-Playing Dialogue Agents
  115. Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning
  116. Tool learning via Inference-time Scaling and Cycle Verifier
  117. Towards Text-Image Interleaved Retrieval
  118. Unleashing LLM Reasoning Capability via Scalable Question Synthesis from Scratch
  119. Unlocking Recursive Thinking of LLMs: Alignment via Refinement
  120. Unveiling the Potential of BERT-family: A New Recipe for Building Scalable, General and Competitive Large Language Models
  121. VideoVista-CulturalLingo: 360° Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension
  122. VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models
  123. When Words Smile: Generating Diverse Emotional Facial Expressions from Text
  124. XIFBench: Evaluating Large Language Models on Multilingual Instruction Following
  125. Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
  126. \mathcalA³: Automatic Alignment Framework for Attributed Text Generation
  127. A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any Translation
  128. A Simple yet Effective Training-free Prompt-free Approach to Chinese Spelling Correction Based on Large Language Models
  129. A Two-Stage Adaptation of Large Language Models for Text Ranking
  130. Achieving Stronger Generation via Simple Contrastive Tuning
  131. Adaptive Feature-based Low-Rank Compression of Large Language Models via Bayesian Optimization
  132. An Empirical Study of CLIP for Text-Based Person Search
  133. Are Bert Family Good Instruction Followers? A Study on Their Potential And Limitations
  134. AutoSurvey: Large Language Models Can Automatically Write Surveys
  135. CMD: a framework for Context-aware Model self-Detoxification
  136. CTC-based Non-autoregressive Textless Speech-to-Speech Translation
  137. Can LLMs Learn Uncertainty on Their Own? Expressing Uncertainty Effectively in A Self-Training Manner
  138. Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
  139. Chinese Spelling Corrector Is Just a Language Learner
  140. Chinese Spoken Named Entity Recognition in Real-world Scenarios: Dataset and Approaches
  141. Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge Alignment
  142. CommonIT: Commonality-Aware Instruction Tuning for Large Language Models via Data Partitions
  143. Concise and Precise Context Compression for Tool-Using Language Models
  144. ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLMs
  145. Context Consistency between Training and Inference in Simultaneous Machine Translation
  146. CopyNE: Better Contextual ASR by Copying Named Entities
  147. Curriculum Consistency Learning for Conditional Sentence Generation
  148. DB-LLM: Accurate Dual-Binarization for Efficient LLMs
  149. DUAL-REFLECT: Enhancing Large Language Models for Reflective Translation through Dual Learning Feedback Mechanisms
  150. DeMPT: Decoding-enhanced Multi-phase Prompt Tuning for Making LLMs Be Better Context-aware Translators
  151. Demonstration Augmentation for Zero-shot In-context Learning
  152. Domain-Aware k-Nearest-Neighbor Knowledge Distillation for Machine Translation
  153. Dynamic Planning for LLM-based Graphical User Interface Automation
  154. EEG-MACS: Manifold Attention and Confidence Stratification for EEG-based Cross-Center Brain Disease Diagnosis under Unreliable Annotations
  155. Efficient Domain Adaptation for Non-Autoregressive Machine Translation
  156. Efficient k-Nearest-Neighbor Machine Translation with Dynamic Retrieval
  157. Enhancing EEG-to-Text Decoding through Transferable Representations from Pre-trained Contrastive EEG-Text Masked Autoencoder
  158. Exploring Reversal Mathematical Reasoning Ability for Large Language Models
  159. From Role-Play to Drama-Interaction: An LLM Solution
  160. GenView: Enhancing View Quality with Pretrained Generative Model for Self-Supervised Learning
  161. Improving Aspect-Based Sentiment Analysis via Tuple-Order Learning
  162. Improving Attributed Text Generation of Large Language Models via Preference Learning
  163. In-Context Learning State Vector with Inner and Momentum Optimization
  164. LLM-Driven Multimodal Opinion Expression Identification
  165. Learnability Matters: Active Learning for Video Captioning
  166. Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?
  167. Medico: Towards Hallucination Detection and Correction with Multi-source Evidence Fusion
  168. MultiSkill: Evaluating Large Multimodal Models for Fine-grained Alignment Skills
  169. Multimodal Reasoning with Multimodal Knowledge Graph
  170. NewTerm: Benchmarking Real-Time New Terms for Large Language Models with Annual Updates
  171. On the Hallucination in Simultaneous Machine Translation
  172. Parameter Competition Balancing for Model Merging
  173. Paying More Attention to Source Context: Mitigating Unfaithful Translations from Large Language Model
  174. Question-guided Knowledge Graph Re-scoring and Injection for Knowledge Graph Question Answering
  175. Recognizing Everything from All Modalities at Once: Grounded Multimodal Universal Information Extraction
  176. Rethinking Negative Instances for Generative Named Entity Recognition
  177. Retrieval and Reasoning on KGs: Integrate Knowledge Graphs into Large Language Models for Complex Question Answering
  178. Revisiting Demonstration Selection Strategies in In-Context Learning
  179. SEER: Self-Aligned Evidence Extraction for Retrieval-Augmented Generation
  180. SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-Reflection
  181. Self-Powered LLM Modality Expansion for Large Speech-Text Models
  182. Semantic Role Labeling from Chinese Speech via End-to-End Learning
  183. Separate the Wheat from the Chaff: Model Deficiency Unlearning via Parameter-Efficient Module Operation
  184. Speech Sense Disambiguation: Tackling Homophone Ambiguity in End-to-End Speech Translation
  185. SpeechEE: A Novel Benchmark for Speech Event Extraction
  186. StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
  187. Synergistic Dual Spatial-aware Generation of Image-to-text and Text-to-image
  188. Take Off the Training Wheels! Progressive In-Context Learning for Effective Alignment
  189. TasTe: Teaching Large Language Models to Translate through Self-Reflection
  190. Temporal Knowledge Question Answering via Abstract Reasoning Induction
  191. The ACM Multimedia 2024 Viual Spatial Description Grand Challenge
  192. Towards Better Chinese Spelling Check for Search Engines: A New Dataset and Strong Baseline
  193. Towards Demonstration-Aware Large Language Models for Machine Translation
  194. Translatotron-V(ison): An End-to-End Model for In-Image Machine Translation
  195. TruthReader: Towards Trustworthy Document Assistant Chatbot with Reliable Attribution
  196. Unsupervised Sign Language Translation and Generation
  197. VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
  198. XNLP: An Interactive Demonstration System for Universal Structured NLP
  199. mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval