PPaperPicks

Jie Zhou

Tencent Inc., WeChat AI, Pattern Recognition Center, Beijing, China

85 papers at tracked venues · 67 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. ArrowGEV: Grounding Events in Video via Learning the Arrow of Time
  2. ComoRAG: A Cognitive-Inspired Memory-Organized RAG for Stateful Long Narrative Reasoning
  3. Figure It Out: Improve the Frontier of Reasoning with Executable Visual States
  4. Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
  5. Joint Optimization of Training Data and Policy in RLHF
  6. ReFreeKV: Towards Threshold-Free KV Cache Compression
  7. Situated Embedding Models for Context-Aware Dense Retrieval
  8. Think Natively: Unlocking Multilingual Reasoning with Consistency-Enhanced Reinforcement Learning
  9. UltraEval-Audio: A Unified Framework for Comprehensive Evaluation of Audio Foundation Models
  10. A Law Reasoning Benchmark for LLM with Tree-Organized Structures including Factum Probandum, Evidence and Experiences
  11. A Self-Denoising Model for Robust Few-Shot Relation Extraction
  12. A Visual Leap in Clip Compositionality Reasoning Through Generation of Counterfactual Sets
  13. APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
  14. AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity
  15. Advancing SMoE for Continuous Domain Adaptation of MLLMs: Adaptive Router and Domain-Specific Loss
  16. An Empirical Study of Many-to-Many Summarization with Large Language Models
  17. ArtFRD: A Fisher-Rao Mixture Metric for Generative Model Aesthetic Evaluation
  18. Beyond Next Token Prediction: Patch-Level Training for Large Language Models
  19. CM-Align: Consistency-based Multilingual Alignment for Large Language Models
  20. Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
  21. ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoning
  22. Continuous Visual Autoregressive Generation via Score Maximization
  23. DRT: Deep Reasoning Translation via Long Chain-of-Thought
  24. DelTA: An Online Document-Level Translation Agent Based on Multi-Level Memory
  25. Dense Retrievers Can Fail on Simple Queries: Revealing The Granularity Dilemma of Embeddings
  26. Efficient Speech Language Modeling via Energy Distance in Continuous Latent Space
  27. Enhancing Cross-Tokenizer Knowledge Distillation with Contextual Dynamical Mapping
  28. FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
  29. From Imitation to Innovation: The Emergence of Ai's Unique Artistic Styles and the Challenge of Copyright Protection
  30. LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning
  31. Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-Experts
  32. LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information
  33. MCID: Multi-aspect Copyright Infringement Detection for Generated Images
  34. MedDiT: A Knowledge-Controlled Diffusion Transformer Framework for Dynamic Medical Image Generation in Virtual Simulated Patient
  35. MiniPLM: Knowledge Distillation for Pre-training Language Models
  36. POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
  37. Personalized Language Model Learning on Text Data Without User Identifiers
  38. PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
  39. Retrieval-Augmented Machine Translation with Unstructured Knowledge
  40. Secret Lies in Color: Enhancing AI-Generated Images Detection with Color Distribution Analysis
  41. THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine Translation
  42. TIU-Bench: A Benchmark for Evaluating Large Multimodal Models on Text-rich Image Understanding
  43. The Essence of Contextual Understanding in Theory of Mind: A Study on Question Answering with Story Characters
  44. The Stochastic Parrot on LLM's Shoulder: A Summative Assessment of Physical Concept Understanding
  45. Understanding LLMs' Fluid Intelligence Deficiency: An Analysis of the ARC Task
  46. WalkVLM: Aid Visually Impaired People Walking by Vision Language Model
  47. AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors
  48. BranchNorm: Robustly Scaling Extremely Deep Transformers
  49. C-LLM: Learn to Check Chinese Spelling Errors Character by Character
  50. CSCD-NS: a Chinese Spelling Check Dataset for Native Speakers
  51. Comments as Natural Logic Pivots: Improve Code Generation via Comment Perspective
  52. Continual Learning with Semi-supervised Contrastive Distillation for Incremental Neural Machine Translation
  53. Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
  54. Enhancing Byzantine-Resistant Aggregations with Client Embedding
  55. Exploring Conditional Variational Mechanism to Pinyin Input Method for Addressing One-to-Many Mappings in Low-Resource Scenarios
  56. Exploring the Benefit of Activation Sparsity in Pre-training
  57. Few-Shot Character Understanding in Movies as an Assessment to Meta-Learning of Theory-of-Mind
  58. Fine-Grained Modeling of Narrative Context: A Coherence Perspective via Retrospective Questions
  59. Generative Multi-Modal Knowledge Retrieval with Large Language Models
  60. Identifying Factual Inconsistencies in Summaries: Grounding LLM Inference via Task Taxonomy
  61. Improving Machine Translation with Large Language Models: A Preliminary Study with Cooperative Decoding
  62. Instruction Position Matters in Sequence Generation with Large Language Models
  63. LCS: A Language Converter Strategy for Zero-Shot Neural Machine Translation
  64. Language Generation with Strictly Proper Scoring Rules
  65. Large Language Models Are Not Robust Multiple Choice Selectors
  66. MAVEN-ARG: Completing the Puzzle of All-in-One Event Understanding Dataset with Event Argument Annotation
  67. Multi-Level Cross-Modal Alignment for Speech Relation Extraction
  68. On Large Language Models' Hallucination with Regard to Known Facts
  69. On Prompt-Driven Safeguarding for Large Language Models
  70. On the token distance modeling ability of higher RoPE attention dimension
  71. Outdated Issue Aware Decoding for Factual Knowledge Editing
  72. Plot Retrieval as an Assessment of Abstract Semantic Association
  73. TasTe: Teaching Large Language Models to Translate through Self-Reflection
  74. Teaching Large Language Models to Translate with Comparison
  75. Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven Agents
  76. ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
  77. Towards Codable Watermarking for Injecting Multi-Bits Information to LLMs
  78. Towards Multiple References Era - Addressing Data Leakage and Limited Reference Diversity in Machine Translation Evaluation
  79. Translatotron-V(ison): An End-to-End Model for In-Image Machine Translation
  80. Tree-of-Reasoning Question Decomposition for Complex Question Answering with Large Language Models
  81. Trust in Internal or External Knowledge? Generative Multi-Modal Entity Linking with Knowledge Retriever
  82. Understanding and Addressing the Under-Translation Problem from the Perspective of Decoding Objective
  83. Unsupervised Information Refinement Training of Large Language Models for Retrieval-Augmented Generation
  84. Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
  85. XAL: EXplainable Active Learning Makes Classifiers Better Low-resource Learners