PPaperPicks

Zeming Liu

33 papers at tracked venues · 24 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation
  2. Beyond Unimodal Shortcuts: MLLMs as Cross-Modal Reasoners for Grounded Named Entity Recognition
  3. Live-Aid: A Large-Scale Dialogue Dataset and Benchmark for Interleaved Multi-party Interactions in Live Streaming
  4. Mem²Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation
  5. PEAP: Proactive Embodied Action Sequence Planning with Joint Understanding of Vision and Audio Perception
  6. PEC-Home: Interpretation of Progressively Elliptical Commands in Smart Homes
  7. DocMEdit: Towards Document-Level Model Editing
  8. Exploring In-Image Machine Translation with Real-World Background
  9. Flow2Code: Evaluating Large Language Models for Flowchart-based Code Generation Capability
  10. GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art
  11. HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices
  12. KwaiChat: A Large-Scale Video-Driven Multilingual Mixed-Type Dialogue Corpus
  13. Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling
  14. Multi-modal Deepfake Detection via Multi-task Audio-Visual Prompt Learning
  15. NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment
  16. PRIM: Towards Practical In-Image Multilingual Machine Translation
  17. RETAIL: Towards Real-world Travel Planning for Large Language Models
  18. ReFF: Reinforcing Format Faithfulness in Language Models Across Varied Tasks
  19. RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models
  20. Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges
  21. STAMPsy: Towards SpatioTemporal-Aware Mixed-Type Dialogues for Psychological Counseling
  22. SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
  23. Self-Reasoning Language Models: Unfold Hidden Reasoning Chains with Few Reasoning Catalyst
  24. Semi-Supervised Clustering Framework for Fine-grained Scene Graph Generation
  25. SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
  26. Stealthy Jailbreak Attacks on Large Language Models via Benign Data Mirroring
  27. ToolSpectrum: Towards Personalized Tool Utilization for Large Language Models
  28. TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments
  29. Weak2Wise: An Automated, Lightweight Framework for Weak-LLM-Friendly Reasoning Synthesis
  30. AppBench: Planning of Multiple APIs from Various APPs for Complex User Instruction
  31. Deterministic Reversible Data Augmentation for Neural Machine Translation
  32. FAME: Towards Factual Multi-Task Model Editing
  33. Medical Dialogue System: A Survey of Categories, Methods, Evaluation and Challenges