PPaperPicks

Yue Zhang

Westlake University, School of Engineering, Hangzhou, China

63 papers at tracked venues · 47 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. AiraXiv: An AI-Driven Open-Access Platform for Human and AI Scientists
  2. AutoFigure-Edit: Generating Editable Scientific Illustrations via Reference-Guided Styling
  3. Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents
  4. DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
  5. LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
  6. ReEfBench: Quantifying the Reasoning Efficiency of LLMs
  7. Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection
  8. Alleviating Hallucinations of Large Language Models through Induced Hallucinations
    NAACL 2025 · Yue Zhang
  9. An Empirical Analysis of Uncertainty in Large Language Model Evaluations
  10. CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmark
  11. Constrain Alignment with Sparse Autoencoders
  12. Contrastive Learning on LLM Back Generation Treebank for Cross-domain Constituency Parsing
  13. CycleResearcher: Improving Automated Research via Automated Review
  14. DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process
  15. Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values
  16. ELICIT: LLM Augmentation Via External In-context Capability
  17. ExploraCoder: Advancing Code Generation for Multiple Unseen APIs via Planning and Chained Exploration
  18. Exploring Model Editing for LLM-based Aspect-Based Sentiment Classification
  19. Glimpse: Enabling White-Box Methods to Use Proprietary Models for Zero-Shot LLM-Generated Text Detection
  20. Human Simulacra: Benchmarking the Personification of Large Language Models
  21. Learning to Reason under Off-Policy Guidance
  22. Lost in Literalism: How Supervised Training Shapes Translationese in LLMs
  23. MCRanker: Generating Diverse Criteria On-the-Fly to Improve Pointwise LLM Rankers
  24. MMQA: Evaluating LLMs with Multi-Table Multi-Hop Complex Questions
  25. Multi-Document Event Extraction Using Large and Small Language Models
  26. NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens
  27. PerSphere: A Comprehensive Framework for Multi-Faceted Perspective Retrieval and Summarization
  28. Personality Alignment of Large Language Models
  29. Reaction Graph: Towards Reaction-Level Modeling for Chemical Reactions with 3D Structures
  30. Reasoning is All You Need for Video Generalization: A Counterfactual Benchmark with Sub-question Evaluation
  31. SPARTUN3D: Situated Spatial Understanding of 3D World in Large Language Model
    ICLR 2025 · Yue Zhang
  32. Task Calibration: Calibrating Large Language Models on Inference Tasks
  33. ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning
  34. Unveiling Attractor Cycles in Large Language Models: A Dynamical Systems View of Successive Paraphrasing
  35. Vision-and-Language Navigation with Analogical Textual Descriptions in LLMs
    EMNLP 2025 · Yue Zhang
  36. A Rationale-centric Counterfactual Data Augmentation Method for Cross-Document Event Coreference Resolution
  37. A Survey on Open Information Extraction from Rule-based Model to Large Language Model
  38. AutoSurvey: Large Language Models Can Automatically Write Surveys
  39. Can Language Models Learn to Skip Steps?
  40. ECON: On the Detection and Resolution of Evidence Conflicts
  41. Empirical Prior for Text Autoencoders
  42. Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
  43. FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models
  44. Gated Slot Attention for Efficient Linear-Time Sequence Modeling
  45. KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models
  46. Knowledge Conflicts for LLMs: A Survey
  47. LLMEval: A Preliminary Study on How to Evaluate Large Language Models
    AAAI 2024 · Yue Zhang
  48. LexMatcher: Dictionary-centric Data Curation for LLM-based Machine Translation
  49. MAGE: Machine-generated Text Detection in the Wild
  50. Narrowing the Gap between Vision and Action in Navigation
    ACM MM 2024 · Yue Zhang
  51. PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
  52. Prefix Text as a Yarn: Eliciting Non-English Alignment in Foundation Language Model
  53. Pro-Woman, Anti-Man? Identifying Gender Bias in Stance Detection
  54. RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation
  55. RAGLAB: A Modular and Research-Oriented Unified Framework for Retrieval-Augmented Generation
  56. Semformer: Transformer Language Models with Semantic Planning
  57. Spotting AI's Touch: Identifying LLM-Paraphrased Spans in Text
  58. Supervised Knowledge Makes Large Language Models Better In-context Learners
  59. Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs
  60. Understanding In-Context Learning from Repetitions
  61. What Have We Achieved on Non-autoregressive Translation?
  62. XAL: EXplainable Active Learning Makes Classifiers Better Low-resource Learners
  63. ZeroStance: Leveraging ChatGPT for Open-Domain Stance Detection via Dataset Generation