PPaperPicks

Siyu Yuan

32 papers at tracked venues · 18 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Metaphor Reasoning is Meta-reasoning
  2. ARIA: Training Language Agents with Intention-driven Reward Aggregation
  3. CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios
  4. Character is Destiny: Can Persona-assigned Language Models Make Personal Choices?
  5. CoSER: Coordinating LLM-Based Persona Simulation of Established Roles
  6. Collaborative Human Activity Recognition with Passive Inter-Body Electrostatic Field
  7. Curse of Knowledge: Your Guidance and Provided Knowledge are biasing LLM Judges in Complex Evaluation
  8. DEEPER Insight into Your User: Directed Persona Refinement for Dynamic Persona Modeling
  9. EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction
    NAACL 2025 · Siyu Yuan
  10. Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles
  11. EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms
    NAACL 2025 · Siyu Yuan
  12. Implicit Reasoning in Transformers is Reasoning through Shortcuts
  13. KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
  14. LLM-Powered Information Extraction for the Dairy Financial Domain: Tackling Data Scarcity and Ambiguity
  15. MultiLingPoT: Boosting Mathematical Reasoning in LLMs through Multilingual Program Integration
  16. ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
  17. Past Meets Present: Creating Historical Analogy with Large Language Models
  18. PunMemeCN: A Benchmark to Explore Vision-Language Models' Understanding of Chinese Pun Memes
  19. Revealing the Barriers of Language Agents in Planning
  20. SELFGOAL: Your Language Agents Already Know How to Achieve High-level Goals
  21. The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided Improvement
  22. ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use
  23. Unlocking Scientific Concepts: How Effective Are LLM-Generated Analogies for Student Understanding and Classroom Practice?
  24. "A good pun is its own reword": Can Large Language Models Understand Puns?
  25. ANALOGYKB: Unlocking Analogical Reasoning of Language Models with A Million-scale Knowledge Base
    ACL 2024 · Siyu Yuan
  26. Boosting Scientific Concepts Understanding: Can Analogy from Teacher Models Empower Student Models?
    EMNLP 2024 · Siyu Yuan
  27. Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional Works
  28. InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews
  29. Light Up the Shadows: Enhance Long-Tailed Entity Grounding with Concept-Guided Vision-Language Models
  30. TaskBench: Benchmarking Large Language Models for Task Automation
  31. TimeArena: Shaping Efficient Multitasking Language Agents in a Time-Aware Simulation
  32. Translate Meanings, Not Just Words: IdiomKB's Role in Optimizing Idiomatic Translation with Language Models