PPaperPicks

Songyang Zhang

ShanghaiTech University, Shanghai, China

24 papers at tracked venues · 17 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains
  2. Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
  3. RouteMoA: Dynamic Routing without Pre-Inference Boosts Efficient Mixture-of-Agents
  4. Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law
  5. CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
  6. Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
  7. LiT: Delving into a Simple Linear Diffusion Transformer for Image Generation
  8. OpenHuEval: Evaluating Large Language Model on Hungarian Specifics
  9. Rethinking Verification for LLM Code Generation: From Generation to Testing
  10. UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
  11. Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
  12. Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
  13. BotChat: Evaluating LLMs' Capabilities of Having Multi-Turn Dialogues
  14. Fake Alignment: Are LLMs Really Aligned Well?
  15. From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
  16. GTA: A Benchmark for General Tool Agents
  17. InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
  18. LLaST: Improved End-to-end Speech Translation System Leveraged by Large Language Models
  19. LawBench: Benchmarking Legal Knowledge of Large Language Models
  20. MMBench: Is Your Multi-modal Model an All-Around Player?
  21. MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
  22. Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
  23. ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
  24. T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step