PPaperPicks

Yuxiao Dong

56 papers at tracked venues · 48 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. KARL: Reinforcement Learning for LLM Agents on Multi-Turn Knowledge-Intensive Agentic Tasks
  2. VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
  3. A Survey of Post-Training Scaling in Large Language Models
  4. AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
  5. AndroidGen: Building an Android Language Agent under Data Scarcity
  6. AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents
  7. CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
  8. CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
  9. Controlling Large Language Model with Latent Action
  10. LVBench: An Extreme Long Video Understanding Benchmark
  11. LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
  12. LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
  13. LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-Context QA
  14. LongReward: Improving Long-context Large Language Models with AI Feedback
  15. LongSafety: Evaluating Long-Context Safety of Large Language Models
  16. LongWriter: Unleashing 10, 000+ Word Generation from Long Context LLMs
  17. MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
  18. SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
  19. SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling
  20. Scaling Speech-Text Pre-training with Synthetic Interleaved Data
  21. SceneGenAgent: Precise Industrial Scene Generation with Coding Agent
  22. T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
  23. TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
  24. VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
  25. VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
  26. WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
  27. AgentBench: Evaluating LLMs as Agents
  28. AgentTuning: Enabling Generalized Agent Abilities for LLMs
  29. AlignBench: Benchmarking Chinese Alignment of Large Language Models
  30. AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models
  31. AutoRE: Document-Level Relation Extraction with Large Language Models
  32. AutoWebGLM: A Large Language Model-based Web Navigating Agent
  33. Black-Box Prompt Optimization: Aligning Large Language Models without Model Training
  34. CharacterGLM: Customizing Social Characters with Large Language Models
  35. ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline
  36. CogAgent: A Visual Language Model for GUI Agents
  37. CogVLM: Visual Expert for Pretrained Language Models
  38. CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion
  39. CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
  40. Generative AI Day
  41. Inf-DiT: Upsampling Any-Resolution Image with Memory-Efficient Diffusion Transformer
  42. LongAlign: A Recipe for Long Context Alignment of Large Language Models
  43. LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
  44. LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question Answering
  45. Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments
  46. NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Queries
  47. OAG-Bench: A Human-Curated Benchmark for Academic Graph Mining
  48. Open-World Semi-Supervised Learning for Node Classification
  49. OpenWebAgent: An Open Toolkit to Enable Web Agents on Large Language Models
  50. Pre-Training and Prompting for Few-Shot Node Classification on Text-Attributed Graphs
  51. ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
  52. RecDCL: Dual Contrastive Learning for Recommendation
  53. Revisiting Parallel Context Windows: A Frustratingly Simple Alternative and Chain-of-Thought Deterioration
  54. SciInstruct: a Self-Reflective Instruction Annotated Dataset for Training Scientific Language Models
  55. TriSampler: A Better Negative Sampling Principle for Dense Retrieval
  56. Understanding Emergent Abilities of Language Models from the Loss Perspective