PPaperPicks

Ge Zhang

ByteDance Inc.

57 papers at tracked venues · 45 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values
  2. CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs
  3. CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization
  4. M3TQA: Massively Multilingual Multitask Table Question Answering
  5. MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models
  6. MMTableBench: A Multi-level Multimodal Benchmark for Reasoning and Layout Complexity in Table QA
  7. MdEval: Massively Multilingual Code Debugging
  8. PRISM: Probabilistic Reward Model with Inherent Structural Modeling
  9. Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
  10. CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models
  11. COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning
  12. Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
  13. Can MLLMs Understand the Deep Implication Behind Chinese Images?
  14. FlexWorld: Progressively Expanding 3D Scenes for Flexible-View Exploration
  15. General-Reasoner: Advancing LLM Reasoning Across All Domains
  16. KARPA: A Training-free Method of Adapting Knowledge Graph as References for Large Language Model's Reasoning Path Aggregation
  17. KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks
  18. KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
  19. LIME: Less Is More for MLLM Evaluation
  20. M2RC-EVAL: Massively Multilingual Repository-level Code Completion Evaluation
  21. MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation
  22. MIO: A Foundation Model on Multimodal Tokens
  23. MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
  24. MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
  25. MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
  26. McEval: Massively Multilingual Code Evaluation
  27. MuPT: A Generative Symbolic Music Pretrained Transformer
  28. OAgents: An Empirical Study of Building Effective Agents
  29. Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language Models
  30. OmniBench: Towards The Future of Universal Omni-Language Models
  31. OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision
  32. OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models
  33. SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models
  34. TableBench: A Comprehensive and Complex Benchmark for Table Question Answering
  35. VAMBA: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
  36. VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
  37. AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
  38. AutoAgents: A Framework for Automatic Agent Generation
  39. CIF-Bench: A Chinese Instruction-Following Benchmark for Evaluating the Generalizability of Large Language Models
  40. ChatMusician: Understanding and Generating Music Intrinsically with LLM
  41. D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
  42. DDK: Distilling Domain Knowledge for Efficient Large Language Models
  43. E2-LLM: Efficient and Extreme Length Extension of Large Language Models
  44. II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models
  45. MAmmoTH2: Scaling Instructions from the Web
  46. MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning
  47. MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training
  48. MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
  49. MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
  50. MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical Language
  51. Massive Editing for Large Language Models via Meta Learning
  52. MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response
  53. OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
  54. RoleAgent: Building, Interacting, and Benchmarking High-quality Role-Playing Agents from Scripts
  55. SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
  56. UniIR: Training and Benchmarking Universal Multimodal Information Retrievers
  57. VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation