PPaperPicks

Conghui He

64 papers at tracked venues · 52 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch
  2. Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's Nature
  3. MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
  4. MoDora: Tree-Based Semi-Structured Document Analysis System
  5. REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once
  6. The Data Frontier for Large Language Models: Selection, Synthesis, and Tools
  7. Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs
  8. A Strategic Coordination Framework of Small LMs Matches Large LMs in Data Synthesis
  9. BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
  10. BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
  11. CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenge
  12. Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
  13. Conical Visual Concentration for Efficient Large Vision-Language Models
  14. Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context Learning
  15. Dataset Distillation with Neural Characteristic Function: A Minmax Perspective
  16. Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
  17. Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration
  18. GRAIT: Gradient-Driven Refusal-Aware Instruction Tuning for Effective Hallucination Mitigation
  19. GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
  20. Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
  21. Harnessing Diversity for Important Data Selection in Pretraining Large Language Models
  22. IPDreamer: Appearance-Controllable 3D Object Generation with Complex Image Prompts
  23. Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching
  24. LEGION: Learning to Ground and Explain for Synthetic Image Detection
  25. LEMMA: Learning from Errors for MatheMatical Advancement in LLMs
  26. LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models
  27. Large Language Models Meet Symbolic Provers for Logical Reasoning Evaluation
  28. Leveraging BEV Paradigm for Ground-to-Aerial Image Synthesis
  29. MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models
  30. MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion
  31. Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models
  32. MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer
  33. Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop Learning
  34. Multi-step Visual Reasoning with Visual Tokens Scaling and Verification
  35. OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
  36. OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
  37. OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
  38. OpenHuEval: Evaluating Large Language Model on Hungarian Specifics
  39. Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
  40. SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
  41. Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact Explanation
  42. Stop Looking for "Important Tokens" in Multimodal Language Models: Duplication Matters More
  43. SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
  44. Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?
  45. UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
  46. Utilize the Flow Before Stepping into the Same River Twice: Certainty Represented Knowledge Flow for Refusal-Aware Instruction Tuning
  47. VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis
  48. VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
  49. Where am I? Cross-View Geo-localization with Natural Language Descriptions
  50. 3D Building Reconstruction from Monocular Remote Sensing Images with Multi-level Supervisions
  51. Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
  52. Cross-View Image Geo-Localization with Panorama-BEV Co-retrieval Network
  53. InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
  54. LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-Training
  55. LOCR: Location-Guided Transformer for Optical Character Recognition
  56. LongWanjuan: Towards Systematic Measurement for Long Text Quality
  57. MMBench: Is Your Multi-modal Model an All-Around Player?
  58. OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
  59. Parrot Captions Teach CLIP to Spot Text
  60. ProtLLM: An Interleaved Protein-Language LLM with Protein-as-Word Pre-Training
  61. SG-BEV: Satellite-Guided BEV Fusion for Cross-View Semantic Segmentation
  62. SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
  63. ShareGPT4V: Improving Large Multi-modal Models with Better Captions
  64. VIGC: Visual Instruction Generation and Correction