PPaperPicks

Baotian Hu

33 papers at tracked venues · 25 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering
  2. ComfyFlow: Benchmarking LLMs for AIGC Workflow Generation
  3. ComfyUI-R1: Exploring Reasoning Models for Workflow Generation
  4. Dynamic Long Context Reasoning over Compressed Memory via End-to-End Reinforcement Learning
  5. Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs
  6. Improving Value-based Process Verifier via Low-Cost Variance Reduction
  7. Learning to Extract Rational Evidence via Reinforcement Learning for Retrieval-Augmented Generation
  8. Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework
  9. LycheeCluster: Efficient Long-Context Inference with Structure-Aware Chunking and Hierarchical KV Indexing
  10. MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation
  11. Structured Episodic Event Memory
  12. ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution
  13. UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity Mixture-of-Experts
  14. WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments
  15. A Unified Agentic Framework for Evaluating Conditional Image Generation
  16. Advancing Temporal Sensitive Question Answering through Progressive Multi-Step Reflection
  17. CMT: A Memory Compression Method for Continual Knowledge Learning of Large Language Models
  18. ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development
  19. FunnelRAG: A Coarse-to-Fine Progressive Retrieval Paradigm for RAG
  20. MeKB-Sim: Personal Knowledge Base-Powered Multi-Agent Simulation
  21. VideoVista-CulturalLingo: 360° Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension
  22. Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge Alignment
  23. Improving Attributed Text Generation of Large Language Models via Preference Learning
  24. In-Context Learning State Vector with Inner and Momentum Optimization
  25. Medico: Towards Hallucination Detection and Correction with Multi-source Evidence Fusion
  26. MultiSkill: Evaluating Large Multimodal Models for Fine-grained Alignment Skills
  27. SEER: Self-Aligned Evidence Extraction for Retrieval-Augmented Generation
  28. SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-Reflection
  29. Separate the Wheat from the Chaff: Model Deficiency Unlearning via Parameter-Efficient Module Operation
  30. Take Off the Training Wheels! Progressive In-Context Learning for Effective Alignment
  31. Temporal Knowledge Question Answering via Abstract Reasoning Induction
  32. TruthReader: Towards Trustworthy Document Assistant Chatbot with Reliable Attribution
  33. VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context