PPaperPicks

Bin Cui

Peking University, School of Electronics Engineering and Computer Science, Beijing, China

87 papers at tracked venues · 80 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch
  2. Data-Centric Perspectives on Agentic Retrieval-Augmented Generation: A Survey
  3. DistVec: Efficient Distributed Machine Learning in Parallel Database Systems
  4. Efficient Content-based Recommendation Model Training via Noise-aware Coreset Selection
  5. Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's Nature
  6. Let's Verify Math Questions Step by Step
  7. LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
  8. Med-R2: Crafting Trustworthy LLM Physicians via Retrieval and Reasoning of Evidence-Based Medicine
  9. PSEO: Optimizing Post-hoc Stacking Ensemble Through Hyperparameter Tuning
  10. QA-GraphRAG: Query-Adaptive Plug-and-Play Retrieval Integration for Graph-based Retrieval-Augmented Generation
  11. Rethinking Text-to-SQL: Dynamic Multi-turn SQL Interaction for Real-world Database Exploration
  12. Text2VectorSQL: Towards a Unified Interface for Vector Search and SQL Queries
  13. Text2sql-Flow: a Robust Sql-Aware Data Augmentation Framework for Text-To-Sql
  14. A-Tune-Online: Efficient and QoS-Aware Online Configuration Tuning for Dynamic Workloads
  15. ADEPT-SQL: A High-performance Text-to-SQL Application for Real-World Enterprise-Level Databases
  16. BRACE: A Benchmark for Robust Audio Caption Quality Evaluation
  17. CFBench: A Comprehensive Constraints-Following Benchmark for LLMs
  18. CaliEX: A Disk-Based Large-Scale GNN Training System with Joint Design of Caching and Execution
  19. CuckooGraph: A Scalable and Space-Time Efficient Data Structure for Large-Scale Dynamic Graphs
  20. DataLab: A Unified Platform for LLM-Powered Business Intelligence
  21. DataSculpt: A Holistic Data Management Framework for Long-Context LLMs Training
  22. Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
  23. Enhancing Unsupervised Sentence Embeddings via Knowledge-Driven Data Augmentation and Gaussian-Decayed Contrastive Learning
  24. Extendible RDMA-Based Remote Memory KV Store with Dynamic Perfect Hashing Index
  25. FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback
  26. Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning
  27. HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation
  28. HourglassSketch: An Efficient and Scalable Framework for Graph Stream Summarization
  29. Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment
  30. Improved Accuracy, Declining Orders: Uncovering the Vicious Competition Trap in Multi-Domain Recommendation
  31. Improving Low-Resource Sequence Labeling with Knowledge Fusion and Contextual Label Explanations
  32. IterComp: Iterative Composition-Aware Feedback Learning from Model Gallery for Text-to-Image Generation
  33. LLMs Are Noisy Oracles! LLM-based Noise-aware Graph Active Learning for Node Classification
  34. LobRA: Multi-tenant Fine-tuning over Heterogeneous Data
  35. MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training
  36. Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
  37. Modeling All-Atom Glycan Structures via Hierarchical Message Passing and Multi-Scale Pre-training
  38. NetMoE: Accelerating MoE Training through Dynamic Sample Placement
  39. PAS: Plug-and-Play Prompt Augmentation System
  40. PQCache: Product Quantization-based KVCache for Long Context LLM Inference
  41. PS-MI: Accurate, Efficient, and Private Data Valuation in Vertical Federated Learning
  42. SiriusBI: A Comprehensive LLM-powered Solution for Data Analytics in Business Intelligence
  43. SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget
  44. StarRec: A Hypergraph-based Framework with Star-Expansion for Multi-Behavior Recommendation
  45. SuperCorrect: Advancing Small LLM Reasoning with Thought Template Distillation and Self-Correction
  46. SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
  47. SysBench: Can LLMs Follow System Message?
  48. ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
  49. Towards Scalable and Deep Graph Neural Networks via Noise Masking
  50. Towards Scalable and Efficient Graph Structure Learning
  51. Training Data Distribution Estimation for Optimized Pre-training Data Management
  52. Training-Free Heterogeneous Graph Condensation via Data Selection
  53. Training-Free and Adaptive Sparse Attention for Efficient Long Video Generation
  54. VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs
  55. Accelerating Scalable Graph Neural Network Inference with Node-Adaptive Propagation
  56. Accelerating Text-to-Image Editing via Cache-Enabled Sparse Diffusion Inference
  57. BIM: Improving Graph Neural Networks with Balanced Influence Maximization
  58. BOURNE: Bootstrapped Self-Supervised Learning Framework for Unified Graph Anomaly Detection
  59. Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models
  60. CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation Models
  61. CodingSketch: A Hierarchical Sketch with Efficient Encoding and Recursive Decoding
  62. Cross-Modal Contextualized Diffusion Models for Text-Guided Visual Generation and Editing
  63. Demystifying Data Management for Large Language Models
  64. Efficient Multi-task LLM Quantization and Serving for Multiple LoRA Adapters
  65. HGAMLP: Heterogeneous Graph Attention MLP with De-Redundancy Mechanism
  66. LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
  67. MFIX: An Efficient and Reliable Index Advisor via Multi-Fidelity Bayesian Optimization
  68. Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
  69. Mitigating Negative Transfer in Cross-Domain Recommendation via Knowledge Transferability Enhancement
  70. Multi- View Teacher with Curriculum Data Fusion for Robust Unsupervised Domain Adaptation
  71. MultiLoRA: Multi-Directional Low Rank Adaptation for Multi-Domain Recommendation
  72. NC-ALG: Graph-Based Active Learning Under Noisy Crowd
  73. NPA: Improving Large-scale Graph Neural Networks with Non-parametric Attention
  74. Newton Sketches: Estimating Node Intimacy in Dynamic Graphs Using Newton's Law of Cooling
  75. OUTRE: An OUT-of-core De-REdundancy GNN Training Framework for Massive Graphs within A Single Machine
  76. On-Device Recommender Systems: A Tutorial on The New-Generation Recommendation Paradigm
  77. Online Detection of Outstanding Quantiles with QuantileFilter
  78. Protein-Ligand Interaction Prior for Binding-aware 3D Molecule Diffusion Models
  79. RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion Models
  80. Scalable Overspeed Item Detection in Streams
  81. Structure-Guided Adversarial Training of Diffusion Models
  82. Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling
  83. Unveiling Vulnerabilities of Contrastive Recommender Systems to Poisoning Attacks
  84. VQGraph: Rethinking Graph Representation Space for Bridging GNNs and MLPs
  85. VideoTetris: Towards Compositional Text-to-Video Generation
  86. VisionEmbedder: Bit-Level-Compact Key-Value Storage with Constant Lookup, Rapid Updates, and Rare Failure
  87. X-former Elucidator: Reviving Efficient Attention for Long Context Language Modeling