PPaperPicks

Ju Fan

31 papers at tracked venues · 31 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Data Agents: Levels, State of the Art, and Open Problems
  2. Revisiting Single-Table Retrieval: An Open Problem Under 360° Stress Tests
  3. Reward-SQL: Boosting Text-to-SQL via Stepwise Execution-Aware Reasoning and Process-Supervised Rewards
  4. TACO: A Benchmark for Open-Domain Text-to-SQL with Ambiguous and Cross-Database Queries
  5. Telescope: A Learned What-If Call for Column Store Selection in HTAP Databases
  6. VecBench: A Controllable Benchmark for Filtered Vector Search: [Experiments & Analysis]
  7. Alpha-SQL: Zero-Shot Text-to-SQL using Monte Carlo Tree Search
  8. Andromeda: Debugging Database Performance Issues with Retrieval-Augmented Large Language Models
  9. AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework
  10. Automatic Database Configuration Debugging using Retrieval-Augmented Language Models
  11. CloudyBench: A Testbed for A Comprehensive Evaluation of Cloud-Native Databases
  12. Data Imputation with Limited Data Redundancy Using Data Lakes
  13. Harnessing Diversity for Important Data Selection in Pretraining Large Language Models
  14. Natural Language to SQL: State of the Art and Open Problems
  15. PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking
  16. TANDEM: Bi-Level Data Mixture Optimization with Twin Networks
  17. Weak-to-Strong Prompts with Lightweight-to-Powerful LLMs for High-Accuracy, Low-Cost, and Explainable Data Transformation
  18. A Multi-Task Learning Framework for Reading Comprehension of Scientific Tabular Data
  19. Auto-Formula: Recommend Formulas in Spreadsheets using Contrastive Learning for Table Representations
  20. ChatPipe: Orchestrating Data Preparation Pipelines by Optimizing Human-ChatGPT Interactions
  21. CodeS: Towards Building Open-source Language Models for Text-to-SQL
  22. Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL
    VLDB 2024 · Ju Fan
  23. Controllable Tabular Data Synthesis Using Diffusion Models
  24. Cost-Effective In-Context Learning for Entity Resolution: A Design Space Exploration
  25. DINGO: Towards Diverse and Fine-Grained Instruction-Following Evaluation
  26. Front Matter
  27. IDE: A System for Iterative Mislabel Detection
  28. Improving Graph Compression for Efficient Resource-Constrained Graph Analytics
  29. MisDetect: Iterative Mislabel Detection using Early Loss
  30. Mitigating Data Scarcity in Supervised Machine Learning Through Reinforcement Learning Guided Data Generation
  31. Representation Learning for Entity Alignment in Knowledge Graph: A Design Space Exploration