PPaperPicks

Tongshuang Wu

Carnegie Mellon University, PA, USA

26 papers at tracked venues · 19 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Behavioral Indicators of Overreliance During Interaction with Conversational Language Models
  2. Evidotes: Integrating Scientific Evidence and Anecdotes to Support Uncertainties Triggered by Peer Health Posts
  3. From Human-Human Collaboration to Human-Agent Collaboration: A Vision, Design Philosophy, and an Empirical Framework for Achieving Successful Partnerships Between Humans and LLM Agents
  4. Not Everyone Wins with LLMs: Behavioral Patterns and Pedagogical Implications for AI Literacy in Programmatic Data Science
  5. RECAP: An End-to-End Platform for Capturing, Replaying, and Analyzing AI-Assisted Programming Interactions
  6. Scaling Collaborative Effort with Agents
  7. What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts
  8. Bidirectional Human-AI Alignment: Emerging Challenges and Opportunities
  9. Checklists Are Better Than Reward Models For Aligning Language Models
  10. Evaluating Mathematical Reasoning Beyond Accuracy
  11. How to Teach Programming in the AI Era? Using LLMs as a Teachable Agent for Debugging (Extended Abstract)
  12. Human Subjects Research in the Age of Generative AI: Opportunities and Challenges of Applying LLM-Simulated Data to HCI Studies
  13. Human-AI Collaboration: How AIs Augment Human Teammates
    ACL 2025 · Sherry Wu
  14. LLMs as Workers in Human-Computational Algorithms? Replicating Crowdsourcing Pipelines with LLMs
    CHI 2025 · Tongshuang Wu
  15. MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrievers
  16. Orbit: A Framework for Designing and Evaluating Multi-objective Rankers
  17. SOTOPIA-S4: a user-friendly system for flexible, customizable, and large-scale social simulation
  18. SPHERE: An Evaluation Card for Human-AI Systems
  19. cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Tree
  20. Better Synthetic Data by Retrieving and Transforming Existing Datasets
  21. Fact-and-Reflection (FaR) Improves Confidence Calibration of Large Language Models
  22. Large Language Models Help Humans Verify Truthfulness - Except When They Are Convincingly Wrong
  23. Selenite: Scaffolding Online Sensemaking with Comprehensive Overviews Elicited from Large Language Models
  24. Synthetic Multimodal Question Generation
  25. Trust and Reliance in Evolving Human-AI Workflows (TREW)
  26. Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia