PPaperPicks

Diyi Yang

65 papers at tracked venues · 46 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
  2. Can LLM-Simulated Practice and Feedback Upskill Human Counselors? A Randomized Study with 90+ Novice Counselors
  3. Future of Work in the Age of LLMs
  4. Generative Interfaces for Language Models
  5. Just-In-Time Objectives: A General Approach for Specialized AI Interactions
  6. Mapping the Spiral of Silence: Surveying Unspoken Opinions in Online Communities
  7. Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment
  8. Verbalizing LLMs' Assumptions About the User to Calibrate Expectations and Reduce Sycophancy
  9. Whose Knowledge Counts? Co-Designing Community-Centered AI Auditing Tools with Educators in Hawai'i
  10. Aligning Language Models with Demonstrated Feedback
  11. Attacking Vision-Language Computer Agents via Pop-ups
  12. Bidirectional Human-AI Alignment: Emerging Challenges and Opportunities
  13. Blackbox Model Provenance via Palimpsestic Membership Inference
  14. Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
  15. Creating General User Models from Computer Use
  16. Culture Cartography: Mapping the Landscape of Cultural Knowledge
  17. Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
  18. Distilling an End-to-End Voice Assistant Without Instruction Training Data
  19. EgoNormia: Benchmarking Physical-Social Norm Understanding
  20. EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking
  21. Helping the Helper : Supporting Peer Counselors via AI-Empowered Practice and Feedback
  22. Human-AI Collaboration: How AIs Augment Human Teammates
  23. Identifying Unlearned Data in LLMs via Membership Inference Attacks
  24. Information Retrieval Induced Safety Degradation in AI Agents
  25. Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
  26. Knoll: Creating a Knowledge Ecosystem for Large Language Models
  27. Mind the Gap: Static and Interactive Evaluations of Large Audio Models
  28. No Preference Left Behind: Group Distributional Preference Optimization
  29. OpenCUA: Open Foundations for Computer-Use Agents
  30. Position: Towards Bidirectional Human-AI Alignment
  31. SPHERE: An Evaluation Card for Human-AI Systems
  32. SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
  33. SWE-smith: Scaling Data for Software Engineering Agents
  34. Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design Prototyping
  35. StorySage: Conversational Autobiography Writing Powered by a Multi-Agent Framework
  36. SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs
  37. The Practice of Online Peer Counseling and the Potential for AI-Powered Support Tools
  38. When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration
  39. Are Large Language Models Consistent over Value-laden Questions?
  40. Benchmarking Machine Translation with Cultural Awareness
  41. CultureBank: An Online Community-Driven Knowledge Base Towards Culturally Aware Language Technologies
  42. DARG: Dynamic Evaluation of Large Language Models via Adaptive Reasoning Graph
  43. Decoding Susceptibility: Modeling Misbelief to Misinformation Through a Computational Approach
  44. Demystifying Verbatim Memorization in Large Language Models
  45. DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks
  46. Grounding Gaps in Language Model Generations
  47. How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
  48. Language Agents: Foundations, Prospects, and Risks
  49. MIDDAG: Where Does Our News Go? Investigating Information Diffusion via Community-Level Information Pathways
  50. Measuring and Addressing Indexical Bias in Information Retrieval
  51. Modeling Gender and Dialect Bias in Automatic Speech Recognition
  52. Multi-Level Feedback Generation with Large Language Models for Empowering Novice Peer Counselors
  53. Perceptions of Language Technology Failures from South Asian English Speakers
  54. Position: A Safe Harbor for AI Evaluation and Red Teaming
  55. PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
  56. Rehearsal: Simulating Conflict to Teach Conflict Resolution
  57. Roleplay-doh: Enabling Domain-Experts to Create LLM-simulated Patients via Eliciting and Adhering to Principles
  58. Semi-Truths: A Large-Scale Dataset of AI-Augmented Images for Evaluating Robustness of AI-Generated Image detectors
  59. Silent Signals, Loud Impact: LLMs for Word-Sense Disambiguation of Coded Dog Whistles
  60. Simulated Misinformation Susceptibility (SMISTS): Enhancing Misinformation Research with Large Language Model Simulations
  61. Social Intelligence Data Infrastructure: Structuring the Present and Navigating the Future
  62. Training Socially Aligned Language Models on Simulated Social Interactions
  63. Understanding Online Discussion Across Difference: Insights from Gun Discourse on Reddit
  64. Unintended Impacts of LLM Alignment on Global Representation
  65. What Makes Digital Support Effective? How Therapeutic Skills Affect Clinical Well-Being