PPaperPicks

Wei Xu

Georgia Institute of Technology, School of Interactive Computing, Atlanta, GA, USA

28 papers at tracked venues · 15 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
  2. GeoRC: A Benchmark for Geolocation Reasoning Chains
  3. Supporting Informed Self-Disclosure: Design Recommendations for Presenting AI-Estimates of Privacy Risks to Users
  4. Third Workshop on Human-Centered Evaluation and Auditing of Language Models: AI Agents-in-the-Loop
  5. CARE: Multilingual Human Preference Learning for Cultural Awareness
  6. CROSSNEWS: A Cross-Genre Authorship Verification and Attribution Benchmark
  7. Generating CAD Code with Vision-Language Models for 3D Designs
  8. How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation
  9. Human-Centered Evaluation and Auditing of Language Models
  10. Measuring, Modeling, and Helping People Account for Privacy Risks in Online Self-Disclosures with AI
  11. On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena
  12. SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?
  13. The Impact of Visual Information in Chinese Characters: Evaluating Large Models' Ability to Recognize and Utilize Radicals
  14. What are Foundation Models Cooking in the Post-Soviet World?
  15. Automatic and Human-AI Interactive Text Generation (with a focus on Text Simplification and Revision)
  16. ChatHF: Collecting Rich Human Feedback from Real-time Conversations
  17. Constrained Decoding for Cross-lingual Label Projection
  18. FactPICO: Factuality Evaluation for Plain Language Summarization of Medical Evidence
  19. GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation
  20. Granular Privacy Control for Geolocation with Vision Language Models
  21. Having Beer after Prayer? Measuring Cultural Bias in Large Language Models
  22. Improving Minimum Bayes Risk Decoding with Multi-Prompt
  23. InfoLossQA: Characterizing and Recovering Information Loss in Text Simplification
  24. MedReadMe: A Systematic Study for Fine-grained Sentence Readability in Medical Domain
  25. Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
  26. NEO-BENCH: Evaluating Robustness of Large Language Models with Neologisms
  27. ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability Assessment
  28. Reducing Privacy Risks in Online Self-Disclosures with Language Models