PPaperPicks

Xing Xie

Microsoft Research Asia, Beijing, China

63 papers at tracked venues · 54 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Can AI Revise Research Papers with Human Review Feedback? An Empirical Study and Benchmark
  2. Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural Alignment
  3. Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations
  4. DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection
  5. From Passive Consumption to Active Interaction: Exploring Interactive LLM Scaffolding to Support Learning Engagement
  6. How Do Human Creators Embrace Human-AI Co-Creation? A Perspective on Human Agency of Screenwriters
  7. HumanLLM: Towards Personalized Understanding and Simulation of Human Nature
  8. IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
  9. Influence-based Online Experience Selection for Effective RLHF
  10. Interaction-Augmented Instruction: Modeling the Synergy of Prompts and Interactions in Human-GenAI Collaboration
  11. MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation
  12. Measuring Human Contribution in AI-Assisted Content Generation
  13. MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
  14. PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models
  15. Adversarial Style Augmentation via Large Language Model for Robust Fake News Detection
  16. Benchmarking and Defending against Indirect Prompt Injection Attacks on Large Language Models
  17. CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds
  18. Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
  19. LLM-powered Multi-agent Framework for Goal-oriented Learning in Intelligent Tutoring System
  20. MoVa: Towards Generalizable Classification of Human Morals and Values
  21. MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
  22. Pretraining Context Compressor for Large Language Models with Embedding-Based Memory
  23. Prompting Generative AI with Interaction-Augmented Instructions
  24. Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
  25. Specify Privacy Yourself: Assessing Inference-Time Personalized Privacy Preservation Ability of Large Vision-Language Models
  26. Towards Better Value Principles for Large Language Model Alignment: A Systematic Evaluation and Enhancement
  27. TrendSim: Simulating Trending Topics in Social Media Under Poisoning Attacks with LLM-based Multi-agent System
  28. Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
  29. Unveiling the Learning Mind of Language Models: A Cognitive Framework and Empirical Study
  30. Value Compass Benchmarks: A Comprehensive, Generative and Self-Evolving Platform for LLMs' Value Evaluation
  31. A Data-Centric Multi-Objective Learning Framework for Responsible Recommendation Systems
  32. A General Framework for Learning from Weak Supervision
  33. Ada-Retrieval: An Adaptive Multi-Round Retrieval Paradigm for Sequential Recommendations
  34. Aligning Language Models for Versatile Text-based Item Retrieval
  35. Aligning Large Language Models for Controllable Recommendations
  36. CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses
  37. CompeteAI: Understanding the Competition Dynamics of Large Language Model-based Agents
  38. CultureLLM: Incorporating Cultural Differences into Large Language Models
  39. CulturePark: Boosting Cross-cultural Understanding in Large Language Models
  40. Denevil: towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning
  41. DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks
  42. Dynamic Evaluation of Large Language Models by Meta Probing Agents
  43. ERBench: An Entity-Relationship based Automatically Verifiable Hallucination Benchmark for Large Language Models
  44. Embedding an Ethical Mind: Aligning Text-to-Image Synthesis via Lightweight Value Optimization
  45. Foundation Model-oriented Robustness: Robust Image Model Evaluation with Pretrained Models
  46. High-Frequency-aware Hierarchical Contrastive Selective Coding for Representation Learning on Text Attributed Graphs
  47. IRGen: Generative Modeling for Image Retrieval
  48. Imprecise Label Learning: A Unified Framework for Learning with Various Imprecise Label Configurations
  49. KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models
  50. MS MARCO Web Search: A Large-scale Information-rich Web Dataset with Millions of Real Click Labels
  51. Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
  52. On the Essence and Prospect: An Investigation of Alignment Approaches for Big Models
  53. On the Vulnerability of Safety Alignment in Open-Access LLMs
  54. PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
  55. Position: TrustLLM: Trustworthiness in Large Language Models
  56. RecAI: Leveraging Large Language Models for Next-Generation Recommender Systems
  57. RecExplainer: Aligning Large Language Models for Explaining Recommendation Models
  58. SpecFormer: Guarding Vision Transformer Robustness via Maximum Singular Value Penalization
  59. Supervised Knowledge Makes Large Language Models Better In-context Learners
  60. The Good, The Bad, and Why: Unveiling Emotions in Generative AI
  61. Towards Optimization and Model Selection for Domain Generalization: A Mixup-guided Solution
  62. Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
  63. Value FULCRA: Mapping Large Language Models to the Multidimensional Spectrum of Basic Human Value