PPaperPicks

Wenxuan Wang

Chinese University of Hong Kong, Department of Computer Science and Engineering, Hong Kong

50 papers at tracked venues · 42 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. A Survey of Deep Learning for Geometry Problem Solving
  2. A Survey of Large Models in Sports
  3. AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor
  4. Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
    ACL 2026 · Wenxuan Wang
  5. ChartEditor: A Reinforcement Learning Framework for Robust Chart Editing
  6. Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
  7. Exploring Attention Attractors in Large Language Models
  8. Identifying the Achilles' Heel: An Iterative Method for Uncovering Factual Errors in Large Language Models
    ACL 2026 · Wenxuan Wang
  9. Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification
  10. JARVIS or Ultron? A Survey on the Safety and Security Threats of Computer-Using Agents
  11. LongMP-Bench: A Benchmark for Multimodal Persona Understanding in Long-Term Dialogues
  12. MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis
  13. Mem-PAL: Towards Memory-based Personalized Dialogue Assistants for Long-term User-Agent Interaction
  14. POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
  15. Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation
  16. Probing Semantic Insensitivity for Inference-Time Backdoor Defense in Multimodal Large Language Model
  17. Social Welfare Function Leaderboard: On the Emergence of LLM Agents as the Welfare Dictator
  18. A Survey of LLM-based Agents in Medicine: How far are we from Baymax?
    ACL 2025 · Wenxuan Wang
  19. AI Sees Your Location - But With A Bias Toward The Wealthy World
  20. Asclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models
  21. Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
    ACL 2025 · Wenxuan Wang
  22. Chain-of-Jailbreak Attack for Image Generation Models via Step by Step Editing
    ACL 2025 · Wenxuan Wang
  23. ChartM3: Benchmarking Chart Editing with Multimodal Instructions
  24. Competing Large Language Models in Multi-Agent Gaming Environments
  25. EAGLE: Expert-Guided Self-Enhancement for Preference Alignment in Pathology Large Vision-Language Model
  26. Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
  27. IntentionESC: An Intention-Centered Framework for Enhancing Emotional Support in Dialogue Systems
  28. Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
  29. Learning to Ask: When LLM Agents Meet Unclear Instruction
    EMNLP 2025 · Wenxuan Wang
  30. MedChain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive Sequence
  31. On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents
  32. QueryAttack: Jailbreaking Aligned Large Language Models Using Structured Non-natural Query Language
  33. Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
  34. Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
  35. ToolSafety: A Comprehensive Dataset for Enhancing Safety in LLM-Based Agent Tool Invocations
  36. Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
  37. Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
  38. Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
  39. VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
  40. VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models
  41. Where Fact Ends and Fairness Begins: Redefining AI Bias Evaluation through Cognitive Biases
  42. All Languages Matter: On the Multilingual Safety of LLMs
    ACL 2024 · Wenxuan Wang
  43. Apathetic or Empathetic? Evaluating LLMs' Emotional Alignments with Humans
  44. Boosting Adversarial Transferability by Block Shuffle and Rotation
  45. GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
  46. LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
  47. New Job, New Gender? Measuring the Social Bias in Image Generation Models
    ACM MM 2024 · Wenxuan Wang
  48. Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language Models
    ACL 2024 · Wenxuan Wang
  49. On the Humanity of Conversational AI: Evaluating the Psychological Portrayal of LLMs
  50. On the Reliability of Psychological Scales on Large Language Models