PPaperPicks

Michael R. Lyu

Chinese University of Hong Kong, Department of Computer Science and Engineering, Hong Kong

29 papers at tracked venues · 21 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level
  2. Identifying the Achilles' Heel: An Iterative Method for Uncovering Factual Errors in Large Language Models
  3. Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification
  4. Understanding Secret Leakage Risks in Code LLMs: A Tokenization Perspective
  5. Asclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models
  6. CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code Reasoning
  7. Competing Large Language Models in Multi-Agent Gaming Environments
  8. C²LEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation
  9. Learning to Ask: When LLM Agents Meet Unclear Instruction
  10. Learning to Reason from Feedback at Test-Time
  11. MedChain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive Sequence
  12. On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents
  13. SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design
  14. UniDebugger: Hierarchical Multi-Agent Framework for Unified Software Debugging
  15. Where Fact Ends and Fairness Begins: Redefining AI Bias Evaluation through Cognitive Biases
  16. All Languages Matter: On the Multilingual Safety of LLMs
  17. Apathetic or Empathetic? Evaluating LLMs' Emotional Alignments with Humans
  18. Beyond Embeddings: The Promise of Visual Table in Visual Reasoning
  19. Curvature-Invariant Adversarial Attacks for 3D Point Clouds
  20. Enhancing Temporal Modeling of Video LLMs via Time Gating
  21. Improving the Adversarial Transferability of Vision Transformers with Virtual Dense Connection
  22. LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
  23. Make Your Home Safe: Time-aware Unsupervised User Behavior Anomaly Detection in Smart Homes via Loss-guided Mask
  24. Making Long-Context Language Models Better Multi-Hop Reasoners
  25. New Job, New Gender? Measuring the Social Bias in Image Generation Models
  26. Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language Models
  27. On the Humanity of Conversational AI: Evaluating the Psychological Portrayal of LLMs
  28. On the Reliability of Psychological Scales on Large Language Models
  29. SPES: Towards Optimizing Performance-Resource Trade-Off for Serverless Functions