PPaperPicks

William Yang Wang

University of California, Santa Barbara, USA

60 papers at tracked venues · 42 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Can Editing LLMs Inject Harm?
  2. Detecting Training Data of Large Language Models via Expectation Maximization
  3. LEDOM: Reverse Language Model
  4. WildSci: Advancing Scientific Reasoning from In-the-Wild Literature
  5. AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge
  6. Aristotle: Mastering Logical Reasoning with A Logic-Complete Decompose-Search-Resolve Framework
  7. BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
  8. CBT-Bench: Evaluating Large Language Models on Assisting Cognitive Behavior Therapy
  9. Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
  10. DebUnc: Improving Large Language Model Agent Communication With Uncertainty Metrics
  11. Disentangling Memory and Reasoning Ability in Large Language Models
  12. Do You Know About My Nation? Investigating Multilingual Language Models' Cultural Literacy Through Factual Knowledge
  13. Dynamic Evaluation for Oversensitivity in LLMs
  14. Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
  15. Gödel Agent: A Self-Referential Agent Framework for Recursively Self-Improvement
  16. How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
  17. Human Bias in the Face of AI: Examining Human Judgment Against Text Labeled as AI Generated
  18. InductionBench: LLMs Fail in the Simplest Complexity Class
  19. Investigating the Transferability of Code Repair for Low-Resource Programming Languages
  20. MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
  21. MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
  22. MuSLR: Multimodal Symbolic Logical Reasoning
  23. REALM: A Dataset of Real-World LLM Use Cases
  24. RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
  25. SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement
  26. Scaling LLM Inference Efficiently with Optimized Sample Compute Allocation
  27. Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling
  28. T2V-Turbo-v2: Enhancing Video Model Post-Training through Data, Reward, and Conditional Guidance Design
  29. TC-Bench: Benchmarking Temporal Compositionality in Conditional Video Generation
  30. Uncovering Factor-Level Preference to Improve Human-Model Alignment
  31. Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoning
  32. VSP: Diagnosing the Dual Challenges of Perception and Reasoning in Spatial Planning Tasks for MLLMS
  33. Weak-to-Strong Jailbreaking on Large Language Models
  34. A Survey on Detection of LLMs-Generated Content
  35. AKEW: Assessing Knowledge Editing in the Wild
  36. BPO: Staying Close to the Behavior LLM Creates Better Online LLM Alignment
  37. DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text
  38. FASTopic: Pretrained Transformer is a Fast, Adaptive, Stable, and Transferable Topic Model
  39. Global Human-guided Counterfactual Explanations for Molecular Properties via Reinforcement Learning
  40. Guiding Instruction-based Image Editing via Multimodal Large Language Models
  41. Hire a Linguist!: Learning Endangered Languages in LLMs with In-Context Linguistic Descriptions
  42. Knowledge of Knowledge: Exploring Known-Unknowns Uncertainty with Large Language Models
  43. LLMRefine: Pinpointing and Refining Large Language Models via Fine-Grained Actionable Feedback
  44. Language Control Diffusion: Efficiently Scaling through Space, Time, and Tasks
  45. Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts
  46. Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts
  47. Mastering Robot Manipulation with Multimodal Prompts through Pretraining and Multi-task Fine-tuning
  48. MultiAgent Collaboration Attack: Investigating Adversarial Attacks in Large Language Model Collaborations via Debate
  49. Multimodal Procedural Planning via Dual Text-Image Prompting
  50. Neuroformer: Multimodal and Multitask Generative Pretraining for Brain Data
  51. Position: AI/ML Influencers Have a Place in the Academic Process
  52. Position: TrustLLM: Trustworthiness in Large Language Models
  53. Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
  54. RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question Answering
  55. T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback
  56. The Knowledge Alignment Problem: Bridging Human and External Knowledge for Large Language Models
  57. Understanding Reasoning Ability of Language Models From the Perspective of Reasoning Paths Aggregation
  58. VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View
  59. Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)
  60. WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences