PPaperPicks

Huazheng Wang

27 papers at tracked venues · 19 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Bridging the Tokenizer Gap: Semantics and Distribution-aware Knowledge Transfer for Unbiased Cross-Tokenizer Distillation
    AAAI 2026 · Huazheng Wang
  2. Erasing Without Remembering: Implicit Knowledge Forgetting in Large Language Models
    ACL 2026 · Huazheng Wang
  3. Example Quality Matters: Multi-Aspects Example Augmentation for Private Library Programming
  4. A Common Pitfall of Margin-based Language Model Alignment: Gradient Entanglement
  5. Design-Based Bandits Under Network Interference: Trade-Off Between Regret and Statistical Inference
  6. Divide, Optimize, Merge: Scalable Fine-Grained Generative Optimization for LLM Agents
  7. Do LVLMs Truly Understand Video Anomalies? Revealing Hallucination via Co-Occurrence Patterns
  8. Efficient and Robust Reinforcement Learning from Human Feedback
    AAAI 2025 · Huazheng Wang
  9. Evaluating and Mitigating Object Hallucination in Large Vision-Language Models: Can They Still See Removed Objects?
  10. FCOM: A Federated Collaborative Online Monitoring Framework via Representation Learning
  11. Provably Efficient Algorithm for Best Scoring Rule Identification in Online Principal-Agent Information Acquisition
  12. SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement
  13. The Ranking Blind Spot: Decision Hijacking in LLM-based Text Ranking
  14. The Threat of PROMPTS in Large Language Models: A System and User Prompt Perspective
  15. TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
  16. Unveiling Internal Reasoning Modes in LLMs: A Deep Dive into Latent Reasoning vs. Factual Shortcuts with Attribute Rate Ratio
  17. Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
  18. Adversarial Attacks on Combinatorial Multi-Armed Bandits
  19. Conversational Dueling Bandits in Generalized Linear Models
  20. MDR: Model-Specific Demonstration Retrieval at Inference Time for In-Context Learning
    NAACL 2024 · Huazheng Wang
  21. PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human Feedback
  22. Pure Exploration in Asynchronous Federated Bandits
  23. RA-PbRL: Provably Efficient Risk-Aware Preference-Based Reinforcement Learning
  24. SSS: Editing Factual Knowledge in Language Models towards Semantic Sparse Space
    ACL 2024 · Huazheng Wang
  25. Stealthy Adversarial Attacks on Stochastic Multi-Armed Bandits
  26. Tree Search-Based Evolutionary Bandits for Protein Sequence Optimization
  27. Video Anomaly Detection via Progressive Learning of Multiple Proxy Tasks