PPaperPicks

Huaimin Wang

National University of Defense Technology (NUDT), State Key Laboratory of Complex and Critical Software Environment, Changsha, China

16 papers at tracked venues · 14 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Divergence or Convergence? A Deep Insight into the Crowd Collaboration and its Productivity in Open Source Software based on Entropy
  2. Are External Contributions Important to Project Productivity in Open Source Software? A Deep Insight based on Issue Entropy
  3. Empowering Large Language Model Agent through Step-Level Self-Critique and Self-Training
  4. Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models
  5. Improving the Continuity of Goal-Achievement Ability via Policy Self-Regularization for Goal-Conditioned Reinforcement Learning
  6. Knowledge Memorization and Rumination for Pre-trained Model-based Class-Incremental Learning
  7. Maintaining Fairness in Logit-based Knowledge Distillation for Class-Incremental Learning
  8. Preference-Strength-Aware Self-Improving Alignment with Generative Preference Models
  9. V-Pilot: A Velocity Vector Control Agent for Fixed-Wing UAVs from Imperfect Demonstrations
  10. VVC-Gym: A Fixed-Wing UAV Reinforcement Learning Environment for Multi-Goal Long-Horizon Problems
  11. Demonstrative Instruction Following in Multimodal LLMs via Integrating Low-Rank Adaptation with Ensemble Learning
  12. Goal-Conditioned On-Policy Reinforcement Learning
  13. Iterative Regularized Policy Optimization with Imperfect Demonstrations
  14. Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
  15. Stabilizing Zero-Shot Prediction: A Novel Antidote to Forgetting in Continual Vision-Language Tasks
  16. Tracing Training Progress: Dynamic Influence Based Selection for Active Learning