PPaperPicks

Zhiqing Sun

11 papers at tracked venues · 10 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
  2. Improve Vision Language Model Chain-of-thought Reasoning
  3. Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for LLM Problem-Solving
  4. Lean-STaR: Learning to Interleave Thinking and Proving
  5. Self-Play Preference Optimization for Language Model Alignment
  6. Aligning Large Multimodal Models with Factually Augmented RLHF
    ACL 2024 · Zhiqing Sun
  7. Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
    NeurIPS 2024 · Zhiqing Sun
  8. Instruction-tuned Language Models are Better Knowledge Learners
  9. LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery
  10. SALMON: Self-Alignment with Instructable Reward Models
    ICLR 2024 · Zhiqing Sun
  11. Visual Chain-of-Thought Prompting for Knowledge-Based Visual Reasoning