P
PaperPicks
Conferences
Zhiqing Sun
11 papers at tracked venues · 10 at CORE A* · active 2024–2025
DBLP profile ↗
ORCID search ↗
Venues
ICLR
×4
ACL
×3
AAAI
×1
ICML
×1
NAACL
×1
NeurIPS
×1
Frequent coauthors
Ruohong Zhang
DBLP profile ↗
ORCID search ↗
×2
Yangzhen Wu
DBLP profile ↗
ORCID search ↗
×1
Haohan Lin
DBLP profile ↗
ORCID search ↗
×1
Yue Wu
DBLP profile ↗
ORCID search ↗
×1
Zhengbao Jiang
DBLP profile ↗
ORCID search ↗
×1
Pingchuan Ma
DBLP profile ↗
ORCID search ↗
×1
Zhenfang Chen
DBLP profile ↗
ORCID search ↗
×1
Papers
Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
NAACL 2025
·
Ruohong Zhang
DBLP profile ↗
ORCID search ↗
Improve Vision Language Model Chain-of-thought Reasoning
ACL 2025
·
Ruohong Zhang
DBLP profile ↗
ORCID search ↗
Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for LLM Problem-Solving
ICLR 2025
·
Yangzhen Wu
DBLP profile ↗
ORCID search ↗
Lean-STaR: Learning to Interleave Thinking and Proving
ICLR 2025
·
Haohan Lin
DBLP profile ↗
ORCID search ↗
Self-Play Preference Optimization for Language Model Alignment
ICLR 2025
·
Yue Wu
DBLP profile ↗
ORCID search ↗
Aligning Large Multimodal Models with Factually Augmented RLHF
ACL 2024
·
Zhiqing Sun
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
NeurIPS 2024
·
Zhiqing Sun
Instruction-tuned Language Models are Better Knowledge Learners
ACL 2024
·
Zhengbao Jiang
DBLP profile ↗
ORCID search ↗
LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery
ICML 2024
·
Pingchuan Ma
DBLP profile ↗
ORCID search ↗
SALMON: Self-Alignment with Instructable Reward Models
ICLR 2024
·
Zhiqing Sun
Visual Chain-of-Thought Prompting for Knowledge-Based Visual Reasoning
AAAI 2024
·
Zhenfang Chen
DBLP profile ↗
ORCID search ↗