PPaperPicks

Zhenguo Li

60 papers at tracked venues · 49 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Empowering Sparse-Input Neural Radiance Fields with Dual-Level Semantic Guidance from Dense Novel Views
  2. MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes
  3. Accelerating Auto-regressive Text-to-Image Generation with Training-free Speculative Jacobi Decoding
  4. Adding Additional Control to One-Step Diffusion with Joint Distribution Matching
  5. Automated Evaluation of Large Vision-Language Models on Self-Driving Corner Cases
  6. Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning
  7. CARTS: Advancing Neural Theorem Proving with Diversified Tactic Calibration and Bias-Resistant Tree Search
  8. Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction Tuning
  9. DAPE V2: Process Attention Score as Feature Map for Length Extrapolation
  10. EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
  11. FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
  12. Forewarned is Forearmed: Harnessing LLMs for Data Synthesis via Failure-induced Exploration
  13. G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model
  14. Getting More Juice Out of Your Data: Hard Pair Refinement Enhances Visual-Language Models Without Extra Data
  15. How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training
  16. How Numerical Precision Affects Arithmetical Reasoning Capabilities of LLMs
  17. Implicit Search via Discrete Diffusion: A Study on Chess
  18. Improved Diffusion-based Generative Model with Better Adversarial Robustness
  19. Jailbreaking as a Reward Misspecification Problem
  20. LiT: Delving into a Simple Linear Diffusion Transformer for Image Generation
  21. MESH - Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
  22. MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control
  23. Masked Diffusion Models as Energy Minimization
  24. Mixture of insighTful Experts (MoTE): The Synergy of Reasoning Chains and Expert Mixtures in Self-Alignment
  25. Multi-Objective One-Shot Pruning for Large Language Models
  26. ProofAug: Efficient Neural Theorem Proving via Fine-grained Proof Structure Analysis
  27. Self-Adjust Softmax
  28. SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
  29. Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
  30. SplatMesh: Interactive 3D Segmentation and Editing Using Mesh-Based Gaussian Splatting
  31. T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation
  32. Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism
  33. Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data
  34. ATG: Benchmarking Automated Theorem Generation for Generative Language Models
  35. Accelerating Diffusion Sampling with Optimized Time Steps
  36. CVT-xRF: Contrastive In-Voxel Transformer for 3D Consistent Radiance Fields from Sparse Inputs
  37. DAPE: Data-Adaptive Positional Encoding for Length Extrapolation
  38. DeepAccident: A Motion and Accident Prediction Benchmark for V2X Autonomous Driving
  39. DetCLIPv3: Towards Versatile Generative Open-Vocabulary Object Detection
  40. DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and Perception
  41. Diffusion of Thought: Chain-of-Thought Reasoning in Diffusion Language Models
  42. Dual Risk Minimization: Towards Next-Level Robustness in Fine-tuning Zero-Shot Models
  43. Enhancing the Power of OOD Detection via Sample-Aware Model Selection
  44. Eyes Closed, Safety on: Protecting Multimodal LLMs via Image-to-Text Transformation
  45. Fast Training of Diffusion Transformer with Extreme Masking for 3D Point Clouds Generation
  46. Forward-Backward Reasoning in Large Language Models for Mathematical Verification
  47. Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis
  48. GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing
  49. GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
  50. Implicit Concept Removal of Diffusion Models
  51. LEGO-Prover: Neural Theorem Proving with Growing Libraries
  52. Large Language Models as Automated Aligners for benchmarking Vision-Language Models
  53. MUSTARD: Mastering Uniform Synthesis of Theorem and Proof Data
  54. MagicDrive: Street View Generation with Diverse 3D Geometry Control
  55. MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
  56. PIXART-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
  57. PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
  58. Proving Theorems Recursively
  59. The Surprising Effectiveness of Skip-Tuning in Diffusion Sampling
  60. Towards Understanding the Working Mechanism of Text-to-Image Diffusion Model