PPaperPicks

Li Yuan

Peking University, School of Electronic and Computer Engineering, Beijing, China

53 papers at tracked venues · 44 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. 360Explorer: Exploring 4D Controllable World in Panoramic Videos
  2. AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety Basin
  3. Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
  4. NeuralGS: Bridging Neural Fields and 3D Gaussian Splatting for Compact 3D Representations
  5. Next Patch Prediction for AutoRegressive Visual Generation
  6. SAFE-QAQ: End-to-End Slow-Thinking Audio-Text Fraud Detection via Reinforcement Learning
  7. AE-NeRF: Augmenting Event-Based Neural Radiance Fields for Non-ideal Conditions and Larger Scenes
  8. Beyond Chemical QA: Evaluating LLM's Chemical Reasoning with Modular Chemical Operations
  9. CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step
  10. Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle
  11. DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses
  12. E-4DGS: High-Fidelity Dynamic Reconstruction from the Multi-view Event Cameras
  13. Epona: Autoregressive Diffusion World Model for Autonomous Driving
  14. Evagaussians: Event Stream Assisted Gaussian Splatting from Blurry Images
  15. GS2E: Gaussian Splatting is an Effective Data Generator for Event Stream Generation
  16. Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuning
  17. HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation
  18. Identity-Preserving Text-to-Video Generation by Frequency Decomposition
  19. ImgEdit: A Unified Image Editing Dataset and Benchmark
  20. LlaVA-CoT: Let Vision Language Models Reason Step-By-Step
  21. MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
  22. MoH: Multi-Head Attention as Mixture-of-Head Attention
  23. Multi-objective Aligned Bidword Generation Model for E-commerce Search Advertising
  24. OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
  25. Orthogonal Subspace Decomposition for Generalizable AI-Generated Image Detection
  26. PiCO: Peer Review in LLMs based on Consistency Optimization
  27. RoomPainter: View-Integrated Diffusion for Consistent Indoor Scene Texturing
  28. UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model Evaluation
  29. WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
  30. Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
  31. ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation
  32. DF40: Toward Next-Generation Deepfake Detection
  33. Fast and Robust Point Cloud Registration with Tree-based Transformer
  34. FreestyleRet: Retrieving Images from Style-Diversified Queries
  35. GraCo: Granularity-Controllable Interactive Segmentation
  36. HiFi-123: Towards High-Fidelity One Image to 3D Content Generation
  37. LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference
  38. LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
  39. Learning Pseudo 3D Guidance for View-Consistent Texturing with 2D Diffusion
  40. Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation
  41. Parallel Vertex Diffusion for Unified Visual Grounding
  42. Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic Prompts
  43. Prompt2Poster: Automatically Artistic Chinese Poster Creation from Prompt Only
  44. QKFormer: Hierarchical Spiking Transformer using Q-K Attention
  45. RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
  46. Regressor-Segmenter Mutual Prompt Learning for Crowd Counting
  47. Repaint123: Fast and High-Quality One Image to 3D Generation with Progressive Controllable Repainting
  48. ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
  49. Spiking Transformer with Experts Mixture
  50. SynSP: Synergy of Smoothness and Precision in Pose Sequences Refinement
  51. Towards Better Seach Query Classification with Distribution-Diverse Multi-Expert Knowledge Distillation in JD Ads Search
  52. VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions
  53. Video-LLaVA: Learning United Visual Representation by Alignment Before Projection