PPaperPicks

Lei Zhang

Hong Kong Polytechnic University, Department of Computing

59 papers at tracked venues · 46 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code
  2. AlignCVC: Aligning Cross-View Consistency for Single-Image-to-3D Generation
  3. BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object Detection
  4. Fast Multi-view Consistent 3D Editing with Video Priors
  5. NebulaSQL: A Large-scale Feature Computation System for Online Recommendation
  6. BurstDeflicker: A Benchmark Dataset for Flicker Removal in Dynamic Scenes
  7. DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow Editing
  8. DP²O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-Resolution
  9. FiVE-Bench: A Fine-Grained Video Editing Benchmark for Evaluating Emerging Diffusion and Rectified Flow Models
  10. Fine-Structure Preserved Real-World Image Super-Resolution Via Transfer Vae Training
  11. FreCaS: Efficient Higher-Resolution Image Generation via Frequency-aware Cascaded Sampling
  12. GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and Generation
  13. Generalized and Efficient 2D Gaussian Splatting for Arbitrary-Scale Super-Resolution
  14. InsViE-1M: Effective Instruction-Based Video Editing with Elaborate Dataset Construction
  15. InstructRestore: Region-Customized Image Restoration with Human Instructions
  16. Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem Solving
  17. Knowledge Regularized Negative Feature Tuning of Vision-Language Models for Out-of-Distribution Detection
  18. LLaVA-MoD: Making LLaVA Tiny via MoE-Knowledge Distillation
  19. MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis
  20. MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
  21. NTIRE 2025 the 2nd Restore Any Image Model (RAIM) in the Wild Challenge
  22. One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution
  23. Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
  24. Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approach
  25. Polyline Path Masked Attention for Vision Transformer
  26. Progressive Rendering Distillation: Adapting Stable Diffusion for Instant Text-to-Mesh Generation without 3D Data
  27. RORem: Training a Robust Object Remover with Human-in-the-Loop
  28. Registration is a Powerful Rotation-Invariance Learner for 3D Anomaly Detection
  29. Reliable and Private Utility Signaling for Data Markets
  30. Spatial-Mamba: Effective Visual State Space Models via Structure-Aware State Fusion
  31. SyncNoise: Geometrically Consistent Noise Prediction for Instruction-based 3D Editing
  32. Toward Generalized Image Quality Assessment: Relaxing the Perfect Reference Quality Assumption
  33. Toward Generalizing Visual Brain Decoding to Unseen Subjects
  34. Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
  35. VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to Rank
  36. A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment
  37. Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding
  38. Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models
  39. General Geometry-Aware Weakly Supervised 3D Object Detection
  40. LAPT: Label-Driven Automated Prompt Tuning for OOD Detection with Vision-Language Models
  41. LIDIA: Precise Liver Tumor Diagnosis on Multi-Phase Contrast-Enhanced CT via Iterative Fusion and Asymmetric Contrastive Learning
  42. Local Differentially Private Heavy Hitter Detection in Data Streams with Bounded Memory
  43. MasterWeaver: Taming Editability and Face Identity for Personalized Text-to-Image Generation
  44. NTIRE 2024 Restore Any Image Model (RAIM) in the Wild Challenge
  45. Neural Super-Resolution for Real-Time Rendering with Radiance Demodulation
  46. One-Step Effective Diffusion Network for Real-World Image Super-Resolution
  47. Open Vocabulary 3D Scene Understanding via Geometry Guided Self-Distillation
  48. Osprey: Pixel Understanding with Visual Instruction Tuning
  49. Pixel-Aware Stable Diffusion for Realistic Image Super-Resolution and Personalized Stylization
  50. Responsible Visual Editing
  51. SSL: A Self-similarity Loss for Improving Generative Image Super-resolution
  52. ScaleDreamer: Scalable Text-to-3D Synthesis with Asynchronous Score Distillation
  53. ScatterFormer: Efficient Voxel Transformer with Scattered Linear Attention
  54. SeeSR: Towards Semantics-Aware Real-World Image Super-Resolution
  55. Self-Supervised Video Desmoking for Laparoscopic Surgery
  56. Source Prompt Disentangled Inversion for Boosting Image Editability with Diffusion Models
  57. TAPTRv2: Attention-based Position Update Improves Tracking Any Point
  58. UniVS: Unified and Universal Video Segmentation with Prompts as Queries
  59. Voxel Mamba: Group-Free State Space Models for Point Cloud based 3D Object Detection