PPaperPicks

Jiankang Deng

56 papers at tracked venues · 48 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Egocentric Co-Pilot: Web-Native Smart-Glasses Agents for Assistive Egocentric AI
  2. Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
  3. UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
  4. ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
  5. "Principal Components" Enable a New Language of Images
  6. Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs
  7. CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
  8. CaricatureBooth: Data-Free Interactive Caricature Generation in a Photo Booth
  9. Deep Gaussian from Motion: Exploring 3D Geometric Foundation Models for Gaussian Splatting
  10. ForCenNet: Foreground-Centric Network for Document Image Rectification
  11. Fractal Calibration for Long-tailed Object Detection
    CVPR 2025 ·
    Konstantinos Panagiotis Alexandridis
  12. Frequency-Guided Diffusion for Training-Free Text-Driven Image Translation
  13. From Attention to Activation: Unraveling the Enigmas of Large Language Models
  14. Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution
  15. HUST: High-Fidelity Unbiased Skin Tone Estimation via Texture Quantization
  16. HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
  17. ImHead: A Large-Scale Implicit Morphable Model for Localized Head Modeling
    ICCV 2025 ·
    Rolandos Alexandros Potamias
  18. RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
  19. Region-based Cluster Discrimination for Visual Representation Learning
  20. S3-Face: SSS-Compliant Facial Reflectance Estimation via Diffusion Priors
  21. ShapeCraft: LLM Agents for Structured, Textured and Interactive 3D Modeling
  22. Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language Generator
  23. Single-view Image to Novel-view Generation for Hand-Object Interactions
  24. UniAttackData+: Unified Physical-Digital Attack Detection+ Challenge
  25. UniViT: Unifying Image and Video Understanding in One Vision Encoder
  26. Unlocking the Potential of Diffusion Priors in Blind Face Restoration
  27. Unsupervised Audio-Visual Segmentation with Modality Alignment
  28. VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
  29. WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild
    CVPR 2025 ·
    Rolandos Alexandros Potamias
  30. $V_{k}D$: Improving Knowledge Distillation Using Orthogonal Projections
  31. 3DGazeNet: Generalizing 3D Gaze Estimation with Weak-Supervision from Synthetic Views
  32. AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic Synthesis
  33. Adaptive Parametric Activation
    ECCV 2024 ·
    Konstantinos Panagiotis Alexandridis
  34. Any-Size-Diffusion: Toward Efficient Text-Driven Synthesis for Any-Size HD Images
  35. Arc2Face: A Foundation Model for ID-Consistent Human Faces
    ECCV 2024 ·
    Foivos Paraperas Papantoniou
  36. Boosting Object Detection with Zero-Shot Day-Night Domain Adaptation
  37. CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-Spoofing
  38. DiffSED: Sound Event Detection with Denoising Diffusion
  39. G3DR: Generative 3D Reconstruction in ImageNet
  40. ID-to-3D: Expressive ID-guided 3D Heads via Score Distillation Sampling
  41. IDAdapter: Learning Mixed Features for Tuning-Free Personalization of Text-to-Image Models
  42. Improving Face Generation Quality and Prompt Following with Synthetic Captions
  43. Monocular Identity-Conditioned Facial Reflectance Reconstruction
  44. Multi-label Cluster Discrimination for Visual Representation Learning
  45. Multilinear Mixture of Experts: Scalable Expert Specialization through Factorization
  46. Neural Sign Actors: A diffusion model for 3D sign language production from text
  47. RWKV-CLIP: A Robust Vision-Language Representation Learner
  48. Rethinking the Domain Gap in Near-infrared Face Recognition
  49. SAGS: Structure-Aware 3D Gaussian Splatting
  50. Self-Adaptive Reality-Guided Diffusion for Artifact-Free Super-Resolution
  51. Three Heads Are Better than One: Complementary Experts for Long-Tailed Semi-supervised Learning
  52. TopoFR: A Closer Look at Topology Alignment on Face Recognition
  53. Unified Physical-Digital Attack Detection Challenge
  54. Unified Physical-Digital Face Attack Detection
  55. VeLoRA: Memory Efficient Training using Rank-1 Sub-Token Projections
  56. WaveFace: Authentic Face Restoration with Efficient Frequency Recovery