PPaperPicks

Wenguan Wang

Zhejiang University, China

40 papers at tracked venues · 34 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation
  2. 3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language Navigation
  3. A Conditional Probability Framework for Compositional Zero-Shot Learning
  4. Cycle-Consistent Learning for Joint Layout-to-Image Generation and Object Detection
  5. DiffVsgg: Diffusion-Driven Online Video Scene Graph Generation
  6. Do as We Do, Not as You Think: the Conformity of Large Language Models
  7. Dual Reciprocal Learning of Language-based Human Motion Understanding and Generation
  8. Gaussian-Based World Model: Gaussian Priors for Voxel-Based Occupancy Prediction and Future Motion Prediction
  9. Hydra-SGG: Hybrid Relation Assignment for One-stage Scene Graph Generation
  10. LOGICZSL: Exploring Logic-induced Representation for Compositional Zero-shot Learning
  11. Learning Clustering-based Prototypes for Compositional Zero-Shot Learning
  12. Learning Human-Object Interaction as Groups
  13. Multi-view Reconstruction via SfM-guided Monocular Depth Estimation
  14. OmniGaze: Reward-inspired Generalizable Gaze Estimation in the Wild
  15. Scene Map-based Prompt Tuning for Navigation Instruction Generation
  16. TAGA: Self-supervised Learning for Template-free Animatable Gaussian Articulated Model
  17. Towards Human-Like Virtual Beings: Simulating Human Behavior in 3D Scenes
  18. UNIALIGN: Scaling Multimodal Alignment within One Unified Model
  19. Underwater Visual SLAM with Depth Uncertainty and Medium Modeling
  20. Clustering Propagation for Universal Medical Image Segmentation
  21. Clustering for Protein Representation Learning
  22. Controllable Navigation Instruction Generation with Chain of Thought Prompting
  23. DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
  24. Facing the Elephant in the Room: Visual Prompt Tuning or Full finetuning?
  25. General and Task-Oriented Video Segmentation
  26. Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion Models
  27. IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object Detection
  28. Interpretable3D: An Ad-Hoc Interpretable Classifier for 3D Point Clouds
  29. LSK3DNet: Towards Effective and Efficient 3D Perception with Large Sparse Kernels
  30. MS2SL: Multimodal Spoken Data-Driven Continuous Sign Language Production
  31. Mutual Learning for Acoustic Matching and Dereverberation via Visual Scene-Driven Diffusion
  32. Navigation Instruction Generation with BEV Perception and Large Language Models
  33. Neural Clustering Based Visual Representation Learning
  34. Nonverbal Interaction Detection
  35. Poly Kernel Inception Network for Remote Sensing Detection
  36. Psychometry: An Omnifit Model for Image Reconstruction from Human Brain Activity
  37. Scene Graph Generation with Role-Playing Large Language Models
  38. Shape2Scene: 3D Scene Representation Learning Through Pre-training on Shape Data
  39. Vision-Language Navigation with Energy-Based Policy
  40. Volumetric Environment Representation for Vision-Language Navigation