PPaperPicks

Mike Zheng Shou

National University of Singapore

69 papers at tracked venues · 59 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. OptMark: Robust Multi-bit Diffusion Watermarking via Inference Time Optimization
  2. Balanced Image Stylization with Style Matching Score
  3. Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach
  4. Can I Trust You? Advancing GUI Task Automation with Action Trust Score
  5. CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
  6. DIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion Models
  7. DOTA: Distributional Test-time Adaptation of Vision-Language Models
  8. DiffSim: Taming Diffusion Models for Evaluating Visual Similarity
  9. DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles
  10. Factorized Learning for Temporally Grounded Video-Language Models
  11. GUI-Narrator: Detecting and Captioning Computer GUI Actions
  12. Grounding Multimodal Large Language Model in GUI World
  13. IDProtector: An Adversarial Noise Encoder to Protect Against ID-Preserving Image Generation
  14. Image Watermarks are Removable using Controllable Regeneration from Clean Noise
  15. Impossible Videos
  16. InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models with Human Feedback
  17. LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer
  18. LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
  19. MP-Mat: A 3D-and-Instance-Aware Human Matting and Editing Framework with Multiplane Representation
  20. MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
  21. OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data
  22. PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer
  23. PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning
  24. ROICtrl: Boosting Instance Control for Visual Generation
  25. ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning
  26. SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost
  27. Show-o2: Improved Native Unified Multimodal Models
  28. Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
  29. ShowUI: One Vision-Language-Action Model for GUI Visual Agent
  30. Sparse Image Synthesis via Joint Latent and RoI Flow
  31. Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
  32. VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
  33. VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
  34. WMAdapter: Adding WaterMark Control to Latent Diffusion Models
  35. macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
  36. Apprenticeship-Inspired Elegance: Synergistic Knowledge Distillation Empowers Spiking Neural Networks for Efficient Single-Eye Emotion Recognition
  37. AssistEditor: Multi-Agent Collaboration for GUI Workflow Automation in Video Creation
  38. AssistGUI: Task-Oriented PC Graphical User Interface Automation
  39. Bootstrapping SparseFormers from Vision Foundation Models
  40. Can Simple Averaging Defeat Modern Watermarks?
  41. Delocate: Detection and Localization for Deepfake Videos with Randomly-Located Tampered Traces
  42. DoFIT: Domain-aware Federated Instruction Tuning with Alleviated Catastrophic Forgetting
  43. DragAnything: Motion Control for Anything Using Entity Representation
  44. DynVideo-E: Harnessing Dynamic NeRF for Large-Scale Motion- and View-Change Human-Centric Video Editing
  45. EvolveDirector: Approaching Advanced Text-to-Image Generation with Large Vision-Language Models
  46. Exocentric-to-Egocentric Video Generation
  47. Free-ATM: Harnessing Free Attention Masks for Representation Learning on Diffusion-Generated Images
  48. GENIXER: Empowering Multimodal Large Language Model as a Powerful Data Generator
  49. L4D-Track: Language-to-4D Modeling Towards 6-DoF Tracking and Shape Reconstruction in 3D Point Cloud Stream
  50. LOVA3: Learning to Visual Question Answering, Asking and Assessment
  51. Learning Video Context as Interleaved Multimodal Sequences
  52. Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal Learning
  53. MAG-Edit: Localized Image Editing in Complex Scenarios via Mask-Based Attention-Adjusted Guidance
  54. MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model
  55. MotionDirector: Motion Customization of Text-to-Video Diffusion Models
  56. One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
  57. Parrot Captions Teach CLIP to Spot Text
  58. Rethinking the Objectives of Vector-Quantized Tokenizers for Image Synthesis
  59. RingID: Rethinking Tree-Ring Watermarking for Enhanced Multi-key Identification
  60. Skinned Motion Retargeting with Dense Geometric Interaction Perception
  61. SparseFormer: Sparse Visual Recognition via Limited Latent Tokens
  62. Tune-an-Ellipse: CLIP Has Potential to Find what you Want
  63. VIT-LENS: Towards Omni-modal Representations
  64. VideoGUI: A Benchmark for GUI Automation from Instructional Videos
  65. VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
  66. VideoLLM-online: Online Video Large Language Model for Streaming Video
  67. VideoSwap: Customized Video Subject Swapping with Interactive Semantic Point Correspondence
  68. Visual Perception by Large Language Model's Weights
  69. X- Adapter: Universal Compatibility of Plugins for Upgraded Diffusion Model