PPaperPicks

Feng Zheng

Southern University of Science and Technology, Department of Computer Science and Engineering, Shenzhen, Guangdong, China

30 papers at tracked venues · 18 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual Scenarios
  2. SCORP: Scene-Consistent Object Refinement via Proxy Generation and Tuning
  3. Transferability of Adversarial Attacks in Video-based MLLMs: A Cross-modal Image-to-Video Approach
  4. $A_{0}$: An Affordance-Aware Hierarchical Model for General Robotic Manipulation
  5. LLplace: Embodied 3D Indoor Layout Synthesis Framework with Large Language Model
  6. Learn 3D VQA Better with Active Selection and Reannotation
  7. LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
  8. MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
  9. MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning
  10. On the Generalization Ability of Next-Token-Prediction Pretraining
  11. OptiScene: LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference Optimization
  12. Sample then Identify: A General Framework for Risk Control and Assessment in Multimodal Large Language Models
  13. Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors
  14. Beyond Prototypes: Semantic Anchor Regularization for Better Representation Learning
  15. Block Image Compressive Sensing with Local and Global Information Interaction
  16. Depth-Aware Concealed Crop Detection in Dense Agricultural Scenes
  17. Fine-grained Analysis of Stability and Generalization for Stochastic Bilevel Optimization
  18. MS2SL: Multimodal Spoken Data-Driven Continuous Sign Language Production
  19. Mutual Learning for Acoustic Matching and Dereverberation via Visual Scene-Driven Diffusion
  20. Negative Label Guided OOD Detection with Pretrained Vision-Language Models
  21. On the Noise Robustness of In-Context Learning for Text Generation
  22. PVUW 2024 Challenge on Complex Video Understanding: Methods and Results
  23. Place Anything into Any Video
  24. Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models
  25. Self-guided Knowledge-Injected Graph Neural Network for Alzheimer's Diseases
  26. The Second Visual Object Tracking Segmentation VOTS2024 Challenge Results
  27. Tuning-Free Image Customization with Image and Text Guidance
  28. Two in One Go: Single-stage Emotion Recognition with Decoupled Subject-context Transformer
  29. Unlocking Memorization in Large Language Models with Dynamic Soft Prompting
  30. Unsupervised Continual Anomaly Detection with Contrastively-Learned Prompt