PPaperPicks

Zhi-Qi Cheng

26 papers at tracked venues · 19 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
  2. Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation Standards
  3. A Video-grounded Dialogue Dataset and Metric for Event-driven Activities
  4. DyRoNet: Dynamic Routing and Low-Rank Adapters for Autonomous Driving Streaming Perception
  5. Emphasizing Discriminative Features for Dataset Distillation in Complex Scenarios
  6. Large Language Model Agents in Finance: A Survey Bridging Research, Practice, and Real-World Deployment
  7. MaxSup: Overcoming Representation Collapse in Label Smoothing
  8. MetaDesigner: Advancing Artistic Typography through AI-Driven, User-Centric, and Multilingual WordArt Synthesis
  9. MotionFollower: Editing Video Motion via Score-Guided Diffusion
  10. POPoS: Improving Efficient and Robust Facial Landmark Detection with Parallel Optimal Position Search
  11. ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding
  12. Securing the Skies: A Comprehensive Survey on Anti-UAV Methods, Benchmarking, and Future Directions
  13. StableAnimator: High-Quality Identity-Preserving Human Image Animation
  14. UCDR-Adapter: Exploring Adaptation of Pre-Trained Vision-Language Models for Universal Cross-Domain Retrieval
  15. Why We Feel: Breaking Boundaries in Emotional Reasoning with Multimodal Large Language Models
  16. BlockGCN: Redefine Topology Awareness for Skeleton-Based Action Recognition
  17. DCPT: Darkness Clue-Prompted Tracking in Nighttime UAVs
  18. Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
  19. FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled Audio
  20. Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions
  21. MotionEditor: Editing Video Motion via Content-Aware Diffusion
  22. Music2P: A Multi-Modal AI-Driven Tool for Simplifying Album Cover Design
  23. ProS: Prompting-to-Simulate Generalized Knowledge for Universal Cross-Domain Retrieval
  24. SHIELD: LLM-Driven Schema Induction for Predictive Analytics in EV Battery Supply Chain Disruptions
    EMNLP 2024 · Zhi-Qi Cheng
  25. SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition
  26. Towards Calibrated Robust Fine-Tuning of Vision-Language Models