PPaperPicks

Yan Zhang

29 papers at tracked venues · 21 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Beyond Counting: Evaluating Abstract and Emotional Reasoning in Vision-Language Models
  2. Failures are Treasures: Constructing a Pedagogical Bridge for Agentic Strategy Distillation
  3. Graph-Driven Domain Co-Adaptation for Cross-Domain Image Quality Assessment
  4. LookFlow: Training-Free and Efficient High-Resolution Image Synthesis via Dynamic Lookahead Guidance Flow
  5. Mitigating Error Accumulation in Knowledge Editing for Multi-Hop Question Answering
  6. SegMem-RAG: Adaptive Memory for Retrieval-Augmented Generation in Open-Ended Knowledge Environments
  7. SynPlay: Large-Scale Synthetic Human Data with Real-World Diversity for Aerial-View Perception
  8. A Theory for Conditional Generative Modeling on Multiple Data Sources
  9. AdaGCRAG: Adaptive Graph-Chunk Retrieval for Lightweight RAG
  10. DPS: Diverse Prototype Selection for Adaptive In-Context Learning
  11. ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability Knowledge
  12. Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
    ACM MM 2025 · Yan Zhang
  13. Joint Modeling of fMRI and EEG Imaging Using Ordinary Differential Equation-Based Hypergraph Neural Networks
    NeurIPS 2025 · Yan Zhang
  14. MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
  15. NTIRE 2025 Challenge on Cross-Domain Few-Shot Object Detection: Methods and Results
  16. NTIRE 2025 Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results
  17. ProtPainter: Draw or Drag Protein via Topology-guided Diffusion
  18. SmallGS: Gaussian Splatting-based Camera Pose Estimation for Small-Baseline Videos
  19. Test-Time Code-Switching for Cross-lingual Aspect Sentiment Triplet Extraction
  20. Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues
    AAAI 2025 · Yan Zhang
  21. Tuning Less, Prompting More: In-Context Preference Learning Pipeline for Natural Language Transformation
  22. When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
  23. CUBE: Causal Intervention-based Counterfactual Explanation for Prediction Models (Extended Abstract)
  24. Explainable Database Management System Configuration Tuning through Counterfactuals
  25. FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
  26. Graph Neural Networks for Learning Equivariant Representations of Neural Networks
  27. Improved Generalization of Weight Space Networks via Augmentations
  28. RELI11D: A Comprehensive Multimodal Human Motion Dataset and Method
  29. Unsupervised Concept Discovery Mitigates Spurious Correlations