PPaperPicks

Yun Fu

Northeastern University, Department of Electrical and Computer Engineering / College of Computer and Information Science, Boston, MA, USA

38 papers at tracked venues · 29 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. ACBQ: Adaptive Cross-Block Quantization of Large Language Models
  2. Arbitrary-Scale 3D Gaussian Super-Resolution
  3. Distorted or Fabricated? A Survey on Hallucination in Video LLMs
  4. From Words to Pixels: A Comprehensive Survey on Large Language Models in Visual Segmentation
  5. Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning
  6. NoisyCausal: A Benchmark for Evaluating Causal Reasoning Under Structured Noise
  7. Revealing the Seen, Imagining the Beyond: A Survey of Image-Grounded Chain-of-Thought Reasoning in Multimodal LLMs
  8. Towards Unified Multimodal Large Language Models: A survey
  9. UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning
  10. Accessing Vision Foundation Models via ImageNet-1K
  11. AdaSports-Traj: Role- and Domain-Aware Adaptation for Multi-Agent Trajectory Modeling in Sports
  12. Cautious Next Token Prediction
  13. D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
  14. MTS-DMAE: Dual-Masked Autoencoder for Unsupervised Multivariate Time Series Representation Learning
  15. Outlier-Aware Post-Training Quantization for Image Super-Resolution
  16. REGEN: Learning Compact Video Embedding with (Re-)Generative Decoder
  17. RealTalk: Realistic Emotion-Aware Lifelike Talking-Head Synthesis
  18. Representation Potentials of Foundation Models for Multimodal Alignment: A Survey
  19. Scale-Free Graph-Language Models
  20. Sports-Traj: A Unified Trajectory Generation Model for Multi-Agent Movement in Sports
  21. The Indra Representation Hypothesis for Multimodal Alignment
  22. Towards Zero-shot 3D Anomaly Localization
  23. VQToken: Neural Discrete Token Representation Learning for Extreme Token Reduction in Video Large Language Models
  24. Ada-VAD: Domain Adaptable Video Anomaly Detection
  25. AdaFormer: Efficient Transformer with Adaptive Token Sparsification for Image Super-resolution
  26. Adapting to Length Shift: FlexiLength Network for Trajectory Prediction
  27. Advancing Vision-Language Models with Adapter Ensemble Strategies
  28. Aligning Out-of-Distribution Web Images and Caption Semantics via Evidential Learning
  29. Consistency and Uncertainty: Identifying Unreliable Responses From Black-Box Vision-Language Models for Selective Visual Question Answering
  30. Don't Judge by the Look: Towards Motion Coherent Video Representation
  31. Efficient Modulation for Vision Networks
  32. LightAvatar: Efficient Head Avatar as Dynamic Neural Light Field
  33. OOSTraj: Out-of-Sight Trajectory Prediction With Vision-Positioning Denoising
  34. Rewrite the Stars
  35. Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement
  36. SkipDiff: Adaptive Skip Diffusion Model for High-Fidelity Perceptual Image Super-resolution
  37. Slicing Vision Transformer for Flexible Inference
  38. α-Former: Local-Feature-Aware (L-FA) Transformer