PPaperPicks

Jingkuan Song

29 papers at tracked venues · 26 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. De-biased Natural Language Egocentric Task Verification via Prototypical Evidence Learning
  2. Debiased Orthogonal Boundary-Driven Efficient Noise Mitigation
  3. FGIM: a Fast Graph-based Indexes Merging Framework for Approximate Nearest Neighbor Search
  4. Hyper-Opinion Vagueness Quantification for Robust Multimodal Learning
  5. Learning to Curate Context: Jointly Optimizing Retrieval and Prediction for Multimodal Social Media Popularity
  6. SINDI: An Efficient Index for Sparse Vector Approximate Maximum Inner Product Search
  7. AICL: Action In-Context Learning for Text-to-Video Generation
  8. FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models
  9. From Observation to Understanding: Front-Door Adjustments with Uncertainty Calibration for Enhancing Egocentric Reasoning in LVLMs
  10. Improving Multimodal Social Media Popularity Prediction via Selective Retrieval Knowledge Augmentation
  11. MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct
  12. OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
  13. PHGC: Procedural Heterogeneous Graph Completion for Natural Language Task Verification in Egocentric Videos
  14. SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
  15. Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
  16. Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approach
  17. VSAG: An Optimized Search Framework for Graph-based Approximate Nearest Neighbor Search
  18. Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization
  19. Any Target Can be Offense: Adversarial Example Generation via Generalized Latent Infection
  20. CoIN: A Benchmark of Continual Instruction Tuning for Multimodel Large Language Models
  21. Counterfactually Augmented Event Matching for De-biased Temporal Sentence Grounding
  22. DePT: Decoupled Prompt Tuning
  23. F³-Pruning: A Training-Free and Generalized Pruning Strategy towards Faster and Finer Text-to-Video Synthesis
  24. MPT: Multi-grained Prompt Tuning for Text-Video Retrieval
  25. MagicVFX: Visual Effects Synthesis in Just Minutes
  26. ProS: Prompting-to-Simulate Generalized Knowledge for Universal Cross-Domain Retrieval
  27. RoScenes: A Large-Scale Multi-view 3D Dataset for Roadside Perception
  28. SI-BiViT: Binarizing Vision Transformers with Spatial Interaction
  29. Unsupervised Cross-Domain Image Retrieval with Semantic-Attended Mixture-of-Experts