PPaperPicks

Liqiang Nie

106 papers at tracked venues · 98 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. D2MoRA: Diversity-Regulated Asymmetric MoE-LoRA Decomposition for Efficient Multi-Task Adaptation
  2. Evolving Sparsity: Leveraging Token Importance Dynamics for Efficient LLM Decoding with Sparse Attention
  3. Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding
  4. Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation
  5. Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning
  6. Parallel Test-Time Scaling for Latent Reasoning Models
  7. PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records
  8. Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
  9. SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation
  10. StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval
  11. TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval
  12. TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMs
  13. 3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds
  14. A Polynomial-time Algorithm for Online Sparse Linear Regression with Improved Regret Bound under Weaker Conditions
  15. A Survey on the Feedback Mechanism of LLM-based AI Agents
  16. AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
  17. Bi-Tuning with Collaborative Information for Controllable LLM-based Sequential Recommendation
  18. Breakthrough Sensor-Limited Single View: Towards Implicit Temporal Dynamics for Time Series Domain Adaptation
  19. CoRe-MMRAG: Cross-Source Knowledge Reconciliation for Multimodal RAG
  20. CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & Sparsification
  21. Curriculum Coarse-to-Fine Selection for High-IPC Dataset Distillation
  22. DKDM: Data-Free Knowledge Distillation for Diffusion Models with Any Architecture
  23. Debiased Curriculum Adaptation for Safe Transfer Learning in Chest X-Ray Classification
  24. Dual-Center Graph Clustering with Neighbor Distribution
  25. Dually Self-Improved Counterfactual Data Augmentation Using Large Language Model
  26. Efficient Safety Alignment of Large Language Models via Preference Re-ranking and Representation-based Reward Modeling
  27. Embodied Crowd Counting
  28. EmoSym: A Symbiotic Framework for Unified Emotional Understanding and Generation via Latent Reasoning
  29. Enhancing Democratic Mediation through Norm-Awareness in Generative Agent Societies
  30. Enhancing GUI Agent with Uncertainty-Aware Self-Trained Evaluator
  31. Enhancing HOI Detection with Contextual Cues from Large Vision-Language Models
  32. FALCON: Resolving Visual Redundancy and Fragmentation in High-Resolution Multimodal Large Language Models via Visual Registers
  33. Fair Deepfake Detectors Can Generalize
  34. GARLIC: GPT-Augmented Reinforcement Learning with Intelligent Control for Vehicle Dispatching
  35. GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent
  36. Gaming for Boundary: Elastic Localization for Frame-Supervised Video Moment Retrieval
  37. Generative Agents for Multimodal Controversy Detection
  38. HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
  39. Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin
  40. LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant
  41. Language-Assisted Debiasing and Smoothing for Foundation Model-Based Semi-Supervised Learning
  42. Less is More: Empowering GUI Agent with Context-Aware Simplification
  43. Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation
  44. Meta Guidance: Incorporating Inductive Biases into Deep Time Series Imputers
  45. Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization
  46. OFFSET: Segmentation-based Focus Shift Revision for Composed Image Retrieval
  47. ORPP: Self-Optimizing Role-playing Prompts to Enhance Language Model Capabilities
  48. Object-Shot Enhanced Grounding Network for Egocentric Video
  49. Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned Policy
  50. PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
  51. R$2$ec: Towards Large Recommender Models with Reasoning
  52. Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
  53. SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
  54. STAR: Learning Diverse Robot Skill Abstractions through Rotation-Augmented Vector Quantization
  55. Social Context-Aware Community-Level Propagation Prediction
  56. Social Debiasing for Fair Multi-Modal LLMs
  57. Spa-Bench: a comprehensive Benchmark for Smartphone Agent Evaluation
  58. Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
  59. Spatial-Temporal Graph Diffusion Policy with Kinematic Modeling for Bimanual Robotic Manipulation
  60. TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
  61. TextSplat: Text-Guided Semantic Fusion for Generalizable Gaussian Splatting
  62. The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
  63. Towards Harmless Multimodal Assistants with Blind Preference Optimization
  64. Train with Perturbation, Infer after Merging: A Two-Stage Framework for Continual Learning
  65. Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
  66. Understanding the Forgetting of (Replay-based) Continual Learning via Feature Learning: Angle Matters
  67. Unified Transferability Metrics for Time Series Foundation Models
  68. UtilGen: Utility-Centric Generative Data Augmentation with Dual-Level Task Adaptation
  69. Weight-Aware Activation Sparsity with Constrained Bayesian Optimization Scheduling for Large Language Models
  70. A Multi-View Clustering Algorithm for Short Text
  71. Attribute-driven Disentangled Representation Learning for Multimodal Recommendation
  72. Boosting Transferability and Discriminability for Time Series Domain Adaptation
  73. Breaking Barriers of System Heterogeneity: Straggler-Tolerant Multimodal Federated Learning via Knowledge Distillation
  74. CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuning
  75. Decision Mamba: A Multi-Grained State Space Model with Self-Evolution Regularization for Offline RL
  76. DiffPerformer: Iterative Learning of Consistent Latent Guidance for Diffusion-Based Human Video Generation
  77. Differential-Perceptive and Retrieval-Augmented MLLM for Change Captioning
  78. Diffusion Facial Forgery Detection
  79. Discriminative Probing and Tuning for Text-to-Image Generation
  80. Distillation Enhanced Generative Retrieval
  81. Explicit Granularity and Implicit Scale Correspondence Learning for Point-Supervised Video Moment Localization
  82. Exploiting the Social-Like Prior in Transformer for Visual Reasoning
  83. Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
  84. Fourier Priors-Guided Diffusion for Zero-Shot Joint Low-Light Enhancement and Deblurring
  85. GPS-Gaussian: Generalizable Pixel-Wise 3D Gaussian Splatting for Real-Time Human Novel View Synthesis
  86. GaussianAvatar: Towards Realistic Human Avatar Modeling from a Single Video via Animatable 3D Gaussians
  87. Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
  88. GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative Decoding
  89. High-Resolution Image Harmonization with Adaptive-Interval Color Transformation
  90. LION : Empowering Multimodal Large Language Model with Dual-Level Visual Knowledge
  91. LLM vs Small Model? Large Language Model Based Text Augmentation Enhanced Personality Detection Model
  92. LRQuant: Learnable and Robust Post-Training Quantization for Large Language Models
  93. Let Me Show You Step by Step: An Interpretable Graph Routing Network for Knowledge-based Visual Question Answering
  94. Mind the Boundary: Coreset Selection via Reconstructing the Decision Boundary
  95. MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models
  96. Multi-Factor Adaptive Vision Selection for Egocentric Video Question Answering
  97. NovaChart: A Large-scale Dataset towards Chart Understanding and Generation of Multimodal Large Language Models
  98. Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks
  99. Revisiting Context Aggregation for Image Matting
  100. Revisiting Unsupervised Temporal Action Localization: The Primacy of High-Quality Actionness and Pseudolabels
  101. RoboMP2: A Robotic Multimodal Perception-Planning Framework with Multimodal Large Language Models
  102. Self-chats from Large Language Models Make Small Emotional Support Chatbot Better
  103. Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image Retrieval
  104. Thoughts to Target: Enhance Planning for Target-driven Conversation
  105. To Err Like Human: Affective Bias-Inspired Measures for Visual Emotion Recognition Evaluation
  106. Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization