PPaperPicks

Rongrong Ji

Xiamen University, Xiamen, China

97 papers at tracked venues · 85 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. ALGOGEN: Tool-Generated Verifiable Traces for Reliable Algorithm Visualization
  2. EMA: An Episodic Memory Agent for Efficient and Selective Memory
  3. Relaxing the Constraints: A Dual-Importance Projection Mechanism for Lifelong Model Editing
  4. Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings
  5. Aigi-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
  6. Automated Fine-Grained Mixture-of-Experts Quantization
  7. Automated Manipulation of Magnetic Microswarms for Temporal Logic Cargo Delivery Tasks in Complex Environments
  8. BAME: Block-Aware Mask Evolution for Efficient N: M Sparse Training
  9. Benchmarking Abstract and Reasoning Abilities Through A Theoretical Perspective
  10. Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference
  11. CPPO: Accelerating the Training of Group Relative Policy Optimization-Based Reasoning Models
  12. DAMamba: Vision State Space Model with Dynamic Adaptive Scan
  13. DS-VLM: Diffusion Supervision Vision Language Model
  14. Determining Layer-wise Sparsity for Large Language Models Through a Theoretical Perspective
  15. Discovering Important Experts for Mixture-of-Experts Models Pruning Through a Theoretical Perspective
  16. Dynamic Low-Rank Sparse Adaptation for Large Language Models
  17. EasyInv: Toward Fast and Better DDIM Inversion
  18. Enhancing Language Model Hypernetworks with Restart: A Study on Optimization
  19. Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
  20. Few-Shot Image Quality Assessment via Adaptation of Vision-Language Models
  21. FlashSloth : Lightning Multimodal Large Language Models via Embedded Visual Compression
  22. FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
  23. From Objects to Events: Unlocking Complex Visual Understanding in Object Detectors Via LLM-guided Symbolic Reasoning
  24. GPT-ReID: Learning Fine-grained Representation with GPT for Text-based Person Retrieval
  25. GS-Bias: Global-Spatial Bias Learner for Single-Image Test-Time Adaptation of Vision-Language Models
  26. Generate Aligned Anomaly: Region-Guided Few-Shot Anomaly Image-Mask Pair Synthesis for Industrial Inspection
  27. HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation
  28. Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive Segmentation
  29. LTD-Bench: Evaluating Large Language Models by Letting Them Draw
  30. Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
  31. MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
  32. MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
  33. OracleFusion: Assisting the Decipherment of Oracle Bone Script with Structurally Constrained Semantic Typography
  34. Routing Experts: Learning to Route Dynamic Experts in Existing Multi-modal Large Language Models
  35. SVFR: A Unified Framework for Generalized Video Face Restoration
  36. Semantic Alignment and Reinforcement for Data-Free Quantization of Vision Transformers
  37. Spotlight Attention: Towards Efficient LLM Generation via Non-linear Hashing-based KV Cache Retrieval
  38. Towards General Visual-Linguistic Face Forgery Detection
  39. Training Long-Context LLMs Efficiently via Chunk-wise Optimization
  40. VISA: Group-wise Visual Token Selection and Aggregation via Graph Summarization for Efficient MLLMs Inference
  41. VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
  42. VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language Model
  43. VTON-HandFit: Virtual Try-on for Arbitrary Hand Pose Guided by Hand Priors Embedding
  44. Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
  45. What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image Segmentation
  46. Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
  47. polybasic Speculative Decoding Through a Theoretical Perspective
  48. γ-MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models
  49. 3D-GRES: Generalized 3D Referring Expression Segmentation
  50. AccDiffusion: An Accurate Method for Higher-Resolution Image Generation
  51. Adaptive Feature Selection for No-Reference Image Quality Assessment by Mitigating Semantic Noise Sensitivity
  52. Adaptive Selection based Referring Image Segmentation
  53. Advancing Multimodal Large Language Models with Quantization-Aware Scale Learning for Efficient Adaptation
  54. AffineQuant: Affine Transformation Quantization for Large Language Models
  55. Aligning and Prompting Everything All at Once for Universal Visual Perception
  56. AnyTrans: Translate AnyText in the Image with Large Scale Models
  57. Autoregressive Queries for Adaptive Tracking with Spatio-Temporal Transformers
  58. CaM: Cache Merging for Memory-efficient LLMs Inference
  59. CamoTeacher: Dual-Rotation Consistency Learning for Semi-supervised Camouflaged Object Detection
  60. Cantor: Inspiring Multimodal Chain-of-Thought of MLLM
  61. Code Membership Inference for Detecting Unauthorized Data Use in Code Pre-trained Language Models
  62. ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
  63. Cross-Modality Perturbation Synergy Attack for Person Re-identification
  64. Deep Instruction Tuning for Segment Anything Model
  65. DiffAgent: Fast and Accurate Text-to-Image API Selection with Large Language Model
  66. DiffuMatting: Synthesizing Arbitrary Objects with Matting-Level Annotation
  67. DiffusionFake: Enhancing Generalization in Deepfake Detection via Guided Stable Diffusion
  68. Director3D: Real-world Camera Trajectory and 3D Scene Generation from Text
  69. Dynamic Sparse No Training: Training-Free Fine-tuning for Sparse LLMs
  70. ERQ: Error Reduction for Post-Training Quantization of Vision Transformers
  71. Enhancing Tampered Text Detection Through Frequency Feature Fusion and Decomposition
  72. Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language Models
  73. Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
  74. Exploring Target Representations for Masked Autoencoders
  75. Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric Regularization
  76. FocSAM: Delving Deeply into Focused Objects in Segmenting Anything
  77. GOI: Find 3D Gaussians of Interest with an Optimizable Open-vocabulary Semantic-space Hyperplane
  78. GraCo: Granularity-Controllable Interactive Segmentation
  79. I2EBench: A Comprehensive Benchmark for Instruction-based Image Editing
  80. Integrating Global Context Contrast and Local Sensitivity for Blind Image Quality Assessment
  81. Learning Image Demoiréing from Unpaired Real Data
  82. Multi-branch Collaborative Learning Network for 3D Visual Grounding
  83. Multimodal Inplace Prompt Tuning for Open-set Object Detection
  84. Outlier-aware Slicing for Post-Training Quantization in Vision Transformer
  85. PortraitBooth: A Versatile Portrait Model for Fast Identity-Preserved Personalization
  86. Prompting to Adapt Foundational Segmentation Models
  87. QueryMatch: A Query-based Contrastive Learning Framework for Weakly Supervised Visual Grounding
  88. RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation
  89. RLE: A Unified Perspective of Data Augmentation for Cross-Spectral Re-Identification
  90. Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation
  91. SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation
  92. StealthDiffusion: Towards Evading Diffusion Forensic Detection through Diffusion Model
  93. TF-FAS: Twofold-Element Fine-Grained Semantic Guidance for Generalizable Face Anti-spoofing
  94. Textual Grounding for Open-Vocabulary Visual Information Extraction in Layout-Diversified Documents
  95. Toward Open-Set Human Object Interaction Detection
  96. UniPTS: A Unified Framework for Proficient Post-Training Sparsity
  97. X-Oscar: A Progressive Framework for High-quality Text-guided 3D Animatable Avatar Generation