PPaperPicks

Tat-Seng Chua

National University of Singapore, Department of Computer Science, Singapore

162 papers at tracked venues · 138 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Bringing Reasoning to Generative Recommendation Through the Lens of Cascaded Ranking
  2. LLaVA-UHD v2: Exploiting Hierarchical Vision Granularity in MLLMs via Inverse Semantic Pyramid
  3. Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models
  4. ThinkTank-ME: A Multi-Expert Framework for Middle East Event Forecasting
  5. Towards Knowledgeable Deep Research: Framework and Benchmark
  6. Verifiable Reasoning for LLM-based Generative Recommendation
  7. A Federated Framework for LLM-based Recommendation
  8. AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
  9. Aligning Large Language Models for Faithful Integrity Against Opposing Argument
  10. AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models
  11. AnyEdit: Edit Any Knowledge Encoded in Language Models
  12. Attend and Enrich: Enhanced Visual Prompt for Zero-Shot Learning
  13. Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
  14. Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMs
  15. Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark
  16. Bridging Jensen Gap for Max-Min Group Fairness Optimization in Recommendation
  17. Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
  18. Coarse-to-Fine Cross-Modality Generation for Enhancing Vehicle Re-Identification with High-Fidelity Synthetic Data
  19. Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
  20. Continual Multimodal Contrastive Learning
  21. Counterfactual Evolution of Multimodal Datasets via Visual Programming
  22. Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
  23. DRC: Enhancing Personalized Image Generation via Disentangled Representation Composition
  24. Decoding in Latent Spaces for Efficient Inference in LLM-based Recommendation
  25. Decoupling Knowledge and Context: An Efficient and Effective Retrieval Augmented Generation Framework via Cross Attention
  26. DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization
  27. EARN: Efficient Inference Acceleration for LLM-based Generative Recommendation by Register Tokens
  28. Efficient Inference for Large Language Model-based Generative Recommendation
  29. EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
  30. Exploring Training and Inference Scaling Laws in Generative Retrieval
  31. External Memory Matters: Generalizable Object-Action Memory for Retrieval-Augmented Long-Term Video Understanding
  32. FACT-AUDIT: An Adaptive Multi-Agent Framework for Dynamic Fact-Checking Evaluation of Large Language Models
  33. FOCoOp: Enhancing Out-of-Distribution Robustness in Federated Prompt Learning for Vision-Language Models
  34. FinIR: The 2nd Workshop on Financial Information Retrieval in the Era of Generative AI
  35. Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment
  36. Fine-tuning Multimodal Large Language Models for Product Bundling
  37. G2S: A General-to-Specific Learning Framework for Temporal Knowledge Graph Forecasting with Large Language Models
  38. Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos
  39. Generative Recommendation Models: Progress and Directions
  40. Generative Recommendation: Towards Personalized Multimodal Content Generation
  41. GraphVideoAgent: Enhancing Long-form Video Understanding with Entity Relation Graphs
  42. Hello Again! LLM-powered Personalized Agent for Long-term Dialogue
  43. Heterogeneous User Modeling for LLM-based Recommendation
  44. How to Enable Effective Cooperation Between Humans and NLP Models: A Survey of Principles, Formalizations, and Beyond
  45. IGD: Token Decisiveness Modeling via Information Gain in LLMs for Personalized Recommendation
  46. International Workshop on Multimodal Generative Search and Recommendation (MMGenSR@CIKM 2025)
  47. JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
  48. Knowledge Boundary of Large Language Models: A Survey
  49. L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language Models
  50. LLM-based Search Assistant with Holistically Guided MCTS for Intricate Information Seeking
  51. LLM2Rec: Large Language Models Are Powerful Embedding Models for Sequential Recommendation
  52. Language Representations Can be What Recommenders Need: Findings and Potentials
  53. Large Generative Models Meet Multimodal Applications (LGM3A)
  54. Large Language Models Empowered Personalized Web Agents
  55. Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual Scene
  56. Length Controlled Generation for Black-box LLMs
  57. LightPROF: A Lightweight Reasoning Framework for Large Language Model on Knowledge Graph
  58. Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
  59. MPO: Multilingual Safety Alignment via Reward Gap Optimization
  60. Measuring What Makes You Unique: Difference-Aware User Modeling for Enhancing LLM Personalization
  61. Media Source Matters More Than Content: Unveiling Political Bias in LLM-Generated Citations
  62. NExT-Mol: 3D Diffusion Meets 1D Language Modeling for 3D Molecule Generation
  63. NExT-Search: Rebuilding User Feedback Ecosystem for Generative AI Search
  64. Neural Causal Graph for Interpretable and Intervenable Classification
  65. Next Phase of Research on Multimodal Foundation Models: From Alignments to Content Generation and Quality Assessment
    ACM MM 2025 · Tat-Seng Chua
  66. On Path to Multimodal Generalist: General-Level and General-Bench
  67. On Reasoning Strength Planning in Large Reasoning Models
  68. Optimize Incompatible Parameters Through Compatibility-aware Knowledge Integration
  69. Order-agnostic Identifier for Large Language Model-based Generative Recommendation
  70. Personalized Generation In Large Model Era: A Survey
  71. Personalized Text Generation with Contrastive Activation Steering
  72. Preference Diffusion for Recommendation
  73. RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness
  74. RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards
  75. Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
  76. SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
  77. STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
  78. Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models
  79. ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
  80. Self-Calibrated Listwise Reranking with Large Language Models
  81. Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment
  82. TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
  83. The Emergence of Abstract Thought in Large Language Models Beyond Any Language
  84. Towards Large Generative Recommendation: A Tokenization Perspective
  85. Towards Modality Generalization: A Benchmark and Prospective Analysis
  86. Towards Semantic Equivalence of Tokenization in Multimodal LLM
  87. Towards Temporal-Aware Multi-Modal Retrieval Augemented Generation in Finance
  88. Towards Unified and Lossless Latent Space for 3D Molecular Latent Diffusion Modeling
  89. Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
  90. Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models
  91. Understanding Accuracy-Fairness Trade-offs in Re-ranking through Elasticity in Economics
  92. Universal Scene Graph Generation
  93. Video Question Answering and Beyond
  94. When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual Reasoners
  95. Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
  96. A Picture Is Worth a Graph: A Blueprint Debate Paradigm for Multimodal Reasoning
  97. A Study of Implicit Ranking Unfairness in Large Language Models
  98. A Survey on Neural Question Generation: Methods, Applications, and Prospects
  99. A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models
  100. A Taxation Perspective for Fair Re-ranking
  101. ALI-Agent: Assessing LLMs' Alignment with Human Values via Agent-based Evaluation
  102. Abductive Ego-View Accident Video Understanding for Safe Driving Perception
  103. Analyzing Temporal Complex Events with Large Language Models? A Benchmark towards Temporal, Long Context Understanding
  104. Ask-before-Plan: Proactive Language Agents for Real-World Planning
  105. Auto-Encoding Morph-Tokens for Multimodal LLM
  106. Beyond Persuasion: Towards Conversational Recommender System with Credible Explanations
  107. Bridging Items and Language: A Transition Paradigm for Large Language Model-Based Recommendation
  108. CIRP: Cross-Item Relational Pre-training for Multimodal Product Bundling
  109. CLAMBER: A Benchmark of Identifying and Clarifying Ambiguous Information Needs in Large Language Models
  110. Can I Trust Your Answer? Visually Grounded Video Question Answering
  111. Causal-driven Large Language Models with Faithful Reasoning for Knowledge Question Answering
  112. Chain-of-Exemplar: Enhancing Distractor Generation for Multimodal Educational Question Generation
  113. Companion Proceedings of the ACM on Web Conference 2024, WWW 2024, Singapore, Singapore, May 13-17, 2024
    WWW 2024 · Tat-Seng Chua
  114. Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
  115. Data-efficient Fine-tuning for LLM-based Recommendation
  116. Denoising Diffusion Recommender Model
  117. Discriminative Probing and Tuning for Text-to-Image Generation
  118. Disentangling Masked Autoencoders for Unsupervised Domain Generalization
  119. Distillation Enhanced Generative Retrieval
  120. Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations
  121. Dysen-VDM: Empowering Dynamics-Aware Text-to-Video Diffusion with LLMs
  122. Fact : Teaching MLLMs with Faithful, Concise and Transferable Rationales
  123. FashionReGen: LLM-Empowered Fashion Report Generation
  124. Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
  125. GOODAT: Towards Test-Time Graph Out-of-Distribution Detection
  126. Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
  127. Hop-based Heterogeneous Graph Transformer
  128. I3: Intent-Introspective Retrieval Conditioned on Instructions
  129. Improving Expressive Power of Spectral Graph Neural Networks with Eigenvalue Correction
  130. Information-Controllable Graph Contrastive Learning for Recommendation
  131. LARP: Language Audio Relational Pre-training for Cold-Start Playlist Continuation
  132. LASO: Language-Guided Affordance Segmentation on 3D Object
  133. LLaVA-UHD: An LMM Perceiving Any Aspect Ratio and High-Resolution Images
  134. Large Language Model Powered Agents for Information Retrieval
  135. Large Language Model Powered Agents in the Web
  136. Learnable Item Tokenization for Generative Recommendation
  137. Learning to Generate Explainable Stock Predictions using Self-Reflective Large Language Models
  138. Leveraging Multimodal Features and Item-level User Feedback for Bundle Construction
  139. MM-Forecast: A Multimodal Approach to Temporal Event Forecasting with Large Language Models
  140. Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
  141. NExT-Chat: An LMM for Chat, Detection and Segmentation
  142. NExT-GPT: Any-to-Any Multimodal LLM
  143. On Generative Agents in Recommendation
  144. On Softmax Direct Preference Optimization for Recommendation
  145. On the Multi-turn Instruction Following for Conversational Web Agents
  146. Plug-and-Play Policy Planner for Large Language Model Powered Dialogue Agents
  147. Proceedings of the ACM on Web Conference 2024, WWW 2024, Singapore, May 13-17, 2024
    WWW 2024 · Tat-Seng Chua
  148. ProtT3: Protein-to-Text Generation for Text-based Protein Understanding
  149. ReactXT: Understanding Molecular "Reaction-ship" via Reaction-Contextualized Molecule-Text Pretraining
  150. STYLE: Improving Domain Transferability of Asking Clarification Questions in Large Language Model Powered Conversational Agents
  151. Search-in-the-Chain: Interactively Enhancing Large Language Models with Search for Knowledge-intensive Tasks
  152. Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image Retrieval
  153. Strength Lies in Differences! Improving Strategy Planning for Non-collaborative Dialogues via Diversified User Simulation
  154. Temporally and Distributionally Robust Optimization for Cold-Start Recommendation
  155. Think Twice Before Trusting: Self-Detection for Large Language Models through Comprehensive Answer Reflection
  156. Towards 3D Molecule-Text Interpretation in Language Models
  157. Towards Human-centered Proactive Conversational Agents
  158. Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching
  159. Towards Neuron Attributions in Multi-Modal Large Language Models
  160. Uplift Modeling for Target User Attacks on Recommender Systems
  161. Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
  162. XNLP: An Interactive Demonstration System for Universal Structured NLP