PPaperPicks

Hao Fei

University of Oxford, Department of Computer Science, Big Data Institute and OATML, Oxford, UK

70 papers at tracked venues · 62 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. DragNeXt: Rethinking Drag-Based Image Editing
  2. Dynamic Emotion and Personality Profiling for Multimodal Deception Detection
  3. Orthogonal Spatial-temporal Distributional Transfer for 4D Generation
  4. Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment
  5. Aristotle: Mastering Logical Reasoning with A Logic-Complete Decompose-Search-Resolve Framework
  6. CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
  7. CLEAR: A Framework Enabling Large Language Models to Discern Confusing Legal Paragraphs
  8. CogMAEC'25: The 1st Workshop on Cognition-oriented Multimodal Affective and Empathetic Computing
    ACM MM 2025 · Hao Fei
  9. Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
  10. David vs. Goliath: Cost-Efficient Financial QA via Cascaded Multi-Agent Reasoning
  11. Derm1M: A Million-Scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology
  12. Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion Knowledge
  13. FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents
  14. From Data Deluge to Data Curation: A Filtering-WoRA Paradigm for Efficient Text-based Person Search
  15. InTriage: Intelligent Telephone Triage in Pre-Hospital Emergency Care
  16. Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
  17. JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
  18. LEAF-Mamba: Local Emphatic and Adaptive Fusion State Space Model for RGB-D Salient Object Detection
  19. Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual Scene
  20. MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
  21. MuSLR: Multimodal Symbolic Logical Reasoning
  22. Multi-Granular Multimodal Clue Fusion for Meme Understanding
  23. On Path to Multimodal Generalist: General-Level and General-Bench
    ICML 2025 · Hao Fei
  24. PhysSplat: Efficient Physics Simulation for 3D Scenes via MLLM-Guided Gaussian Splatting
  25. The ACM Multimedia 2025 Grand Challenge of Avatar-based Multimodal Empathetic Conversation
  26. The ACM Multimedia 2025 Grand Challenge of Multimodal Conversational Aspect-based Sentiment Analysis
  27. Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
  28. Towards Semantic Equivalence of Tokenization in Multimodal LLM
  29. Universal Scene Graph Generation
  30. VEGAS: Towards Visually Explainable and Grounded Artificial Social Intelligence
  31. ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
  32. VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models
  33. VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
  34. Watch Out Your Album! On the Inadvertent Privacy Memorization in Multi-Modal Large Language Models
  35. When Words Smile: Generating Diverse Emotional Facial Expressions from Text
  36. Where, What, Why: Towards Explainable Driver Attention Prediction
  37. A Survey of Ontology Expansion for Conversational Understanding
  38. Actively Learn from LLMs with Uncertainty Propagation for Generalized Category Discovery
  39. ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
  40. De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
  41. Divide and Conquer: Legal Concept-guided Criminal Court View Generation
  42. Dysen-VDM: Empowering Dynamics-Aware Text-to-Video Diffusion with LLMs
    CVPR 2024 · Hao Fei
  43. EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot
    ACL 2024 · Hao Fei
  44. Faithful Logical Reasoning via Symbolic Chain-of-Thought
  45. From Multimodal LLM to Human-level AI: Modality, Instruction, Reasoning and Beyond
    ACM MM 2024 · Hao Fei
  46. Guided Knowledge Generation with Language Models for Commonsense Reasoning
  47. Harnessing Holistic Discourse Features and Triadic Interaction for Sentiment Quadruple Extraction in Dialogues
  48. I3: Intent-Introspective Retrieval Conditioned on Instructions
  49. Improving Expressive Power of Spectral Graph Neural Networks with Eigenvalue Correction
  50. LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning
  51. MMLSCU: A Dataset for Multi-modal Multi-domain Live Streaming Comment Understanding
  52. Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
  53. NExT-GPT: Any-to-Any Multimodal LLM
  54. OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
  55. PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis
  56. ProtT3: Protein-to-Text Generation for Text-based Protein Understanding
  57. RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation
  58. Recognizing Everything from All Modalities at Once: Grounded Multimodal Universal Information Extraction
  59. Reverse Multi-Choice Dialogue Commonsense Inference with Graph-of-Thought
  60. Revisiting Structured Sentiment Analysis as Latent Dependency Graph Parsing
  61. Self-Adaptive Fine-grained Multi-modal Data Augmentation for Semi-supervised Muti-modal Coreference Resolution
  62. SpeechEE: A Novel Benchmark for Speech Event Extraction
  63. Synergistic Dual Spatial-aware Generation of Image-to-text and Text-to-image
  64. Synergizing Large Language Models and Pre-Trained Smaller Models for Conversational Intent Discovery
  65. The ACM Multimedia 2024 Viual Spatial Description Grand Challenge
  66. Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
  67. Unified Generative and Discriminative Training for Multi-modal Large Language Models
  68. Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
    ICML 2024 · Hao Fei
  69. Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
    NeurIPS 2024 · Hao Fei
  70. XNLP: An Interactive Demonstration System for Universal Structured NLP
    ACL 2024 · Hao Fei