PPaperPicks

Gao Huang

Tsinghua University, Department of Automation, BNRist, Beijing, China

60 papers at tracked venues · 46 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
  2. SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
  3. Vision Transformers Are Circulant Attention Learners
  4. 4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language Models
  5. ART: Anonymous Region Transformer for Variable Multi-Layer Transparent Image Generation
  6. Absolute Zero: Reinforced Self-play Reasoning with Zero Data
  7. Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
  8. CODA: Repurposing Continuous VAEs for Discrete Tokenization
  9. CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning
  10. DTOS: Dynamic Time Object Sensing with Large Multimodal Model
  11. DenseGrounding: Improving Dense Language-Vision Semantics for Ego-centric 3D Visual Grounding
  12. Differential Transformer
  13. DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints
  14. Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
  15. DyMoDreamer: World Modeling with Dynamic Modulation
  16. Dynamic Diffusion Transformer
  17. EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance
  18. Everything to the Synthetic: Diffusion-driven Test-time Adaptation via Synthetic-Domain Alignment
  19. GridMix: Exploring Spatial Modulation for Neural Fields in PDE Modeling
  20. HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
  21. How Far Is Video Generation from World Model: A Physical Law Perspective
  22. IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance
  23. Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials
  24. Model Surgery: Modulating LLM's Behavior Via Simple Parameter Editing
  25. ProxyTransformation: Preshaping Point Cloud Manifold With Proxy Attention For 3D Visual Grounding
  26. Towards Understanding Text Hallucination of Diffusion Models via Local Generation Bias
  27. UltraDP: Generalizable Carotid Ultrasound Scanning with Force-Aware Diffusion Policy
  28. Video Perception Models for 3D Scene Synthesis
  29. A Unified Interaction Control Framework for Safe Robotic Ultrasound Scanning with Human-Intention-Aware Compliance
  30. ADDP: Learning General Representations for Image Recognition and Generation with Alternating Denoising Diffusion Process
  31. AdaNAT: Exploring Adaptive Policy for Token-Based Image Generation
  32. Agent Attention: On the Integration of Softmax and Linear Attention
  33. Boosting LLM Agents with Recursive Contemplation for Effective Deception Handling
  34. Bridging the Divide: Reconsidering Softmax and Linear Attention
  35. COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video Editing
  36. Cardiac Copilot: Automatic Probe Guidance for Echocardiography with World Model
  37. DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution
  38. Demystify Mamba in Vision: A Linear Attention Perspective
  39. DyFADet: Dynamic Feature Aggregation for Temporal Action Detection
  40. Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation
  41. ENAT: Rethinking Spatial-temporal Interactions in Token-based Image Synthesis
  42. Efficient Diffusion Transformer with Step-Wise Dynamic Attention Mediators
  43. ExpeL: LLM Agents Are Experiential Learners
  44. Exploring Temporal Feature Correlation for Efficient and Stable Video Semantic Segmentation
  45. Exploring Text-to-Motion Generation with Human Preference
  46. GRA: Detecting Oriented Objects Through Group-Wise Rotating and Attention
  47. GSVA: Generalized Segmentation via Multimodal Large Language Models
  48. Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering
  49. LLaVA-UHD: An LMM Perceiving Any Aspect Ratio and High-Resolution Images
  50. Learning 1D Causal Visual Representation with De-focus Attention Networks
  51. Mask Grounding for Referring Image Segmentation
  52. Prompt-Free Diffusion: Taking "Text" Out of Text-to-Image Diffusion Models
  53. PsychoGAT: A Novel Psychological Measurement Paradigm through Interactive Fiction Games with LLM Agents
  54. Rethinking the Architecture Design for Efficient Generic Event Boundary Detection
  55. Revisiting Non-Autoregressive Transformers for Efficient Image Synthesis
  56. Segment3D: Learning Fine-Grained Class-Agnostic 3D Segmentation Without Manual Labels
  57. SimPro: A Simple Probabilistic Framework Towards Realistic Long-Tailed Semi-Supervised Learning
  58. Smooth Diffusion: Crafting Smooth Latent Spaces in Diffusion Models
  59. Structure-aware World Model for Probe Guidance via Large-scale Self-supervised Pre-train
  60. Training an Open-Vocabulary Monocular 3D Detection Model without 3D Data