PPaperPicks

Yan Lu

Microsoft Research Asia, Beijing, China

45 papers at tracked venues · 41 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. CSPO: Alleviating Reward Ambiguity for Structured Table-to-LaTeX Generation
  2. Closing the Modality Reasoning Gap for Speech Large Language Models
  3. InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
  4. SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
  5. Toward Natural and Companionable Virtual Agents via Cross-Temporal Emotional Modeling
  6. When Systems Take Initiative: A Design Framework for Adaptive, Mixed-initiative Database Querying
  7. Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video
  8. DLF: Extreme Image Compression with Dual-Generative Latent Fusion
  9. Deciphering Functions of Neurons in Vision-Language Models
  10. Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
  11. FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
  12. FuncGenFoil: Airfoil Generation and Editing Model in Function Space
  13. I2VGuard: Safeguarding Images against Misuse in Diffusion-based Image-to-Video Models
  14. Image as a World: Generating Interactive World from Single Image via Panoramic Video Generation
  15. Omnidirectional 3D Scene Reconstruction from Single Image
  16. One-Step Diffusion-Based Image Compression with Semantic Distillation
  17. PICD: Versatile Perceptual Image Compression with Diffusion Rendering
  18. PRING: Rethinking Protein-Protein Interaction Prediction from Pairs to Graphs
  19. Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
  20. SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
  21. STAR: A Benchmark for Astronomical Star Fields Super-Resolution
  22. SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation
  23. Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning
  24. StreamGS: Online Generalizable Gaussian Splatting Reconstruction for Unposed Image Streams
  25. Towards Anytime Retrieval: A Benchmark for Anytime Person Re-Identification
  26. Towards Practical Real-Time Neural Video Compression
  27. TrInk: Ink Generation with Transformer Network
  28. UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic Grasping
  29. VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
  30. Arbitrary-Scale Video Super-resolution Guided by Dynamic Context
  31. Breaking through the learning plateaus of in-context learning in Transformer
  32. Convert and Speak: Zero-shot Accent Conversion with Minimum Supervision
  33. Diffusion Model with Cross Attention as an Inductive Bias for Disentanglement
  34. Generative Latent Coding for Ultra-Low Bitrate Image Compression
  35. Hierarchical Intra-Modal Correlation Learning for Label-Free 3D Semantic Segmentation
  36. Implicit Motion Function
  37. Long-Term Temporal Context Gathering for Neural Video Compression
  38. Mask-Based Modeling for Neural Radiance Fields
  39. MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators
  40. MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
  41. Neural Video Compression with Feature Modulation
  42. QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic Decomposition
  43. Slot-VLM: Object-Event Slots for Video-Language Modeling
  44. Text Grouping Adapter: Adapting Pre-Trained Text Detector for Layout Analysis
  45. Unifying Multi-Modal Uncertainty Modeling and Semantic Alignment for Text-to-Image Person Re-identification