PPaperPicks

Zhaoxiang Zhang

Chinese Academy of Sciences, Institute of Automation, Center for Research on Intelligent Perception and Computing (CRIPAC), National Laboratory of Pattern Recognition (NLPR), Beijing, China

59 papers at tracked venues · 49 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. AdaField: Generalizable Surface Pressure Modeling with Physics-Informed Pre-training and Flow-Conditioned Adaptation
  2. CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization
  3. AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
  4. C2KD: Cross-layer and Cross-head Knowledge Distillation for Small Language Model-based Recommendation
  5. Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
  6. CityGaussianV2: Efficient and Geometrically Accurate Reconstruction for Large-Scale Scenes
  7. DexVLG: Dexterous Vision-Language-Grasp Model at Scale
  8. DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving
  9. DrivingGPT: Unifying Driving World Modeling and Planning with Multi-Modal Autoregressive Transformers
  10. End-to-End Driving with Online Trajectory Evaluation via BEV World Model
  11. Enhancing End-to-End Autonomous Driving with Latent World Model
  12. FIRM: Flexible Interactive Reflection ReMoval
  13. FlexDrive: Toward Trajectory Flexibility in Driving Scene Gaussian Splatting Reconstruction and Rendering
  14. FreeSim: Toward Free-viewpoint Camera Simulation in Driving Scenes
  15. FreeVS: Generative View Synthesis on Free Driving Trajectory
  16. KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
  17. LayerAnimate: Layer-Level Control for Animation
  18. M2RC-EVAL: Massively Multilingual Repository-level Code Completion Evaluation
  19. MCOP: Multi-UAV Collaborative Occupancy Prediction
  20. MIO: A Foundation Model on Multimodal Tokens
  21. MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
  22. MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
  23. McEval: Massively Multilingual Code Evaluation
  24. OmniBench: Towards The Future of Universal Omni-Language Models
  25. OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models
  26. Reconstructive Visual Instruction Tuning
  27. Ross3d: Reconstructive Visual Instruction Tuning With 3D-Awareness
  28. SceneX: Procedural Controllable Large-Scale Scene Generation
  29. TC-Light: Temporally Coherent Generative Rendering for Realistic World Transfer
  30. Top-Down Guidance for Learning Object-Centric Representations
  31. UIPro: Unleashing Superior Interaction Capability for GUI Agents
  32. CSOT: Cross-scan Object Transfer for Semi-Supervised LiDAR Object Detection
  33. CityGaussian: Real-Time High-Quality Large-Scale Scene Rendering with Gaussians
  34. Compositional Inversion for Stable Diffusion Models
  35. Continual Forgetting for Pre-Trained Vision Models
  36. Driving Into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving
  37. DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model
  38. Enhancing Visual Continual Learning with Language-Guided Supervision
  39. Expanding Scene Graph Boundaries: Fully Open-Vocabulary Scene Graph Generation via Visual-Concept Alignment and Retention
  40. Fully Data-Driven Pseudo Label Estimation for Pointly-Supervised Panoptic Segmentation
  41. General Geometry-Aware Weakly Supervised 3D Object Detection
  42. Generative Active Learning for Image Synthesis Personalization
  43. HardMo: A Large-Scale Hardcase Dataset for Motion Capture
  44. MaterialSeg3D: Segmenting Dense Materials from 2D Priors for 3D Assets
  45. MemoNav: Working Memory Model for Visual Navigation
  46. MixSup: Mixed-grained Supervision for Label-efficient LiDAR-based 3D Object Detection
  47. Monocular Occupancy Prediction for Scalable Indoor Scenes
  48. OneTrack: Demystifying the Conflict Between Detection and Tracking in End-to-End 3D Trackers
  49. Open Vocabulary 3D Scene Understanding via Geometry Guided Self-Distillation
  50. OpenSatMap: A Fine-grained High-resolution Satellite Dataset for Large-scale Map Construction
  51. PanoOcc: Unified Occupancy Representation for Camera-based 3D Panoptic Segmentation
  52. Point-Supervised Panoptic Segmentation via Estimating Pseudo Labels from Learnable Distance
  53. RCL: Reliable Continual Learning for Unified Failure Detection
  54. Robust Depth Enhancement via Polarization Prompt Fusion Tuning
  55. RoleAgent: Building, Interacting, and Benchmarking High-quality Role-Playing Agents from Scripts
  56. RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models
  57. StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation Framework
  58. VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector Quantization
  59. Voxel Mamba: Group-Free State Space Models for Point Cloud based 3D Object Detection