Publications
Selected papers are listed first. Equal contribution/Corresponding author is marked with *.
Selected Publications
NegAS: Negative Label Guided Attention and Scoring for Out-of-Distribution Object Detection with Vision-Language Models
GDPO-SR: Group Direct Preference Optimization for One-Step Generative Image Super-Resolution
Mamba-ISTD: Multi-context Hierarchical Aggregation Mamba for Robust Infrared Small Target Detection
ADIL: Adaptive Dual Imitation Learning for Real-time Object Detection in Remote Sensing Images
DPO-SR: Direct Perceptual Preference Optimization for Real-World Image Super-Resolution
DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow Editing
Fine-structure Preserved Real-world Image Super-resolution via Transfer VAE Training
Dual-Rate Dynamic Teacher for Source-Free Domain Adaptive Object Detection
Exploring Dynamic Transformer for Efficient Object Tracking
Open Vocabulary 3D Scene Understanding via Geometry Guided Self-distillation
UniVS: Unified and Universal Video Segmentation with Prompts as Queries
SeeSR: Towards Semantics-Aware Real-World Image Super-Resolution
FPR: False Positive Rectification for Weakly Supervised Semantic Segmentation
One-to-Few Label Assignment for End-to-End Dense Detection
DynaMask: Dynamic Mask Selection for Instance Segmentation
SIM: Semantic-aware Instance Mask Generation for Box-Supervised Instance Segmentation
MDQE: Mining Discriminative Query Embeddings to Segment Occluded Instances on Challenging Videos
MSF: Motion-guided Sequential Fusion for Efficient 3D Object Detection from Point Cloud Sequences
One-stage Visual Relationship Referring with Transformers and Adaptive Message Passing
A Dual Weighting Label Assignment Scheme for Object Detection
Class-balanced Pixel-level Self-labeling for Domain Adaptive Semantic Segmentation
Voxel Set Transformer: A Set-to-Set Approach to 3D Object Detection from Point Clouds
Category Dictionary Guided Unsupervised Domain Adaptation for Object Detection
Spatial Feature Calibration and Temporal Fusion for Effective One-stage Video Instance Segmentation
LST-Net: Learning a Convolutional Neural Network with a Learnable Sparse Transform
Dynamic Anchor Feature Selection for Single-Shot Object Detection
