PromptDet: Expand Your Detector Vocabulary with Uncurated Images

Last update: Dec 20, 2022

Overview

PromptDet: Expand Your Detector Vocabulary with Uncurated Images

Introduction

The goal of this work is to establish a scalable pipeline for expanding an object detector towards novel/unseen categories, using zero manual annotations. To achieve that, we make the following four contributions: (i) in pursuit of generalisation, we propose a two-stage open-vocabulary object detector that categorises each box proposal by a classifier generated from the text encoder of a pre-trained visual-language model; (ii) To pair the visual latent space (from RPN box proposal) with that of the pre-trained text encoder, we propose the idea of regional prompt learning to optimise a couple of learnable prompt vectors, converting the textual embedding space to fit those visually object-centric images; (iii) To scale up the learning procedure towards detecting a wider spectrum of objects, we exploit the available online resource, iteratively updating the prompts, and later self-training the proposed detector with pseudo labels generated on a large corpus of noisy, uncurated web images. The self-trained detector, termed as PromptDet, significantly improves the detection performance on categories for which manual annotations are unavailable or hard to obtain, e.g. rare categories. Finally, (iv) to validate the necessity of our proposed components, we conduct extensive experiments on the challenging LVIS and MS-COCO dataset, showing superior performance over existing approaches with fewer additional training images and zero manual annotations whatsoever.

Training framework

Prerequisites

MMDetection version 2.16.0.
Please see get_started.md for installation and the basic usage of MMDetection.

Inference

./tools/dist_test.sh configs/promptdet/promptdet_mask_rcnn_r50_fpn_sample1e-3_mstrain_1x_lvis_v1.py work_dirs/promptdet_mask_rcnn_r50_fpn_sample1e-3_mstrain_1x_lvis_v1.pth 4 --eval bbox segm

Train

To be updated.

Models

For your convenience, we provide the following trained models (PromptDet) with mask AP.

Model	Epochs	Scale Jitter	Input Size	AP_novel	APc	AP_f	AP	Config	Download
PromptDet_R_50_FPN_1x	12	640~800	800x800	19.0	18.5	25.8	21.4	config	google / baidu
PromptDet_R_50_FPN_6x	72	100~1280	800x800	21.4	23.3	29.3	25.3	config	google / baidu

[0] All results are obtained with a single model and without any test time data augmentation such as multi-scale, flipping and etc..
[1] Refer to more details in config files in config/promptdet/.
[2] Extraction code of baidu netdisk: promptdet.

Acknowledgement

Thanks MMDetection team for the wonderful open source project!

Citation

If you find PromptDet useful in your research, please consider citing:

@inproceedings{feng2022promptdet,
    title={PromptDet: Expand Your Detector Vocabulary with Uncurated Images},
    author={Feng, Chengjian and Zhong, Yujie and Jie, Zequn and Chu, Xiangxiang and Ren, Haibing and Wei, Xiaolin and Xie, Weidi and Ma, Lin},
    journal={arXiv preprint arXiv:2203.16513},
    year={2022}
}

PromptDet: Expand Your Detector Vocabulary with Uncurated Images

Related tags

Overview

PromptDet: Expand Your Detector Vocabulary with Uncurated Images

Introduction

Training framework

Prerequisites

Inference

Train

Models

Acknowledgement

Citation

Owner

Implement face detection, and age and gender classification, and emotion classification.

A numpy-based implementation of RANSAC for fundamental matrix and homography estimation. The degeneracy updating and local optimization components are included and optional.

Few-Shot Object Detection via Association and DIscrimination

This repository is an official implementation of the paper MOTR: End-to-End Multiple-Object Tracking with TRansformer.

Real Time Object Detection and Classification using Yolo Algorithm.

BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization

Official PyTorch implementation of the paper: DeepSIM: Image Shape Manipulation from a Single Augmented Training Sample

PyTorch code accompanying our paper on Maximum Entropy Generators for Energy-Based Models

Interactive Visualization to empower domain experts to align ML model behaviors with their knowledge.

Temporal Segment Networks (TSN) in PyTorch

🕵 Artificial Intelligence for social control of public administration

Code for: Gradient-based Hierarchical Clustering using Continuous Representations of Trees in Hyperbolic Space. Nicholas Monath, Manzil Zaheer, Daniel Silva, Andrew McCallum, Amr Ahmed. KDD 2019.

Codes for building and training the neural network model described in Domain-informed neural networks for interaction localization within astroparticle experiments.

Fashion Recommender System With Python

Evaluation toolkit of the informative tracking benchmark comprising 9 scenarios, 180 diverse videos, and new challenges.

Final project code: Implementing MAE with downscaled encoders and datasets, for ESE546 FA21 at University of Pennsylvania

Angular & Electron desktop UI framework. Angular components for native looking and behaving macOS desktop UI (Electron/Web)

PyTorch implementation of PNASNet-5 on ImageNet

DeepStruc is a Conditional Variational Autoencoder which can predict the mono-metallic nanoparticle from a Pair Distribution Function.

PyTorch implementation of deep GRAph Contrastive rEpresentation learning (GRACE).