GLIP: Grounded Language-Image Pre-training

Updates

12/06/2021: GLIP paper on arxiv https://arxiv.org/abs/2112.03857. Code and Model are under internal review and will release soon. Stay tuned!

11/23/2021: Project page built.

Introduction

This repository is the project page for GLIP, containing necessary instructions to reproduce the results presented in the paper. This paper presents a grounded language-image pre-training (GLIP) model for learning object-level, language-aware, and semantic-rich visual representations. GLIP unifies object detection and phrase grounding for pre-training. The unification brings two benefits: 1) it allows GLIP to learn from both detection and grounding data to improve both tasks and bootstrap a good grounding model; 2) GLIP can leverage massive image-text pairs by generating grounding boxes in a self-training fashion, making the learned representation semantic-rich. In our experiments, we pre-train GLIP on 27M grounding data, including 3M human-annotated and 24M web-crawled image-text pairs. The learned representations demonstrate strong zero-shot and few-shot transferability to various object-level recognition tasks.

When directly evaluated on COCO and LVIS (without seeing any images in COCO during pre-training), GLIP achieves 49.8 AP and 26.9 AP, respectively, surpassing many supervised baselines.
After fine-tuned on COCO, GLIP achieves 60.8 AP on val and 61.5 AP on test-dev, surpassing prior SoTA.
When transferred to 13 downstream object detection tasks, a few-shot GLIP rivals with a fully-supervised Dynamic Head.

Supervised baselines on COCO object detection: Faster-RCNN w/ ResNet50 (40.2) or ResNet101 (42.0) from Detectron2, and DyHead w/ Swin-Tiny (49.7).

Citations

Please consider citing this paper if you use the code:

@inproceedings{harold_GLIP2021,
      title={Grounded Language-Image Pre-training},
      author={Liunian Harold Li* and Pengchuan Zhang* and Haotian Zhang* and Jianwei Yang and Chunyuan Li and Yiwu Zhong and Lijuan Wang and Lu Yuan and Lei Zhang and Jenq-Neng Hwang and Kai-Wei Chang and Jianfeng Gao},
      year={2021},
      booktitle={arXiv preprint arXiv:2112.03857},
}

GLIP: Grounded Language-Image Pre-training

Related tags

Overview

GLIP: Grounded Language-Image Pre-training

Updates

Introduction

Citations

Owner

Microsoft

A Python training and inference implementation of Yolov5 helmet detection in Jetson Xavier nx and Jetson nano

TorchMultimodal is a PyTorch library for training state-of-the-art multimodal multi-task models at scale.

Face recognize and crop them

Adaptive Attention Span for Reinforcement Learning

FairyTailor: Multimodal Generative Framework for Storytelling

UFT - Universal File Transfer With Python

OpenGAN: Open-Set Recognition via Open Data Generation

ProMP: Proximal Meta-Policy Search

Real-time LIDAR-based Urban Road and Sidewalk detection for Autonomous Vehicles 🚗

OpenFed: A Comprehensive and Versatile Open-Source Federated Learning Framework

Scaling Vision with Sparse Mixture of Experts

Neural Motion Learner With Python

Google AI Open Images - Object Detection Track: Open Solution

[NeurIPS2021] Code Release of Learning Transferable Perturbations

Benchmark for evaluating open-ended generation

Snscrape-jsonl-urls-extractor - Extracts urls from jsonl produced by snscrape

A model that attempts to learn and benefit from data collected on card counting.

Public scripts, services, and configuration for running a smart home K3S network cluster

Seq2seq - Sequence to Sequence Learning with Keras

High frequency AI based algorithmic trading module.