A concise but complete implementation of CLIP with various experimental improvements from recent papers

Last update: Dec 26, 2022

Overview

x-clip (wip)

A concise but complete implementation of CLIP with various experimental improvements from recent papers

Install

$ pip install x-clip

Usage

import torch
from x_clip import CLIP

clip = CLIP(
    dim_text = 512,
    dim_image = 512,
    dim_latent = 512,
    num_text_tokens = 10000,
    text_enc_depth = 6,
    text_seq_len = 256,
    text_heads = 8,
    num_visual_tokens = 512,
    visual_enc_depth = 6,
    visual_image_size = 256,
    visual_patch_size = 32,
    visual_heads = 8,
    use_all_token_embeds = True   # whether to use fine-grained contrastive learning (FILIP)
)

text = torch.randint(0, 10000, (4, 256))
images = torch.randn(4, 3, 256, 256)
mask = torch.ones_like(text).bool()

loss = clip(text, images, text_mask = mask, return_loss = True)
loss.backward()

Citations

@misc{radford2021learning,
    title   = {Learning Transferable Visual Models From Natural Language Supervision}, 
    author  = {Alec Radford and Jong Wook Kim and Chris Hallacy and Aditya Ramesh and Gabriel Goh and Sandhini Agarwal and Girish Sastry and Amanda Askell and Pamela Mishkin and Jack Clark and Gretchen Krueger and Ilya Sutskever},
    year    = {2021},
    eprint  = {2103.00020},
    archivePrefix = {arXiv},
    primaryClass = {cs.CV}
}

@misc{yao2021filip,
    title   = {FILIP: Fine-grained Interactive Language-Image Pre-Training}, 
    author  = {Lewei Yao and Runhui Huang and Lu Hou and Guansong Lu and Minzhe Niu and Hang Xu and Xiaodan Liang and Zhenguo Li and Xin Jiang and Chunjing Xu},
    year    = {2021},
    eprint  = {2111.07783},
    archivePrefix = {arXiv},
    primaryClass = {cs.CV}
}

Comments

Model forward outputs to text/image similarity score

Any insight on how to take the image/text embeddings (or nominal model forward output) to achieve a simple similarity score as done in the huggingface implementation? HF example here

In the original paper I see the dot products of the image/text encoder outputs were used, but here I was having troubles with the dimensions on the outputs.

opened by paulcjh 12

Using different encoders in CLIP

Hi, I am wondering if it was possible to use different encoders in CLIP ? For images not using vit but resnet for example. And is it possible to replace the text encoder by a features encoder for example ? If I have a vector of features for a given image and I want to use x-clip how should I do that ? I have made a code example that doesnt seems to work, here is what I did:

import torch
from x_clip import CLIP
import torch.nn as nn
from torchvision import models

class Image_Encoder(torch.nn.Module):
    #output size is (bs,512)
    def __init__(self):
        super(Image_Encoder, self).__init__()
        self.model_pre = models.resnet18(pretrained=False)
        self.base=nn.Sequential(*list(self.model_pre.children()))
        self.base[0]=nn.Conv2d(3, 64, kernel_size=(7, 7), stride=(2, 2), padding=(3, 3), bias=False)
        self.resnet=self.base[:-1]

    def forward(self, x):
        out=self.resnet(x).squeeze()
        return out


class features_encoder(torch.nn.Module):
    #output size is (bs,512)
    def __init__(self):
        super(features_encoder, self).__init__()
        self.model =nn.Linear(2048,512)

    def forward(self, x):
        out=self.model(x)
        return out

images_encoder=Image_Encoder()
features_encoder=features_encoder()

clip = CLIP(
    image_encoder = images_encoder,
    text_encoder = features_encoder,
    dim_image = 512,
    dim_text = 512,
    dim_latent = 512
)

features= torch.randn(4,2048)
images = torch.randn(4, 3, 256, 256)

loss = clip(features, images, return_loss = True)
loss.backward()

but I got the following error : forward() takes 2 positional arguments but 3 were given

Thanks

opened by ethancohen123 8

Visual ssl with channels different than 3

Hi, seems to be a bug when trying to use visual ssl with a different number of channel than 3 . I think the error came from the visual ssl type ~row 280 here:

#send a mock image tensor to instantiate parameters self.forward(torch.randn(1, 3, image_size, image_size))

opened by ethancohen123 4

Allow other types of visual SSL when initiating CLIP

In the following code as part of CLIP.__init__

        if use_visual_ssl:
            if visual_ssl_type == 'simsiam':
                ssl_type = SimSiam
            elif visual_ssl_type == 'simclr':
                ssl_type = partial(SimCLR, temperature = simclr_temperature)
            else:
                raise ValueError(f'unknown visual_ssl_type')

            self.visual_ssl = ssl_type(
                self.visual_transformer,
                image_size = visual_image_size,
                hidden_layer = visual_ssl_hidden_layer
            )

the visual self-supervised learning is hardcoded. I would suggest changing this to accept the visual SSL module as an argument when instantiating CLIP to allow flexibility in the same manner as it does for the image encoder and text encoder.

Example:

barlow = BarlowTwins(augmentatation_fns)
clip = CLIP(..., visual_ssl=barlow)

opened by Froskekongen 4

Extract Text and Image Latents

Hi, in the current implementation we can only extract text and image embedding (by set return_encodings=True) which are obtained before applying latent linear layers. Isn't it better to add an option to extract latent embeddings? Another importance of this is that with the current code, it is impossible to extract the similarity matrix between a batch of images and a batch of text.

opened by mmsamiei 2

NaN with mock data

Hi lucidrains,

Try this and it will NaN within 100 steps (latest Github code). The loss looks fine before NaN.

import torch
torch.backends.cudnn.allow_tf32 = True
torch.backends.cuda.matmul.allow_tf32 = True    
torch.backends.cudnn.benchmark = True

import random
import numpy as np
seed = 42
random.seed(seed)
np.random.seed(seed)
torch.manual_seed(seed)
torch.cuda.manual_seed_all(seed)

num_text_tokens = 10000
batch_sz = 12
text_seq_len = 256
visual_image_size = 256

# mock data

data_sz = 1000
all_text = torch.randint(0, num_text_tokens, (data_sz, text_seq_len)).cuda()
all_images = torch.randn(data_sz, 3, visual_image_size, visual_image_size).cuda()

text = torch.zeros((batch_sz, text_seq_len), dtype=torch.long).cuda()
images = torch.zeros((batch_sz, 3, visual_image_size, visual_image_size)).cuda()

##########################################################################################

import wandb
import datetime
wandb.init(project="Test", name=datetime.datetime.today().strftime('%Y-%m-%d-%H-%M-%S'), save_code=False)

from x_clip import CLIP

clip = CLIP(
    dim_text = 512,
    dim_image = 512,
    dim_latent = 512,
    num_text_tokens = num_text_tokens,
    text_enc_depth = 6,
    text_seq_len = text_seq_len,
    text_heads = 8,
    visual_enc_depth = 6,
    visual_image_size = visual_image_size,
    visual_patch_size = 32,
    visual_heads = 8,
    use_all_token_embeds = False,           # whether to use fine-grained contrastive learning (FILIP)
    decoupled_contrastive_learning = True,  # use decoupled contrastive learning (DCL) objective function, removing positive pairs from the denominator of the InfoNCE loss (CLOOB + DCL)
    extra_latent_projection = True,         # whether to use separate projections for text-to-image vs image-to-text comparisons (CLOOB)
    use_visual_ssl = True,                  # whether to do self supervised learning on iages
    visual_ssl_type = 'simclr',             # can be either 'simclr' or 'simsiam', depending on using DeCLIP or SLIP
    use_mlm = False,                        # use masked language learning (MLM) on text (DeCLIP)
    text_ssl_loss_weight = 0.05,            # weight for text MLM loss
    image_ssl_loss_weight = 0.05            # weight for image self-supervised learning loss
).cuda()

optimizer = torch.optim.Adam(clip.parameters(), lr=1e-4, betas=(0.9, 0.99))

for step in range(999999):
    for i in range(batch_sz):
        data_id = random.randrange(0, data_sz - 1)
        text[i] = all_text[data_id]
        images[i] = all_images[data_id]

    loss = clip(
        text,
        images,
        freeze_image_encoder = False,   # whether to freeze image encoder if using a pretrained image net, proposed by LiT paper
        return_loss = True              # needs to be set to True to return contrastive loss
    )
    clip.zero_grad()
    loss.backward()
    torch.nn.utils.clip_grad_norm_(clip.parameters(), 1.0)
    optimizer.step()

    now_loss = loss.item()
    wandb.log({"loss": now_loss}, step = step)
    print(step, now_loss)

    if 'nan' in str(now_loss):
        break

opened by BlinkDL 1

Unable to train to convergence (small dataset)

Hi nice work with x-clip. Hoping to play around with it and eventually combine it into your DALLE2 work.

Currently having some trouble training on roughly 30k image-text pairs. Loss eventually goes negative and starts producing Nan's. I've dropped learning rate down (1e-4) and I'm clipping gradients (max_norm=0.5).

Any thoughts on what are sane training params/configs on such a small dataset using x-clip?

opened by jacobwjs 9

Releases(0.12.0)

0.12.0(Dec 2, 2022)

null
Source code(tar.gz)
Source code(zip)
0.11.0(Oct 16, 2022)

null
Source code(tar.gz)
Source code(zip)
0.10.0(Sep 14, 2022)

null
Source code(tar.gz)
Source code(zip)
0.9.0(Aug 17, 2022)

null
Source code(tar.gz)
Source code(zip)
0.8.4(Aug 5, 2022)

null
Source code(tar.gz)
Source code(zip)
0.8.3(Aug 4, 2022)

null
Source code(tar.gz)
Source code(zip)
0.8.2(Aug 3, 2022)

null
Source code(tar.gz)
Source code(zip)
0.8.1(Aug 3, 2022)

null
Source code(tar.gz)
Source code(zip)
0.8.0(Aug 3, 2022)

null
Source code(tar.gz)
Source code(zip)
0.7.4(Aug 3, 2022)

null
Source code(tar.gz)
Source code(zip)
0.7.3(Aug 3, 2022)

null
Source code(tar.gz)
Source code(zip)
0.7.2(Aug 3, 2022)

null
Source code(tar.gz)
Source code(zip)
0.7.1(Jul 30, 2022)

null
Source code(tar.gz)
Source code(zip)
v0.7.0(Jun 23, 2022)

null
Source code(tar.gz)
Source code(zip)
v0.6.1(May 24, 2022)

null
Source code(tar.gz)
Source code(zip)
0.6.1(May 24, 2022)

Source code(tar.gz)
Source code(zip)
0.6.0(May 24, 2022)

Source code(tar.gz)
Source code(zip)
0.5.1(Apr 29, 2022)

Source code(tar.gz)
Source code(zip)
0.5.0(Apr 15, 2022)

Source code(tar.gz)
Source code(zip)
0.4.6(Apr 13, 2022)

Source code(tar.gz)
Source code(zip)
0.4.5(Apr 13, 2022)

Source code(tar.gz)
Source code(zip)
0.4.4(Apr 13, 2022)

Source code(tar.gz)
Source code(zip)
0.4.3(Apr 12, 2022)

Source code(tar.gz)
Source code(zip)
0.4.2(Apr 12, 2022)

Source code(tar.gz)
Source code(zip)
0.4.1(Apr 12, 2022)

Source code(tar.gz)
Source code(zip)
0.4.0(Apr 6, 2022)

Source code(tar.gz)
Source code(zip)
0.3.0(Mar 1, 2022)

Source code(tar.gz)
Source code(zip)
0.2.4(Mar 1, 2022)

Source code(tar.gz)
Source code(zip)
0.2.3(Feb 5, 2022)

Source code(tar.gz)
Source code(zip)
0.2.2(Jan 27, 2022)

Source code(tar.gz)
Source code(zip)

Owner

Phil Wang

Working with Attention. It's all we need

GitHub Repository

Pywonderland - A tour in the wonderland of math with python.

A Tour in the Wonderland of Math with Python A collection of python scripts for drawing beautiful figures and animating interesting algorithms in math

4.1k Jan 03, 2023

Photographic Image Synthesis with Cascaded Refinement Networks - Pytorch Implementation

Photographic Image Synthesis with Cascaded Refinement Networks-Pytorch (https://arxiv.org/abs/1707.09405) This is a Pytorch implementation of cascaded

63 Mar 27, 2022

Offical code for the paper: "Growing 3D Artefacts and Functional Machines with Neural Cellular Automata" https://arxiv.org/abs/2103.08737

Growing 3D Artefacts and Functional Machines with Neural Cellular Automata Video of more results: https://www.youtube.com/watch?v=-EzztzKoPeo Requirem

51 Jan 01, 2023

Code for the Weighted, Accelerated and Restarted Primal-dual algorithm. This algorithm achieves stable linear convergence for reconstruction from undersampled noisy measurements under an approximate sharpness condition. See the paper for details.

WARPd Code for the Weighted, Accelerated and Restarted Primal-dual algorithm. This algorithm achieves stable linear convergence for reconstruction fro

1 Apr 08, 2022

Using machine learning to predict undergrad college admissions.

College-Prediction Project- Overview: Many have tried, many have failed. Few trailblazers are ambitious enought to chase acceptance into the top 15 un

1 Jan 05, 2022

Classification Modeling: Probability of Default

Credit Risk Modeling in Python Introduction: If you've ever applied for a credit card or loan, you know that financial firms process your information

2 Nov 07, 2022

A data-driven maritime port simulator

PySeidon - A Data-Driven Maritime Port Simulator 🌊 Extendable and modular software for maritime port simulation. This software uses entity-component

6 Apr 10, 2022

Official codebase for Decision Transformer: Reinforcement Learning via Sequence Modeling.

Decision Transformer Lili Chen*, Kevin Lu*, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas†, and Igor M

1.4k Jan 07, 2023

Learning to Predict Gradients for Semi-Supervised Continual Learning

Learning to Predict Gradients for Semi-Supervised Continual Learning Code for project: "Learning to Predict Gradients for Semi-Supervised Continual Le

2 Mar 05, 2022

[NeurIPS 2021] Galerkin Transformer: a linear attention without softmax

[NeurIPS 2021] Galerkin Transformer: linear attention without softmax Summary A non-numerical analyst oriented explanation on Toward Data Science abou

159 Dec 20, 2022

A particular navigation route using satellite feed and can help in toll operations & traffic managemen

How about adding some info that can quanitfy the stress on a particular navigation route using satellite feed and can help in toll operations & traffic management The current analysis is on the satel

1 Feb 14, 2022

Surrogate- and Invariance-Boosted Contrastive Learning (SIB-CL)

Surrogate- and Invariance-Boosted Contrastive Learning (SIB-CL) This repository contains all source code used to generate the results in the article "

3 Jul 23, 2022

Hierarchical probabilistic 3D U-Net, with attention mechanisms (—𝘈𝘵𝘵𝘦𝘯𝘵𝘪𝘰𝘯 𝘜-𝘕𝘦𝘵, 𝘚𝘌𝘙𝘦𝘴𝘕𝘦𝘵) and a nested decoder structure with deep supervision (—𝘜𝘕𝘦𝘵++).

Hierarchical probabilistic 3D U-Net, with attention mechanisms (—𝘈𝘵𝘵𝘦𝘯𝘵𝘪𝘰𝘯 𝘜-𝘕𝘦𝘵, 𝘚𝘌𝘙𝘦𝘴𝘕𝘦𝘵) and a nested decoder structure with deep supervision (—𝘜𝘕𝘦𝘵++). Built in TensorFlow 2.5. Configured for vox

32 Dec 08, 2022

A concise but complete implementation of CLIP with various experimental improvements from recent papers

Related tags

Overview

x-clip (wip)

Install

Usage

Citations

Comments

Model forward outputs to text/image similarity score

Using different encoders in CLIP

Visual ssl with channels different than 3

Allow other types of visual SSL when initiating CLIP

Extract Text and Image Latents

NaN with mock data

Unable to train to convergence (small dataset)

Releases(0.12.0)

0.12.0(Dec 2, 2022)

0.11.0(Oct 16, 2022)

0.10.0(Sep 14, 2022)

0.9.0(Aug 17, 2022)

0.8.4(Aug 5, 2022)

0.8.3(Aug 4, 2022)

0.8.2(Aug 3, 2022)

0.8.1(Aug 3, 2022)

0.8.0(Aug 3, 2022)

0.7.4(Aug 3, 2022)

0.7.3(Aug 3, 2022)

0.7.2(Aug 3, 2022)

0.7.1(Jul 30, 2022)

v0.7.0(Jun 23, 2022)

v0.6.1(May 24, 2022)

0.6.1(May 24, 2022)

0.6.0(May 24, 2022)

0.5.1(Apr 29, 2022)

0.5.0(Apr 15, 2022)

0.4.6(Apr 13, 2022)

0.4.5(Apr 13, 2022)

0.4.4(Apr 13, 2022)

0.4.3(Apr 12, 2022)

0.4.2(Apr 12, 2022)

0.4.1(Apr 12, 2022)

0.4.0(Apr 6, 2022)

0.3.0(Mar 1, 2022)

0.2.4(Mar 1, 2022)

0.2.3(Feb 5, 2022)

0.2.2(Jan 27, 2022)

Owner

Phil Wang

Pywonderland - A tour in the wonderland of math with python.

Photographic Image Synthesis with Cascaded Refinement Networks - Pytorch Implementation

Offical code for the paper: "Growing 3D Artefacts and Functional Machines with Neural Cellular Automata" https://arxiv.org/abs/2103.08737

Code for the Weighted, Accelerated and Restarted Primal-dual algorithm. This algorithm achieves stable linear convergence for reconstruction from undersampled noisy measurements under an approximate sharpness condition. See the paper for details.

Using machine learning to predict undergrad college admissions.

Classification Modeling: Probability of Default

A data-driven maritime port simulator

Official codebase for Decision Transformer: Reinforcement Learning via Sequence Modeling.

Learning to Predict Gradients for Semi-Supervised Continual Learning

[NeurIPS 2021] Galerkin Transformer: a linear attention without softmax

A particular navigation route using satellite feed and can help in toll operations & traffic managemen

Surrogate- and Invariance-Boosted Contrastive Learning (SIB-CL)

Hierarchical probabilistic 3D U-Net, with attention mechanisms (—𝘈𝘵𝘵𝘦𝘯𝘵𝘪𝘰𝘯 𝘜-𝘕𝘦𝘵, 𝘚𝘌𝘙𝘦𝘴𝘕𝘦𝘵) and a nested decoder structure with deep supervision (—𝘜𝘕𝘦𝘵++).

Real-Time-Student-Attendence-System - Real Time Student Attendence System

This is the official code for the paper "Ad2Attack: Adaptive Adversarial Attack for Real-Time UAV Tracking".

Implementation of H-Transformer-1D, Hierarchical Attention for Sequence Learning

[ICLR'21] Counterfactual Generative Networks

This repository contains the code for the binaural-detection model used in the publication arXiv:2111.04637

An open source library for face detection in images. The face detection speed can reach 1000FPS.

This is an official implementation of CvT: Introducing Convolutions to Vision Transformers.