Google Landmark Recogntion and Retrieval 2021 Solutions

Last update: Nov 25, 2022

Related tags

Overview

Google Landmark Recogntion and Retrieval 2021 Solutions

In this repository you can find solution and code for Google Landmark Recognition 2021 and Google Landmark Retrieval 2021 competitions (both in top-100).

Brief Summary

My solution is based on the latest modeling from the previous competition and strong post-processing based on re-ranking and using side models like detectors. I used single RTX 3080, EfficientNet B0 and only competition data for training.

Model and loss function

I used the same model and loss as the winner team of the previous competition as a base. Since I had only single RTX 3080, I hadn't enough time to experiment with that and change it. The only things I managed to test is Subcenter ArcMarginProduct as the last block of model and ArcFaceLossAdaptiveMargin loss function, which has been used by the 2nd place team in the previous year. Both those things gave me a signifact score boost (around 4% on CV and 5% on LB).

Setting up the training and validation

Optimizing and scheduling

Optimizer - Ranger (lr=0.003)
Scheduler - CosineAnnealingLR (T_max=12) + 1 epoch Warm-Up

Training stages

I found the best perfomance in training for 15 epochs and 5 stages:

(1-3) - Resize to image size, Horizontal Flip
(4-6) - Resize to bigger image size, Random Crop to image size, Horizontal Flip
(7-9) - Resize to bigger image size, Random Crop to image size, Horizontal Flip, Coarse Dropout with one big square (CutMix)
(10-12) - Resize to bigger image size, Random Crop to image size, Horizontal Flip, FMix, CutMix, MixUp
(13-15) - Resize to bigger image size, Random Crop to image size, Horizontal Flip

I used default Normalization on all the epochs.

Validation scheme

Since I hadn't enough hardware, this became my first competition where I wasn't able to use a K-fold validation, but at least I saw stable CV and CV/LB correlation at the previous competitions, so I used simple stratified train-test split in 0.8, 0.2 ratio. I also oversampled all the samples up to 5 for each class.

Inference and Post-Processing:

Change class to non-landmark if it was predicted more than 20 times .
Using pretrained YoloV5 for detecting non-landmark images. All classes are used, boxes with confidence < 0.5 are dropped. If total area of boxes is greater than total_image_area / 2.7, the sample is marked as non-landmark. I tried to use YoloV5 for cleaning the train dataset as well, but it only decreased a score.
Tuned post-processing from this paper, based on the cosine similarity between train and test images to non-landmark ones.
Higher image size for extracting embeddings on inference.
Also using public train dataset as an external data for extracting embeddings.

Didn't work for me

Knowledge Distillation
Resnet architectures (on average they were worse than effnets)
Adding an external non-landmark class to training from 2019 test dataset
Train binary non-landmark classifier

Transfer Learning on the full dataset and Label Smoothing should be useful here, but I didn't have time to test it.

Google Landmark Recogntion and Retrieval 2021 Solutions

Related tags

Overview

Google Landmark Recogntion and Retrieval 2021 Solutions

Brief Summary

Model and loss function

Setting up the training and validation

Optimizing and scheduling

Training stages

Validation scheme

Inference and Post-Processing:

Didn't work for me

Owner

Vadim Timakin

Deep Learning as a Cloud API Service.

Semi-supervised Learning for Sentiment Analysis

An Empirical Investigation of Model-to-Model Distribution Shifts in Trained Convolutional Filters

git git《Transformer Meets Tracker: Exploiting Temporal Context for Robust Visual Tracking》(CVPR 2021) GitHub:git2] 《Masksembles for Uncertainty Estimation》(CVPR 2021) GitHub:git3]

Densely Connected Search Space for More Flexible Neural Architecture Search (CVPR2020)

[CVPR 2021] A Peek Into the Reasoning of Neural Networks: Interpreting with Structural Visual Concepts

Implementation of Wasserstein adversarial attacks.

Algorithmic trading with deep learning experiments

PyTorch version of Stable Baselines, reliable implementations of reinforcement learning algorithms.

Code base for NeurIPS 2021 publication titled Kernel Functional Optimisation (KFO)

[CVPR'21] FedDG: Federated Domain Generalization on Medical Image Segmentation via Episodic Learning in Continuous Frequency Space

OpenMMLab Semantic Segmentation Toolbox and Benchmark.

A repository for interferometer controller code.

PyTorch implementation of "Transparency by Design: Closing the Gap Between Performance and Interpretability in Visual Reasoning"

Instance-level Image Retrieval using Reranking Transformers

This repository contains the source code and data for reproducing results of Deep Continuous Clustering paper

A python bot to move your mouse every few seconds to appear active on Skype, Teams or Zoom as you go AFK. 🐭 🤖

use machine learning to recognize gesture on raspberrypi

[NeurIPS 2021] SSUL: Semantic Segmentation with Unknown Label for Exemplar-based Class-Incremental Learning

This project intends to use SVM supervised learning to determine whether or not an individual is diabetic given certain attributes.