Official code repository for the work: "The Implicit Values of A Good Hand Shake: Handheld Multi-Frame Neural Depth Refinement"

Last update: Dec 14, 2022

Related tags

Deep Learning HNDR

Overview

Handheld Multi-Frame Neural Depth Refinement

This is the official code repository for the work: The Implicit Values of A Good Hand Shake: Handheld Multi-Frame Neural Depth Refinement .

If you use parts of this work, or otherwise take inspiration from it, please considering citing our paper:

@article{chugunov2021implicit,
  title={The Implicit Values of A Good Hand Shake: Handheld Multi-Frame Neural Depth Refinement},
  author={Chugunov, Ilya and Zhang, Yuxuan and Xia, Zhihao and Zhang, Cecilia and Chen, Jiawen and Heide, Felix},
  journal={arXiv preprint arXiv:2111.13738},
  year={2021}
}

Requirements:

Developed using PyTorch 1.10.0 on Linux x64 machine
Condensed package requirements are in \requirements.txt. Note that this contains the package versions at the time of publishing, if you update to, for example, a newer version of PyTorch you will need to watch out for changes in class/function calls

Data:

Download data from this Google Drive link and unpack into the \data folder
Each folder corresponds to a scene [castle, eagle, elephant, frog, ganesha, gourd, rocks, thinker] and contains four files.
- model.pt is the frozen, trained MLP corresponding to the scene
- frame_bundle.npz is the recorded bundle data (images, depth, and poses)
- reprojected_lidar.npy is the merged LiDAR depth baseline as described in the paper
- snapshot.mp4 is a video of the recorded snapshot for visualization purposes

An explanation of the format and contents of the frame bundles (frame_bundle.npz) is given in an interactive format in \0_data_format.ipynb. We recommend you go through this jupyter notebook before you record your own bundles or otherwise manipulate the data.

Project Structure:

HNDR
  ├── checkpoints  
  │   └── // folder for network checkpoints
  ├── data  
  │   └── // folder for recorded bundle data
  ├── utils  
  │   ├── dataloader.py  // dataloader class for bundle data
  │   ├── neural_blocks.py  // MLP blocks and positional encoding
  │   └── utils.py  // miscellaneous helper functions (e.g. grid/patch sample)
  ├── 0_data_format.ipynb  // interactive tutorial for understanding bundle data
  ├── 1_reconstruction.ipynb  // interactive tutorial for depth reconstruction
  ├── model.py  // the learned implicit depth model
  │             // -> reproject points, query MLP for offsets, visualization
  ├── README.md  // a README in the README, how meta
  ├── requirements.txt  // frozen package requirements
  ├── train.py  // wrapper class for arg parsing and setting up training loop
  └── train.sh  // example script to run training

Reconstruction:

The jupyter notebook \1_reconstruction.ipynb contains an interactive tutorial for depth reconstruction: loading a model, loading a bundle, generating depth.

Training:

The script \train.sh demonstrates a basic call of \train.py to train a model on the gourd scene data. It contains the arguments

checkpoint_path - path to save model and tensorboard checkpoints
device - device for training [cpu, cuda]
bundle_path - path to the bundle data

For other training arguments, see the argument parser section of \train.py.

Best of luck,
Ilya

Official code repository for the work: "The Implicit Values of A Good Hand Shake: Handheld Multi-Frame Neural Depth Refinement"

Related tags

Overview

Handheld Multi-Frame Neural Depth Refinement

Requirements:

Data:

Project Structure:

Reconstruction:

Training:

Owner

Simulating an AI playing 2048 using the Expectimax algorithm

The official code repo of "HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection"

Capsule endoscopy detection DACON challenge

Codes for AAAI 2022 paper: Context-aware Health Event Prediction via Transition Functions on Dynamic Disease Graphs

This is a project based on retinaface face detection, including ghostnet and mobilenetv3

FrankMocap: A Strong and Easy-to-use Single View 3D Hand+Body Pose Estimator

Code repository for "Reducing Underflow in Mixed Precision Training by Gradient Scaling" presented at IJCAI '20

This repository contains the code for the binaural-detection model used in the publication arXiv:2111.04637

YOLOv5 Series Multi-backbone, Pruning and quantization Compression Tool Box.

Detecting drunk people through thermal images using Deep Learning (CNN)

[ACM MM 2021] Diverse Image Inpainting with Bidirectional and Autoregressive Transformers

The 1st Place Solution of the Facebook AI Image Similarity Challenge (ISC21) : Descriptor Track.

Turi Create simplifies the development of custom machine learning models.

A modular, research-friendly framework for high-performance and inference of sequence models at many scales

[ICCV2021] Official Pytorch implementation for SDGZSL (Semantics Disentangling for Generalized Zero-Shot Learning)

Official Pytorch implementation of "Beyond Static Features for Temporally Consistent 3D Human Pose and Shape from a Video", CVPR 2021

EGNN - Implementation of E(n)-Equivariant Graph Neural Networks, in Pytorch

Lightweight stereo matching network based on MobileNetV1 and MobileNetV2

Image processing in Python

Hunt down social media accounts by username across social networks