A novel Engagement Detection with Multi-Task Training (ED-MTT) system

Last update: Nov 11, 2022

Related tags

Overview

ED-MTT

A novel Engagement Detection with Multi-Task Training (ED-MTT) system which minimizes MSE and triplet loss together to determine the engagement level of students in an e-learning environment. You can check the colab notebook bellow for detailed explanatoins about data loading and code execution.

Introduction & Problem Definition

With the Covid-19 outbreak, the online working and learning environments became essential in our lives. For this reason, automatic analysis of non-verbal communication becomes crucial in online environments.

Engagement level is a type of social signal that can be predicted from facial expression and body pose. To this end, we propose an end-to-end deep learning-based system that detects the engagement level of the subject in an e-learning environment.

The engagement level feedback is important because:

Make aware students of their performance in classes.
Will help instructors to detect confusing or unclear parts of the teaching material.

Model Architecture

The proposed system first extracts features with OpenFace, then aggregates frames in a window for calculating feature statistics as additional features. Finally, uses Bi-LSTM for generating vector embeddings from input sequences. In this system, we introduce a triplet loss as an auxiliary task and design the system as a multi-task training framework by taking inspiration from, where self-supervised contrastive learning of multi-view facial expressions was introduced. To the best of our knowledge, this is a novel approach in engagement detection literature. The key novelty of this work is the multi-task training framework using triplet loss together with Mean Squared Error (MSE). The main contributions of this paper are as follows:

Multi-task training with triplet and MSE losses introduces an additional regularization and reduces over-fitting due to very small sample size.
Using triplet loss mitigates the label reliability problem since it measures relative similarity between samples.
A system with lightweight feature extraction is efficient and highly suitable for real-life applications.

Dataset

We evaluate the performance of ED-MTT on a publicly available ``Engagement in The Wild'' dataset which is comprised of separated training and validation sets.

The dataset is comprised of 78 subjects (25 females and 53 males) whose ages are ranged from 19 to 27. Each subject is recorded while watching an approximately 5 minutes long stimulus video of a Korean Language lecture.

Results

We compare the performance of ED-MTT with 9 different works from the state-of-the-art which will be reviewed in the rest of this section. Our results show that ED-MTT outperforms these state-of-the-art methods with at least a 5.74% improvement on MSE.

Repository structure

ED-MTT
│   README.md
│   Engagement_Labels.txt
|   ED-MTT.ipynb

└───code
│   │   dataloader.py
|   |   model.py
|   |   train.py
|   |   test.py
│   │   fix_path.py
|   |   utils.py
|   |   requirements.txt

└───configs
    │   batchnorm_default.yaml
    │   sweep.yaml

Running the Code

To train the experiments and manage the experiments, we used PyTorch Lightning together with Weights&Biases. All the detailed explonations to;

Load data and pre-trained weights,
Train the model from scratch,
Manage expriments and hyper-parameter search with wandb,
Reproduce the results presented in the paper,

are shown in ED-MTT.ipynb colab notebook.

A novel Engagement Detection with Multi-Task Training (ED-MTT) system

Related tags

Overview

ED-MTT

Introduction & Problem Definition

Model Architecture

Dataset

Results

Repository structure

Running the Code

Owner

Onur Çopur

PyTorch code for training MM-DistillNet for multimodal knowledge distillation

PyTorch implementation for Partially View-aligned Representation Learning with Noise-robust Contrastive Loss (CVPR 2021)

Bayes-Newton—A Gaussian process library in JAX, with a unifying view of approximate Bayesian inference as variants of Newton's algorithm.

PyTorch/GPU re-implementation of the paper Masked Autoencoders Are Scalable Vision Learners

[NeurIPS 2021] Official implementation of paper "Learning to Simulate Self-driven Particles System with Coordinated Policy Optimization".

PASTRIE: A Corpus of Prepositions Annotated with Supersense Tags in Reddit International English

HandTailor: Towards High-Precision Monocular 3D Hand Recovery

Age and Gender prediction using Keras

Code for the AAAI-2022 paper: Imagine by Reasoning: A Reasoning-Based Implicit Semantic Data Augmentation for Long-Tailed Classification

3D position tracking for soccer players with multi-camera videos

[ACMMM 2021 Oral] Enhanced Invertible Encoding for Learned Image Compression

Invert and perturb GAN images for test-time ensembling

Study of human inductive biases in CNNs and Transformers.

The implementation of "Bootstrapping Semantic Segmentation with Regional Contrast".

Source code of NeurIPS 2021 Paper ''Be Confident! Towards Trustworthy Graph Neural Networks via Confidence Calibration''

I decide to sync up this repo and self-critical.pytorch. (The old master is in old master branch for archive)

Pytorch implementation of the paper: "SAPNet: Segmentation-Aware Progressive Network for Perceptual Contrastive Image Deraining"

Depth-Aware Video Frame Interpolation (CVPR 2019)

Api for getting bin info and getting encrypted card details for adyen.

UI2I via StyleGAN2 - Unsupervised image-to-image translation method via pre-trained StyleGAN2 network