Multimodel deepfake detection
-
Updated
Aug 14, 2026 - Jupyter Notebook
Multimodel deepfake detection
Tone.me helps users improve their pronunciation in Mandarin
In this code, we have used common and well-known datasets such as the Toronto dataset available on Kaggle to create a sentiment analysis model from human voice. This model is designed based on the Bert model and is called Hubert.
Adaptive Speech Monitoring for Instruction-Critical Environments
Multimodal Model which take text audio and video to predict the turn taking. That is, to predict whether the speaker in a discussion will change.
Generates section wise topics and transcription for lecture videos and helps to control the lecture video playback based on generated topic-wise timestamps.
Monitor driving instructor speech with real-time keyword spotting and emotion classification to measure instruction quality.
Fine-tuned Wav2Vec2 for Hindi ASR with NLLB-200 translation to English — modular training/eval pipeline, ~30% WER
PyTorch implementation of MixCap: A Multimodal Video Captioning model fusing BLIP-2 (Visual) & Wav2Vec2 (Audio). Features a novel Dual-Target MixUp strategy for low-resource training.
Low-resource multimodal hate speech detection leveraging acoustic and textual representations for robust moderation in Telugu.
To associate your repository with the wave2vec2 topic, visit your repo's landing page and select "manage topics."