# Voice Sentinel

## Title
CAPSTONE PROJECT · COMPUTER SCIENCE · ASHESI UNIVERSITY · 2026
VOICE SENTINEL
AI-Powered Deepfake Audio Detection System for Cybersecurity Defence
STUDENT
Nana Kwaku Afriyie Ampadu-Boateng
SUPERVISOR
Dr. Govindha Yeluripati
DEPARTMENT
Computer Science — Ashesi University
Audio Deepfake Detection

## Deepfake Landscape
INTRODUCTION
The Deepfake Landscape: A New Era of Synthetic Media
What are deepfakes?
| Point |
| --- |
| AI-synthesized digital forgeries that convincingly mimic a person's **voice or likeness** — indistinguishable to the human ear. |
| Powered by **GANs**, **Diffusion Models**, **Voice Conversion** &amp; neural TTS architectures. |
| Platforms — ElevenLabs, HeyGen, VALL-E — require **only seconds** of reference audio and no technical skill. |
| Democratised access has transformed synthetic media from an academic curiosity into a ubiquitous cybersecurity threat. |
CIA Triad Impact
| Label | Color | Description |
| --- | --- | --- |
| Confidentiality | #EF4444 | Bypasses voice biometrics; impersonates authorised users to extract sensitive information. |
| Integrity | #F59E0B | Injects fabricated audio, falsifying evidence, market data, and official communications. |
| Availability | #22C55E | Disinformation campaigns disrupt access to trusted information channels. |
**Why this matters for AI in cybersecurity:** Machine learning introduces novel attack surfaces — model inversion, adversarial evasion, and data poisoning — that traditional firewalls cannot mitigate. Audio deepfakes exploit all three CIA dimensions simultaneously.
01
VOICE SENTINEL · ASHESI UNIVERSITY · 2026

## Why Audio Deepfakes?
THE PROBLEM
Why Focus on Audio Deepfakes?
The scale of the threat
| Number | Label | Source | Color |
| --- | --- | --- | --- |
| 73% | Human detection rate of audio deepfakes | Kimberly et al., 2023 | #EF4444 |
| $25M+ | Lost in a single deepfake voice fraud incident | Elmisery et al., 2025 | #F59E0B |
Why audio is the harder problem
| Challenge | Description |
| --- | --- |
| No visual cues | Video fakes have blinking/boundary artifacts; audio fakes have nothing to see. |
| Underresearched | Most detection work targets video; audio lags significantly behind (Dhabe et al., 2024). |
| African accent gap | Models trained on Western datasets misclassify genuine African voices as synthetic, leaving millions unprotected (Müller et al.). |
| Compression artifacts | WhatsApp voice notes hide traditional forensic signals. |
VOICE SENTINEL · ASHESI UNIVERSITY · 2026
02

## Objectives
DESIGN GOALS
VOICE SENTINEL: Project Objectives
| Number | ID | Title | Description | CIA Tags | Color |
| --- | --- | --- | --- | --- | --- |
| 01 | OBJ 1 | Secure Ingestion & Feature Extraction | Establish an efficient and secure audio ingestion and feature extraction pipeline for diverse audio inputs. | Confidentiality · Availability | #4FC3F7 |
| 02 | OBJ 2 | Ensemble Detection Engine | Deploy an advanced ensemble machine learning model for accurate and generalisable deepfake audio detection. | Integrity | #00BCD4 |
| 03 | OBJ 3 | Reporting System | Design and implement an intuitive user interface and comprehensive reporting system for diverse stakeholders. | Availability · Integrity | #22C55E |
VOICE SENTINEL · ASHESI UNIVERSITY · 2026
03

## Audio Features
FEATURE ENGINEERING
Detection Approach: What We Listen For
Audio deepfakes betray themselves through spectro-temporal inconsistencies invisible to the human ear — but detectable by ML.
| Feature Name | What it captures | Why it discriminates deepfakes |
| --- | --- | --- |
| MFCCs | Timbre & spectral envelope | AI speech shows unnatural smoothness — human voices have micro-variations from breath, emotion, and vocal-tract physiology. |
| Mel Spectrogram | Time-frequency fingerprint | Deepfakes produce overly regular frequency patterns; real voices show organic harmonic variation. |
| Log Energy | Frame-level energy contour | Human breathing creates natural energy dips between words; TTS engines generate unnaturally uniform energy levels. |
| Spectral Centroid | Brightness & harmonic content | Synthetic speech compresses harmonic brightness into an unrealistically stable range vs. natural voice dynamics. |
Feature
What It Captures
Why It Discriminates Deepfakes
VOICE SENTINEL · ASHESI UNIVERSITY · 2026
04

## Feature Extraction
TOOLING
Feature Extraction: Librosa & Wav2Vec2
Two complementary tools cover the full acoustic feature space — traditional signal processing meets deep learned representations.
Librosa
Traditional DSP
Python library for audio & music analysis.
Extracts:
| Item |
| --- |
| MFCCs — timbre, vocal-tract shape |
| Mel Spectrograms — time-frequency map |
| Log Energy — breathing patterns |
| Spectral Centroid — harmonic brightness |
**Why Librosa:** Handcrafted acoustic features correspond directly to the physical properties where deepfake artifacts manifest. Deterministic and highly interpretable — backbone of the **Random Forest** and **CNN** models.
Wav2Vec2
Self-Supervised DL
Facebook AI transformer trained on 960 h of real speech (LibriSpeech).
Extracts:
| Item |
| --- |
| 1024-dim contextual embeddings |
| Prosody, naturalness, phonetic context |
| Learned representation of authentic speech |
**Why Wav2Vec2:** Has internalised the statistical distribution of genuine human voice. Synthetic audio deviates measurably in its latent space — a powerful learned signal for the **TCN** and **CNN-LSTM** models.
VOICE SENTINEL · ASHESI UNIVERSITY · 2026
05

## 5-Model Ensemble
COMPONENT B: DETECTION ENGINE
The 5-Model Ensemble Architecture
Each model targets a different spectro-temporal dimension of deepfake artifacts. The GBT meta-classifier learns the combined prediction space.
Model
Input
What It Detects
| Abbreviation | Model | Input | What it detects |
| --- | --- | --- | --- |
| RF | Random Forest | Librosa acoustics | Non-linear feature interactions; robust baseline on handcrafted features |
| CNN | Conv. Neural Network | Mel Spectrograms | Spatial frequency patterns & texture regularities in spectrograms |
| TCN | Temporal Conv. Network | Wav2Vec2 + acoustics | Long-range temporal dependencies via dilated causal convolutions |
| TSSD | Time-Domain Synth. Speech | Raw waveform | Artifacts in the raw signal that frequency-domain methods miss |
| CNN-LSTM | CNN-LSTM Hybrid | MFCC sequences | Local spectral patterns (CNN) combined with voice evolution memory (LSTM) |
Gradient Boosting Meta-Classifier:
Receives all 5 model confidence scores, learns their combined prediction space, and produces a final verdict with a calibrated confidence score.
VOICE SENTINEL · ASHESI UNIVERSITY · 2026
07

## Architecture
WEB APPLICATION
VOICE SENTINEL: System Architecture & Workflow
| Step | Title | Technology |
| --- | --- | --- |
| 01 | User Uploads Audio | HTML, CSS & Vanilla JS |
| 02 | Ingestion & Storage | FastAPI · Nginx · SQL |
| 03 | Feature Extraction | Librosa + Wav2Vec2 |
| 04 | 5-Model Inference | CNN · TCN · RF · TSSD · CNN-LSTM |
| 05 | GBT Meta-Learner | Gradient Boosting |
| 06 | Gemini Report | Gemini 2.5 Flash |
Stakeholder-tailored reporting
| Point |
| --- |
| **General users** — plain-language verdict with Gemini explanation: <em>"why was this audio flagged?"</em> |
| **Security professionals** — full forensic report: 5 confidence scores, identified artifact locations, synthesis attack vectors, and temporal anomaly breakdown. |
Infrastructure
| Point |
| --- |
| **Frontend:** HTML, CSS &amp; Vanilla JS deployed on Vercel |
| **Backend:** FastAPI on DigitalOcean Droplet |
| **Microservices:** Each model runs in an isolated containerised environment — fault-tolerant and scalable |
| **Database:** SQL — users, analyses, forensic reports |
| **Live:** voice-sentinel.vercel.app |
VOICE SENTINEL · ASHESI UNIVERSITY · 2026
08

## Conclusion
SUMMARY
Conclusion
Scan to interact with the system
![qrImage](https://d6yvfl55smr7u.cloudfront.net/assets/b8imng3e-1774963132770-screenshot-2026-03-31-at-1-18-26-pm.png)
VOICE SENTINEL · ASHESI UNIVERSITY · 2026
10

## References
BIBLIOGRAPHY
References
| Authors | Title |
| --- | --- |
| Müller, N. M., Kawa, P., Choong, W., & Todisco, M. | Replay Attacks Against Audio Deepfake Detection. |
| Dhabe, P., Choudhary, N., Vidhale, A., Singh, V., Gupta, S., & Agrawal, A. | Advanced Sequential Modeling for DeepFake Audio Identification. |
| Kimberly, T., Bray, S. D., Davies, T., & Ma, L. | Warning: Humans cannot reliably detect speech deepfakes. |
| Elmisery, A. M., Sertovic, M., Zayin, A., & Ahmad, M. | Cyber Threats in Financial Transactions — Addressing the Dual Challenge of AI and Quantum Computing. |
| Muhly, F., Chizzonic, E., & Leo, P. | AI-deepfake scams and the importance of a holistic communication security strategy. |
VOICE SENTINEL · ASHESI UNIVERSITY · 2026
11