Skip to content

Latest commit

 

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Speech Emotion Recognition (SER) AI System

Computer Engineering Capstone | Completed May 2025 | Security updates apply

Welcome. As humans and Artificial Intelligence continue to collaborate, the demand for AI to understand audio cues and perceive emotions will increase. This project demonstrates that capability through deep learning-based speech emotion recognition.

Introduction

Demo

Watch a live demonstration of the SER system processing audio files and predicting emotions in real-time:

demo.mp4

The demo achieves 67% accuracy (4 out of 6 correct) on sample audio files spanning six emotions across RAVDESS and CREMA-D datasets.

Installation

Dependencies Required

Instructions

To test and run this application yourself, clone the repository and run the Streamlit app. You will have access to the trained model that I used for my final capstone presentation.

  1. Clone this repository
git clone https://github.com/amvdevlab/Speech-Emotion-Recognition-project.git
cd SER-Capstone-Project
  1. Install dependencies
pip install -r requirements.txt
  1. Launch the Streamlit web interface
streamlit run src/streamlit_app.py

The application will open in your browser. Upload a .wav audio file to get emotion predictions.

Troubleshooting

If you encounter issues with librosa installation on macOS, you may need to install ffmpeg first:

brew install ffmpeg

See requirements.txt for complete list of dependencies.

Advanced Configuration (Optional)

The application uses sensible default paths. To customize paths, set these environment variables before running:

export SER_DATA_PATH="/custom/path/to/processed/data"
export SER_INPUT_PATH="/custom/path/to/raw/audio"

Why This Matters

Traditional AI systems process words but miss the emotional context that humans naturally perceive in speech. A frustrated customer and a satisfied customer might use the same words, but their tone reveals their true emotional state. As AI becomes more integrated into human-facing applications, this emotional blindness creates a significant gap in human-computer interaction.

This Speech Emotion Recognition system addresses that gap by analyzing audio features that humans subconsciously use to interpret emotion: pitch variations, speaking rate, voice intensity, and spectral patterns. By training on diverse emotional speech datasets, the model learns to classify eight distinct emotional states from raw audio input.

The implications span multiple industries: Customer service centers can route frustrated callers to specialized agents. Mental health platforms can monitor patient emotional states during telehealth sessions. Educational technology can detect student confusion or engagement during online learning. Automotive systems can adjust responses based on driver stress levels. As voice interfaces become ubiquitous, emotion recognition transforms AI from merely functional to emotionally aware.

Tech Stack

  • Python 3.13
  • PyTorch CNN-LSTM model
  • Streamlit App
  • Training Datasets:
    • RAVDESS
    • CREMA-D
    • EmoDB
  • Apple Silicon (M3 Pro) GPU acceleration
  • 250 training epochs with Adam optimizer

Key Achievements

  • Multi-dataset training: Unified 3 emotion datasets (9,417 audio samples) with different labeling schemes and file structures into a unified training pipeline
  • Efficient architecture: CNN-LSTM model with ~293K parameters, trainable on laptop hardware in approximately 1 hour
  • 250-epoch training: Achieved 92.7% loss reduction (1.66 → 0.1247) with steady convergence
  • Production deployment: Streamlit web application enabling real-time emotion prediction from user-uploaded audio files
  • Cross-platform compatibility: Apple Silicon (M3 Pro) MPS support for accelerated training on modern Mac hardware
  • Diverse speaker demographics: Trained on American English actors, German speakers, varied age ranges (20-74 years)
  • Modular codebase: Separated concerns (models, datasets, loaders, inference) enabling easy extension to new datasets or model architectures

Results

Training Performance

The model demonstrated strong convergence over 250 epochs, achieving a 92.7% loss reduction (1.66 → 0.1247). The training curve exhibits three distinct phases characteristic of successful deep learning:

  1. Rapid Learning (Epochs 1-60): Exponential loss decrease as the model learns fundamental emotion-discriminative features
  2. Refinement Phase (Epochs 60-150): Gradual improvement with diminishing returns as the model fine-tunes decision boundaries
  3. Convergence (Epochs 150-250): Stable plateau around loss ~0.13 with healthy noise, indicating the model has reached optimal capacity
loss_curves

The smooth decay without upticks at later epochs suggests the model is generalizing rather than memorizing training samples.

Model Evaluation

Overall Performance:

  • Estimated Test Accuracy: ~65-70%
  • Dataset: RAVDESS evaluation subset (1,440 samples)
  • 8 emotion classes evaluated (model trained primarily on 7 core emotions)

Per-Emotion Performance

Emotion Samples Correct Accuracy Key Observations
Angry 192 190 98.96% ✓ Strongest performer
Calm 192 131 68.23% Often confused with Angry (56 samples)
Disgust 192 109 56.77% Frequently misclassified as Angry (69 samples)
Fearful 192 131 68.23% Solid performance with minimal confusion
Happy 192 104 54.17% Major confusion with Angry (78 samples)
Neutral 96 61 63.54% Moderate performance despite fewer samples
Sad 192 132 68.75% Good discrimination from other emotions
Surprised 192 0 0% ⚠️ Complete failure - see analysis below

Confusion Matrix Analysis

confusion_matrix

Strengths:

  • Angry emotion: Near-perfect classification (98.96%), indicating strong feature learning for high-arousal negative emotions
  • Fearful & Sad: Solid performance (~68%), suggesting the model distinguishes negative emotion valences effectively
  • Diagonal dominance: Most emotions show strong diagonal values, indicating correct classifications outnumber errors

Critical Issues:

  1. Surprised Class Failure (0% accuracy)

    • All 192 "surprised" samples misclassified
    • Primary confusions: Angry (95), Fearful (29), Sad (23), Happy (22)
    • Likely cause: "Surprised" may not exist in training data (CREMA-D and EmoDB don't include this emotion), or label mapping error between datasets
  2. Angry Bias

    • Model over-predicts "Angry" across multiple emotions:
      • 78 Happy samples → Angry
      • 69 Disgust samples → Angry
      • 58 Fearful samples → Angry
      • 57 Sad samples → Angry
      • 56 Calm samples → Angry
    • Possible causes: Class imbalance in training data, or anger's distinct high-energy acoustic features dominate learned representations
  3. Calm vs. Angry Confusion

    • 56 calm samples misclassified as angry (29% error rate)
    • May indicate difficulty distinguishing low-arousal from high-arousal speech when prosodic cues are subtle

Emotion Clusters:

  • High performers: Angry (98.96%)
  • Mid-tier performers: Fearful, Sad, Calm (68%)
  • Struggling emotions: Neutral (63%), Happy, Disgust (54-57%)
  • Failed class: Surprised (0%)

Key Findings

What Worked:

  • CNN-LSTM architecture successfully learned discriminative features for high-arousal emotions (angry, fearful, sad)
  • Strong performance on negative emotions suggests the model captures arousal and valence dimensions
  • Training convergence was smooth, validating the model capacity and regularization strategy

What Needs Improvement:

  • Dataset alignment: The "surprised" class failure indicates a mismatch between training labels and evaluation expectations. RAVDESS includes 8 emotions, but CREMA-D (6 emotions) and EmoDB (7 emotions) may not include "surprised" consistently
  • Class balancing: The angry bias suggests either imbalanced training distribution or that anger's acoustic features (shouting, high pitch variation) are overrepresented
  • Validation methodology: Training on 100% of data prevents true generalization assessment. The confusion matrix likely reflects performance on training data, explaining the surprisingly low accuracy despite low loss

Recommendations for Production:

  • Remove "surprised" from prediction classes or retrain with balanced surprised samples
  • Implement weighted loss function to address class imbalance
  • Add train/validation/test split to measure true generalization
  • Consider post-processing to reduce false "angry" predictions (confidence thresholding)

Project Structure

├── LICENSE
├── README.md
├── dataset
│   ├── __init__.py
│   └── ser_dataset.py
├── loaders
│   ├── __init__.py
│   └── dataloader_factory.py
├── models
│   ├── __init__.py
│   └── cnn_lstm.py
├── requirements.txt
├── saved_models
│   └── ser_model_final.pth
├── src
│   ├── batch_plot_loss.py
│   ├── evaluate_model.py
│   ├── label_mapping.txt
│   ├── plot_loss_curve.py
│   ├── predict.py
│   ├── preprocess_audio.py
│   ├── streamlit_app.py
│   ├── train_model.py
│   └── training_log.txt
└── test and demo audio
    ├── 1001 DFA DIS (disgust).wav
    ├── 1001 DFA HAP XX (happy).wav
    ├── 1001 DFA NEU XX (neutral).wav
    ├── CREMA-D 1001 DFA ANG (angry).wav
    ├── CREMA-D 1001 DFA FEA XX (fear).wav
    └── CREMA-D 1001 DFA SAD (sad).wav

Program Overview

The SER system operates in two distinct phases: offline preprocessing and real-time inference. During preprocessing, raw WAV audio files are converted to mel-spectrogram representations using librosa. Each audio sample is resampled to 16kHz, normalized to unit amplitude, and transformed into a 128-band mel-spectrogram. The spectrograms are saved as NumPy arrays for fast loading during training.

Training combines all three datasets through a custom DataLoader that handles each dataset's unique filename conventions and emotion label schemes. The CNN-LSTM model processes fixed-size 128×128 spectrograms, with convolutional layers extracting spatial frequency patterns and LSTM layers capturing temporal dynamics. Training uses CrossEntropyLoss with Adam optimization over 250 epochs.

For inference, the Streamlit web application accepts user-uploaded WAV files, applies identical preprocessing (with padding or truncation to 128 frames), and feeds the spectrogram through the trained model. The output layer produces probability scores across emotion classes, with the highest-scoring emotion displayed alongside corresponding emoji feedback.

Data Pipeline & Program Flow

The SER system operates through two distinct pipelines: Training Pipeline (offline) and Inference Pipeline (real-time).

Training Pipeline


External Entities:
    ┌──────────────────────┐            ┌────────────────────────────┐
    │   Dataset Sources    │            │     ML Engineer / System   │
    │ (RAVDESS, CREMA-D,   │            │     (controls training)    │
    │  Emodb, etc.)        │            └────────────────────────────┘
    └──────────┬───────────┘                            ▲
               │ raw audio files (.wav)                 │ training control signals
               ▼                                        │ (hyperparams, start/stop)
    ┌─────────────────────────────┐                     │
    │     Audio Pre-processor     │◄────────────────────┘
    │ • Resample to 16 kHz        │
    │ • Normalize amplitude       │
    │ • Extract Mel-Spectrogram   │
    │ • Resize/Pad → 128×128      │
    └──────────────┬──────────────┘
                   │ processed spectrograms + labels
                   ▼
    ┌──────────────────────────────────────┐
    │        Data Preparation & Loading    │
    │ • Label mapping & unification        │
    │ • Dataset splitting (train/val)      │
    │ • Batch creation (mini-batches)      │
    └──────────────────────┬───────────────┘
                           │ training batches (spectrograms + labels)
                           ▼
    ┌──────────────────────────────────────┐
    │        Model Training Engine         │
    │   (CNN → LSTM → Dense Classifier)    │
    │ • Forward pass → predictions         │
    │ • Compute Cross-Entropy Loss         │
    │ • Backpropagation                    │
    │ • Update weights (Adam optimizer)    │
    └──────────────────────┬───────────────┘
                           │ updated model parameters
                           ▼
    ┌─────────────────────────────┐
    │     Trained Model Storage   │ ───►  Persistent Model Artifact
    │   (final trained weights)   │       (ready for inference/deployment)
    └─────────────────────────────┘

Feedback / Monitoring Flows (optional at this level):
    Validation loss & metrics ───►  ML Engineer / System

Key Components:

  • preprocess_audio.py: Converts raw WAV files to mel-spectrograms, saves as .npy for fast loading
  • dataloader_factory.py: Handles multi-dataset loading with unified labels across RAVDESS, CREMA-D, EmoDB
  • ser_dataset.py: Dataset-specific parsing logic for different filename conventions and emotion labels
  • train_model.py: Main training loop with loss tracking and model checkpointing
  • models/cnn_lstm.py: CNN-LSTM hybrid architecture (293K parameters)

Inference Pipeline

External Entity                  ┌───────────────────────────────┐
    User      ──── uploads ───►  │          Web Interface        │
                                 │     (Streamlit application)   │
                                 └───────────────┬───────────────┘
                                                 │ audio file (.wav)
                                                 ▼
                                 ┌───────────────────────────────┐
                                 │        File Receiver          │
                                 └───────────────┬───────────────┘
                                                 │ raw audio data
                                                 ▼
                                 ┌───────────────────────────────┐   ┌──────────────┐
                                 │      Audio Pre-processor      │──►│ Invalid File │
                                 │   • Resample to 16 kHz        │no │  Notification│
                                 │   • Extract Mel-Spectrogram   │   └──────────────┘
                                 │   • Normalize & Resize/Pad    │
                                 │        → 1×128×128 tensor     │
                                 └───────────────┬───────────────┘
                                                 │ mel-spectrogram tensor
                                                 ▼
                                 ┌───────────────────────────────┐
                                 │       Emotion Classifier      │
                                 │     (CNN → LSTM → Dense)      │
                                 └───────────────┬───────────────┘
                                                 │ emotion class probabilities (7 classes)
                                                 ▼
                                 ┌───────────────────────────────┐
                                 │       Result Formatter        │
                                 │   • argmax → predicted label  │
                                 │   • Softmax probabilities     │
                                 └───────────────┬───────────────┘
                                                 │ prediction + confidence
                                                 ▼
                                 ┌───────────────────────────────┐
                                 │      Web Interface Display    │
                                 │   (emoji + label + probs)     │
                                 └───────────────┬───────────────┘
                                                 ▲
                                                 │ displays result
                                 ┌───────────────┴───────────────┐
                                 │             User              │
                                 └───────────────────────────────┘

Data Stores (if needed at this level – optional):
    ┌──────────────────────┐
    │ Temporary Audio File │  ← used only during processing
    └──────────────────────┘

Key Components:

  • streamlit_app.py: Web interface for file uploads with security validation
  • predict.py: Inference logic with identical preprocessing pipeline as training
  • saved_models/ser_model_final.pth: Pre-trained weights (~1.2MB file)

Data Transformations

Stage Input Output Dimensions
Raw Audio WAV file Time-series waveform Variable length, 16kHz sample rate
Preprocessing Waveform Mel-spectrogram (dB scale) 128 mel bins × variable width
Model Input Mel-spectrogram Padded/cropped tensor 128 × 128 × 1 (batch × height × width × channels)
CNN Output Tensor Feature maps 32 × 32 × 32 (after 2 pooling layers)
LSTM Output Flattened features Temporal encoding 64-dimensional hidden state
Classifier Output Hidden state Logits 7 emotion classes
Final Prediction Softmax probabilities Emotion label Single class (0-6)

Pipeline Optimizations

Training Phase:

  • Pre-computed spectrograms eliminate redundant on-the-fly processing (60% speedup)
  • MPS backend leverages Apple Silicon GPU acceleration (2.5× faster than CPU)
  • Batch size 16 optimized for 18GB unified memory on M3 Pro

Inference Phase:

  • Model loaded once at startup, reused for all predictions
  • Preprocessing identical to training ensures consistent feature distribution
  • Sub-100ms inference latency per audio file

Training Pipeline

The training process utilizes three emotion speech datasets totaling 9,417 audio samples spanning multiple speakers, ages, and recording conditions. Each dataset contributes different strengths:

  • RAVDESS (Ryerson Audio-Visual Database of Emotional Speech and Song) provides 1,440 files from 24 professional actors speaking emotionally-neutral statements with eight different emotions. The high-quality studio recordings offer clean training signals.

  • CREMA-D (Crowd-sourced Emotional Multimodal Actors Dataset) contributes 7,442 clips from 91 actors with six emotion categories, providing demographic diversity across age (20-74 years), ethnicity, and gender.

  • EmoDB (Berlin Database of Emotional Speech) adds 535 German-language utterances from 10 actors, testing the model's ability to generalize across languages despite training on emotion-specific acoustic features rather than linguistic content.

Audio Preprocessing

All audio undergoes standardized preprocessing before training. Raw WAV files are loaded with librosa and resampled to 16kHz (balancing speech intelligibility with computational efficiency). Amplitude normalization scales each sample to [-1, 1] range, removing loudness variations across different recording equipment.

Mel-spectrogram extraction applies Short-Time Fourier Transform with 128 mel-frequency bands, mimicking human auditory perception by allocating more resolution to lower frequencies where speech formants reside. The resulting power spectrograms are converted to decibel scale (logarithmic) to match human loudness perception. Final dimensions are 128 mel bins by variable width (depending on audio duration), saved as .npy files in the Processed_Data directory.

Model Architecture

  • Architecture: Convolutional Neural Network (CNN) - Long Short-Term Memory (LSTM) hybrid model
  • Classes: 8 emotion classes (angry, bored, disgust, fearful, happy, neutral, sad)
  • Input: Mel-spectrograms (128 mels × variable width, resized to 128×128)
  • Sample Rate: 16kHz

Training Configuration & Results

Training Configuration:

  • Epochs: 250
  • Batch Size: 16 samples
  • Optimizer: Adam (learning rate = 0.001)
  • Loss Function: CrossEntropyLoss
  • Hardware: Apple M3 Pro MacBook (MPS backend)
  • Training Duration: Approximately 1 hour
  • Model Parameters: ~293,159 trainable parameters
  • Dataset Size: 9,417 audio samples

Training Results:

  • Initial Loss (Epoch 1): 1.6623
  • Final Loss (Epoch 250): 0.1247
  • Loss Reduction: 92.7%
  • Convergence Pattern: Steep improvement through epoch 60, gradual refinement through epoch 150, plateau with minor fluctuations through epoch 250

The training exhibited healthy convergence without signs of catastrophic overfitting, though lack of a validation set prevents measuring true generalization performance. The loss curve and detailed results are presented in the Results section above.

Model Optimization

This section documents the optimization strategies employed to train a production-ready model on laptop hardware within capstone project constraints.

Hardware Acceleration

Apple Silicon GPU Support:

  • Leveraged Metal Performance Shaders (MPS) backend for M3 Pro GPU acceleration
  • Training time: ~1 hour for 250 epochs (vs. ~2.5 hours on CPU alone)
  • Memory usage: ~500MB during training with batch size 16
  • Unified memory architecture (18GB) eliminates CPU-GPU data transfer overhead

Preprocessing Optimization:

  • Offline spectrogram computation: All 9,417 audio files pre-processed and saved as .npy files
  • Eliminated redundant on-the-fly audio loading during training
  • Result: 60% reduction in training time, consistent batch loading speeds
  • Trade-off: 2-3GB additional disk storage for preprocessed data

Hyperparameter Tuning

Optimizer Configuration:

  • Algorithm: Adam optimizer (adaptive learning rates per parameter)
  • Learning rate: 0.001 (PyTorch default, suitable for small-to-medium datasets)
  • No learning rate scheduling: Loss converged smoothly without requiring decay
  • Gradient clipping: Not needed (gradients remained stable throughout training)

Training Parameters:

  • Epochs: 250 (selected based on loss convergence plateau around epoch 150-200)
  • Batch size: 16 (balanced between memory constraints and gradient stability)
  • Loss function: CrossEntropyLoss (standard for multi-class classification)
  • No weight decay: Model capacity (293K params) appropriate for dataset size

Architecture Choices:

  • Model size: ~293K parameters (lightweight by modern standards)
  • Rationale: Trainable on laptop hardware, fast inference (<100ms per file)
  • Trade-off: Smaller capacity vs. training accessibility and deployment simplicity

Audio Processing Optimizations

Sampling Strategy:

  • 16kHz sample rate: Adequate for speech (vs. 44.1kHz for music)
  • Reduces file size by ~60% compared to standard audio sampling
  • Preserves critical speech features (formants, pitch) while improving computational efficiency

Spectrogram Configuration:

  • 128 mel bands: Balance between frequency resolution and computation
  • Log-scale (dB): Matches human auditory perception, improves gradient flow
  • Fixed 128×128 input: Enables batch processing and GPU optimization

Normalization Strategy:

  • Amplitude normalization: Scales each audio to [-1, 1] range
  • Removes loudness variations from different recording equipment
  • Per-sample normalization: Prevents dataset-specific biases

Padding/Cropping Logic:

  • Short clips (<128 frames): Zero-padding on right side (preserves temporal structure)
  • Long clips (>128 frames): Center-cropping (preserves emotional content vs. edge-cropping)

Performance Trade-offs

Optimization Benefit Trade-off Impact
Pre-computed spectrograms 60% faster training 2-3GB disk space ✅ High ROI
16kHz sampling Smaller files, faster processing Less high-frequency detail ✅ Acceptable for speech
128×128 fixed size Batch processing, GPU efficiency Information loss for long audio ⚠️ Potential quality impact
Lightweight model (293K params) Fast training, low latency Lower capacity than transformers ⚠️ Limits max accuracy
MPS GPU acceleration 2.5× training speedup Mac-specific, limits portability ✅ Justified for development
No data augmentation Simpler pipeline, faster iteration Reduced robustness to noise ⚠️ Future improvement needed
100% training data (no validation) Maximizes learning data Cannot measure true generalization ❌ Critical oversight

What Was NOT Optimized

Validation Strategy:

  • Issue: No train/validation split implemented
  • Impact: Cannot measure true generalization or detect overfitting reliably
  • Reason: Prioritized maximizing training data over validation methodology
  • Future fix: Implement 80/10/10 train/validation/test split with stratified sampling

Data Augmentation:

  • Issue: No augmentation techniques applied (time stretching, pitch shifting, noise injection)
  • Impact: Model may struggle with real-world audio variations beyond training distribution
  • Reason: Computational constraints, scope limitations of capstone project
  • Future fix: Add background noise, speed perturbation, SpecAugment masking

Learning Rate Scheduling:

  • Issue: Fixed learning rate throughout training
  • Impact: Potentially missed finer convergence in later epochs
  • Reason: Loss decreased smoothly without scheduling necessity
  • Future fix: Implement ReduceLROnPlateau or cosine annealing for optimal convergence

Cross-Dataset Validation:

  • Issue: Evaluation performed only on RAVDESS samples
  • Impact: Unknown performance on CREMA-D and EmoDB held-out data
  • Reason: Time constraints during capstone development
  • Future fix: Evaluate separately on each dataset, measure cross-dataset transfer learning

Optimization Results Summary

Training Efficiency:

  • Total training time: ~1 hour (9,417 samples, 250 epochs)
  • Throughput: ~39 samples/second during training
  • Loss reduction: 92.7% (1.6623 → 0.1247)

Model Deployment:

  • Inference latency: <100ms per audio file on M3 Pro
  • Model checkpoint size: 1.2MB (highly portable)
  • Memory footprint: ~50MB during inference

Hardware Utilization:

  • GPU acceleration: 2.5× speedup vs. CPU
  • Memory efficiency: 500MB training, 50MB inference
  • Disk usage: 2-3GB preprocessed spectrograms + 1.2MB model

These optimizations enabled training a competitive SER model on consumer laptop hardware, demonstrating that effective ML engineering can overcome resource constraints through thoughtful architectural choices and preprocessing strategies.

Challenges & Solutions

Challenge 1: Cross-Dataset Label Inconsistencies

Problem: Each dataset uses different emotion labeling conventions. RAVDESS encodes emotions as two-digit numbers in filenames (01 = neutral, 03 = happy), CREMA-D uses three-letter codes (ANG, HAP, NEU), and EmoDB uses single German letters (W = angry, F = happy). Without standardization, the model couldn't learn consistent emotion representations.

Solution: Implemented dataset-specific parsing logic in ser_dataset.py with emotion mapping dictionaries that translate each dataset's labels to a unified string format (angry, happy, neutral, etc.). A secondary label_to_index mapping converts these strings to integers (0-7) for PyTorch training.

Challenge 2: Variable-Length Audio Handling

Problem: Audio samples ranged from 1 to 5 seconds, producing mel-spectrograms with inconsistent widths (128 mels × 50-500 time frames). Neural networks require fixed-size inputs, but cropping loses information while padding wastes computation.

Solution: During inference, implemented dynamic padding (zero-padding right side) for short clips and center-cropping for long clips to achieve uniform 128×128 dimensions. Training phase uses native variable-length handling through PyTorch's dynamic batching, though this introduces some inefficiency.

Challenge 3: Model Architecture Selection

Problem: Modern SER systems use transformer-based models (Wav2Vec2, HuBERT) pre-trained on thousands of hours of audio, requiring high-end GPU resources and extensive computational budgets beyond capstone project scope.

Solution: Selected CNN-LSTM hybrid architecture as a proven baseline (established in SER literature 2017-2020) that balances accuracy with trainability. CNNs extract local spectral features (formants, pitch) while LSTMs model temporal prosody (rhythm, intonation). With ~293K parameters, the model trains in approximately 1 hour on laptop hardware while achieving strong convergence.

Challenge 4: Limited Training Resources

Problem: Training on personal laptop (Apple M3 Pro) with shared memory imposed constraints on batch size, model capacity, and training duration compared to cloud GPU setups.

Solution: Leveraged Apple's MPS (Metal Performance Shaders) backend for GPU acceleration, used modest batch size (16) to fit memory, and preprocessed all audio offline to .npy files, eliminating redundant on-the-fly spectrogram computation. Pre-saving spectrograms reduced training time by approximately 60% compared to real-time audio loading.

Reflections on Improvements

Given additional time and resources, several improvements would strengthen the system:

Architecture Enhancements

Adding an attention mechanism after the LSTM layer would allow the model to focus on emotionally salient moments (e.g., pitch spikes indicating anger) rather than treating all time steps equally. Ensemble methods combining CNN-LSTM with traditional machine learning (SVM on hand-crafted acoustic features like jitter and shimmer) could capture complementary patterns.

Training Methodology

The most critical oversight was training on 100% of available data without a held-out validation set. This prevents measuring true generalization and risks overfitting going undetected. Implementing 80/20 train-validation split with early stopping would provide reliable performance metrics. Additionally, data augmentation techniques (adding background noise, time stretching, pitch shifting) would improve robustness to real-world audio conditions.

Model Evaluation

The most critical oversight was training on 100% of available data without a held-out validation set. This prevents measuring true generalization and risks overfitting going undetected. Implementing 80/20 train-validation split with early stopping would provide reliable performance metrics. Additionally, data augmentation techniques (adding background noise, time stretching, pitch shifting) would improve robustness to real-world audio conditions.

Evaluation Methodology Flaw: The evaluation script (evaluate_model.py) only tested the model on RAVDESS samples, despite training on all three datasets (RAVDESS, CREMA-D, EmoDB). More critically, since no train/test split was implemented, the evaluation likely measured performance on samples the model had already seen during training rather than true generalization ability. This explains the disconnect between the low training loss (0.1247) and modest accuracy (~65-70%). A proper evaluation would:

  • Implement stratified train/validation/test splits across all datasets before training
  • Evaluate separately on held-out samples from each dataset to measure cross-dataset generalization
  • Report per-class performance across all three datasets, not just RAVDESS
  • Test cross-dataset transfer learning (e.g., train on RAVDESS+CREMA-D, test on EmoDB) to validate that learned features capture language-agnostic emotional acoustics

Beyond loss curves, comprehensive evaluation should include per-class accuracy, confusion matrices on held-out test data, and cross-dataset testing (train on RAVDESS, test on CREMA-D) to measure transfer learning. Recording prediction confidence scores would identify when the model is uncertain, enabling fallback strategies in production.

Deployment Considerations

The current Streamlit app processes pre-recorded files, but real-world applications need streaming audio support. Implementing sliding-window inference over live microphone input with sub-second latency would enable interactive use cases. Model quantization (converting to INT8) could reduce inference time by 4x for mobile deployment.

Dataset Expansion

All training data consists of acted emotional speech recorded in studio conditions. Testing on in-the-wild audio (podcast clips, phone calls, customer service recordings) would reveal performance gaps. Additionally, current datasets are primarily English/German; expanding to multilingual data would test whether learned features generalize across languages.

Sources

Datasets

The three audio datasets used in this project are available on Kaggle:

PyTorch


Security Notice

January 2026 Update: Post-presentation security review identified and addressed the following improvements:

  • Enhanced Streamlit file upload validation (file size limits, format verification, secure temporary file handling)
  • Added weights_only=True parameter to PyTorch model loading for protection against pickle exploits
  • Replaced hardcoded absolute paths with environment variables for cross-platform compatibility
  • Added comprehensive input validation and error handling

These improvements enhance security and portability without affecting the model performance or results presented in the capstone demonstration.

About

A Speech Emotion Recognition System that I built for my Capstone Project using PyTorch and academic emotion audio databases.

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages