Computer Engineering Capstone | Completed May 2025 | Security updates apply
Welcome. As humans and Artificial Intelligence continue to collaborate, the demand for AI to understand audio cues and perceive emotions will increase. This project demonstrates that capability through deep learning-based speech emotion recognition.
Watch a live demonstration of the SER system processing audio files and predicting emotions in real-time:
demo.mp4
The demo achieves 67% accuracy (4 out of 6 correct) on sample audio files spanning six emotions across RAVDESS and CREMA-D datasets.
To test and run this application yourself, clone the repository and run the Streamlit app. You will have access to the trained model that I used for my final capstone presentation.
- Clone this repository
git clone https://github.com/amvdevlab/Speech-Emotion-Recognition-project.git
cd SER-Capstone-Project
- Install dependencies
pip install -r requirements.txt
- Launch the Streamlit web interface
streamlit run src/streamlit_app.py
The application will open in your browser. Upload a .wav audio file to get emotion predictions.
If you encounter issues with librosa installation on macOS, you may need to install ffmpeg first:
brew install ffmpeg
See requirements.txt for complete list of dependencies.
The application uses sensible default paths. To customize paths, set these environment variables before running:
export SER_DATA_PATH="/custom/path/to/processed/data"
export SER_INPUT_PATH="/custom/path/to/raw/audio"
Traditional AI systems process words but miss the emotional context that humans naturally perceive in speech. A frustrated customer and a satisfied customer might use the same words, but their tone reveals their true emotional state. As AI becomes more integrated into human-facing applications, this emotional blindness creates a significant gap in human-computer interaction.
This Speech Emotion Recognition system addresses that gap by analyzing audio features that humans subconsciously use to interpret emotion: pitch variations, speaking rate, voice intensity, and spectral patterns. By training on diverse emotional speech datasets, the model learns to classify eight distinct emotional states from raw audio input.
The implications span multiple industries: Customer service centers can route frustrated callers to specialized agents. Mental health platforms can monitor patient emotional states during telehealth sessions. Educational technology can detect student confusion or engagement during online learning. Automotive systems can adjust responses based on driver stress levels. As voice interfaces become ubiquitous, emotion recognition transforms AI from merely functional to emotionally aware.
- Python 3.13
- PyTorch CNN-LSTM model
- Streamlit App
- Training Datasets:
- RAVDESS
- CREMA-D
- EmoDB
- Apple Silicon (M3 Pro) GPU acceleration
- 250 training epochs with Adam optimizer
- Multi-dataset training: Unified 3 emotion datasets (9,417 audio samples) with different labeling schemes and file structures into a unified training pipeline
- Efficient architecture: CNN-LSTM model with ~293K parameters, trainable on laptop hardware in approximately 1 hour
- 250-epoch training: Achieved 92.7% loss reduction (1.66 → 0.1247) with steady convergence
- Production deployment: Streamlit web application enabling real-time emotion prediction from user-uploaded audio files
- Cross-platform compatibility: Apple Silicon (M3 Pro) MPS support for accelerated training on modern Mac hardware
- Diverse speaker demographics: Trained on American English actors, German speakers, varied age ranges (20-74 years)
- Modular codebase: Separated concerns (models, datasets, loaders, inference) enabling easy extension to new datasets or model architectures
The model demonstrated strong convergence over 250 epochs, achieving a 92.7% loss reduction (1.66 → 0.1247). The training curve exhibits three distinct phases characteristic of successful deep learning:
- Rapid Learning (Epochs 1-60): Exponential loss decrease as the model learns fundamental emotion-discriminative features
- Refinement Phase (Epochs 60-150): Gradual improvement with diminishing returns as the model fine-tunes decision boundaries
- Convergence (Epochs 150-250): Stable plateau around loss ~0.13 with healthy noise, indicating the model has reached optimal capacity
The smooth decay without upticks at later epochs suggests the model is generalizing rather than memorizing training samples.
Overall Performance:
- Estimated Test Accuracy: ~65-70%
- Dataset: RAVDESS evaluation subset (1,440 samples)
- 8 emotion classes evaluated (model trained primarily on 7 core emotions)
| Emotion | Samples | Correct | Accuracy | Key Observations |
|---|---|---|---|---|
| Angry | 192 | 190 | 98.96% | ✓ Strongest performer |
| Calm | 192 | 131 | 68.23% | Often confused with Angry (56 samples) |
| Disgust | 192 | 109 | 56.77% | Frequently misclassified as Angry (69 samples) |
| Fearful | 192 | 131 | 68.23% | Solid performance with minimal confusion |
| Happy | 192 | 104 | 54.17% | Major confusion with Angry (78 samples) |
| Neutral | 96 | 61 | 63.54% | Moderate performance despite fewer samples |
| Sad | 192 | 132 | 68.75% | Good discrimination from other emotions |
| Surprised | 192 | 0 | 0% |
Complete failure - see analysis below |
Strengths:
- Angry emotion: Near-perfect classification (98.96%), indicating strong feature learning for high-arousal negative emotions
- Fearful & Sad: Solid performance (~68%), suggesting the model distinguishes negative emotion valences effectively
- Diagonal dominance: Most emotions show strong diagonal values, indicating correct classifications outnumber errors
Critical Issues:
-
Surprised Class Failure (0% accuracy)
- All 192 "surprised" samples misclassified
- Primary confusions: Angry (95), Fearful (29), Sad (23), Happy (22)
- Likely cause: "Surprised" may not exist in training data (CREMA-D and EmoDB don't include this emotion), or label mapping error between datasets
-
Angry Bias
- Model over-predicts "Angry" across multiple emotions:
- 78 Happy samples → Angry
- 69 Disgust samples → Angry
- 58 Fearful samples → Angry
- 57 Sad samples → Angry
- 56 Calm samples → Angry
- Possible causes: Class imbalance in training data, or anger's distinct high-energy acoustic features dominate learned representations
- Model over-predicts "Angry" across multiple emotions:
-
Calm vs. Angry Confusion
- 56 calm samples misclassified as angry (29% error rate)
- May indicate difficulty distinguishing low-arousal from high-arousal speech when prosodic cues are subtle
Emotion Clusters:
- High performers: Angry (98.96%)
- Mid-tier performers: Fearful, Sad, Calm (68%)
- Struggling emotions: Neutral (63%), Happy, Disgust (54-57%)
- Failed class: Surprised (0%)
What Worked:
- CNN-LSTM architecture successfully learned discriminative features for high-arousal emotions (angry, fearful, sad)
- Strong performance on negative emotions suggests the model captures arousal and valence dimensions
- Training convergence was smooth, validating the model capacity and regularization strategy
What Needs Improvement:
- Dataset alignment: The "surprised" class failure indicates a mismatch between training labels and evaluation expectations. RAVDESS includes 8 emotions, but CREMA-D (6 emotions) and EmoDB (7 emotions) may not include "surprised" consistently
- Class balancing: The angry bias suggests either imbalanced training distribution or that anger's acoustic features (shouting, high pitch variation) are overrepresented
- Validation methodology: Training on 100% of data prevents true generalization assessment. The confusion matrix likely reflects performance on training data, explaining the surprisingly low accuracy despite low loss
Recommendations for Production:
- Remove "surprised" from prediction classes or retrain with balanced surprised samples
- Implement weighted loss function to address class imbalance
- Add train/validation/test split to measure true generalization
- Consider post-processing to reduce false "angry" predictions (confidence thresholding)
├── LICENSE
├── README.md
├── dataset
│ ├── __init__.py
│ └── ser_dataset.py
├── loaders
│ ├── __init__.py
│ └── dataloader_factory.py
├── models
│ ├── __init__.py
│ └── cnn_lstm.py
├── requirements.txt
├── saved_models
│ └── ser_model_final.pth
├── src
│ ├── batch_plot_loss.py
│ ├── evaluate_model.py
│ ├── label_mapping.txt
│ ├── plot_loss_curve.py
│ ├── predict.py
│ ├── preprocess_audio.py
│ ├── streamlit_app.py
│ ├── train_model.py
│ └── training_log.txt
└── test and demo audio
├── 1001 DFA DIS (disgust).wav
├── 1001 DFA HAP XX (happy).wav
├── 1001 DFA NEU XX (neutral).wav
├── CREMA-D 1001 DFA ANG (angry).wav
├── CREMA-D 1001 DFA FEA XX (fear).wav
└── CREMA-D 1001 DFA SAD (sad).wav
The SER system operates in two distinct phases: offline preprocessing and real-time inference. During preprocessing, raw WAV audio files are converted to mel-spectrogram representations using librosa. Each audio sample is resampled to 16kHz, normalized to unit amplitude, and transformed into a 128-band mel-spectrogram. The spectrograms are saved as NumPy arrays for fast loading during training.
Training combines all three datasets through a custom DataLoader that handles each dataset's unique filename conventions and emotion label schemes. The CNN-LSTM model processes fixed-size 128×128 spectrograms, with convolutional layers extracting spatial frequency patterns and LSTM layers capturing temporal dynamics. Training uses CrossEntropyLoss with Adam optimization over 250 epochs.
For inference, the Streamlit web application accepts user-uploaded WAV files, applies identical preprocessing (with padding or truncation to 128 frames), and feeds the spectrogram through the trained model. The output layer produces probability scores across emotion classes, with the highest-scoring emotion displayed alongside corresponding emoji feedback.
The SER system operates through two distinct pipelines: Training Pipeline (offline) and Inference Pipeline (real-time).
External Entities:
┌──────────────────────┐ ┌────────────────────────────┐
│ Dataset Sources │ │ ML Engineer / System │
│ (RAVDESS, CREMA-D, │ │ (controls training) │
│ Emodb, etc.) │ └────────────────────────────┘
└──────────┬───────────┘ ▲
│ raw audio files (.wav) │ training control signals
▼ │ (hyperparams, start/stop)
┌─────────────────────────────┐ │
│ Audio Pre-processor │◄────────────────────┘
│ • Resample to 16 kHz │
│ • Normalize amplitude │
│ • Extract Mel-Spectrogram │
│ • Resize/Pad → 128×128 │
└──────────────┬──────────────┘
│ processed spectrograms + labels
▼
┌──────────────────────────────────────┐
│ Data Preparation & Loading │
│ • Label mapping & unification │
│ • Dataset splitting (train/val) │
│ • Batch creation (mini-batches) │
└──────────────────────┬───────────────┘
│ training batches (spectrograms + labels)
▼
┌──────────────────────────────────────┐
│ Model Training Engine │
│ (CNN → LSTM → Dense Classifier) │
│ • Forward pass → predictions │
│ • Compute Cross-Entropy Loss │
│ • Backpropagation │
│ • Update weights (Adam optimizer) │
└──────────────────────┬───────────────┘
│ updated model parameters
▼
┌─────────────────────────────┐
│ Trained Model Storage │ ───► Persistent Model Artifact
│ (final trained weights) │ (ready for inference/deployment)
└─────────────────────────────┘
Feedback / Monitoring Flows (optional at this level):
Validation loss & metrics ───► ML Engineer / System
Key Components:
preprocess_audio.py: Converts raw WAV files to mel-spectrograms, saves as.npyfor fast loadingdataloader_factory.py: Handles multi-dataset loading with unified labels across RAVDESS, CREMA-D, EmoDBser_dataset.py: Dataset-specific parsing logic for different filename conventions and emotion labelstrain_model.py: Main training loop with loss tracking and model checkpointingmodels/cnn_lstm.py: CNN-LSTM hybrid architecture (293K parameters)
External Entity ┌───────────────────────────────┐
User ──── uploads ───► │ Web Interface │
│ (Streamlit application) │
└───────────────┬───────────────┘
│ audio file (.wav)
▼
┌───────────────────────────────┐
│ File Receiver │
└───────────────┬───────────────┘
│ raw audio data
▼
┌───────────────────────────────┐ ┌──────────────┐
│ Audio Pre-processor │──►│ Invalid File │
│ • Resample to 16 kHz │no │ Notification│
│ • Extract Mel-Spectrogram │ └──────────────┘
│ • Normalize & Resize/Pad │
│ → 1×128×128 tensor │
└───────────────┬───────────────┘
│ mel-spectrogram tensor
▼
┌───────────────────────────────┐
│ Emotion Classifier │
│ (CNN → LSTM → Dense) │
└───────────────┬───────────────┘
│ emotion class probabilities (7 classes)
▼
┌───────────────────────────────┐
│ Result Formatter │
│ • argmax → predicted label │
│ • Softmax probabilities │
└───────────────┬───────────────┘
│ prediction + confidence
▼
┌───────────────────────────────┐
│ Web Interface Display │
│ (emoji + label + probs) │
└───────────────┬───────────────┘
▲
│ displays result
┌───────────────┴───────────────┐
│ User │
└───────────────────────────────┘
Data Stores (if needed at this level – optional):
┌──────────────────────┐
│ Temporary Audio File │ ← used only during processing
└──────────────────────┘
Key Components:
streamlit_app.py: Web interface for file uploads with security validationpredict.py: Inference logic with identical preprocessing pipeline as trainingsaved_models/ser_model_final.pth: Pre-trained weights (~1.2MB file)
| Stage | Input | Output | Dimensions |
|---|---|---|---|
| Raw Audio | WAV file | Time-series waveform | Variable length, 16kHz sample rate |
| Preprocessing | Waveform | Mel-spectrogram (dB scale) | 128 mel bins × variable width |
| Model Input | Mel-spectrogram | Padded/cropped tensor | 128 × 128 × 1 (batch × height × width × channels) |
| CNN Output | Tensor | Feature maps | 32 × 32 × 32 (after 2 pooling layers) |
| LSTM Output | Flattened features | Temporal encoding | 64-dimensional hidden state |
| Classifier Output | Hidden state | Logits | 7 emotion classes |
| Final Prediction | Softmax probabilities | Emotion label | Single class (0-6) |
Training Phase:
- Pre-computed spectrograms eliminate redundant on-the-fly processing (60% speedup)
- MPS backend leverages Apple Silicon GPU acceleration (2.5× faster than CPU)
- Batch size 16 optimized for 18GB unified memory on M3 Pro
Inference Phase:
- Model loaded once at startup, reused for all predictions
- Preprocessing identical to training ensures consistent feature distribution
- Sub-100ms inference latency per audio file
The training process utilizes three emotion speech datasets totaling 9,417 audio samples spanning multiple speakers, ages, and recording conditions. Each dataset contributes different strengths:
-
RAVDESS (Ryerson Audio-Visual Database of Emotional Speech and Song) provides 1,440 files from 24 professional actors speaking emotionally-neutral statements with eight different emotions. The high-quality studio recordings offer clean training signals.
-
CREMA-D (Crowd-sourced Emotional Multimodal Actors Dataset) contributes 7,442 clips from 91 actors with six emotion categories, providing demographic diversity across age (20-74 years), ethnicity, and gender.
-
EmoDB (Berlin Database of Emotional Speech) adds 535 German-language utterances from 10 actors, testing the model's ability to generalize across languages despite training on emotion-specific acoustic features rather than linguistic content.
All audio undergoes standardized preprocessing before training. Raw WAV files are loaded with librosa and resampled to 16kHz (balancing speech intelligibility with computational efficiency). Amplitude normalization scales each sample to [-1, 1] range, removing loudness variations across different recording equipment.
Mel-spectrogram extraction applies Short-Time Fourier Transform with 128 mel-frequency bands, mimicking human auditory perception by allocating more resolution to lower frequencies where speech formants reside. The resulting power spectrograms are converted to decibel scale (logarithmic) to match human loudness perception. Final dimensions are 128 mel bins by variable width (depending on audio duration), saved as .npy files in the Processed_Data directory.
- Architecture: Convolutional Neural Network (CNN) - Long Short-Term Memory (LSTM) hybrid model
- Classes: 8 emotion classes (angry, bored, disgust, fearful, happy, neutral, sad)
- Input: Mel-spectrograms (128 mels × variable width, resized to 128×128)
- Sample Rate: 16kHz
Training Configuration:
- Epochs: 250
- Batch Size: 16 samples
- Optimizer: Adam (learning rate = 0.001)
- Loss Function: CrossEntropyLoss
- Hardware: Apple M3 Pro MacBook (MPS backend)
- Training Duration: Approximately 1 hour
- Model Parameters: ~293,159 trainable parameters
- Dataset Size: 9,417 audio samples
Training Results:
- Initial Loss (Epoch 1): 1.6623
- Final Loss (Epoch 250): 0.1247
- Loss Reduction: 92.7%
- Convergence Pattern: Steep improvement through epoch 60, gradual refinement through epoch 150, plateau with minor fluctuations through epoch 250
The training exhibited healthy convergence without signs of catastrophic overfitting, though lack of a validation set prevents measuring true generalization performance. The loss curve and detailed results are presented in the Results section above.
This section documents the optimization strategies employed to train a production-ready model on laptop hardware within capstone project constraints.
Apple Silicon GPU Support:
- Leveraged Metal Performance Shaders (MPS) backend for M3 Pro GPU acceleration
- Training time: ~1 hour for 250 epochs (vs. ~2.5 hours on CPU alone)
- Memory usage: ~500MB during training with batch size 16
- Unified memory architecture (18GB) eliminates CPU-GPU data transfer overhead
Preprocessing Optimization:
- Offline spectrogram computation: All 9,417 audio files pre-processed and saved as
.npyfiles - Eliminated redundant on-the-fly audio loading during training
- Result: 60% reduction in training time, consistent batch loading speeds
- Trade-off: 2-3GB additional disk storage for preprocessed data
Optimizer Configuration:
- Algorithm: Adam optimizer (adaptive learning rates per parameter)
- Learning rate: 0.001 (PyTorch default, suitable for small-to-medium datasets)
- No learning rate scheduling: Loss converged smoothly without requiring decay
- Gradient clipping: Not needed (gradients remained stable throughout training)
Training Parameters:
- Epochs: 250 (selected based on loss convergence plateau around epoch 150-200)
- Batch size: 16 (balanced between memory constraints and gradient stability)
- Loss function: CrossEntropyLoss (standard for multi-class classification)
- No weight decay: Model capacity (293K params) appropriate for dataset size
Architecture Choices:
- Model size: ~293K parameters (lightweight by modern standards)
- Rationale: Trainable on laptop hardware, fast inference (<100ms per file)
- Trade-off: Smaller capacity vs. training accessibility and deployment simplicity
Sampling Strategy:
- 16kHz sample rate: Adequate for speech (vs. 44.1kHz for music)
- Reduces file size by ~60% compared to standard audio sampling
- Preserves critical speech features (formants, pitch) while improving computational efficiency
Spectrogram Configuration:
- 128 mel bands: Balance between frequency resolution and computation
- Log-scale (dB): Matches human auditory perception, improves gradient flow
- Fixed 128×128 input: Enables batch processing and GPU optimization
Normalization Strategy:
- Amplitude normalization: Scales each audio to [-1, 1] range
- Removes loudness variations from different recording equipment
- Per-sample normalization: Prevents dataset-specific biases
Padding/Cropping Logic:
- Short clips (<128 frames): Zero-padding on right side (preserves temporal structure)
- Long clips (>128 frames): Center-cropping (preserves emotional content vs. edge-cropping)
| Optimization | Benefit | Trade-off | Impact |
|---|---|---|---|
| Pre-computed spectrograms | 60% faster training | 2-3GB disk space | ✅ High ROI |
| 16kHz sampling | Smaller files, faster processing | Less high-frequency detail | ✅ Acceptable for speech |
| 128×128 fixed size | Batch processing, GPU efficiency | Information loss for long audio | |
| Lightweight model (293K params) | Fast training, low latency | Lower capacity than transformers | |
| MPS GPU acceleration | 2.5× training speedup | Mac-specific, limits portability | ✅ Justified for development |
| No data augmentation | Simpler pipeline, faster iteration | Reduced robustness to noise | |
| 100% training data (no validation) | Maximizes learning data | Cannot measure true generalization | ❌ Critical oversight |
Validation Strategy:
- Issue: No train/validation split implemented
- Impact: Cannot measure true generalization or detect overfitting reliably
- Reason: Prioritized maximizing training data over validation methodology
- Future fix: Implement 80/10/10 train/validation/test split with stratified sampling
Data Augmentation:
- Issue: No augmentation techniques applied (time stretching, pitch shifting, noise injection)
- Impact: Model may struggle with real-world audio variations beyond training distribution
- Reason: Computational constraints, scope limitations of capstone project
- Future fix: Add background noise, speed perturbation, SpecAugment masking
Learning Rate Scheduling:
- Issue: Fixed learning rate throughout training
- Impact: Potentially missed finer convergence in later epochs
- Reason: Loss decreased smoothly without scheduling necessity
- Future fix: Implement ReduceLROnPlateau or cosine annealing for optimal convergence
Cross-Dataset Validation:
- Issue: Evaluation performed only on RAVDESS samples
- Impact: Unknown performance on CREMA-D and EmoDB held-out data
- Reason: Time constraints during capstone development
- Future fix: Evaluate separately on each dataset, measure cross-dataset transfer learning
Training Efficiency:
- Total training time: ~1 hour (9,417 samples, 250 epochs)
- Throughput: ~39 samples/second during training
- Loss reduction: 92.7% (1.6623 → 0.1247)
Model Deployment:
- Inference latency: <100ms per audio file on M3 Pro
- Model checkpoint size: 1.2MB (highly portable)
- Memory footprint: ~50MB during inference
Hardware Utilization:
- GPU acceleration: 2.5× speedup vs. CPU
- Memory efficiency: 500MB training, 50MB inference
- Disk usage: 2-3GB preprocessed spectrograms + 1.2MB model
These optimizations enabled training a competitive SER model on consumer laptop hardware, demonstrating that effective ML engineering can overcome resource constraints through thoughtful architectural choices and preprocessing strategies.
Problem: Each dataset uses different emotion labeling conventions. RAVDESS encodes emotions as two-digit numbers in filenames (01 = neutral, 03 = happy), CREMA-D uses three-letter codes (ANG, HAP, NEU), and EmoDB uses single German letters (W = angry, F = happy). Without standardization, the model couldn't learn consistent emotion representations.
Solution: Implemented dataset-specific parsing logic in ser_dataset.py with emotion mapping dictionaries that translate each dataset's labels to a unified string format (angry, happy, neutral, etc.). A secondary label_to_index mapping converts these strings to integers (0-7) for PyTorch training.
Problem: Audio samples ranged from 1 to 5 seconds, producing mel-spectrograms with inconsistent widths (128 mels × 50-500 time frames). Neural networks require fixed-size inputs, but cropping loses information while padding wastes computation.
Solution: During inference, implemented dynamic padding (zero-padding right side) for short clips and center-cropping for long clips to achieve uniform 128×128 dimensions. Training phase uses native variable-length handling through PyTorch's dynamic batching, though this introduces some inefficiency.
Problem: Modern SER systems use transformer-based models (Wav2Vec2, HuBERT) pre-trained on thousands of hours of audio, requiring high-end GPU resources and extensive computational budgets beyond capstone project scope.
Solution: Selected CNN-LSTM hybrid architecture as a proven baseline (established in SER literature 2017-2020) that balances accuracy with trainability. CNNs extract local spectral features (formants, pitch) while LSTMs model temporal prosody (rhythm, intonation). With ~293K parameters, the model trains in approximately 1 hour on laptop hardware while achieving strong convergence.
Problem: Training on personal laptop (Apple M3 Pro) with shared memory imposed constraints on batch size, model capacity, and training duration compared to cloud GPU setups.
Solution: Leveraged Apple's MPS (Metal Performance Shaders) backend for GPU acceleration, used modest batch size (16) to fit memory, and preprocessed all audio offline to .npy files, eliminating redundant on-the-fly spectrogram computation. Pre-saving spectrograms reduced training time by approximately 60% compared to real-time audio loading.
Given additional time and resources, several improvements would strengthen the system:
Adding an attention mechanism after the LSTM layer would allow the model to focus on emotionally salient moments (e.g., pitch spikes indicating anger) rather than treating all time steps equally. Ensemble methods combining CNN-LSTM with traditional machine learning (SVM on hand-crafted acoustic features like jitter and shimmer) could capture complementary patterns.
The most critical oversight was training on 100% of available data without a held-out validation set. This prevents measuring true generalization and risks overfitting going undetected. Implementing 80/20 train-validation split with early stopping would provide reliable performance metrics. Additionally, data augmentation techniques (adding background noise, time stretching, pitch shifting) would improve robustness to real-world audio conditions.
The most critical oversight was training on 100% of available data without a held-out validation set. This prevents measuring true generalization and risks overfitting going undetected. Implementing 80/20 train-validation split with early stopping would provide reliable performance metrics. Additionally, data augmentation techniques (adding background noise, time stretching, pitch shifting) would improve robustness to real-world audio conditions.
Evaluation Methodology Flaw: The evaluation script (evaluate_model.py) only tested the model on RAVDESS samples, despite training on all three datasets (RAVDESS, CREMA-D, EmoDB). More critically, since no train/test split was implemented, the evaluation likely measured performance on samples the model had already seen during training rather than true generalization ability. This explains the disconnect between the low training loss (0.1247) and modest accuracy (~65-70%). A proper evaluation would:
- Implement stratified train/validation/test splits across all datasets before training
- Evaluate separately on held-out samples from each dataset to measure cross-dataset generalization
- Report per-class performance across all three datasets, not just RAVDESS
- Test cross-dataset transfer learning (e.g., train on RAVDESS+CREMA-D, test on EmoDB) to validate that learned features capture language-agnostic emotional acoustics
Beyond loss curves, comprehensive evaluation should include per-class accuracy, confusion matrices on held-out test data, and cross-dataset testing (train on RAVDESS, test on CREMA-D) to measure transfer learning. Recording prediction confidence scores would identify when the model is uncertain, enabling fallback strategies in production.
The current Streamlit app processes pre-recorded files, but real-world applications need streaming audio support. Implementing sliding-window inference over live microphone input with sub-second latency would enable interactive use cases. Model quantization (converting to INT8) could reduce inference time by 4x for mobile deployment.
All training data consists of acted emotional speech recorded in studio conditions. Testing on in-the-wild audio (podcast clips, phone calls, customer service recordings) would reveal performance gaps. Additionally, current datasets are primarily English/German; expanding to multilingual data would test whether learned features generalize across languages.
The three audio datasets used in this project are available on Kaggle:
January 2026 Update: Post-presentation security review identified and addressed the following improvements:
- Enhanced Streamlit file upload validation (file size limits, format verification, secure temporary file handling)
- Added
weights_only=Trueparameter to PyTorch model loading for protection against pickle exploits - Replaced hardcoded absolute paths with environment variables for cross-platform compatibility
- Added comprehensive input validation and error handling
These improvements enhance security and portability without affecting the model performance or results presented in the capstone demonstration.