Speech Emotion Recognition
An audio emotion recognition system: ZCR, RMS, and MFCC features extracted with Librosa, classified by a TensorFlow/Keras model into six emotions, and served through a Streamlit interface.
- Role
- Machine Learning Developer
- Year
- 2024
- Fields
- AI / ML / Data / Optimization
- Stack
- Python
- TensorFlow
- Keras
- Librosa
- NumPy
- Pandas
- Streamlit
What shipped
Signal-level features
ZCR, RMS, and MFCC extraction with Librosa captures how something is said, not what.
Six-class classifier
A TensorFlow/Keras model separates neutral, happy, sad, angry, fear, and disgust.
Streamlit interface
The full pipeline is usable through a simple app, from audio in to emotion out.
Reproducible data handling
NumPy and Pandas keep feature and dataset handling consistent end to end.
What changed
- Classifies speech into six emotion categories.
- Feature extraction, training, and inference share one Python stack.
- The Streamlit interface makes the model usable without any code.
Process details
Reading feeling from a waveform
Emotion hides in signal shape, not words. The model had to learn from raw audio features — zero-crossing rate, RMS energy, MFCCs — and separate six classes that overlap even for human listeners.
The pipeline had to stay honest end to end: consistent feature extraction, a model that generalizes beyond its training recordings, and an interface that lets anyone test it without touching Python.
Features first, then a classifier, then a face for it
Features
ZCR, RMS, and MFCC extraction
Librosa extracts the signal features that carry prosody — energy, rhythm, spectral shape — into a representation a model can learn from.
Model
A TensorFlow/Keras classifier for six emotions
A Keras network maps extracted features to six classes: neutral, happy, sad, angry, fear, and disgust.
Interface
Streamlit as the product surface
A Streamlit app wraps the pipeline so recordings can be classified interactively — no code required.