2025 · Speech
ASR lab
RNN, GRU, BiLSTM, and attention ASR on AfriSpeech-200 (Twi) and LibriSpeech, reported in WER and CER.
Automatic speech recognition pipelines in PyTorch and Torchaudio. Models include vanilla RNNs, attention-based RNNs, GRUs, and bidirectional LSTMs. Audio preprocessing uses MFCC and spectrogram features. One track trains on AfriSpeech-200 (Twi) with CTC loss; another benchmarks on LibriSpeech.
- CTC training on AfriSpeech-200 Twi
- MFCC and spectrogram feature extraction
- WER and CER across recurrent families