m-fakhry/DSAI-456-Speech

Vue

21

0 commits

updated Sep 27, 2026

See the code

README

Speech Recognition (DSAI 456) - 2026-2027

Repository for the Speech Recognition undergraduate course (DSAI 456) for the 2026-2027 academic year at Zewail City University.

This year the course

  • transitions the focus from traditional statistical ASR systems to modern neural speech architectures.
  • adds a class project.

Previous Offerings


Logistics

CourseSpeech Recognition - DSAI 456
Webpagehttps://github.com/m-fakhry/DSAI-456-Speech
InstructorProf. Mohamed Ghalwash (mghalwash@zewailcity.edu.eg)
Structure2-hour lecture (Mon 8-10) and 2-hour lab (Mon 10-12, Tue 12-2, Tue 4-6)
TAsEng. Ahmed Aamer
CommunicationMoodle or Email or office hours. no phone, no whatsapp
Lab PolicyAssignments, quizzes, and project milestones
Book"Speech and Language Processing", Jurafsky and Martin, 3rd Edition, 2025
SupplementaryHuggingFace Audio Course
ObjectiveProvide students with the theory and practical skills to build, adapt, and evaluate modern speech systems (recognition and synthesis), and to read the current literature critically
PrerequisitesDeep Learning
Tools/APIsWASP, librosa, HuggingFace Transformers + PEFT. Optional: openSmile, torchaudio, NeMo or ESPnet

Course Learning Outcomes

CLO #OutcomeStatement
1Analyze Speech SignalsExplain the mathematical foundations of digital audio, including sampling, quantization, and the transformation of signals from the time domain to the frequency domain
2Extract Acoustic FeaturesImplement and evaluate feature extraction pipelines, specifically log-mel spectrograms and Mel-Frequency Cepstral Coefficients (MFCCs), for speech processing tasks
3Model Temporal SequencesApply dynamic programming over sequences, specifically the Forward and Viterbi algorithms, to align audio with text, using GMM-HMM systems as the classical case
4Develop and Adapt Neural ASR SystemsDesign and implement modern End-to-End speech recognition architectures, including Connectionist Temporal Classification (CTC), transducer, and Encoder-Decoder frameworks, and fine-tune pretrained models for new languages and domains under realistic label budgets
5Build Generative Speech SystemsExplain neural audio codecs and token-based synthesis, and implement speech generation systems
6Evaluate & Appraise ResearchEvaluate speech systems empirically against baselines, including accuracy-latency trade-offs, and critically appraise contemporary research in speech and audio AI, including its limitations and ethical implications

Lectures

Papers in the Paper column are required reading. Individual papers are not graded, but the midterm and the final each include one question asking you to appraise an assigned paper - to take a position, predict a result, or propose the experiment that would settle a disagreement between two of them. Summaries will not earn marks.

Please note that the syllabus content is subject to change throughout the semester. Topics may be added or removed based on the instructor's discretion, student progress, and available time. Your feedback and participation will inform these adjustments to ensure alignment with course goals and schedule constraints.

WeekDateTopicContentsPaperCLOLectureAssignment
109-21IntroGeneral introduction to the course. WER/CER metrics.Browse Open ASR Leaderboard as a benchmark6Lecture 1Assignment 1
209-28FoundationsFormants, quantization, framing, F0 and intensity, spectrogram, sampling.Ch. 16.1-16.21Lecture 2Assignment 2
310-05Spectral Front EndDFT/FFT, windowing and spectral leakage, time-frequency resolution, STFT and spectrogram, mel filterbank, log-mel. MFCCCh.1, 2
410-12Alignment & DecodingThe alignment problem, HMM, forward algorithm, Viterbi, GMM as a density model.Rabiner (1989), A Tutorial on Hidden Markov Models3Project proposal
510-19CTCBlank symbol, collapse function, summing over alignments, forward-backward, prefix beam search, shallow LM fusionGraves et al. (2006), Connectionist Temporal Classification3, 4
610-26Encoders & TransducersCNN, BiLSTM, Transformer. RNN-T (encoder + prediction + joint network), Conformer/FastConformer, subsampling, chunked attention and causalityGraves (2012), Sequence Transduction with RNNs; Gulati et al. (2020), Conformer4
711-02Self-Supervised LearningWhy labels are the bottleneck, wav2vec 2.0 (masking, quantized targets, contrastive loss), HuBERT (offline clustering, masked prediction), WavLM, low-resource fine-tuningBaevski et al. (2020), wav2vec 2.0; Hsu et al. (2021), HuBERT4Data + evaluation protocol + pretrained baseline
811-09Midterm1, 2, 3, 4
911-16Weak Supervision at ScaleAttention encoder-decoder, cross-attention, Whisper's multitask token format, hallucination and repetition, long-form chunking, the scale-vs-curation debateRadford et al. (2023), Whisper; Puvvada et al. (2024), Less is more4, 6
1011-23Neural Audio CodecsVQ-VAE, straight-through estimator, codebook collapse, residual vector quantization, EnCodec, semantic vs acoustic tokensDéfossez et al. (2022), EnCodec5Working system + training curves
1111-30TTS as Language ModellingVALL-E's two-stage AR + NAR recipe, zero-shot voice cloning from a 3-second prompt, instability and duration control, flow-matching alternatives, vocoders, MOS and speaker similarityWang et al. (2023), VALL-E5, 6
1212-07Speech LLMsSpeech encoder + projector + LLM, what is frozen and what is trained, prompt-conditioned ASR, spoken QA and audio understanding, why a large LLM does not automatically lower WERTang et al. (2023), SALMONN; Qwen3-Omni5, 6
1312-14Project DemosTeam presentations and evaluation-4, 5, 6Final report due (all teams)
1412-21Project DemosTeam presentations and evaluation-4, 5, 6
1512-28Project DemosTeam presentations and evaluation-4, 5, 6
16Finalall

Class Project

  • Teams of 3-4 propose a novel speech application and build a working system around it.
  • Arabic data is highly preferred. Modern and dialectal Arabic, code-switched Arabic-English, and Egyptian speech in particular are underserved by current models, and a project that improves on a pretrained baseline there is doing something that has not already been done a hundred times in English. Arabic is not a hard requirement, but a non-Arabic project has to carry its weight through a stronger idea.

  • The idea matters as much as the implementation. A project is novel if it targets a task, a user group, a dialect, or a setting that existing systems handle badly, or if it combines components in a way the literature has not. It is not novel if it fine-tunes a standard model on a standard benchmark and reports the expected number. You do not need a research contribution - you need a reason the thing you built should exist.

  • Every project must have: a defined task and evaluation metric, a held-out test set the team did not tune on, a pretrained baseline, a system that improves on that baseline, honest failure analysis, and a working demo.

  • The final report must open with a positioning section: what already exists for this task, where it falls short for your case, and what you did differently. Cite real systems and papers, not a general description of the field. How you weight research depth against engineering polish is up to the team and should be stated in the proposal.

  • Teams are encouraged, but not required, to find a mentor or a real user, who is someone has the problem and will tell you whether your output is any good. A linguist, a teacher, a clinician, a call-centre supervisor. Their feedback is worth more than another epoch of training.

  • Milestones

    WeekDeliverable
    4Proposal: the idea, why it is needed, dataset plan, evaluation metric
    7Data (with consent documentation where applicable), evaluation protocol, pretrained baseline results
    10Working system with training curves and error analysis
    13Final report, due for all teams regardless of demo slot; demos run weeks 13-15

Grading Policy

TopicPercentageNotes
Lab Assignments20%Graded in lab with your TA
Lab Quizzes10%Weeks 6 and 12
Class Project20%Distributed across the four milestones
Midterm10%Covers weeks 1-7, including questions on the assigned papers
Final40%Includes question on the assigned papers

Course Instructions

Principle: deadlines are firm.

Submissions

  • All assignments must be uploaded to the Moodle system before the deadline, even if the assignment has already been graded in the lab.
  • Assignment grades depend on discussing your work with your TA. There are no extensions on these discussions: if you miss the discussion for an assignment, you lose its grade.
  • Any assignment involving model training must report the hardware used and the wall-clock training time. "It did not converge" is a result; report it with evidence.

Excuses and make-up tasks

  • Medical excuses must be submitted within one week of the excused task (assignment, quiz, or midterm). Excuses submitted after that window will not be considered.
  • Make-up tasks cover the material taught up to the date of the make-up, not the material of the original task.

Grade petitions

  • Once coursework grades (assignment, quiz, etc.) are posted, you have one week to raise an issue. If you believe there is an error in your grade, email me and CC your TA with:

    1. A clear and detailed explanation of exactly why you are petitioning the grade.
    2. Any relevant and approved documentation supporting your request, if the petition concerns a missed assignment, quiz, or exam.
  • No grade adjustments will be considered after this deadline, and petitions that do not follow the format above will not be reviewed.


Policies

Use of AI tools. You may use AI assistants for debugging, visualization, and boilerplate. You may not use them to produce your implementations, your error analysis, or your written analysis. Running a pretrained speech model is the subject of this course and is always allowed; having a model write your analysis of it is not. Disclose any AI assistance in your submission.

Voice data, consent, and cloning. Any recording of another person requires their informed consent, including consent for how the recording will be stored and who will hear it. Do not synthesize any identifiable person's voice without their written permission, and never for content they did not agree to say. Project datasets involving human subjects need a consent statement at the week 7 milestone. This is a course requirement, not a formality: it is the same standard the field is currently failing to meet.

Academic integrity. Standard Zewail City policy applies. Project code must be your team's own; pretrained models, third-party datasets, and libraries are permitted and must be cited with their licence.


Resources

m-fakhry/DSAI-456-Speech

Vue

21

0 commits

updated Sep 27, 2026

See the code

README

Speech Recognition (DSAI 456) - 2026-2027

Repository for the Speech Recognition undergraduate course (DSAI 456) for the 2026-2027 academic year at Zewail City University.

This year the course

  • transitions the focus from traditional statistical ASR systems to modern neural speech architectures.
  • adds a class project.

Previous Offerings


Logistics

CourseSpeech Recognition - DSAI 456
Webpagehttps://github.com/m-fakhry/DSAI-456-Speech
InstructorProf. Mohamed Ghalwash (mghalwash@zewailcity.edu.eg)
Structure2-hour lecture (Mon 8-10) and 2-hour lab (Mon 10-12, Tue 12-2, Tue 4-6)
TAsEng. Ahmed Aamer
CommunicationMoodle or Email or office hours. no phone, no whatsapp
Lab PolicyAssignments, quizzes, and project milestones
Book"Speech and Language Processing", Jurafsky and Martin, 3rd Edition, 2025
SupplementaryHuggingFace Audio Course
ObjectiveProvide students with the theory and practical skills to build, adapt, and evaluate modern speech systems (recognition and synthesis), and to read the current literature critically
PrerequisitesDeep Learning
Tools/APIsWASP, librosa, HuggingFace Transformers + PEFT. Optional: openSmile, torchaudio, NeMo or ESPnet

Course Learning Outcomes

CLO #OutcomeStatement
1Analyze Speech SignalsExplain the mathematical foundations of digital audio, including sampling, quantization, and the transformation of signals from the time domain to the frequency domain
2Extract Acoustic FeaturesImplement and evaluate feature extraction pipelines, specifically log-mel spectrograms and Mel-Frequency Cepstral Coefficients (MFCCs), for speech processing tasks
3Model Temporal SequencesApply dynamic programming over sequences, specifically the Forward and Viterbi algorithms, to align audio with text, using GMM-HMM systems as the classical case
4Develop and Adapt Neural ASR SystemsDesign and implement modern End-to-End speech recognition architectures, including Connectionist Temporal Classification (CTC), transducer, and Encoder-Decoder frameworks, and fine-tune pretrained models for new languages and domains under realistic label budgets
5Build Generative Speech SystemsExplain neural audio codecs and token-based synthesis, and implement speech generation systems
6Evaluate & Appraise ResearchEvaluate speech systems empirically against baselines, including accuracy-latency trade-offs, and critically appraise contemporary research in speech and audio AI, including its limitations and ethical implications

Lectures

Papers in the Paper column are required reading. Individual papers are not graded, but the midterm and the final each include one question asking you to appraise an assigned paper - to take a position, predict a result, or propose the experiment that would settle a disagreement between two of them. Summaries will not earn marks.

Please note that the syllabus content is subject to change throughout the semester. Topics may be added or removed based on the instructor's discretion, student progress, and available time. Your feedback and participation will inform these adjustments to ensure alignment with course goals and schedule constraints.

WeekDateTopicContentsPaperCLOLectureAssignment
109-21IntroGeneral introduction to the course. WER/CER metrics.Browse Open ASR Leaderboard as a benchmark6Lecture 1Assignment 1
209-28FoundationsFormants, quantization, framing, F0 and intensity, spectrogram, sampling.Ch. 16.1-16.21Lecture 2Assignment 2
310-05Spectral Front EndDFT/FFT, windowing and spectral leakage, time-frequency resolution, STFT and spectrogram, mel filterbank, log-mel. MFCCCh.1, 2
410-12Alignment & DecodingThe alignment problem, HMM, forward algorithm, Viterbi, GMM as a density model.Rabiner (1989), A Tutorial on Hidden Markov Models3Project proposal
510-19CTCBlank symbol, collapse function, summing over alignments, forward-backward, prefix beam search, shallow LM fusionGraves et al. (2006), Connectionist Temporal Classification3, 4
610-26Encoders & TransducersCNN, BiLSTM, Transformer. RNN-T (encoder + prediction + joint network), Conformer/FastConformer, subsampling, chunked attention and causalityGraves (2012), Sequence Transduction with RNNs; Gulati et al. (2020), Conformer4
711-02Self-Supervised LearningWhy labels are the bottleneck, wav2vec 2.0 (masking, quantized targets, contrastive loss), HuBERT (offline clustering, masked prediction), WavLM, low-resource fine-tuningBaevski et al. (2020), wav2vec 2.0; Hsu et al. (2021), HuBERT4Data + evaluation protocol + pretrained baseline
811-09Midterm1, 2, 3, 4
911-16Weak Supervision at ScaleAttention encoder-decoder, cross-attention, Whisper's multitask token format, hallucination and repetition, long-form chunking, the scale-vs-curation debateRadford et al. (2023), Whisper; Puvvada et al. (2024), Less is more4, 6
1011-23Neural Audio CodecsVQ-VAE, straight-through estimator, codebook collapse, residual vector quantization, EnCodec, semantic vs acoustic tokensDéfossez et al. (2022), EnCodec5Working system + training curves
1111-30TTS as Language ModellingVALL-E's two-stage AR + NAR recipe, zero-shot voice cloning from a 3-second prompt, instability and duration control, flow-matching alternatives, vocoders, MOS and speaker similarityWang et al. (2023), VALL-E5, 6
1212-07Speech LLMsSpeech encoder + projector + LLM, what is frozen and what is trained, prompt-conditioned ASR, spoken QA and audio understanding, why a large LLM does not automatically lower WERTang et al. (2023), SALMONN; Qwen3-Omni5, 6
1312-14Project DemosTeam presentations and evaluation-4, 5, 6Final report due (all teams)
1412-21Project DemosTeam presentations and evaluation-4, 5, 6
1512-28Project DemosTeam presentations and evaluation-4, 5, 6
16Finalall

Class Project

  • Teams of 3-4 propose a novel speech application and build a working system around it.
  • Arabic data is highly preferred. Modern and dialectal Arabic, code-switched Arabic-English, and Egyptian speech in particular are underserved by current models, and a project that improves on a pretrained baseline there is doing something that has not already been done a hundred times in English. Arabic is not a hard requirement, but a non-Arabic project has to carry its weight through a stronger idea.

  • The idea matters as much as the implementation. A project is novel if it targets a task, a user group, a dialect, or a setting that existing systems handle badly, or if it combines components in a way the literature has not. It is not novel if it fine-tunes a standard model on a standard benchmark and reports the expected number. You do not need a research contribution - you need a reason the thing you built should exist.

  • Every project must have: a defined task and evaluation metric, a held-out test set the team did not tune on, a pretrained baseline, a system that improves on that baseline, honest failure analysis, and a working demo.

  • The final report must open with a positioning section: what already exists for this task, where it falls short for your case, and what you did differently. Cite real systems and papers, not a general description of the field. How you weight research depth against engineering polish is up to the team and should be stated in the proposal.

  • Teams are encouraged, but not required, to find a mentor or a real user, who is someone has the problem and will tell you whether your output is any good. A linguist, a teacher, a clinician, a call-centre supervisor. Their feedback is worth more than another epoch of training.

  • Milestones

    WeekDeliverable
    4Proposal: the idea, why it is needed, dataset plan, evaluation metric
    7Data (with consent documentation where applicable), evaluation protocol, pretrained baseline results
    10Working system with training curves and error analysis
    13Final report, due for all teams regardless of demo slot; demos run weeks 13-15

Grading Policy

TopicPercentageNotes
Lab Assignments20%Graded in lab with your TA
Lab Quizzes10%Weeks 6 and 12
Class Project20%Distributed across the four milestones
Midterm10%Covers weeks 1-7, including questions on the assigned papers
Final40%Includes question on the assigned papers

Course Instructions

Principle: deadlines are firm.

Submissions

  • All assignments must be uploaded to the Moodle system before the deadline, even if the assignment has already been graded in the lab.
  • Assignment grades depend on discussing your work with your TA. There are no extensions on these discussions: if you miss the discussion for an assignment, you lose its grade.
  • Any assignment involving model training must report the hardware used and the wall-clock training time. "It did not converge" is a result; report it with evidence.

Excuses and make-up tasks

  • Medical excuses must be submitted within one week of the excused task (assignment, quiz, or midterm). Excuses submitted after that window will not be considered.
  • Make-up tasks cover the material taught up to the date of the make-up, not the material of the original task.

Grade petitions

  • Once coursework grades (assignment, quiz, etc.) are posted, you have one week to raise an issue. If you believe there is an error in your grade, email me and CC your TA with:

    1. A clear and detailed explanation of exactly why you are petitioning the grade.
    2. Any relevant and approved documentation supporting your request, if the petition concerns a missed assignment, quiz, or exam.
  • No grade adjustments will be considered after this deadline, and petitions that do not follow the format above will not be reviewed.


Policies

Use of AI tools. You may use AI assistants for debugging, visualization, and boilerplate. You may not use them to produce your implementations, your error analysis, or your written analysis. Running a pretrained speech model is the subject of this course and is always allowed; having a model write your analysis of it is not. Disclose any AI assistance in your submission.

Voice data, consent, and cloning. Any recording of another person requires their informed consent, including consent for how the recording will be stored and who will hear it. Do not synthesize any identifiable person's voice without their written permission, and never for content they did not agree to say. Project datasets involving human subjects need a consent statement at the week 7 milestone. This is a course requirement, not a formality: it is the same standard the field is currently failing to meet.

Academic integrity. Standard Zewail City policy applies. Project code must be your team's own; pretrained models, third-party datasets, and libraries are permitted and must be cited with their licence.


Resources

Languages

Vue

100.0%