Skip to content
AI.info

AI.info

Speech and audio

Sound as measurement — sampling, features, recognition — and the decisions people build on top of it.

  1. Speech and Audio AI as Measurement and Decision Engineering
  2. Sound Waves, Amplitude, Phase, and Decibels
  3. Sampling, Aliasing, Quantization, and Bit Depth
  4. Microphones, Channels, Gain, Clipping, and Calibration
  5. Time–Frequency Analysis: Fourier Transform, STFT, and Windowing
  6. Spectrograms, Mel Filterbanks, MFCCs, and Learned Front Ends
  7. Pitch, Loudness, Timbre, and Auditory Perception
  8. Audio Units, Segmentation, Labels, and Timestamps
  9. Audio Dataset Design, Rights, Consent, and Provenance
  10. Audio Augmentation, Mixing, Room Simulation, and Leakage
  11. Voice Activity Detection, Endpointing, and Segmentation
  12. Speech Enhancement and Noise Suppression
  13. Dereverberation, Echo Cancellation, and Acoustic Feedback
  14. Microphone Arrays, Beamforming, and Spatial Audio
  15. Source Separation: Masks, Waveforms, and Permutation
  16. Overlapped Speech, Meeting Capture, and Multi-Speaker Attribution
  17. Music Source Separation and Stem Quality
  18. Neural Audio Codecs, Quantization, and Audio Tokens
  19. Automatic Speech Recognition as a Product System
  20. Acoustic Units, Phonemes, Characters, and Subwords
  21. CTC and Monotonic Alignment
  22. Attention Encoder–Decoder Speech Recognition
  23. Transducers and Streaming Recognition
  24. Conformer and Transformer Speech Encoders
  25. Self-Supervised Speech Representation Learning
  26. Multilingual, Low-Resource, Accented, and Code-Switched ASR
  27. ASR Decoding, Language Models, and Contextual Biasing
  28. Text Normalization, Punctuation, Timestamps, and Confidence
  29. Long-Form ASR, Chunking, and Non-Speech Hallucinations
  30. ASR Evaluation Beyond Word Error Rate
  31. ASR Serving, Latency, Privacy, and Monitoring
  32. Speaker Embeddings, Verification, and Identification
  33. Speaker Diarization and Overlap Attribution
  34. Keyword Spotting, Wake Words, and Command Recognition
  35. Spoken Language, Accent, and Paralinguistic Analysis
  36. Speech Health Signals and High-Stakes Boundaries
  37. Replay, Synthetic Speech, and Anti-Spoofing
  38. Voice Biometrics, Privacy, and De-Identification
  39. Text-to-Speech as a Product and the Text Front End
  40. TTS Acoustic Models, Alignment, Duration, and Prosody
  41. Neural Vocoders and End-to-End Speech Synthesis
  42. Multi-Speaker, Multilingual, and Zero-Shot Voice Cloning
  43. Voice Conversion and Speech-to-Speech Translation
  44. TTS Evaluation, Consent, Provenance, and Deepfake Response
  45. Audio Tagging, Sound Event Detection, and Acoustic Scene Analysis
  46. Audio Foundation Models and Audio–Text Representation
  47. Audio Captioning, Retrieval, and Question Answering
  48. Music Information Retrieval: Pitch, Beat, Chords, and Transcription
  49. Text-to-Audio and Music Generation
  50. Bioacoustics and Environmental Monitoring
  51. Real-Time Voice Agents, Duplex Interaction, and Turn-Taking
  52. Audio System Evaluation, Robustness, Fairness, and Deployment
  53. Capstone: Design and Defend a Speech & Audio AI System