AI.info
Speech and audio
Sound as measurement — sampling, features, recognition — and the decisions people build on top of it.
- Speech and Audio AI as Measurement and Decision Engineering
- Sound Waves, Amplitude, Phase, and Decibels
- Sampling, Aliasing, Quantization, and Bit Depth
- Microphones, Channels, Gain, Clipping, and Calibration
- Time–Frequency Analysis: Fourier Transform, STFT, and Windowing
- Spectrograms, Mel Filterbanks, MFCCs, and Learned Front Ends
- Pitch, Loudness, Timbre, and Auditory Perception
- Audio Units, Segmentation, Labels, and Timestamps
- Audio Dataset Design, Rights, Consent, and Provenance
- Audio Augmentation, Mixing, Room Simulation, and Leakage
- Voice Activity Detection, Endpointing, and Segmentation
- Speech Enhancement and Noise Suppression
- Dereverberation, Echo Cancellation, and Acoustic Feedback
- Microphone Arrays, Beamforming, and Spatial Audio
- Source Separation: Masks, Waveforms, and Permutation
- Overlapped Speech, Meeting Capture, and Multi-Speaker Attribution
- Music Source Separation and Stem Quality
- Neural Audio Codecs, Quantization, and Audio Tokens
- Automatic Speech Recognition as a Product System
- Acoustic Units, Phonemes, Characters, and Subwords
- CTC and Monotonic Alignment
- Attention Encoder–Decoder Speech Recognition
- Transducers and Streaming Recognition
- Conformer and Transformer Speech Encoders
- Self-Supervised Speech Representation Learning
- Multilingual, Low-Resource, Accented, and Code-Switched ASR
- ASR Decoding, Language Models, and Contextual Biasing
- Text Normalization, Punctuation, Timestamps, and Confidence
- Long-Form ASR, Chunking, and Non-Speech Hallucinations
- ASR Evaluation Beyond Word Error Rate
- ASR Serving, Latency, Privacy, and Monitoring
- Speaker Embeddings, Verification, and Identification
- Speaker Diarization and Overlap Attribution
- Keyword Spotting, Wake Words, and Command Recognition
- Spoken Language, Accent, and Paralinguistic Analysis
- Speech Health Signals and High-Stakes Boundaries
- Replay, Synthetic Speech, and Anti-Spoofing
- Voice Biometrics, Privacy, and De-Identification
- Text-to-Speech as a Product and the Text Front End
- TTS Acoustic Models, Alignment, Duration, and Prosody
- Neural Vocoders and End-to-End Speech Synthesis
- Multi-Speaker, Multilingual, and Zero-Shot Voice Cloning
- Voice Conversion and Speech-to-Speech Translation
- TTS Evaluation, Consent, Provenance, and Deepfake Response
- Audio Tagging, Sound Event Detection, and Acoustic Scene Analysis
- Audio Foundation Models and Audio–Text Representation
- Audio Captioning, Retrieval, and Question Answering
- Music Information Retrieval: Pitch, Beat, Chords, and Transcription
- Text-to-Audio and Music Generation
- Bioacoustics and Environmental Monitoring
- Real-Time Voice Agents, Duplex Interaction, and Turn-Taking
- Audio System Evaluation, Robustness, Fairness, and Deployment
- Capstone: Design and Defend a Speech & Audio AI System