AI.info
Natural language processing
Text as engineering — corpora, rights, tokens, and evaluation that survives contact with real language.
- Natural Language Processing as Communication Engineering
- Corpus Design: Domains, Documents, Rights, and Sampling
- Unicode, Normalization, and the Boundaries of Text
- Segmentation: Sentences, Words, Subwords, and Bytes
- Vocabulary Design and Tokenization Diagnostics
- Morphology and Lexical Structure
- Syntax: Constituents, Dependencies, and Parsing
- Semantics: Meaning, Ambiguity, and Lexical Relations
- Pragmatics, Discourse, and Coreference
- Sparse Text Representations: Counts, N-Grams, and TF–IDF
- Static Word Embeddings and Distributional Meaning
- Contextual Representations and Transformer Encoders
- Language Modeling and Sequence Probability
- Topic Models and Unsupervised Text Discovery
- Text Classification and Intent Detection
- Multilabel, Hierarchical, and Long-Document Classification
- Sequence Labeling: POS, Named Entities, and Span Boundaries
- Relations, Events, and Slot Extraction
- Sentiment, Emotion, Stance, and Aspect Analysis
- Lexical Search, Indexing, and Ranking
- Dense Retrieval and Neural Reranking
- Question Answering and Reading Comprehension
- Summarization and Content Condensation
- Machine Translation, Alignment, and Meaning Preservation
- Multilingual and Cross-Lingual NLP
- Semantic Parsing and Structured Prediction
- Dialogue Systems and Conversational State
- Large Language Models in the NLP Toolkit
- NLP Evaluation: Metrics, Human Judgment, and Behavioral Tests
- Error Analysis, Calibration, and Abstention
- Robustness, Bias, Privacy, and Domain Shift
- Production NLP Systems and Monitoring
- Capstone: Design, Evaluate, and Defend an End-to-End NLP System