Skip to content
AI.info

AI.info

Natural language processing

Text as engineering — corpora, rights, tokens, and evaluation that survives contact with real language.

  1. Natural Language Processing as Communication Engineering
  2. Corpus Design: Domains, Documents, Rights, and Sampling
  3. Unicode, Normalization, and the Boundaries of Text
  4. Segmentation: Sentences, Words, Subwords, and Bytes
  5. Vocabulary Design and Tokenization Diagnostics
  6. Morphology and Lexical Structure
  7. Syntax: Constituents, Dependencies, and Parsing
  8. Semantics: Meaning, Ambiguity, and Lexical Relations
  9. Pragmatics, Discourse, and Coreference
  10. Sparse Text Representations: Counts, N-Grams, and TF–IDF
  11. Static Word Embeddings and Distributional Meaning
  12. Contextual Representations and Transformer Encoders
  13. Language Modeling and Sequence Probability
  14. Topic Models and Unsupervised Text Discovery
  15. Text Classification and Intent Detection
  16. Multilabel, Hierarchical, and Long-Document Classification
  17. Sequence Labeling: POS, Named Entities, and Span Boundaries
  18. Relations, Events, and Slot Extraction
  19. Sentiment, Emotion, Stance, and Aspect Analysis
  20. Lexical Search, Indexing, and Ranking
  21. Dense Retrieval and Neural Reranking
  22. Question Answering and Reading Comprehension
  23. Summarization and Content Condensation
  24. Machine Translation, Alignment, and Meaning Preservation
  25. Multilingual and Cross-Lingual NLP
  26. Semantic Parsing and Structured Prediction
  27. Dialogue Systems and Conversational State
  28. Large Language Models in the NLP Toolkit
  29. NLP Evaluation: Metrics, Human Judgment, and Behavioral Tests
  30. Error Analysis, Calibration, and Abstention
  31. Robustness, Bias, Privacy, and Domain Shift
  32. Production NLP Systems and Monitoring
  33. Capstone: Design, Evaluate, and Defend an End-to-End NLP System