Skip to content
AI.info

Natural language processing

Sentiment, Emotion, Stance, and Aspect Analysis

Separate sentiment, emotion, stance, target, holder, intensity, and aspect while avoiding unsupported inferences about people.

By the end you can

Example

“Positive” about what, according to whom?

A product review can hold several evaluations at once, and a document-level label compresses them badly. The constructs underneath are hard enough that trained people disagree about them. Three judges classified the same 270 tweets as sarcastic, positive or negative. They agreed unanimously on 135 of them. Credit only those unanimous judgements and accuracy over the whole test set falls to 43.33%.

  • “The screen is beautiful, but the battery is unacceptable” carries opposite sentiments about two aspects of one product. SemEval-2014 Task 4 gave that case a polarity value of its own, conflict, alongside positive, negative and neutral. In the restaurant training set alone, 91 aspect terms carry it.
  • “Great, another delayed train” uses positive wording sarcastically to criticise, and that costs more than it looks. A 2011 study by González-Ibáñez and colleagues assembled 900 sarcastic, 900 positive and 900 negative self-labelled tweets, sampled 270 of them, and gave the three-way call to three human judges. Overall agreement was 50%, Fleiss' kappa 0.4788, mean accuracy 62.59%. Their best classifier scored 57.41%. The abstract does not soften it: “Perhaps unsurprisingly, neither the human judges nor the machine learning techniques perform very well.”
  • “Analysts welcomed the plan, while unions opposed it” attributes different stances to different groups.
  • “I am worried the update may fail” expresses emotion and uncertainty, not an observed failure.
  • “The article calls the policy reckless” reports another source’s stance rather than the author’s own.
  • “This drug is aggressive against the tumor” contains evaluative vocabulary without a human emotion.

Comparison

Related tasks with different labels

Using “sentiment” as a catch-all creates invalid interpretations. The distance between two of these constructs has been measured rather than asserted. On the SemEval-2016 stance data, Mohammad and colleagues built an oracle. It is handed the gold sentiment label of every tweet, and it may then pick the best sentiment-to-stance mapping for each target. Perfect sentiment, converted as favourably as it can be converted, reaches F-macroT 53.1 and F-microT 57.2. The shared task's winning stance system reached F-macroT 56.0 and F-microT 67.8. A system trained to predict stance beats flawless sentiment labels. The authors say why: “This shows that even though sentiment can play a key role in detecting stance, sentiment alone is not sufficient.”

So the four constructs below are not four names for one judgement. Sentiment evaluates a target as positive, negative, neutral or graded, and can be mixed. Emotion identifies expressed or described affect; there the holder matters, and expression differs from experience. Stance determines support, opposition or uncertainty toward a proposition or entity that may never appear in the text. Subjectivity separates opinion and evaluation from more factual reporting, clause by clause, with quotation complicating who owns the view.

FigureComparison · 4 columns

Sentiment or polarity

Evaluate a target as positive, negative, neutral, or graded.

  • Requires a target
  • Can be mixed
  • Domain-sensitive words
  • Useful for reviews and feedback

Emotion

Identify expressed or described affect categories or dimensions.

  • Holder matters
  • Expression differs from experience
  • Cultural variation
  • Avoid diagnosis

Stance

Determine support, opposition, or uncertainty toward a proposition or entity.

  • Target may be absent from text
  • Can differ from sentiment
  • Attribution matters
  • Useful for claims and debates

Subjectivity

Distinguish opinion, evaluation, or perspective from more factual reporting.

  • Not identical to falsehood
  • Can be clause-level
  • Quotation complicates ownership
  • Guideline-dependent

Visual

An opinion frame has several slots

A single polarity label discards most of this structure. It discards who expresses or is attributed the evaluation, which entity or feature is evaluated, and which words support the label. It discards the direction and the degree. It discards the context too: negation, modality, time, comparison, irony and quoted source.

FigureHierarchy · 5 levels
  • Holder

    Who expresses, reports, or is attributed the evaluation?

    • Target or aspect

      Which entity, feature, proposition, or event is evaluated?

      • Expression and evidence

        Which words or constructions support the label?

        • Orientation and intensity

          What direction and degree are expressed?

          • Context

            Negation, modality, time, comparison, irony, and quoted source.

Analogy

A restaurant scorecard with separate dishes

A diner rates service, price, dessert and noise independently rather than assigning one number to the entire evening. The aspects reveal where praise and criticism coexist.

A scorecard reads as one diner’s settled preference, and language evaluation is not always that. Reviews can quote others, joke, compare, or strategically present a stance.

Aspect analysis preserves the target and evidence behind an evaluation.

Key idea

Language cues do not justify psychological surveillance

A message that contains anger words does not establish a clinical condition, a stable personality or future behavior. Models can also misread dialect, reclaimed language, humor, or culturally specific expression.

Limit outputs to the annotated construct, disclose uncertainty, and avoid high-consequence inferences without valid evidence and governance.

It is worth knowing how little of that restraint is imposed from outside. European law prohibits AI systems used “to infer emotions of a natural person in the areas of workplace and education institutions”, except for medical or safety reasons. That is Article 5(1)(f) of the AI Act, Regulation (EU) 2024/1689, and it has applied since 2 February 2025. Recital 44 gives the reason in the legislature's own words: “serious concerns about the scientific basis of AI systems aiming to identify or infer emotions”, and their “limited reliability, the lack of specificity and the limited generalisability”. That is a regulator agreeing with this section.

Then read how far the rule reaches. On 29 July 2025 the European Commission published guidelines on prohibited AI practices, and they state: “An AI system inferring emotions from written text (content/sentiment analyses) to define the style or the tone of a certain article is not based on biometric data and therefore does not fall within the scope of the prohibition.” The prohibition is built on biometric data. A text emotion classifier reading employee messages sits outside it. The scientific concerns Recital 44 records apply to it in full; Article 5(1)(f) does not. What stops that deployment is the team's own restraint, and nothing else in this paragraph.

Classify the linguistic expression you defined, not an imagined inner state.

The same subtask scored 84.01% on restaurants and 74.55% on laptops

“Unpredictable” can be negative for a vehicle, positive for a thriller, and merely descriptive in a scientific report. “Positive” in medicine can indicate a finding rather than approval.

Domain adaptation needs labeled examples, target-aware evaluation, and phrase-level context. Blindly importing a general lexicon can produce confident category errors.

SemEval-2014 Task 4 put a number on the gap by holding everything else fixed. It released 7,686 hand-annotated review sentences in two domains: 3,841 restaurant sentences (3,041 training, 800 test) and 3,845 laptop sentences (3,045 and 800). One scheme covered both, with four aspect-term polarity values — positive, negative, neutral, and conflict for terms evaluated both ways at once. Same extraction subtask, same competitors, same year. Best F1 was 84.01% on restaurants and 74.55% on laptops. The task report states it plainly: “Overall, the systems achieved significantly higher scores (+10%) in the restaurants domain, compared to laptops.” Nothing changed but the things being reviewed.

Polarity belongs to an expression–target–domain relationship, not to a word forever.

Case

Phrase-level sentiment exists because 215,154 phrases from film reviews were labelled by hand

Phrase-level sentiment became measurable because someone paid for it by hand. The Sentiment Treebank, built by Socher and colleagues and published in 2013, carries “fine grained sentiment labels for 215,154 phrases in the parse trees of 11,855 sentences”. Those sentences come from a 10,662-sentence corpus of movie reviews. Each one was parsed, and every resulting phrase was labelled through Amazon Mechanical Turk. Two hundred thousand judgements bought sentiment at phrase level in exactly one domain: film criticism. Nothing in that corpus says what “unpredictable” means in a maintenance log. The ten-point restaurant-to-laptop drop in SemEval-2014 Task 4 is what the distance between two domains costs when someone measures it.

Steps

Build an opinion annotation scheme

The scheme should preserve disagreement rather than force annotators to infer unsupported intent. Disagreement here is not a rounding error. Three trained annotators worked over the same 210 sentences of world-press text under the MPQA scheme, and in 2005 Wiebe and colleagues reported what they found: “In the 210 sentences in the annotation study, the annotators A, M, and S respectively marked 311, 352 and 249 expressive subjective elements.” One text, three trained people, a 41% spread in how many opinion spans it contains. Average pairwise agreement was 0.72. It rose to 0.80 for medium-or-higher intensity spans and 0.88 for high-or-extreme ones. Agreement on the explicit private-state and speech-event anchors — the ones with a word you can point at — was higher, at 0.82. The annotators rated 89% of editorials as difficult and 73% of objective-topic articles as easy.

Read those numbers as a design brief for the five steps. Define the construct and the unit: document, sentence, clause, holder–target pair, or aspect. Mark holder, target and evidence, separating quoted, author and user-attributed opinions. That is the part that reached 0.82 rather than 0.72, because it is anchored to visible words. Specify context rules for negation, comparison, modality, mixed sentiment, irony and uncertainty. Include cannot-determine labels for ambiguity and insufficient context. A scheme that forces one answer on the sentences in the editorial 89% does not produce agreement; it hides the 311-versus-249 disagreement inside a single number. Then audit group and domain behavior across dialect, language, topic, source and consequence slices.

FigureProcess · 5 steps
  1. 1. Define the construct and unit

    Choose document, sentence, clause, holder–target pair, or aspect.

  2. 2. Mark holder, target, and evidence

    Separate quoted, author, and user-attributed opinions.

  3. 3. Specify context rules

    Cover negation, comparison, modality, mixed sentiment, irony, and uncertainty.

  4. 4. Include cannot-determine labels

    Use ambiguity and insufficient-context options where appropriate.

  5. 5. Audit group and domain behavior

    Review dialect, language, topic, source, and consequence slices.

Evaluate structure and downstream use

Document accuracy can hide wrong targets, holders or evidence. Report aspect, holder, span, orientation, intensity and attribution errors where the schema includes them.

SemEval-2016 Task 6 shows what one headline score conceals. It annotated 4,870 English tweets for stance towards six US targets — 2,914 training and 1,249 test instances across five targets in Task A, plus 707 Donald Trump tweets in Task B — and 19 teams entered. Now split the test set by whether the opinion expressed in the tweet is aimed at the target of stance or at some other entity. One system becomes two. The winning system scores F-macroT 59.7 on the tweets whose opinion is directed at the target and 35.4 on those directed at another entity (F-microT 72.5 against 44.5). The organisers' own SVM scores 62.5 against 37.9 (75.0 against 43.0). Their abstract records it: “However, systems found it markedly more difficult to infer stance towards the target of interest from tweets that express opinion towards another entity.” A single average conceals the split between the cases where the target is on the surface and the cases where it has to be reasoned about.

The same effect appears when one paper reports two accuracies. Socher and colleagues report that their model “pushes the state of the art in single sentence positive/negative classification from 80% up to 85.4%”. Separately, “The accuracy of predicting fine-grained sentiment labels for all phrases reaches 80.7%, an improvement of 9.7% over bag of features baselines”. One number is a document-level binary call. The other is a label on every phrase in a parse tree. Report whichever one the product actually consumes, and say which it is.

For monitoring or aggregation, examine prevalence, calibration, sampling bias, and who chooses to submit feedback. A dashboard summarizes recorded language, not the complete population’s feelings. In the US that has been more than methodological advice since 21 October 2024, when the FTC's rule on consumer reviews and testimonials, 16 CFR Part 465, took effect. Under §465.7(b) it is an unfair or deceptive practice “For a business to materially misrepresent, expressly or by implication, that the consumer reviews of one or more of the products or services it sells displayed in a portion of its website or platform dedicated in whole or in part to receiving and displaying consumer reviews represent most or all the reviews submitted to the website or platform when reviews are being suppressed (i.e., not displayable) based upon their ratings or their negative sentiment.” The rule leaves room only for sentiment-blind withholding criteria: confidential information, abusive or obscene content, suspected fakes, off-topic reviews. Note what the rule has isolated. The offence is not a wrong polarity label. It is a filter upstream of every polarity label, selecting on the very quantity the dashboard then averages.

Opinion metrics inherit both linguistic ambiguity and participation bias.

Position

The language identifier decides whose opinion is in the number

Everything a sentiment project argues about happens after the messages have arrived: which construct, which schema, which model, which slice of the error report. The measurement in this lesson sits before all of that, in the step nobody on the project owns.

langid.py is a widely used open-source language identifier, the kind of component added to a pipeline in one line. Blodgett and colleagues ran it over demographically aligned Twitter data and reported in 2016: “Of the AA-aligned tweets, 13.2% were classified by langid.py as non-English; in contrast, 7.6% of white-aligned tweets were classified as such”. Then they sampled the rejections, 50 per run, and hand-checked them: “Of these 300 tweets, only 3 could be unambiguously identified as written in a language other than English”. So what that component rejected was, almost entirely, not another language. At nearly twice the rate, it was removing one variety of English.

Put a component like that at the front of a collection pipeline and follow where those messages go. That is the part no downstream metric can catch. They are absent from the training data. If the evaluation set is drawn along the same collection path — and it usually is, because it comes out of the same pipe — they are absent from that too. They are absent from the denominator of the dashboard, which this lesson has already warned summarises recorded language rather than a population’s feelings. A sentiment classifier can then be re-annotated, recalibrated and re-evaluated to any standard you like, and every number it reports can improve without one of those messages ever coming back. The error was committed by a component with no sentiment metric attached to it. That is why no sentiment metric will find it.

The commercial version of the same structure has a price. From late 2015 through November 2019, reviews arrived at fashionnova.com and a filter stood in front of them. In the Federal Trade Commission's own words: “According to the Commission’s proposed complaint, from late 2015 through November 2019, Fashion Nova had four- and five-star reviews automatically posted to its website but did not approve for posting or publish lower-starred, more negative reviews.” The FTC's press release describes hundreds of thousands of suppressed reviews. On 25 January 2022 the Commission announced a settlement: Fashion Nova, LLC paid $4,200,000 and accepted a 20-year order requiring it to display all submitted product reviews, with only sentiment-blind exclusions. For four years, any average rating or sentiment summary computed from that page was arithmetically correct on the reviews it had. The reviews it did not have had been removed for their sentiment — the identical shape as langid.py, with a commercial motive in place of a training-data artefact.

Hence the position: no accuracy figure for the sentiment model can tell you whether the sentiment number is right. What decides that is the collection path — who was admitted, by which component, at what unequal rate. The authors state the consequence in two verbs — “dialect speakers’ opinions may be mischaracterized under social media sentiment analysis or omitted altogether” — and the second one is the dangerous one. A mischaracterisation shows up as an error. An omission shows up as nothing at all.

Of 300 tweets the filter rejected as non-English, three could be unambiguously identified as another language.

Rewrite a sentiment claim into a defensible analysis

Take a claim such as “Customers hate the new interface.” Specify the sampled source, the time, the aspect, the holder, the label definition, the uncertainty and the missing population. For that last one, name the component that did the excluding — as langid.py did at 13.2% against 7.6%, and as the filter on fashionnova.com did for four years.

Then design a model output that returns target, evidence, orientation and cannot-determine status. Say what the product must not conclude. Not a stance: perfect sentiment labels reached only F-macroT 53.1 against the stance system's 56.0. Not a diagnosis: Recital 44 of Regulation (EU) 2024/1689 records “serious concerns about the scientific basis of AI systems aiming to identify or infer emotions”. And not a claim about a population, because the denominator was assembled by components that no sentiment metric evaluates.

A responsible result narrows the claim to what the language evidence actually supports.

Key takeaways