After the images of project 2, another kind of non-tabular data: text. In this project you will build, end to end, a sentiment classifier — positive or negative — for social-media opinions about a home delivery service, using classical NLP: text cleaning, bag-of-words and TF-IDF representations, and the models from module 4, with the multinomial Naive Bayes from 04-06 in the leading role (back there we promised it to you "for text"; today it delivers). To keep the project self-contained, the dataset is fictional but realistic: around 40 opinions about "RepartoYa", a delivery service in the style of MercaFresh, built right in the code. In a real case they would be thousands of rows from a public CSV — at the end of the project we tell you where to get them.

Contents

  1. Problem definition and target metric
  2. Building the dataset
  3. Exploring the text
  4. Cleaning and normalization
  5. Representation: from words to numbers
  6. Modeling: Naive Bayes, logistic regression and linear SVM
  7. Evaluation with stratified cross-validation
  8. Interpretation: which words carry weight
  9. Limits of the classical approach
  10. Conclusions

Problem definition and target metric

  • Problem: binary classification of short texts — is the opinion positive (1) or negative (0)?
  • Business use: prioritizing customer support (answering angry customers first) and measuring brand health on social media.
  • Metric: with balanced classes like here, accuracy plus the confusion matrix; in a real deployment you would care especially about the recall of the negative class (no furious customer slips through), the same cost-based reasoning as in 06-04 "ROC curve and AUC".
  • Success criterion: clearly beat random guessing (50%) and, above all, keep the model interpretable: we must be able to see which words convince it.

Building the dataset

With 40 examples you do not train a product: you learn a workflow. We build them balanced and with the real vices of social-media text (caps, URLs, mentions, emojis, excessive punctuation):

import pandas as pd

positives = [
    "My order arrived in 20 minutes, amazing service!!",
    "The fruit was super fresh and the delivery guy was lovely 😍",
    "@RepartoYa you guys are the best, everything perfect as always",
    "Really happy with the delivery, on time and everything well packed",
    "Will definitely order again, good prices and super fast delivery",
    "The frozen fish arrived in impeccable condition, great job",
    "Customer service was a 10, they fixed my problem right away",
    "How wonderful to do the shopping without leaving home 👏",
    "Complete order with zero mistakes, that is how it is done",
    "The app is super convenient and the delivery very fast",
    "Excellent as always, not a single complaint in six months",
    "The delivery guy carried the bags up five flights, what a legend",
    "Everything fresh, everything on time, highly recommended https://repartoya.example",
    "They warned me about the delay and gave me free shipping, well handled",
    "The substitutes they picked were better than my original order",
    "One-hour delivery slot and they nailed it, perfect",
    "Excellent quality on the meat, I will be ordering again",
    "Very good service, the website works great",
    "Five stars, fast and nothing broken",
    "The best online supermarket I have tried, honestly",
]

negatives = [
    "Two hours late and nobody answers the phone 😡",
    "@RepartoYa you delivered me EXPIRED yogurts, disgraceful",
    "Incomplete order again, 5 products missing",
    "The milk arrived warm and the ice cream was soup, awful",
    "Impossible to reach customer service, terrible service",
    "They charged me twice and I am still waiting for the refund",
    "Torn bags and crushed eggs, a total disaster",
    "They cancelled my order without warning, never ordering again",
    "The delivery guy left my groceries at the wrong building, unbelievable",
    "Overripe fruit and stale bread, what a disappointment https://complaint.example",
    "45 minutes on the phone for nothing, horrible experience",
    "Third time it arrives late, they never learn",
    "The app crashes at checkout, impossible to finish the order",
    "They swapped my salmon for crab sticks, seriously??",
    "Everything overpriced and the delivery late on top of it, not worth it",
    "I ordered on sale and they charged me full price, misleading",
    "Nobody replies on social media, terrible support",
    "The order arrived frozen when it should have been fresh, very bad",
    "An hour waiting at the door because they could not find the street",
    "Disappointing from start to finish, I do not recommend it 👎",
]

df = pd.DataFrame({
    "text": positives + negatives,
    "sentiment": [1] * len(positives) + [0] * len(negatives),
})
print(df.shape, "| balance:", df["sentiment"].mean())

Methodological transparency: in a real project this step would be pd.read_csv(...) over thousands of labeled opinions (labeled by hand or by rules such as a review's star rating). Everything that follows works identically with 40 rows or with 40,000 — only the statistical solidity of the results changes, as 02-04 "Statistical inference" taught you: with small samples, wide intervals and humility.

Exploring the text

Text EDA starts by counting things:

df["n_words"] = df["text"].str.split().str.len()
print(df.groupby("sentiment")["n_words"].describe().round(1))

from collections import Counter
neg_words = Counter(" ".join(df[df.sentiment == 0]["text"]).lower().split())
print(neg_words.most_common(10))

Clues are already visible: the negatives are full of "late", "terrible", "nobody", "again"; the positives of "fast", "perfect", "fresh". The classifier we are about to build will formalize exactly this intuition.

Cleaning and normalization

Raw social-media text carries noise with no sentiment signal (URLs, mentions) and superficial variants of the same thing ("Amazing" vs. "amazing"). Cleaning is the textual equivalent of 03-01 "Data cleaning":

import re

def clean_text(text):
    t = text.lower()                     # 1. lowercase
    t = re.sub(r"https?://\S+", " ", t)  # 2. URLs out
    t = re.sub(r"@\w+", " ", t)          # 3. mentions out
    t = re.sub(r"[^\w\s]", " ", t)       # 4. punctuation and emojis out
    t = re.sub(r"\s+", " ", t).strip()   # 5. collapse multiple spaces
    return t

df["text_clean"] = df["text"].map(clean_text)
print(df[["text", "text_clean"]].head(3).to_string())

Every step is a decision with a cost: by removing emojis we lose 😡 and 😍, which are the purest sentiment signal there is (an obvious improvement would be converting them into tokens like emoji_angry before cleaning). Tokenization — chopping the text into words — will be done for us by the vectorizer, splitting on non-alphanumerics; for this simple approach that is enough.

Representation: from words to numbers

No model from module 4 eats text: they eat numeric matrices. The classical solution is bag-of-words: every distinct word in the vocabulary becomes a column, and every text becomes a row with the number of times it uses each word. You lose the order ("not good at all" and "good not at all" look identical) but you gain a table on which everything you have learned just works.

CountVectorizer builds that matrix; TfidfVectorizer improves it by weighting: TF-IDF multiplies the word's frequency in the document (TF) by a factor that penalizes words appearing in almost every document (IDF). So "order", which is everywhere, weighs little, while "expired", rare but revealing, weighs a lot.

from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer

# sklearn's built-in "english" stopword list removes negations,
# which is fatal for sentiment: we define our own
stopwords_en = [
    "the", "a", "an", "and", "to", "of", "in", "it", "is", "was",
    "i", "my", "me", "at", "on", "for", "with", "they", "them", "that",
    "this", "have",
]

vectorizer = TfidfVectorizer(
    stop_words=stopwords_en,
    ngram_range=(1, 2),   # unigrams and bigrams
    min_df=1,
)
X_demo = vectorizer.fit_transform(df["text_clean"])
print(X_demo.shape)  # (40, n_terms): sparse matrix

Two parameters deserve an explanation:

  • ngram_range=(1, 2): besides single words, it uses consecutive pairs (bigrams). It is the cheap way to recover some word order: "not recommend" as a term of its own distinguishes what "not" and "recommend" taken separately would confuse.
  • stop_words: filler words ("the", "of", "that") that only add noise. Careful: with sentiment you must prune with caution — "not" must never be a stopword here.

The resulting matrix is sparse (almost all zeros): each text uses a few dozen terms out of a vocabulary of hundreds. sklearn handles it efficiently without ever converting it to dense.

Modeling: Naive Bayes, logistic regression and linear SVM

Three classics ideally suited to high-dimensional sparse matrices, each in its own Pipeline (vectorizer included, so the vocabulary is learned only from the training portion of each fold — the anti-leak discipline from 06-01):

from sklearn.pipeline import Pipeline
from sklearn.naive_bayes import MultinomialNB
from sklearn.linear_model import LogisticRegression
from sklearn.svm import LinearSVC

def pipe(model):
    return Pipeline([
        ("vec", TfidfVectorizer(stop_words=stopwords_en, ngram_range=(1, 2))),
        ("clf", model),
    ])

models = {
    "MultinomialNB": pipe(MultinomialNB()),
    "Logistic": pipe(LogisticRegression(max_iter=1000)),
    "Linear SVM": pipe(LinearSVC()),
}

MultinomialNB is the reunion promised in 04-06 "Naive Bayes": it models the probability of each word given the class (what is the probability of seeing "expired" in a negative opinion?) and combines them all with Bayes' theorem from 02-05, assuming independence between words — a false assumption that works surprisingly well on text. The logistic regression from 04-02 and the linear SVM from 04-04 instead learn one weight per term, which makes them highly interpretable.

Evaluation with stratified cross-validation

With 40 examples, a single train/test split would be a lottery. We use the stratified CV from 06-03 "Cross-validation", which also keeps the 50/50 class balance in every fold:

from sklearn.model_selection import StratifiedKFold, cross_val_score, cross_val_predict
from sklearn.metrics import confusion_matrix, classification_report

cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)

for name, model in models.items():
    acc = cross_val_score(model, df["text_clean"], df["sentiment"], cv=cv)
    print(f"{name}: {acc.mean():.2f} ± {acc.std():.2f}")

# Aggregated confusion matrix for the best one
y_pred = cross_val_predict(models["MultinomialNB"],
                           df["text_clean"], df["sentiment"], cv=cv)
print(confusion_matrix(df["sentiment"], y_pred))
print(classification_report(df["sentiment"], y_pred,
                            target_names=["negative", "positive"]))
Model Accuracy (CV 5) Comment
MultinomialNB ≈ 0.85–0.95 Excellent with little data: its strong bias protects it
Logistic ≈ 0.80–0.90 Very interpretable; it will need more data to shine
Linear SVM ≈ 0.80–0.90 Similar; it tends to win once there are thousands of examples

With 8 examples per test fold, every single hit moves the metric by 12.5%: the deviations are enormous and no ranking among these three is conclusive. The honest conclusion is the opposite one: all three clearly beat random guessing, a sign that the lexical signal of sentiment is strong. With thousands of examples the differences would indeed become measurable.

Interpretation: which words carry weight

The great virtue of the classical approach: it can be audited. We train the logistic regression on everything and look at its coefficients per term (positive = pushes toward positive sentiment):

model = models["Logistic"].fit(df["text_clean"], df["sentiment"])
vec = model.named_steps["vec"]
weights = pd.Series(model.named_steps["clf"].coef_[0],
                    index=vec.get_feature_names_out())

print("Most POSITIVE terms:\n", weights.nlargest(10).round(2))
print("\nMost NEGATIVE terms:\n", weights.nsmallest(10).round(2))

You should see "fast", "perfect", "fresh" at the top and "late", "terrible", "nobody", "not" at the bottom. This serves three purposes: validating that the model learns sentiment rather than coincidences ("if 'flights' came out as a negative term, I would get suspicious"), explaining individual predictions to the business, and detecting bias — if a proper name or a demographic group showed up with a strong weight, you would be looking at the proxy problem from 08-04 "Ethical and privacy considerations".

Limits of the classical approach

Honesty is mandatory here: bag-of-words does not understand language, it counts words.

  • Negation and context: "not bad at all" is positive; to a bag of words it is "not" + "bad" = extremely negative. Bigrams mitigate this, they do not cure it.
  • Irony and sarcasm: "great, late again 👏" defeats any word counter.
  • Closed vocabulary: a word never seen in training contributes nothing, even if it is "atrocious".

The transformers mentioned in 07-04 "Deep neural networks (Deep Learning)" and among the frameworks in 08-01 (Hugging Face) solve much of this: they read the whole sentence, understand context and come pretrained with the language already learned. In exchange: higher compute cost, less direct interpretability and more infrastructure. In professional practice, TF-IDF + a linear model remains an excellent baseline and is often enough — and it is always the first rung against which any transformer should be measured.

Common Mistakes and Tips

  • Vectorizing before splitting. Fitting the TfidfVectorizer on all the texts leaks vocabulary (and IDF frequencies!) from the test set into training. Inside the Pipeline, it is impossible to get wrong.
  • Removing "not" as a stopword. Generic stopword lists include negations; in sentiment they are the most informative words there are. Always review the list for your problem.
  • Over-cleaning. Emojis, sustained caps ("EXPIRED") and repeated punctuation ("!!!") carry emotional signal. Before deleting them, ask yourself whether they could be features.
  • Believing an accuracy computed on 40 examples. Always report the deviation across folds and, with small samples, prefer qualitative conclusions ("it beats random") over rankings decided by tenths of a point.
  • Lazy labels. If you label reviews by stars (1–2 = negative, 4–5 = positive), document what you do with the 3-star ones: including them mislabeled pollutes more than discarding them.

Challenges to go further

  1. Scale up to a real dataset. Replace the 40 sentences with thousands: the classic IMDb movie-review dataset, the Multilingual Amazon Reviews corpus, or other public review datasets available on Hugging Face Datasets (if you ever work with Spanish-language data, the TASS corpora from the SEPLN sentiment workshops are the reference). Hint: with thousands of examples, rerun the three-model comparison and check whether the SVM overtakes Naive Bayes; add min_df=5 to prune rarities.
  2. Lemmatization with spaCy. Install spacy and the en_core_web_sm model, and replace each word with its lemma ("arrived" → "arrive", "charged" → "charge") before vectorizing. Hint: it shrinks the vocabulary by grouping variants; measure whether it improves the CV or just slows things down — it does not always pay off.
  3. A transformer in three lines. Try Hugging Face's pipeline("sentiment-analysis", model=...) with an English model (search for "sentiment" on their hub) over your 40 sentences and compare its hits with your TF-IDF's, looking especially at the sentences with negation or sarcasm. Hint: it needs no training — it comes pretrained; also note how long it takes per sentence compared to your linear model.

Conclusion

Third project done: you have turned free text into a matrix that the models from module 4 know how to digest — cleaning with regular expressions, bag-of-words, TF-IDF with n-grams —, you have reunited with MultinomialNB on the very ground we promised you in 04-06, you have evaluated with the stratified CV that the small sample demanded, and you have audited the model word by word, which is the superpower of the classical approach over blacker boxes. And you have learned to declare limits: negation and sarcasm mark the border where transformers take over, at their cost. Notice the pattern repeating across these three projects: the representation of the data matters as much as the algorithm. In the next project the challenge will lie not in the representation but in the distribution: detecting fraud when 99% of transactions are legitimate and accuracy becomes a professional liar.

Machine Learning Course

Module 1: Introduction to Machine Learning

Module 2: Foundations of Statistics and Probability

Module 3: Data Preprocessing

Module 4: Supervised Machine Learning Algorithms

Module 5: Unsupervised Machine Learning Algorithms

Module 6: Model Evaluation and Validation

Module 7: Advanced Techniques and Optimization

Module 8: Model Implementation and Deployment

Module 9: Hands-On Projects

Module 10: Additional Resources

© Copyright 2026. All rights reserved