You now command the recurrent machinery: cells with memory (LSTM/GRU), stacked and bidirectional layers. But one obstacle stands before applying it to TecnoMarket's reviews: neural networks only process numbers, and a review is text. This lesson builds that bridge — tokenization, vocabulary, padding and, above all, embeddings — and applies it to the flagship project of natural language processing (NLP): sentiment analysis. Following the course methodology, we'll first build a prototype with the public IMDB dataset (movie reviews) and then read it through the TecnoMarket lens: classifying customer reviews and feeding the customer-service dashboard. We'll close with an overview of other NLP tasks with RNNs and their limits, which motivate the evolution towards attention.
Contents
- The problem: networks eat numbers, not words
- Tokenization and vocabulary
- Different lengths: padding and truncation
- Embeddings: from indices to meaning
- The Keras Embedding layer
- Prototype: sentiment analysis on IMDB (Embedding + LSTM)
- The TecnoMarket reading: the customer-service dashboard
- Overview: other NLP tasks with RNNs and their limits
The problem: networks eat numbers, not words
Everything you've trained so far received numbers: pixels between 0 and 255 (MNIST, 02-05), sales figures, synthetic sequences. "The battery dies fast" is not a tensor. The standard pipeline for turning text into network input has three stages:
graph LR
A["Raw text<br>'The battery dies fast'"] --> B["Tokenization<br>['the','battery','dies','fast']"]
B --> C["Vocabulary indices<br>[5, 214, 89, 47]"]
C --> D["Padding<br>[5, 214, 89, 47, 0, 0, 0, 0]"]
D --> E["Embedding<br>8 × 32 matrix of real numbers"]
E --> F["LSTM → classification"]
Let's go stage by stage.
Tokenization and vocabulary
Tokenizing means chopping the text into units (tokens): most intuitively, words. Then a vocabulary is built: a dictionary assigning each token an integer index, normally ordered by frequency (low indices = the most common words).
from tensorflow.keras.layers import TextVectorization
import numpy as np
reviews = [
"the battery runs out fast",
"fast shipping and excellent product",
"the product arrived broken",
"excellent quality for the price",
]
vectorizer = TextVectorization(
max_tokens=1000, # maximum vocabulary size
output_sequence_length=8 # fixed output length (automatic padding/truncation)
)
vectorizer.adapt(reviews) # learns the vocabulary FROM the texts
print(vectorizer.get_vocabulary()[:8])
# ['', '[UNK]', 'the', 'product', 'fast', 'excellent', 'and', 'arrived']
print(vectorizer(["fast shipping and excellent product"]).numpy())
# [[ 4 9 6 5 3 0 0 0]]Details that matter:
- Index 0 is reserved for padding and 1 for
[UNK](unknown): any word not in the vocabulary (typos, rare words) gets mapped to[UNK]. Capping the vocabulary (e.g., to the 10,000 most frequent words) is a trade-off: fewer parameters in exchange for losing rare words. - Real tokenization also handles lowercasing, accents and punctuation (
TextVectorizationapplies a default standardization). Modern models use subwords (splitting "unbreakable" into "un-break-able"), but the underlying idea is the same.
Different lengths: padding and truncation
Reviews come in different lengths, but a tensor is rectangular: every row must be the same size. The solution is to fix a length L:
- Shorter sequences → get filled with zeros (padding).
- Longer sequences → get cut (truncation).
Choosing L is a trade-off: too short loses the end of long reviews (where the conclusion usually lives!); too long wastes compute on zeros. A practical rule: look at the 90th-95th percentile of the actual lengths in your corpus.
Embeddings: from indices to meaning
The indices [5, 214, 89, 47] are already numbers, but arbitrary numbers: the fact that "fridge" is 214 and "refrigerator" is 611 says nothing about their similarity. Two options for handing them to the network:
| One-hot | Embedding | |
|---|---|---|
| Representation of a word | Vocabulary-sized vector with a single 1 (e.g., 10,000 dims) | Short dense vector of real numbers (e.g., 32-300 dims) |
| Dimension with a 10,000-word vocab | 10,000 | 32-300 (you choose) |
| Captures similarity? | No: every pair of words is equidistant | Yes: words used in similar ways end up with nearby vectors |
| Is it learned? | No, it's fixed | Yes, by backpropagation like any other weight |
An embedding is a lookup table: a matrix of shape (vocabulary, dimension) where row i is the vector for word i. Those vectors start out random and are learned during training: if "broken" and "defective" appear in similar contexts and predict the same label, backpropagation pushes their vectors closer together — the blame assignment (02-03) reaches all the way down to each word's own representation.
Does this ring a bell? It's the same idea as in 03-04: there, a CNN turned each product image into a feature vector and we measured product similarity with cosine similarity. A word embedding is exactly that for language: a space where geometry encodes meaning. With well-trained embeddings, cosine(v_fridge, v_refrigerator) is high, cosine(v_fridge, v_rug) is low, and famous regularities show up, such as v_king − v_man + v_woman ≈ v_queen.
# Cosine similarity between rows of an embedding matrix (recap from 03-04)
def cosine_similarity(a, b):
return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))The Keras Embedding layer
In Keras, the table is just another layer:
from tensorflow.keras import layers
emb = layers.Embedding(
input_dim=10000, # vocabulary size
output_dim=32, # dimension of each word vector
mask_zero=True # optional: flags the padding (index 0) so the LSTM skips it
)
# Input: (batch, L) integer indices
# Output: (batch, L, 32) a 32-number vector per word- Parameters:
10000 × 32 = 320,000. It's usually the biggest layer in the model — another reason to cap the vocabulary. - Look at the output shape:
(batch, steps, features)— exactly the 3D format the recurrent layers of 04-01 expect. Embedding and LSTM snap together like Lego bricks. mask_zero=Truepropagates a "mask" so the downstream layers don't process the padding zeros as if they were words.
Prototype: sentiment analysis on IMDB (Embedding + LSTM)
Course methodology: prototype on a public dataset first. IMDB ships 50,000 movie reviews, labeled positive (1) or negative (0), already tokenized into indices — perfect for focusing on the model. It's a textbook many-to-one task.
import numpy as np
from tensorflow import keras
from tensorflow.keras import layers
VOCAB = 10000 # we keep the 10,000 most frequent words
L = 200 # fixed sequence length
# 1. Load and prepare
(x_train, y_train), (x_test, y_test) = keras.datasets.imdb.load_data(num_words=VOCAB)
x_train = keras.preprocessing.sequence.pad_sequences(x_train, maxlen=L)
x_test = keras.preprocessing.sequence.pad_sequences(x_test, maxlen=L)
print(x_train.shape) # (25000, 200): 25,000 reviews of 200 indices
# 2. Model: Embedding -> LSTM -> decision
model = keras.Sequential([
layers.Input(shape=(L,)),
layers.Embedding(VOCAB, 32, mask_zero=True), # (200,) -> (200, 32)
layers.LSTM(64), # reads the review, summarizes into 64 numbers
layers.Dense(1, activation="sigmoid") # probability of "positive"
])
model.compile(optimizer="adam",
loss="binary_crossentropy", # BCE: binary classification (02-04)
metrics=["accuracy"])
model.summary()
# Embedding: 320,000 params | LSTM: 4·64·(32+64+1) = 24,832 | Dense: 65
# 3. Train and evaluate
history = model.fit(x_train, y_train, epochs=4, batch_size=64,
validation_split=0.2)
print(model.evaluate(x_test, y_test))
# Typical: ~86-88% test accuracyHow to read the results, as we did with MNIST in 02-05:
- The final sigmoid gives the probability of a positive review; with BCE as the loss, it's the standard binary-classification recipe from 02-04.
- Watch the curves: on IMDB, around epoch 3-4 the training accuracy keeps climbing but validation accuracy stalls or drops — the model is starting to memorize. Stop there (the systematic anti-overfitting techniques arrive in 05-04).
- 87% from such a small model is remarkable: chance would give 50%. The remaining errors are usually reviews with irony, complex negations or mixed sentiment — precisely long, subtle dependencies.
The TecnoMarket reading: the customer-service dashboard
Second step of the methodology: carry it over to TecnoMarket's (fictional) data. The customer-service team receives hundreds of reviews and messages a day; the goal is to prioritize them automatically.
Changes with respect to the prototype:
- Your own data: the store's real reviews with its own labels. Here the model goes from binary to three classes:
positive,negativeandurgent(mentions damage, fraud or legal deadlines). Only the model's head needs changing:Dense(3, activation="softmax")with acategorical_crossentropyloss — the activation/loss-per-problem-type guide from 02-02 and 02-04 tells you exactly what to touch. - A vocabulary of your own: learned with
TextVectorization.adapt()over TecnoMarket's corpus ("doesn't cool", "cracked screen", "refund" need to be in the vocabulary). - Confidence threshold and human review queue, the same deployment architecture we set up for product photos in 03-04:
def route_review(text, model, vectorizer, threshold=0.80):
"""Classifies a review and decides its destination on the dashboard."""
x = vectorizer([text])
probs = model.predict(x, verbose=0)[0] # e.g. [0.05, 0.15, 0.80]
label = ["positive", "negative", "urgent"][int(np.argmax(probs))]
confidence = float(probs.max())
if confidence < threshold:
return {"destination": "human_review", "suggestion": label,
"confidence": confidence}
if label == "urgent":
return {"destination": "priority_inbox", "confidence": confidence}
return {"destination": f"archive_{label}", "confidence": confidence}The principle is the same as in 03-04: the model automates the clear-cut cases and routes the doubtful ones to people. In customer service this is not optional: misclassifying an urgent review (a fridge dripping onto a power strip) costs far more than one extra manual review. That's why, beyond accuracy, what matters most here is not letting urgencies slip through, even at the cost of some false positives.
Overview: other NLP tasks with RNNs and their limits
Sentiment analysis is many-to-one, but for years RNNs covered almost all of NLP:
| Task | Topology | TecnoMarket example |
|---|---|---|
| NER (named entity recognition) | Synchronized many-to-many | Flag which tokens in each review are a product, a brand or an order number, to link it with the catalog |
| Machine translation | Seq2seq (offset many-to-many) | Translate reviews from the international storefronts into English for the support team |
| Summarization | Seq2seq | Condense 40 reviews of a product into three sentences for the internal product sheet |
The seq2seq scheme deserves two conceptual lines: an encoder (an LSTM) reads the entire source sentence and compresses it into its final state; a decoder (another LSTM) generates the target sentence word by word starting from that state. Elegant... and with an obvious bottleneck: the whole source sentence, whether 5 or 60 words long, must fit into a single state vector. With long sentences, the decoder "forgets" the beginning, no matter how LSTM it is. On top of that, recurrence forces word-by-word processing in order, with no parallelization.
The historical solution was to let the decoder "look back" at all the encoder positions and decide which one to attend to at each step: the attention mechanism, which ended up removing recurrence entirely and giving rise to transformers. We won't develop it here: it's the subject of lesson 05-05. Text generation with RNNs (writing character by character) isn't practiced now either: it's guided project 07-02.
Common Mistakes and Tips
- Learning the vocabulary (or the vectorizer) from test data.
adapt()must run only on the training set; otherwise, information from the test set leaks into the model and the evaluation comes out inflated. It's the same data-leakage principle we'll see with scalers in 04-04. - Truncating on the wrong side.
pad_sequencestruncates from the beginning by default (truncating='pre'). In reviews, the conclusion usually sits at the end, so that default is often the right one; but decide consciously, not by accident. - An oversized vocabulary. 100,000 words × 128 dims = 12.8M parameters in the Embedding alone, most of them for words seen 2 or 3 times (impossible to learn well). Trim down to the frequent ones and let
[UNK]absorb the rest. - Forgetting
mask_zero=True(or an equivalent) with heavy padding. If half of each sequence is zeros and the LSTM processes them as words, you dilute the signal. With the mask, the network skips them. - Evaluating with accuracy alone on imbalanced classes. If 90% of reviews are positive, a model that always says "positive" scores 90% accuracy and is worthless. Look at the per-class confusion matrix, especially the recall of
urgent.
Exercises
- Pipeline by hand: with the vocabulary
{'': 0, '[UNK]': 1, 'product': 2, 'excellent': 3, 'broken': 4, 'arrived': 5}and fixed lengthL = 6, convert the review "the product arrived broken" into a sequence of indices (the word "the" is not in the vocabulary). What shape will the output have after anEmbedding(6, 4)? - Parameter count: compute the total parameters of the lesson's IMDB model (Embedding 10,000×32, LSTM(64) with a 32-dim input, Dense(1) over 64). What percentage belongs to the Embedding?
- TecnoMarket design: the dashboard classifies into positive/negative/urgent with a 0.80 threshold. For review X the model returns
[0.42, 0.45, 0.13]. What does theroute_reviewfunction do, and why is that the correct behavior?
Solutions
- Tokens:
['the', 'product', 'arrived', 'broken']→ indices[1, 2, 5, 4]("the" →[UNK]= 1). Padded to 6 (filled at the end):[1, 2, 5, 4, 0, 0]. AfterEmbedding(6, 4): shape(6, 4)— 6 time steps with a 4-number vector each (with batch:(1, 6, 4)). - Embedding:
10,000 × 32 = 320,000. LSTM:4·64·(32+64+1) = 24,832. Dense:64 + 1 = 65. Total:344,897. The Embedding accounts for320,000 / 344,897 ≈ 93%of the parameters — which is why vocabulary size dominates model size. - The most probable class is
negative(0.45), but the maximum confidence (0.45) falls well below the 0.80 threshold, so it returnsdestination: human_reviewwithsuggestion: negative. That's correct: the model is essentially undecided between positive and negative (0.42 vs. 0.45) — probably a mixed or ironic review — and deciding automatically would be a coin toss; a person resolves the case in seconds.
Conclusion
You've closed the bridge between language and networks: tokenization and vocabulary to chop and index, padding to make things rectangular, and embeddings — learned dense vectors where geometric closeness captures semantic similarity, like the image embeddings of 03-04 — to give the indices meaning. With Embedding + LSTM you've built an 87% sentiment classifier on IMDB and carried it over to TecnoMarket's dashboard with three classes and the confidence-threshold-plus-human-review setup from 03-04. You've also seen the map of NLP tasks (NER, seq2seq for translation and summarization) and the single-vector bottleneck that motivated attention and transformers (05-05). One flagship project of the module remains: the TecnoMarket sequences that aren't words but dated numbers — forecasting daily demand, where time series, sliding windows and a classic mistake or two we'll learn to dodge are waiting for us.
Deep Learning Course
Module 1: Introduction to Deep Learning
- What is Deep Learning?
- History and evolution of Deep Learning
- Applications of Deep Learning
- Basic concepts of neural networks
- Setting up the work environment
Module 2: Neural Network Fundamentals
- Perceptron and Multilayer Perceptron
- Activation functions
- Forward and backward propagation
- Optimization and loss functions
- Your first complete neural network
Module 3: Convolutional Neural Networks (CNN)
- Introduction to CNNs
- Convolutional and pooling layers
- Popular CNN architectures
- CNN applications in image recognition
Module 4: Recurrent Neural Networks (RNN)
- Introduction to RNNs
- LSTM and GRU
- RNN applications in natural language processing
- Sequences and time series
Module 5: Advanced Deep Learning Techniques
- Generative Adversarial Networks (GAN)
- Autoencoders
- Transfer Learning
- Regularization and improvement techniques
- Attention mechanisms and Transformers
Module 6: Tools and Frameworks
- Introduction to TensorFlow
- Introduction to PyTorch
- Framework comparison
- Development environments and additional resources
- Saving, loading and deploying models
Module 7: Hands-On Projects
- Image classification with CNNs
- Text generation with RNNs
- Anomaly detection with Autoencoders
- Building a GAN for image generation
- Fine-tuning a pretrained model
