Closing the CNN module, we left a door open: convolutional networks rule TecnoMarket's catalog images, but two of its most valuable problems — analyzing customer reviews and forecasting daily demand — are not images. They are sequences: data where order matters and where each element depends on the ones before it. In this lesson you'll discover why dense networks and CNNs fall short with this kind of data, and meet the architecture designed specifically for it: the recurrent neural network (RNN). You'll understand its central idea (a hidden state that acts as memory), its mathematical formulation with a numerical example worked by hand, its typical topologies and its first Keras implementation, plus its great limitation, which will motivate the next lesson.
Contents
- Sequential data: when order is everything
- Why dense networks and CNNs don't model order
- The idea of recurrence: a hidden state as memory
- Unrolling an RNN through time
- The RNN formula with a numerical example by hand
- Topologies: one-to-many, many-to-one, many-to-many
- SimpleRNN in Keras on a synthetic sequence
- The limitation: short-term memory
Sequential data: when order is everything
Sequential data is an ordered collection of elements where position and context matter. Shuffle the elements and you destroy the information.
Examples at TecnoMarket:
| Data | Elements of the sequence | What happens if you shuffle it? |
|---|---|---|
| Customer review | Words: "not", "bad", "at", "all" | "bad all at not" is gibberish; worse still: "not bad at all" (positive) and "not good at all" (negative) differ by a single word |
| Daily demand | Sales per day: 120, 135, 128, 410... | You lose the trend, the weekly seasonality and the Black Friday spike |
| A user's clicks | home → phones → iPhone → cart | The order distinguishes "bought after comparing" from random browsing |
| Call-center audio | Sound samples over time | Scrambled speech is just noise |
Look closely at the first example: in English, "not bad at all" is praise. A model that merely counts words ("not", "bad" → negative!) will fail systematically. We need a model that reads in order and remembers what it has read.
Why dense networks and CNNs don't model order
You already know two families of architectures. Let's see why neither is quite the right fit:
- Dense network (like the 784-128-64-10 from MNIST):
- Expects a fixed-size input. Reviews come in different lengths (5 words or 200) and a sales series grows every day.
- Treats each input position independently: nothing in the architecture says that word 3 comes after word 2. To the network, the input is a "bag" of numbers.
- If you train a dense network to spot "fast shipping" at positions 1-2, it won't recognize it at positions 7-8: it would have to learn the pattern at every position separately (the same parameter-explosion problem we saw in 03-01 with images).
- CNNs:
- They solve part of the problem: their shared-weight filters detect a local pattern wherever it appears (in fact, 1D CNNs for text exist). But their "field of view" is local and fixed-size: a filter of size 3 sees 3 consecutive words.
- A long dependency — "The TV I bought two months ago, which arrived with a cracked screen, doesn't work" — requires connecting "TV" with "doesn't work" 15 words apart. Stacking many convolutional layers widens the view, but in a rigid and expensive way.
What we need is a mechanism that processes the elements one by one, in order, carrying along a summary of everything seen so far. That is exactly what recurrence is.
The idea of recurrence: a hidden state as memory
An RNN processes the sequence step by step with a single cell (a small neural network block) applied over and over. At each time step t, the cell receives two things:
- The current input
x_t(today's word, today's sales figure). - The hidden state
h_{t-1}: a vector summarizing everything processed up to the previous step.
And it produces a new hidden state h_t, which is passed on to the next step. The hidden state is the network's memory: a compressed, fixed-size summary of the whole sequence read so far.
Think of TecnoMarket's purchasing committee from module 1, but reading a review aloud word by word: after each word, the committee updates a mental note ("so far it sounds positive... careful, they said 'but'... now they're complaining about the shipping"). That mental note is h_t. By the end of the review, the note summarizes the full opinion.
Two key properties:
- The same cell at every step: the weights do not change from one time step to the next. It's the exact parallel of CNNs: where a CNN shares weights across space (the same filter sweeps the whole image), an RNN shares weights across time (the same cell sweeps the whole sequence). That's why an RNN can process sequences of any length with a fixed number of parameters.
- The state has a fixed size: whether you feed it 10-word or 300-word reviews,
h_tis always a vector of, say, 64 numbers. The network learns to decide what deserves a place in that summary.
Unrolling an RNN through time
Although the RNN is a single cell with a loop, to understand it (and to train it) it helps to "unroll" it through time, as if it were a deep network with one layer per step:
graph LR
subgraph "RNN unrolled through time"
x1["x₁: 'not'"] --> C1[RNN cell]
x2["x₂: 'bad'"] --> C2[RNN cell]
x3["x₃: 'at'"] --> C3[RNN cell]
x4["x₄: 'all'"] --> C4[RNN cell]
h0["h₀ = 0"] --> C1
C1 -- "h₁" --> C2
C2 -- "h₂" --> C3
C3 -- "h₃" --> C4
C4 -- "h₄" --> S["Output: sentiment"]
end
Important points about the diagram:
- The four "RNN cell" boxes are the same cell with the same weights, drawn four times.
- The initial state
h₀is usually a vector of zeros: at the start, the network has read nothing. - Information flows left to right along the chain of states: for "not" (step 1) to influence the final decision, its effect must survive through
h₁ → h₂ → h₃ → h₄.
That last point, which now looks like a detail, will turn out to be the key to the limitation we'll meet at the end.
The RNN formula with a numerical example by hand
The basic RNN cell (the one Keras calls SimpleRNN) computes:
Where:
x_t: input vector at stept.h_{t-1}: previous hidden state.Wx: input→state weight matrix (learned).Wh: state→state weight matrix (learned). This is the recurrence matrix: it decides how much of the memory is kept, and how.b: bias (learned).tanh: the activation you already know from 02-02; it keeps the state between -1 and 1, which stops the memory from growing out of control step after step.
Let's do the computation by hand with tiny dimensions: an input of 1 number, a state of 2 numbers.
Suppose these already-trained weights:
And the input sequence x = [1, 0, 1] (three time steps). We start from h₀ = [0, 0].
Step 1 (x₁ = 1):
Step 2 (x₂ = 0; all the information comes from memory):
z₂ = Wx·0 + Wh·h₁ = [0.8·0.462 + (-0.2)·(-0.291), 0.1·0.462 + 0.6·(-0.291)] = [0.428, -0.129] h₂ = tanh([0.428, -0.129]) = [0.404, -0.128]
Step 3 (x₃ = 1):
z₃ = Wx·1 + Wh·h₂ = [0.5 + 0.8·0.404 + (-0.2)·(-0.128), -0.3 + 0.1·0.404 + 0.6·(-0.128)] = [0.849, -0.336] h₃ = tanh([0.849, -0.336]) = [0.690, -0.324]
Notice two things: at step 2, even though the input was 0, the state did not empty out — the memory term Wh·h₁ kept the information from step 1 alive (attenuated: from 0.462 down to 0.404). And at step 3, the output is not the same as at step 1 even though the input is identical (x=1): the accumulated context changes the result. That is exactly what a dense network cannot do.
In numpy, the same computation generalized:
import numpy as np
def simple_rnn_step_by_step(x_seq, Wx, Wh, b):
"""Runs a SimpleRNN by hand over a sequence.
x_seq: (steps, input_dim), returns all the states h_t."""
h = np.zeros(Wh.shape[0]) # h0 = vector of zeros
states = []
for x_t in x_seq: # one time step per iteration
h = np.tanh(Wx @ x_t + Wh @ h + b) # the cell's formula
states.append(h)
return np.array(states)
Wx = np.array([[0.5], [-0.3]])
Wh = np.array([[0.8, -0.2], [0.1, 0.6]])
b = np.zeros(2)
x = np.array([[1.0], [0.0], [1.0]]) # 3-step sequence
print(simple_rnn_step_by_step(x, Wx, Wh, b))
# [[ 0.462 -0.291]
# [ 0.404 -0.128]
# [ 0.690 -0.324]]The for loop is literal: an RNN is a loop over time with the same weights on every pass.
Parameter count
With input dimension d and state dimension u, the cell has d·u (Wx) + u·u (Wh) + u (b) parameters. For d=1, u=2: 2 + 4 + 2 = 8. Note that this does not depend on the sequence length: 8 parameters handle 3 steps just as well as 3,000, thanks to the weights shared across time.
Topologies: one-to-many, many-to-one, many-to-many
Depending on how many inputs and outputs the sequence has, there are several ways to use an RNN:
| Topology | Input | Output | TecnoMarket example |
|---|---|---|---|
| Many-to-one | Sequence | One value | Read a whole review → one label (positive/negative) |
| One-to-many | One value | Sequence | A product photo (CNN vector) → generate its description word by word |
| Many-to-many (synchronized) | Sequence | Sequence of the same length | For each word of a review, tag whether it is a product name or not |
| Many-to-many (offset / seq2seq) | Sequence | Sequence of a different length | Translate a review from one language into another |
In this module we'll mostly work with many-to-one: it's the natural shape of sentiment analysis (04-03) and of windowed demand forecasting (04-04): many observations go in, one prediction comes out. Text generation (one-to-many) will be developed as a project in 07-02, and we'll cover seq2seq conceptually in 04-03.
SimpleRNN in Keras on a synthetic sequence
Let's verify that an RNN really learns a temporal dependency. Synthetic task: given a sequence of 10 numbers, predict the sum of the last 3. A task that is impossible without knowing which position each number occupies.
import numpy as np
from tensorflow import keras
from tensorflow.keras import layers
# 1. Generate the synthetic dataset
rng = np.random.default_rng(42)
X = rng.uniform(0, 1, size=(2000, 10, 1)) # 2000 sequences, 10 steps, 1 value per step
y = X[:, -3:, 0].sum(axis=1) # target: sum of the last 3 steps
# 2. Model: many-to-one SimpleRNN + regression output
model = keras.Sequential([
layers.Input(shape=(10, 1)), # (steps, features per step)
layers.SimpleRNN(16), # returns only the LAST state h_10 (16 numbers)
layers.Dense(1) # from the final summary to the prediction
])
model.compile(optimizer="adam", loss="mse", metrics=["mae"])
model.summary()
# 3. Train
history = model.fit(X, y, epochs=30, batch_size=32,
validation_split=0.2, verbose=0)
print(f"Final validation MAE: {history.history['val_mae'][-1]:.3f}")
# Typical: ~0.05 (on sums around 1.5, a very small error)Reading keys for beginners:
- The input shape of an RNN in Keras is always
(samples, steps, features). Here: 2000 sequences, 10 steps each, 1 number per step. That third axis is confusing at first: even if each step is a single number, you must declare it. SimpleRNN(16)creates a cell with a 16-dimensional state. Parameters:1·16 + 16·16 + 16 = 288(check it in thesummary()).- By default the layer returns only the last state
h_10: perfect for many-to-one. (Thereturn_sequences=Trueparameter, which returns all the states, we'll use in 04-02 to stack layers.) - The network learns on its own to "ignore" the first 7 numbers and "remember" the last 3: nobody told it which positions matter.
The limitation: short-term memory
Thought experiment: change the previous task to "predict the sum of the first 3 numbers of a 100-step sequence". Now the relevant information must survive 97 steps inside the hidden state... and the SimpleRNN will fail.
Why? Training an RNN uses backpropagation through time (BPTT): conceptually, it is the same backpropagation from 02-03 applied to the unrolled network — the "blame assignment" travels backwards along the chain of cells, step by step. And here an old acquaintance reappears: at each backward step, the gradient gets multiplied by Wh (and by the derivative of the tanh, which is at most 1). After many steps:
- If those factors are < 1, the gradient vanishes exponentially: the blame never reaches the first steps, so the network cannot learn long dependencies. It's the vanishing gradient from 02-03, but across time instead of across layers.
- If they are > 1, the gradient explodes and training becomes unstable.
Practical consequence: a SimpleRNN remembers well for about 5-10 steps; beyond that, its memory blurs away. For TecnoMarket's reviews ("The TV I bought... [20 words] ...doesn't work") or for demand with weekly and yearly seasonality, that is not enough.
In 03-03 we saw that ResNet's skip connections created "shortcuts" so the gradient could flow through very deep networks. Sequences need their own version of that idea: a cell with protected memory. That is exactly what the LSTM and GRU of the next lesson do.
Common Mistakes and Tips
- Feeding data with the wrong shape. The most frequent beginner error:
expected ndim=3, found ndim=2. Keras demands(samples, steps, features); if each step is a scalar, add the third axis withX.reshape(n, steps, 1)orX[..., np.newaxis]. - Believing there is a different cell for each time step. No: it is the same cell (same
Wx,Wh,b) applied in a loop. The unrolled view is just a way of drawing it. That's why the parameter count does not depend on sequence length. - Using an RNN when order doesn't matter. If your features are independent (customer age, region, average spend), a dense network is simpler and works better. An RNN only adds value when there is genuine sequential dependency.
- Expecting long memory from a SimpleRNN. If your relevant dependency sits more than ~10 steps away, don't fight the SimpleRNN: it's a structural limitation, not a hyperparameter one. The solution arrives in 04-02.
- Forgetting that
tanhsaturates. If the state fills up with ±1 values, the gradients along that path tend to zero (we saw this in 02-02). It's another face of the same short-memory problem.
Exercises
- By hand: with the example weights (
Wx = [0.5, -0.3]ᵀ,Wh = [[0.8, -0.2], [0.1, 0.6]],b = 0), computeh₁andh₂for the sequencex = [0, 1]. Before computing, reason it out: what willh₁be, and why? - Parameter count: how many parameters does
SimpleRNN(32)have with inputs of 8 features per step? And if the sequence grows from 20 to 200 time steps? - Topology classification: classify these TecnoMarket cases as one-to-many, many-to-one or many-to-many: (a) predicting tomorrow's sales from the last 30 daily sales figures; (b) tagging, word by word, which parts of a review mention shipping; (c) generating an automatic apology reply from the complaint's category.
Solutions
- With
x₁ = 0andh₀ = [0,0]:z₁ = Wx·0 + Wh·[0,0] = [0,0], soh₁ = tanh([0,0]) = [0, 0]— with no input and no prior memory, the state remains empty. Withx₂ = 1:z₂ = Wx·1 + Wh·[0,0] = [0.5, -0.3], henceh₂ = [0.462, -0.291](same as step 1 of the example, because the previous context was null). Wx: 8·32 = 256,Wh: 32·32 = 1024,b: 32. Total: 1312 parameters. With 200 steps instead of 20: exactly the same 1312 — the weights are shared across time; sequence length adds no parameters (though it does add compute cost and more trouble for the gradient).- (a) Many-to-one: 30 values go in, 1 prediction comes out. (b) Synchronized many-to-many: one label per word. (c) One-to-many: one category goes in, a sequence of words comes out (we'll practice this kind of generation in project 07-02).
Conclusion
TecnoMarket's reviews and demand are sequences, and sequences need an architecture that respects order: the RNN achieves it with a single cell applied step by step, dragging along a hidden state that works as memory — weights shared across time, just as CNNs share them across space. You've seen its formula (h_t = tanh(Wx·x_t + Wh·h_{t-1} + b)) working by hand, its topologies and a SimpleRNN in Keras learning a real temporal dependency. But also its Achilles' heel: the gradient traveling backwards through time vanishes, and long-term memory vanishes with it. In the next lesson we'll meet the cells that solved this problem and have powered real-world sequence applications for decades: LSTM and GRU.
Deep Learning Course
Module 1: Introduction to Deep Learning
- What is Deep Learning?
- History and evolution of Deep Learning
- Applications of Deep Learning
- Basic concepts of neural networks
- Setting up the work environment
Module 2: Neural Network Fundamentals
- Perceptron and Multilayer Perceptron
- Activation functions
- Forward and backward propagation
- Optimization and loss functions
- Your first complete neural network
Module 3: Convolutional Neural Networks (CNN)
- Introduction to CNNs
- Convolutional and pooling layers
- Popular CNN architectures
- CNN applications in image recognition
Module 4: Recurrent Neural Networks (RNN)
- Introduction to RNNs
- LSTM and GRU
- RNN applications in natural language processing
- Sequences and time series
Module 5: Advanced Deep Learning Techniques
- Generative Adversarial Networks (GAN)
- Autoencoders
- Transfer Learning
- Regularization and improvement techniques
- Attention mechanisms and Transformers
Module 6: Tools and Frameworks
- Introduction to TensorFlow
- Introduction to PyTorch
- Framework comparison
- Development environments and additional resources
- Saving, loading and deploying models
Module 7: Hands-On Projects
- Image classification with CNNs
- Text generation with RNNs
- Anomaly detection with Autoencoders
- Building a GAN for image generation
- Fine-tuning a pretrained model
