In the previous lesson we discovered that the SimpleRNN has the memory of a goldfish: the gradient traveling backwards through time vanishes and, with it, the ability to learn dependencies more than a few steps apart. This lesson presents the two cells that solved that problem and that underpin most real-world sequence applications: the LSTM (Long Short-Term Memory, 1997) and its modern simplification, the GRU (Gated Recurrent Unit, 2014). You'll understand the LSTM piece by piece — its cell state and its three gates — see its parallel with ResNet's skip connections, count its parameters, and verify experimentally in Keras how an LSTM solves a long-memory task where the SimpleRNN fails. We'll close with two architecture tools you'll use constantly: stacked layers (return_sequences) and bidirectional layers.

Contents

  1. The long-term memory problem
  2. The LSTM cell: the cell state as a conveyor belt
  3. The three gates, one by one
  4. The gradient highway: the ResNet parallel
  5. GRU: the simplification that is almost always enough
  6. SimpleRNN / LSTM / GRU comparison
  7. Counting the parameters of an LSTM
  8. Keras experiment: SimpleRNN vs. LSTM on long memory
  9. Stacking recurrent layers and Bidirectional

The long-term memory problem

Read this review, typical of the kind TecnoMarket receives:

"I bought this refrigerator three months ago. Checkout was smooth, the courier was punctual, the packaging arrived flawless and installation was straightforward... but since last week it doesn't cool."

To classify it correctly, the model must connect "refrigerator" (word 3) with "doesn't cool" (words 30+), leaping over a long list of positive remarks. A SimpleRNN rewrites its entire hidden state at every step (h_t = tanh(...)): each new word tramples the previous ones, and after 25 steps of logistics compliments, the "refrigerator" from the beginning has washed out. And during training the mirror image happens: the gradient of the error ("you said positive and it was negative") vanishes before reaching the first steps, so the network never even learns that it should have remembered.

The LSTM's solution is not "remember harder", but something more elegant: separate the memory from the processing and put access to that memory under the control of learned gates.

The LSTM cell: the cell state as a conveyor belt

An LSTM maintains two state vectors instead of one:

  • h_t — the hidden state, as in the SimpleRNN: each step's "working output".
  • C_t — the cell state: the long-term memory.

The classic image for C_t is a conveyor belt running the whole length of the sequence: information rides it from one step to the next almost untouched — only two gentle operations (a multiplication and an addition) modify it. Nothing rewrites it wholesale. Alongside the belt sit three gates: small mechanisms deciding what gets taken off the belt, what gets placed onto it, and what gets looked up.

A gate is simply a mini-layer with a sigmoid activation:

gate = σ(W · [h_{t-1}, x_t] + b)

The sigmoid (02-02) produces values between 0 and 1 for each component of the memory: 0 means "block the flow" and 1 means "let everything through". It plays the same role as a faucet — the course's shower-faucet analogy —: the network learns how far to open each faucet depending on the context (h_{t-1} and x_t), not with a fixed value.

graph LR
    subgraph "LSTM cell at step t"
        Cprev["C_{t-1}<br>(belt: memory)"] --> MUL1(("×"))
        F["Forget gate f_t<br>σ: what do I erase?"] --> MUL1
        MUL1 --> ADD(("+"))
        I["Input gate i_t<br>σ: what do I write?"] --> MUL2(("×"))
        CAND["Candidate C̃_t<br>tanh: new content"] --> MUL2
        MUL2 --> ADD
        ADD --> Cnew["C_t<br>(updated belt)"]
        Cnew --> TANH["tanh"]
        TANH --> MUL3(("×"))
        O["Output gate o_t<br>σ: what do I reveal?"] --> MUL3
        MUL3 --> Hnew["h_t<br>(hidden state)"]
    end
    X["x_t and h_{t-1}"] -.feed.-> F
    X -.-> I
    X -.-> CAND
    X -.-> O

The three gates, one by one

  1. Forget gate: what do I erase from the belt?

f_t = σ(Wf · [h_{t-1}, x_t] + bf)

It looks at the current input and the context and decides which components of C_{t-1} to keep (values near 1) and which to empty (near 0). It's applied by multiplying: f_t ⊙ C_{t-1} (⊙ = component-wise).

TecnoMarket intuition: if the review says "On the other hand, the second product...", it pays to forget the details of the first product to make room.

  1. Input gate: what do I write onto the belt?

It works in tandem with a candidate for new memory:

i_t  = σ(Wi · [h_{t-1}, x_t] + bi)       # how much to write (0-1)
C̃_t = tanh(Wc · [h_{t-1}, x_t] + bc)    # what content to write (-1 to 1)

The candidate C̃_t proposes new information (it is, in fact, the SimpleRNN formula); the gate i_t decides which parts of that proposal deserve to enter the memory. The complete belt update is:

C_t = f_t ⊙ C_{t-1}  +  i_t ⊙ C̃_t
      (what I keep)      (what I add)

Intuition: on reading "refrigerator", the input gate writes "the product is a refrigerator" onto the belt; during the compliments to the courier, i_t can stay nearly closed and f_t nearly open: the belt passes through intact.

  1. Output gate: what do I look up from the belt?

o_t = σ(Wo · [h_{t-1}, x_t] + bo)
h_t = o_t ⊙ tanh(C_t)

The hidden state h_t (what the cell "shows" to the outside world and to the next step) is a filtered view of the memory: the output gate decides which parts of the belt are relevant right now. Not everything memorized is needed at every step: you can remember that the product is a refrigerator without mentioning it until the "doesn't cool" arrives.

Flow summary: I forget what's obsolete (f_t), I incorporate the relevant new material (i_t ⊙ C̃_t), and I expose the useful part of my memory (o_t). All the gate weights are learned by backpropagation, exactly like the rest of the network: nobody hand-programs what to forget.

The gradient highway: the ResNet parallel

Look at the belt equation again: C_t = f_t ⊙ C_{t-1} + i_t ⊙ C̃_t. The previous memory reaches the present through an addition, not by passing through a tanh and a multiplication by Wh as in the SimpleRNN. Sound familiar? It's the same move as ResNet's skip connections (03-03): there, output = F(x) + x created a shortcut so the gradient could cross dozens of layers in space; here, the belt creates a shortcut so the gradient can cross dozens of steps in time. As long as the forget gate stays reasonably open (f_t ≈ 1), the gradient flows backwards along the belt almost unattenuated: a temporal gradient highway.

It's a recurring idea in deep learning worth pinning down: when the gradient dies along the way (vanishing, 02-03), the fix is usually to give it a direct additive path. ResNet did it in depth, the LSTM in time, and the residual connections in transformers (05-05) repeat the pattern.

GRU: the simplification that is almost always enough

The GRU asks: do we really need two states and three gates? Its answer: merge C_t and h_t into a single state h_t, and use two gates:

  • Update gate (z_t): it merges the forgetting and writing roles into one. The update is an interpolation:

    h_t = (1 - z_t) ⊙ h_{t-1} + z_t ⊙ h̃_t
    

    If z_t ≈ 0, the state passes through intact (memory preserved); if z_t ≈ 1, it is replaced by the new candidate. Note that it is still an additive path: the gradient highway is preserved.

  • Reset gate (r_t): when computing the candidate h̃_t = tanh(W · [r_t ⊙ h_{t-1}, x_t]), it decides how much past context to use when proposing new content. With r_t ≈ 0, the candidate ignores the past and "starts fresh" locally.

Fewer gates means ~25% fewer parameters and somewhat faster training, with performance usually comparable to the LSTM's. In practice: try GRU when the dataset is small or compute time matters; try LSTM when the task demands very fine-grained memory or data is plentiful. There is no universal winner — it's an empirical decision.

SimpleRNN / LSTM / GRU comparison

Aspect SimpleRNN LSTM GRU
States 1 (h) 2 (h and C) 1 (h)
Gates 0 3 (forget, input, output) 2 (update, reset)
Parameters (input d, units u) u(d+u+1) 4·u(d+u+1) 3·u(d+u+1)
Effective memory ~5-10 steps Hundreds of steps Hundreds of steps
Cost per step Low High (×4) Medium (×3)
Vanishing across time Severe Mitigated (additive belt) Mitigated (interpolation)
When to use it Very short sequences, teaching demos Long, complex dependencies, large datasets Default choice: good quality/cost balance

Counting the parameters of an LSTM

An LSTM computes 4 blocks with the same structure (the 3 gates + the candidate), each with its input matrix, its recurrent matrix and its bias. With input dimension d and u units:

parameters = 4 × (d·u + u·u + u) = 4·u·(d + u + 1)

Example: LSTM(64) with inputs of 32 features per step:

4 × 64 × (32 + 64 + 1) = 4 × 64 × 97 = 24,832 parameters

Exactly 4 times what an equivalent SimpleRNN(64) would have (6,208), and as always with recurrent layers, independent of sequence length. Verify it with model.summary(), as we did with CNNs in 03-02.

Keras experiment: SimpleRNN vs. LSTM on long memory

Let's design a task that demands long memory: in a sequence of 50 random numbers, the goal is to recover the sum of the first 2 values. All the useful information sits at the beginning; the remaining 48 steps are noise the network must cross without forgetting.

import numpy as np
from tensorflow import keras
from tensorflow.keras import layers

# Synthetic long-memory dataset
rng = np.random.default_rng(0)
STEPS = 50
X = rng.uniform(0, 1, size=(5000, STEPS, 1))
y = X[:, :2, 0].sum(axis=1)          # target: sum of the FIRST TWO steps

def train(recurrent_layer, name):
    model = keras.Sequential([
        layers.Input(shape=(STEPS, 1)),
        recurrent_layer,
        layers.Dense(1)
    ])
    model.compile(optimizer="adam", loss="mse", metrics=["mae"])
    h = model.fit(X, y, epochs=25, batch_size=64,
                  validation_split=0.2, verbose=0)
    print(f"{name:>10}: validation MAE = {h.history['val_mae'][-1]:.3f}")

train(layers.SimpleRNN(32), "SimpleRNN")
train(layers.LSTM(32),      "LSTM")

Typical results (they will vary slightly on your machine):

 SimpleRNN: validation MAE = 0.39   # ≈ always predicting the mean: it has NOT learned
      LSTM: validation MAE = 0.05   # memory intact after 48 steps of noise

Interpretation: the sum of two uniforms in [0,1] has mean 1.0, and a "dumb" predictor that always said 1.0 would score an MAE of ~0.4 — exactly where the SimpleRNN lands. The gradient never managed to reach the first steps, so the network gave up and predicts the mean. The LSTM, thanks to the belt, writes the first two values into C_t, keeps the forget gate open for 48 steps, and retrieves them at the end. This miniature experiment is the reason almost nobody uses SimpleRNN in production.

Stacking recurrent layers and Bidirectional

Stacked layers: return_sequences=True

Just as we stacked Conv2D layers to go from edges to textures and objects (03-01), we can stack recurrent layers so the second one processes more abstract representations of the sequence. The technical detail: a recurrent layer by default returns only its last state, but the next layer needs a full sequence as input. The fix is return_sequences=True on every layer except the last:

model = keras.Sequential([
    layers.Input(shape=(50, 1)),
    layers.LSTM(64, return_sequences=True),   # returns all 50 states: (50, 64)
    layers.LSTM(32),                          # returns only the last one: (32,)
    layers.Dense(1)
])

Mnemonic: return_sequences=True = "hand the whole movie to the next layer"; False (the default) = "keep only the final frame".

Bidirectional: reading in both directions

In many tasks, future context helps too: in "my order hasn't arrived yet", the "yet" softens the complaint, but it comes last. Bidirectional duplicates the layer: one copy reads the sequence left to right and another right to left, and their outputs are concatenated (doubling the output dimension and the parameters):

model = keras.Sequential([
    layers.Input(shape=(50, 8)),
    layers.Bidirectional(layers.LSTM(32)),    # output: 64 numbers (32 + 32)
    layers.Dense(1, activation="sigmoid")
])

Use it when you have the complete sequence before deciding (classifying an already-written review: yes). Avoid it when you predict the future in real time (tomorrow's demand: the "reverse direction" would read data that doesn't exist yet). We'll come back to this distinction in 04-04.

Common Mistakes and Tips

  • Stacking recurrent layers without return_sequences=True. The expected ndim=3, found ndim=2 error on the second LSTM is almost always this: the first layer returned only the last state. Enable it on every recurrent layer except the last (if the problem is many-to-one).
  • Confusing h_t with C_t. The hidden state is the working view (what comes out of the cell); the cell state is the LSTM's internal memory. Keras manages both automatically: you only see h_t.
  • Picking LSTM "because it's the good one" without comparing. With short sequences or little data, a GRU (or even a dense network over aggregate statistics) can perform just as well at lower cost. Always compare against the simple option, as we will with the baseline in 04-04.
  • Using Bidirectional for future prediction. It's a conceptual information leak: in production you won't have the "future" to read backwards. Reserve it for complete, already-available sequences.
  • Thinking the LSTM eliminates vanishing entirely. It mitigates it enormously, but with sequences of thousands of steps it suffers too. For those extremes, attention mechanisms emerged — we'll mention them in 04-03 and develop them in 05-05.

Exercises

  1. Gates in action: for the review "The refrigerator arrived fast and well packaged, but it doesn't cool", describe what you would expect the forget and input gates to do (open/closed, what they write or erase) when processing "refrigerator", "fast" and "but".
  2. Parameter count: compute the parameters of (a) LSTM(128) with an input of 64 features per step; (b) the equivalent GRU using the formula 3·u·(d+u+1); (c) by what percentage does the GRU reduce the parameters?
  3. Architecture diagnosis: this model throws an error when built — explain why and fix it: Sequential([Input(shape=(30, 4)), LSTM(64), LSTM(32), Dense(1)]).

Solutions

  1. On "refrigerator": input gate wide open (writes the review's subject onto the belt), forget gate open (there's nothing prior to erase). On "fast": input moderate (records a positive signal about shipping), forget almost fully open (keeps "refrigerator"). On "but": it's a discourse pivot — the forget gate may partially close over the positive-sentiment components (what follows probably contradicts them) while preserving the subject, and the input gate gets ready to write the new assessment. (This is an idealized reading: in a real LSTM the components aren't that legible, but the mechanism is this one.)
  2. (a) 4·128·(64+128+1) = 4·128·193 = 98,816. (b) 3·128·193 = 74,112. (c) 1 − 74,112/98,816 = 25% reduction, the 3/4 ratio you'd expect from having one gate fewer.
  3. The first LSTM(64) by default returns only the last state (a 2D tensor (batch, 64)), but the second LSTM needs a 3D sequence. Fix: LSTM(64, return_sequences=True) on the first layer.

Conclusion

The LSTM cures the SimpleRNN's short memory by separating the memory (the cell state, a conveyor belt crossing the sequence along an additive path) from the processing, and governing access with three learned gates: forget, input and output. That additive path is the same medicine against vanishing as ResNet's skip connections, applied to time instead of depth. The GRU compresses the idea into two gates and a single state, with 25% fewer parameters and almost always comparable results. You've verified it empirically: faced with a 48-step dependency, the SimpleRNN gives up and the LSTM solves it. With the composition tools (return_sequences, Bidirectional) you now have the complete recurrent arsenal. In the next lesson we'll put it to work on the first of the module's two flagship projects: teaching a network to read TecnoMarket's customer reviews — which first demands solving a new problem: turning text into numbers.

© Copyright 2026. All rights reserved