TecnoMarket's second big sequence project works not with words but with dated numbers: the daily sales. Forecasting the next few days' demand decides how much stock to order, how much staff to add in logistics and which promotions to launch — getting it wrong costs money in both directions (stockouts or overflowing warehouses). In this lesson you'll turn demand forecasting into a supervised learning problem using sliding windows, learn to normalize without committing the classic data-leakage mistake, build the indispensable naive baseline, train an LSTM on a synthetic TecnoMarket sales series (with weekly seasonality and Black Friday-style spikes) and evaluate honestly in real units. We'll finish with multi-step forecasting and its risks, and with the criteria for deciding when an LSTM pays off against the classical methods.
Contents
- TecnoMarket's demand as a time series
- Components of a series: trend, seasonality, noise
- Sliding windows: from series to supervised dataset
- Normalization without data leakage
- The naive baseline: the mandatory yardstick
- An LSTM on TecnoMarket's daily sales
- Evaluation in real units
- Multi-step forecasting: iterative vs. direct
- When does an LSTM pay off against the classical methods?
TecnoMarket's demand as a time series
A time series is a sequence of observations taken at regular intervals: units sold each day, orders per hour, visits per week. It differs from the text sequences of 04-03 in three practical respects:
- The values are continuous (not vocabulary indices): there is no Embedding; the end problem is regression (MSE/MAE from 02-04), not classification.
- The "future" genuinely exists: the model will be used to predict values that haven't happened yet, which imposes strict rules on what data it may see during training (and rules out
Bidirectional, as we warned in 04-02). - Time has calendar structure: Mondays look like Mondays, November doesn't look like August.
Components of a series: trend, seasonality, noise
Almost any business series can be thought of as the sum of three components:
| Component | What it is | In TecnoMarket's sales |
|---|---|---|
| Trend | The long-term underlying direction | Sustained growth: the store gains customers every month |
| Seasonality | A pattern repeating with a fixed period | Weekly: weekend peaks; yearly: Black Friday, Christmas, back-to-school |
| Noise | Unexplainable irregular variation | A rainy day, a viral tweet, pure chance |
Let's generate a synthetic series of 2 years of daily sales with these three components plus Black Friday-style events. Working with synthetic data has a teaching advantage: we know the truth the model is supposed to discover.
import numpy as np
rng = np.random.default_rng(7)
DAYS = 730 # 2 years
t = np.arange(DAYS)
trend = 200 + 0.15 * t # from ~200 to ~310 units/day
weekday = t % 7 # 0 = Monday ... 6 = Sunday
seasonal = np.where(weekday >= 5, 60, 0) + 10 * np.sin(2*np.pi*t/7)
noise = rng.normal(0, 15, DAYS) # random variation (σ = 15)
sales = trend + seasonal + noise
# Black Friday-style spikes: days 328 and 693 (late November each year)
for bf in (328, 693):
sales[bf-3:bf+2] *= rng.uniform(2.2, 2.8) # a few days at roughly x2.5
sales = np.round(np.clip(sales, 0, None)) # whole units, never negative
print(sales[:10]) # e.g. [214. 199. 217. 208. 224. 288. 273. 216. 204. 223.]Plot the series (plt.plot(sales)) and you'll see the three layers: the gentle slope, the weekly sawtooth pattern and two spectacular needles on the Black Fridays. The module's question: can an LSTM learn all of this just by looking at past values?
Sliding windows: from series to supervised dataset
A supervised network needs (input X, target y) pairs, but a series is a single strip of numbers. The universal trick is the sliding window: use the last W values as input and the next value as target, then slide the window along the whole series.
Numerical example with the mini-series [10, 12, 15, 14, 18, 21, 19] and window W = 3:
| Window (X) | Target (y) |
|---|---|
| [10, 12, 15] | 14 |
| [12, 15, 14] | 18 |
| [15, 14, 18] | 21 |
| [14, 18, 21] | 19 |
From 7 values you get 7 − 3 = 4 supervised examples. Each row is an exam question: "having seen these 3 days, what happened the next one?". It's a many-to-one task (04-01): many steps go in, one value comes out.
def make_windows(series, window):
"""Turns a 1D series into supervised (X, y) with a sliding window."""
X, y = [], []
for i in range(len(series) - window):
X.append(series[i : i + window]) # W consecutive values
y.append(series[i + window]) # the value immediately after
return np.array(X), np.array(y)
WINDOW = 28 # 4 weeks: enough to capture the weekly pattern several times overChoosing W = 28 is no accident: the window must be longer than the seasonal period you want to capture (here, 7 days) — ideally containing it several times. With W = 3, the model could never "see" that today is Saturday by comparing it with the previous Saturday.
Normalization without data leakage
As with the MNIST pixels (we divided by 255 in 02-05), networks train better with inputs in small ranges; sales between 200 and 800 are worth scaling. But in time series the classic mistake lies in wait: normalizing using statistics of the WHOLE series.
Why is that serious? The maximum of the full series includes the future Black Fridays. If you scale with it, every training example carries an embedded hint about the future ("sales will reach X") that you won't have in production. That is data leakage: the evaluation will come out optimistic and the model will disappoint in the real world. It's the same principle as the train-only adapt() from 04-03.
The golden rule has two parts:
- Split first, and split by time: in time series you NEVER shuffle to split train/test — the test set must be the final stretch (the "future" relative to the train set), because that's how the model will be used.
- Fit the scaler on the train set only and apply it (without refitting) to the test set.
# 1. Temporal split: first 80% for training, last 20% for evaluation
split = int(DAYS * 0.8) # day 584
train_series, test_series = sales[:split], sales[split:]
# 2. Statistics from the train set ONLY
mean, std = train_series.mean(), train_series.std()
train_norm = (train_series - mean) / std # the test set is scaled with the SAME
test_norm = (test_series - mean) / std # mean and std from the train set
# 3. Windows over each stretch, already normalized
X_train, y_train = make_windows(train_norm, WINDOW)
X_test, y_test = make_windows(test_norm, WINDOW)
# 4. Third axis for the LSTM: (samples, steps, features)
X_train = X_train[..., np.newaxis]
X_test = X_test[..., np.newaxis]
print(X_train.shape, X_test.shape) # (556, 28, 1) (118, 28, 1)Keep mean and std: they are part of the model. In production, every prediction gets de-normalized with them (pred * std + mean) to return to real units.
The naive baseline: the mandatory yardstick
Before training anything sophisticated, work out what a trivial rule would achieve. In time series, two canonical baselines:
- Naive: "tomorrow I'll sell the same as today" →
ŷ_t = y_{t-1}. - Seasonal naive: "tomorrow I'll sell the same as a week ago" →
ŷ_t = y_{t-7}. In a series with a strong weekly pattern, it is surprisingly hard to beat.
# Baselines evaluated on the same stretch the LSTM will be evaluated on
actuals = test_series[WINDOW:] # targets in real units
mae_naive = np.abs(actuals - test_series[WINDOW-1:-1]).mean() # yesterday
mae_seasonal = np.abs(actuals - test_series[WINDOW-7:-7]).mean() # 7 days ago
print(f"Naive MAE: {mae_naive:.1f} u. | Seasonal MAE: {mae_seasonal:.1f} u.")
# Typical with this series: naive ≈ 35 u., seasonal ≈ 17 u.This is the most important rule of the lesson: a deep learning model that doesn't beat the seasonal baseline contributes nothing, no matter how many layers it has. The baseline is to time series what the "50% chance" was to IMDB: the honest floor of comparison.
An LSTM on TecnoMarket's daily sales
With the data shaped (samples, 28, 1), the model is a direct application of 04-02:
from tensorflow import keras
from tensorflow.keras import layers
model = keras.Sequential([
layers.Input(shape=(WINDOW, 1)),
layers.LSTM(64, return_sequences=True), # first layer: passes the whole sequence on
layers.LSTM(32), # second layer: summarizes into 32 numbers
layers.Dense(1) # regression: no activation (02-02)
])
model.compile(optimizer="adam", loss="mse", metrics=["mae"])
history = model.fit(X_train, y_train,
epochs=40, batch_size=32,
validation_split=0.15, # the LAST 15% of the train set (Keras doesn't shuffle here)
shuffle=True, # shuffling already-built windows IS valid
verbose=0)A subtle and important nuance: once the windows are built, they can be shuffled among themselves for training (each window is a self-contained example); what you never shuffle is the series before splitting train/test.
Evaluation in real units
The MAE Keras reports is in normalized units — nobody at the logistics meeting understands "0.21". De-normalize:
pred_norm = model.predict(X_test, verbose=0).ravel()
pred = pred_norm * std + mean # back to units/day
actuals = y_test * std + mean
mae_lstm = np.abs(actuals - pred).mean()
print(f"LSTM MAE: {mae_lstm:.1f} units/day")
print(f"vs. naive: {mae_naive:.1f} | vs. seasonal: {mae_seasonal:.1f}")
# Typical result: LSTM ≈ 13-15 u. — beats the seasonal (~17), and the naive (~35) by a mileHow to read it in business terms:
- An MAE of ~14 units on average sales of ~300 is a relative error of ~5%: enough to size stock with a reasonable margin.
- The improvement over the seasonal baseline (~17 → ~14) looks modest, but that's where the value lives: the weekly pattern came for free with the baseline; the LSTM adds the trend and the interactions. Plot
predagainstactuals: you'll see it tracks the weekly sawtooth and the slope well. - What about the Black Fridays? If they fall in the test stretch, the LSTM will smooth them out or arrive late: a yearly spike appears 1-2 times in the whole training set, not enough to learn it from the series alone. In practice this is solved by adding calendar features as extra inputs (day of the week, "it's a Black Friday campaign": the third axis of
Xgoes from 1 to several features per step) — the(samples, steps, features)format from 04-01 was ready for this from the start.
Multi-step forecasting: iterative vs. direct
Logistics doesn't just want tomorrow: they want the next 7 days. Two strategies:
| Strategy | How it works | Advantages | Risks |
|---|---|---|---|
| Iterative | Predict t+1, feed that prediction into the window and predict t+2, and so on 7 times | A single model; flexible horizon | Error accumulation: each step rests on predictions, not data; the error grows with the horizon |
| Direct | One model trained to predict the vector of 7 values at once (Dense(7)), or one model per horizon |
Each horizon is optimized on real data; no error feedback | More training; doesn't reuse the t+1 prediction for t+2 |
def predict_iterative(model, last_window_norm, steps):
"""Iterative multi-step forecasting: feeds each prediction back in."""
window = last_window_norm.copy() # (WINDOW,)
future = []
for _ in range(steps):
p = model.predict(window[np.newaxis, :, np.newaxis], verbose=0)[0, 0]
future.append(p)
window = np.append(window[1:], p) # slide: the prediction moves in
return np.array(future) * std + mean # back to real units
print(predict_iterative(model, test_norm[-WINDOW:], steps=7))The risk of the iterative strategy is best understood through an image: it's a game of telephone — each prediction inherits and amplifies the errors of the previous one. For 2-3 steps it usually holds up; for 14-30 days, the direct strategy (or classical models with explicit seasonality) is usually safer. Always evaluate the MAE per horizon (how far off am I at t+1? and at t+7?): you'll see the degradation curve with your own eyes.
When does an LSTM pay off against the classical methods?
Deep learning is not always the answer in time series. Classical statistical methods (moving averages, exponential smoothing, ARIMA/SARIMA) have been forecasting demand for decades:
| Situation | Best first bet |
|---|---|
| A single short series (< 1-2 years), regular pattern | Classical (SARIMA, exponential smoothing): simpler, interpretable and often just as accurate |
| Hundreds/thousands of related series (one per TecnoMarket product) | LSTM or other DL models: a single global model learns patterns shared across products |
| Many external variables (price, promotions, calendar, weather) | DL: it incorporates per-step features naturally |
| Complex non-linear relationships (promotion × day × category) | DL |
| You need to explain the forecast to the business | Classical (or at the very least, accompany the DL with a clear baseline) |
| Very scarce or very noisy data | Classical or the seasonal baseline; DL will overfit |
The methodological honesty of this lesson — baseline first, temporal split, evaluation in real units — holds in both worlds, and it's what separates a serious prototype from a demo.
Common Mistakes and Tips
- Fitting the scaler on the whole series. The classic data-leakage mistake: the test set gets contaminated with statistics from the future and the metrics inflate.
fitthe scaler on the train set only, always. - Splitting train/test by shuffling. In time series, the test set must be the final temporal stretch. If you shuffle before splitting, the model "trains on the future" and the evaluation is fiction. (Shuffling the already-built training windows, on the other hand, is fine.)
- Not computing the baseline. Without a baseline you don't know whether your MAE of 14 is good or embarrassing. The seasonal naive takes one line to compute and is the first number you should write down.
- A window shorter than the seasonality. With
W < 7on a weekly series, the model can't compare "today" with "a week ago". Window ≥ 2-4 seasonal periods. - Reporting the error in normalized units. "MAE 0.21" informs no decision. De-normalize and speak in units/day (or revenue): that's what the business can judge.
- Trusting iterative multi-step forecasts at long horizons. The error accumulates step by step. Measure the MAE per horizon and cut off where it stops being useful.
Exercises
-
Windows by hand: with the series
[100, 110, 105, 120, 125, 118, 130]andW = 4, write out all the resulting (X, y) pairs. How many supervised examples do you get from a series ofNvalues with windowW? -
Spot the leak: this code contains two methodology errors — find them and fix them:
mean, std = sales.mean(), sales.std() norm = (sales - mean) / std X, y = make_windows(norm, 28) X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, shuffle=True) -
Baseline and verdict: on the test stretch, the seasonal baseline achieves an MAE of 17.2 units and your LSTM 16.8. Training the LSTM takes 20 minutes and the baseline is one line of code. What would you recommend to the TecnoMarket team, and what would you try before discarding the LSTM?
Solutions
- Pairs:
([100,110,105,120] → 125),([110,105,120,125] → 118),([105,120,125,118] → 130). In general,N − Wexamples: here7 − 4 = 3. - Error 1: the scaler is fitted on the whole series (
sales.mean()) — it must be computed on the training stretch only. Error 2:train_test_splitwithshuffle=Trueshuffles before splitting, mixing future into the train set — the split must be temporal (first 80% / last 20%) and happen before building the windows (if it happens afterwards there are, in addition, windows straddling the train/test boundary). Correct version: split the series by temporal index, normalize with train statistics, and build the windows inside each stretch. - With an improvement of 0.4 units (~2%), the LSTM barely justifies its cost: I'd recommend deploying the seasonal baseline today as the reference system. Before discarding the LSTM I would try: adding calendar features (day of the week, holidays, campaigns — where DL can make a real difference), widening the window, and training one global model over many products at once instead of a single series. If after that the improvement is still marginal, the simple method wins: that's the honest conclusion the baseline makes possible.
Conclusion
You've turned TecnoMarket's demand forecasting into a deep learning problem end to end: decomposing the series (trend + seasonality + noise), transforming it into a supervised dataset with sliding windows, normalizing without data leakage (scaler fitted on train only, temporal split with no shuffling), demanding a naive baseline before celebrating anything, training a stacked LSTM that beats that baseline, and evaluating in units the business understands — plus multi-step forecasting with its error accumulation and the criteria for choosing between an LSTM and classical methods. This closes the recurrent networks module: you now know how to model order, memory and time, in text and in numbers alike. The next module changes the question: instead of classifying or predicting, can a network create — generate promotional images, compress representations, leverage pretrained models? We start with generative adversarial networks (GANs).
Deep Learning Course
Module 1: Introduction to Deep Learning
- What is Deep Learning?
- History and evolution of Deep Learning
- Applications of Deep Learning
- Basic concepts of neural networks
- Setting up the work environment
Module 2: Neural Network Fundamentals
- Perceptron and Multilayer Perceptron
- Activation functions
- Forward and backward propagation
- Optimization and loss functions
- Your first complete neural network
Module 3: Convolutional Neural Networks (CNN)
- Introduction to CNNs
- Convolutional and pooling layers
- Popular CNN architectures
- CNN applications in image recognition
Module 4: Recurrent Neural Networks (RNN)
- Introduction to RNNs
- LSTM and GRU
- RNN applications in natural language processing
- Sequences and time series
Module 5: Advanced Deep Learning Techniques
- Generative Adversarial Networks (GAN)
- Autoencoders
- Transfer Learning
- Regularization and improvement techniques
- Attention mechanisms and Transformers
Module 6: Tools and Frameworks
- Introduction to TensorFlow
- Introduction to PyTorch
- Framework comparison
- Development environments and additional resources
- Saving, loading and deploying models
Module 7: Hands-On Projects
- Image classification with CNNs
- Text generation with RNNs
- Anomaly detection with Autoencoders
- Building a GAN for image generation
- Fine-tuning a pretrained model
