Closing the previous module, a warning was left hanging in the air: nearly every tool we have seen so far — intervals, tests, regression, ANOVA — rests on a silent assumption, namely that the observations are independent of one another. But when the data are the sales of each week, one after another, that assumption blows up: Christmas week looks like last year's Christmas, one Monday looks like another Monday, and a good week tends to be followed by another good one. In this lesson you will learn to work with data ordered in time: to decompose them into trend, seasonality and noise, to smooth them with moving averages, to build simple forecasts, and to measure how wrong your forecasts are. The case that will guide the whole lesson is a real one at NovaMarket: Marta needs an e-commerce sales forecast to size the Christmas campaign — how much stock to order, how many delivery drivers to add, how much advertising budget to allocate.

Contents

  1. What a time series is and why it breaks independence
  2. The components of a series: trend, seasonality, cycle and irregular
  3. Additive vs multiplicative decomposition
  4. Moving averages: smoothing to see the trend
  5. Seasonal indices with NovaMarket's online sales
  6. Simple exponential smoothing
  7. Forecasting: from the naïve method to trend + seasonality
  8. Evaluating the forecast: MAE and MAPE
  9. Autocorrelation: the series' memory

What a time series is and why it breaks independence

A time series is a sequence of observations of the same variable taken at consecutive — and usually equally spaced — points in time: daily sales, weekly complaints, quarterly revenue. Back in the lesson on types of data we distinguished cross-sectional data (many units at a single moment: the 42 stores today) from longitudinal data (the same unit over time: NovaMarket's online sales quarter by quarter). This lesson is the machinery we promised for the latter.

Why can't the tools of Modules 5 and 6 simply be applied as they stand? Because in a time series:

  • Observations resemble their neighbors. If this week's online sales were high, next week's most likely will be too. That is dependence, and t or z tests assume exactly the opposite.
  • Order matters. In a sample of receipts you could shuffle the data and lose nothing; in a series, shuffling destroys the main information (the evolution).
  • The mean may not exist as a useful concept. What does "average quarterly sales" mean for a series growing 10% a year? It is a number that describes neither the distant past nor the present.

The practical consequence: instead of estimating one fixed parameter, the goal becomes understanding the temporal structure (which part is trend, which part is seasonal, which part is noise) and using it to forecast.

The components of a series

The classic way to think about a series is to decompose it into four components:

Component Symbol What it is Example at NovaMarket
Trend \(T_t\) Long-run underlying movement, upward or downward The e-commerce business has grown steadily since its launch
Seasonality \(S_t\) Pattern that repeats with a fixed, known period (week, month, quarter, year) The fourth quarter is always above the rest (Christmas); Mondays bring a complaints spike (32 of the 100 weekly complaints)
Cycle \(C_t\) Long oscillations with no fixed period, tied to the economy In recession years the basket gets cheaper and the average transaction value falls
Irregular \(I_t\) What is left over: unpredictable noise A heat wave that sends drinks soaring for one particular week

Two important nuances:

  • Seasonality ≠ cycle. Seasonality has a fixed period (every fourth quarter, every Monday); the cycle does not (an economic expansion can last 4 years or 8). With short series, like the ones we will handle, the cycle can barely be separated from the trend and is usually treated together with it (the trend-cycle component).
  • Seasonality is not "bad". It is enormously valuable information: knowing that Q4 sells 29% more than a normal quarter (we will compute this shortly) is exactly what Marta needs to size the campaign.

Additive vs multiplicative decomposition

How do the components combine? There are two classic models:

\[ \text{Additive:} \quad y_t = T_t + S_t + I_t \]

\[ \text{Multiplicative:} \quad y_t = T_t \times S_t \times I_t \]

Choice criterion (always look at the plot of the series):

  • If the seasonal swings have constant amplitude in euros even as the series grows (every Christmas adds a fixed "+1 million"), the model is additive: seasonality is measured in the units of the series.
  • If the swings grow with the level of the series (every Christmas the series rises "by 29%", which is more euros the bigger the business gets), the model is multiplicative: seasonality is measured as a percentage or index.

In business sales the second is the usual case — the Christmas effect scales with the size of the business — so we will work the NovaMarket case in its multiplicative version. The computational logic is identical in both: estimate the trend, compare the series against it (subtracting in the additive model, dividing in the multiplicative one), and whatever repeats periodically is the seasonality.

Moving averages: smoothing to see the trend

The simplest tool for estimating the trend is the moving average: replace each data point with the average of a window of \(k\) observations centered on it. Averaging cancels out the noise (and the seasonality, if the window spans a full period), leaving the underlying signal.

\[ MA_k(t) = \frac{y_{t-\frac{k-1}{2}} + \dots + y_t + \dots + y_{t+\frac{k-1}{2}}}{k} \]

A worked example by hand: weekly e-commerce sales

NovaMarket's online sales over the last 7 weeks (€ thousands):

Week 1 2 3 4 5 6 7
Sales 310 295 330 315 350 340 365

The series is climbing, but in fits and starts: was week 2 a "bad" week or just noise? Let's compute the moving average of order 3 (the average of each week with the one before and the one after):

  • \(MA_3(2) = \dfrac{310 + 295 + 330}{3} = \dfrac{935}{3} = 311.7\)
  • \(MA_3(3) = \dfrac{295 + 330 + 315}{3} = \dfrac{940}{3} = 313.3\)
  • \(MA_3(4) = \dfrac{330 + 315 + 350}{3} = \dfrac{995}{3} = 331.7\)
  • \(MA_3(5) = \dfrac{315 + 350 + 340}{3} = \dfrac{1005}{3} = 335.0\)
  • \(MA_3(6) = \dfrac{350 + 340 + 365}{3} = \dfrac{1055}{3} = 351.7\)

The smoothed series — 311.7, 313.3, 331.7, 335.0, 351.7 — tells a clean story: steady growth of about €10,000 per week, without the sawtooth of the raw data. Practical observations:

  • The endpoints are lost: there is no moving average for weeks 1 and 7 (no neighbors on one side). The larger \(k\) is, the more is lost.
  • The larger \(k\), the smoother the result, but the slower it reacts to genuine changes. It is the same bias–responsiveness dilemma that will reappear in exponential smoothing.
  • If there is seasonality, the window must cover a full period: for quarterly data, order 4; for daily data with a weekly pattern, order 7. With an even order (4, 12) the moving average lands "between" two periods and must be centered by averaging two consecutive moving averages — you will see this in the next example.

Seasonal indices with NovaMarket's online sales

On to Marta's case. The e-commerce quarterly sales for the last two years (€ millions):

Quarter Q1 2024 Q2 2024 Q3 2024 Q4 2024 Q1 2025 Q2 2025 Q3 2025 Q4 2025
Sales 4.2 4.0 3.8 6.0 4.6 4.4 4.2 6.6

You can see it with the naked eye: every Q4 spikes (Christmas) and every Q3 sags (summer), on top of a rising trend. Let's quantify it.

Step 1. Moving average of order 4 (a full year, so the seasonality cancels out):

  • \(t=1\text{–}4:; (4.2+4.0+3.8+6.0)/4 = 4.50\)
  • \(t=2\text{–}5:; (4.0+3.8+6.0+4.6)/4 = 4.60\)
  • \(t=3\text{–}6:; (3.8+6.0+4.6+4.4)/4 = 4.70\)
  • \(t=4\text{–}7:; (6.0+4.6+4.4+4.2)/4 = 4.80\)
  • \(t=5\text{–}8:; (4.6+4.4+4.2+6.6)/4 = 4.95\)

Step 2. Centering (each order-4 average falls "between" two quarters; they are averaged in consecutive pairs to line them up with a specific quarter):

Quarter Sales \(y_t\) Centered moving average \(T_t\) Ratio \(y_t / T_t\)
Q3 2024 3.8 \((4.50+4.60)/2 = 4.55\) \(3.8/4.55 = 0.835\)
Q4 2024 6.0 \((4.60+4.70)/2 = 4.65\) \(6.0/4.65 = 1.290\)
Q1 2025 4.6 \((4.70+4.80)/2 = 4.75\) \(4.6/4.75 = 0.968\)
Q2 2025 4.4 \((4.80+4.95)/2 = 4.875\) \(4.4/4.875 = 0.903\)

Step 3. Seasonal indices. With more years, the ratios for each quarter would be averaged; here we have one per quarter and take them directly (checking that they sum to ≈ 4, i.e., that they average 1 — here they sum to 3.996, a negligible adjustment):

Quarter Seasonal index Reading
Q1 0.968 3.2% below a "normal" quarter
Q2 0.903 9.7% below
Q3 0.835 16.5% below (the summer effect)
Q4 1.290 29% above (the Christmas effect)

That 1.29 is gold for Marta: whatever the level of the business, the fourth quarter sells around 29% more than the trend. The deseasonalized series (\(y_t / S_t\)) answers another common question: was Q4 good beyond the fact that it was Christmas? For instance, Q4 2025 deseasonalized: \(6.6 / 1.29 = 5.12\), in line with the trend — it was a good Christmas, but not an exceptional one.

Simple exponential smoothing

The moving average gives every data point in the window the same weight and ignores everything outside it. Simple exponential smoothing does something more elegant: it averages the entire history, but with weights that decay geometrically — the recent past weighs heavily, the distant past almost nothing. Its recursive formula could hardly be simpler:

\[ S_t = \alpha , y_t + (1-\alpha) , S_{t-1} \]

where \(S_t\) is the smoothed value at \(t\) and \(\alpha \in (0,1)\) is the smoothing constant: with a high \(\alpha\) (0.7) the smoothing reacts quickly but smooths little; with a low \(\alpha\) (0.1) it smooths heavily but is slow to catch on to changes. Values between 0.1 and 0.3 are a common starting point.

Two steps by hand with the weekly sales from before, taking \(\alpha = 0.3\) and starting from \(S_1 = y_1 = 310\):

  • \(S_2 = 0.3 \times 295 + 0.7 \times 310 = 88.5 + 217.0 = 305.5\)
  • \(S_3 = 0.3 \times 330 + 0.7 \times 305.5 = 99.0 + 213.9 = 312.9\)
  • \(S_4 = 0.3 \times 315 + 0.7 \times 312.9 = 94.5 + 219.0 = 313.5\)

As a forecast, simple smoothing proposes \(\hat{y}_{t+1} = S_t\): "tomorrow, like today's smoothed value". It is suited to series with no clear trend or seasonality (for example, the weekly demand for a mature Pantry product). For series with trend and seasonality there are extensions (Holt's double smoothing, Holt-Winters triple smoothing) that we will not develop: take away the idea that they add analogous equations for the slope and for the seasonal indices.

Forecasting: from the naïve method to trend + seasonality

The naïve method: the bar to clear

The naïve forecast is: "the next value will equal the last one observed"; in its seasonal version: "next Q4 will equal the last Q4". It sounds like a joke, but it is the honest benchmark: any sophisticated method that fails to beat the naïve one is adding nothing.

  • Seasonal naïve for Q4 2026: \(\hat{y} = 6.6\) million (what Q4 2025 sold).

Forecasting with trend + seasonality

A better method combines what we have already computed:

  1. Estimate the trend from the centered moving averages: they went from 4.55 to 4.875 in 3 quarters, i.e., they grow about \((4.875-4.55)/3 \approx 0.11\) million per quarter.
  2. Extrapolate the trend to the quarter being forecast. The last centered average (4.875) corresponds to Q2 2025; there are 6 quarters to Q4 2026: \(T \approx 4.875 + 6 \times 0.11 = 5.53\).
  3. Apply the seasonal index: \(\hat{y}_{\text{Q4 2026}} = 5.53 \times 1.29 \approx 7.1\) million €.

The forecast for Marta: about €7.1 million in Q4 2026, against the naïve method's 6.6. Half a million of difference that translates directly into stock, delivery drivers and campaign budget. It should come with a range attached — the confidence intervals lesson taught us why a bare number is a reckless promise — and for that we need to measure the method's typical forecasting error, which is exactly the next section.

Evaluating the forecast: MAE and MAPE

How do you know whether a method forecasts well? You forecast a period you already know (held back on purpose) and compare forecasts with reality. The two most widely used metrics:

\[ MAE = \frac{1}{n}\sum_{t=1}^{n} |y_t - \hat{y}t| \qquad \qquad MAPE = \frac{100}{n}\sum{t=1}^{n} \frac{|y_t - \hat{y}_t|}{y_t} \]

  • MAE (mean absolute error): "on average, I am off by this many euros". In the units of the series, easy to communicate.
  • MAPE (mean absolute percentage error): "on average, I am off by this percentage". Comparable across series of different scales; it breaks down when actual values are close to zero.

Let's compare both methods by forecasting 2025 using only the information available at the end of 2024:

Quarter Actual Seasonal naïve \(\lvert e \rvert\) naïve Trend + seasonality \(\lvert e \rvert\) T+S
Q1 2025 4.6 4.2 0.4 4.50 0.10
Q2 2025 4.4 4.0 0.4 4.35 0.05
Q3 2025 4.2 3.8 0.4 4.10 0.10
Q4 2025 6.6 6.0 0.6 6.50 0.10
  • Naïve: \(MAE = (0.4+0.4+0.4+0.6)/4 = 0.45\) million; \(MAPE = (8.7% + 9.1% + 9.5% + 9.1%)/4 \approx 9.1%\).
  • Trend + seasonality: \(MAE = (0.10+0.05+0.10+0.10)/4 \approx 0.09\) million; \(MAPE \approx 1.8%\).

The structured method cuts the error from 9% to under 2%: it does clear the naïve bar, and comfortably. That 1.8% MAPE is, moreover, the honest range to attach to the forecast: €7.1 million ± roughly 2%, provided the pattern holds.

Autocorrelation: the series' memory

Autocorrelation is the correlation of the series with itself shifted by \(k\) periods (the shift is called the lag): the coefficient \(r_k\) measures, on the same −1 to +1 scale as the Pearson coefficient you already know, how much \(y_t\) resembles \(y_{t-k}\).

Reading it is an express diagnostic of the temporal structure:

Observed pattern What it indicates
\(r_1\) high and positive, decaying slowly across \(r_2, r_3, \dots\) Trend: each value drags the next along
Autocorrelation spike at the lag of the period (\(r_4\) for quarters, \(r_7\) for daily data) Seasonality
All \(r_k\) small, with no pattern A series without memory: the past does not help predict

Two examples at NovaMarket:

  • In the daily complaints, \(r_7 = 0.81\): every Monday resembles previous Mondays — the numerical signature of the Monday spike (32 of the 100 weekly complaints) we uncovered with the chi-square goodness-of-fit test.
  • In the quarterly online sales, \(r_4\) is high: the signature of the annual seasonality we have just quantified.

Autocorrelation has a second, quality-control use: the residuals of a good forecasting model should show no autocorrelation (if they do, there is structure left unexploited). On this idea rests an entire family of models — the ARIMA models — which model the series directly from its autocorrelations; they are the "next level" after this course, and here it is enough that the name rings a bell: when your software offers you an ARIMA, you now know what question it answers.

Common Mistakes and Tips

  • Comparing periods without deseasonalizing. "Q3 sales fell 30% versus Q2 — crisis!" No: Q3 is always 16.5% below. Compare with the same quarter of the previous year, or deseasonalize before you panic (or celebrate).
  • Applying t, z or ordinary regression to time-ordered data without a second thought. Dependence between observations invalidates p-values computed as if they were independent: they tend to come out "too significant".
  • Extrapolating the trend too far. Extending the line 2 quarters ahead is reasonable; extending it 5 years is science fiction. It is the same extrapolation danger we saw in regression, made worse because business trends change.
  • Confusing the cycle with seasonality. If the period is not fixed and known, it is not seasonality and cannot be captured with seasonal indices.
  • Reaching for the most complex method by default. Always compute the naïve method's error first; if your method does not beat it on MAE/MAPE over data it has not seen, do not use it.
  • Evaluating the error on the same data used to fit the model. That is cheating at solitaire: hold back the most recent periods and forecast them "blind", as we did with 2025.

Exercises

Exercise 1

Units of turrón (the Spanish Christmas nougat) sold at Madrid-Centro over the last 5 weeks, in thousands: 12, 18, 15, 21, 24. Compute the moving average of order 3 and comment on what it shows.

Exercise 2

Using the seasonal indices computed in the lesson (Q1: 0.968; Q2: 0.903; Q3: 0.835; Q4: 1.290) and the extrapolated trend, Marta asks you for the online sales forecast for Q3 2026. The last centered moving average was 4.875 (at Q2 2025) and the trend grows by 0.11 million per quarter. Compute the forecast and compare it with the seasonal naïve one (Q3 2025 = 4.2).

Exercise 3

Two methods forecast last week's daily deliveries. Method A's absolute errors were 12, 8, 15, 10, 5 orders; method B's were 2, 3, 40, 2, 3. Compute each method's MAE, decide which one you would keep, and explain the important nuance that looking at the errors one by one reveals.

Solutions

Exercise 1.

  • \(MA_3(2) = (12+18+15)/3 = 15.0\)
  • \(MA_3(3) = (18+15+21)/3 = 18.0\)
  • \(MA_3(4) = (15+21+24)/3 = 20.0\)

The smoothed series (15, 18, 20) shows clean growth of about 2,500 units per week: the turrón Christmas season is taking off. Common mistake: computing the moving average "backward" (weeks 1-2-3 assigned to week 3) — that is valid for real-time monitoring, but the centered version estimates the trend better; whichever convention you use, state it and stick to it.

Exercise 2.

From Q2 2025 to Q3 2026 there are 5 quarters: trend \(\approx 4.875 + 5 \times 0.11 = 5.425\). Applying the Q3 index: \(\hat{y} = 5.425 \times 0.835 \approx 4.53\) million. The seasonal naïve would say 4.2 (repeat Q3 2025); the structured method adds the trend growth: about €330,000 more. Common mistake: forgetting to apply the seasonal index and handing in 5.43 as the forecast — you would overestimate the summer by more than 19%; the correct order is always trend first, index second.

Exercise 3.

\(MAE_A = (12+8+15+10+5)/5 = 10\) orders; \(MAE_B = (2+3+40+2+3)/5 = 10\) orders. They tie on MAE, but they are very different methods: A errs moderately and consistently; B nails almost every day but failed spectacularly once (a public holiday the method doesn't account for?). For planning delivery drivers day by day, B is better if its blind spot gets fixed; if large misses are costly (a breached SLA — remember that 13% of deliveries already fall outside the SLA), A's stability may be worth more. Moral: the MAE summarizes, but it is no substitute for looking at the distribution of the errors — the same lesson the mean learned from the median back in Module 2.

Conclusion

You have learned how to handle data once time enters the picture: decomposing a series into trend, seasonality, cycle and noise; choosing between additive and multiplicative decomposition based on the amplitude of the swings; smoothing with moving averages and exponential smoothing; building a forecast from trend and seasonal indices that clearly beat the naïve method (a MAPE of 1.8% versus 9.1%); and reading autocorrelation as the series' "memory", with ARIMA models flagged as the next level. Marta now has her number for Christmas: about €7.1 million in Q4 2026.

But notice that the entire lesson has revolved around a single variable moving through time. In practice, NovaMarket's questions are rarely univariate: a store's sales depend simultaneously on its floor area, the foot traffic on its street and the neighborhood's income; a Club Nova member's churn depends on their age, their spend and their activity. Handling several variables at once — multiple regression, logistic regression, principal components, clusters — is the territory of Multivariate Analysis.

© Copyright 2026. All rights reserved