In the previous lesson we defined Machine Learning as the discipline that enables computers to learn from data. But this idea wasn't born yesterday: it has more than seventy years of history behind it, with periods of euphoria, decades of disappointment (the famous "AI winters") and a spectacular renaissance in recent years. Knowing this evolution matters for two reasons: you will understand why ML is exploding precisely now (and not in 1960, when many of the ideas already existed), and you will gain the perspective to be swayed neither by hype nor by skepticism. In this lesson we will not explain the technical workings of the algorithms we mention; that comes in modules 4, 5 and 7.

Contents

  1. The origins: Turing and the founding question (1950)
  2. The first wave of enthusiasm: the perceptron and the pioneers (1950-1970)
  3. The AI winters (1974-1980 and 1987-1993)
  4. The statistical boom of the 1990s
  5. Big data and the industrialization of ML (2000-2012)
  6. The deep learning revolution (2012-2017)
  7. The current era: transformers and generative AI
  8. Timeline of milestones
  9. Why is it exploding now? The three ingredients

The origins: Turing and the founding question (1950)

In 1950, the British mathematician Alan Turing published the paper Computing Machinery and Intelligence, in which he posed the question "can machines think?" and proposed his famous Turing test: if, in a written conversation, you cannot tell whether your counterpart is a person or a machine, the machine is exhibiting intelligent behavior.

The most visionary idea in the paper, for our course, is a different one: Turing suggested that instead of programming an "adult mind" with all its rules, it would be more practical to build a "child mind" and educate it with examples. In other words, he described the essence of Machine Learning decades before it was technically possible.

Other founding milestones of the decade:

  • 1956 – The Dartmouth Conference: John McCarthy coins the term Artificial Intelligence. It is considered the official birth certificate of AI as a research field.
  • 1959 – Arthur Samuel: creates a program that plays checkers and improves by playing against itself. Samuel coins the term Machine Learning and defines it as "the field of study that gives computers the ability to learn without being explicitly programmed" — the intuitive definition we saw in the previous lesson.

The first wave of enthusiasm: the perceptron and the pioneers (1950-1970)

In 1958, the psychologist Frank Rosenblatt unveiled the perceptron, a machine inspired (very loosely) by the neurons of the brain, capable of learning to classify simple patterns from examples. It was the direct ancestor of today's neural networks (which we will study in lesson 04-07).

The press of the day went wild: it was even reported that the perceptron would be "the embryo of a computer able to walk, talk, see, write and be conscious of its own existence". This pattern — inflated expectations followed by disappointment — will repeat itself several times in our story.

In 1969, Marvin Minsky and Seymour Papert published the book Perceptrons, proving mathematically that the simple perceptron had severe limitations (it could not solve certain basic problems). The book was so influential that funding for neural networks all but vanished.

The AI winters (1974-1980 and 1987-1993)

An AI winter is a period in which, after broken promises, funding for and interest in the field collapse. There were two major winters:

  • First winter (1974-1980): the US and UK governments drastically cut investment after highly critical reports (such as the 1973 Lighthill report) which found that actual results were nowhere near what had been promised.
  • Second winter (1987-1993): the 80s saw a revival thanks to expert systems (programs with thousands of rules written by hand by human specialists: exactly the "traditional" approach we saw in the previous lesson). They worked in very narrow domains, but were extremely expensive to maintain and brittle outside their niche. When the market realized this, the bubble burst.

A lesson for the present: the winters did not come because the ideas were bad, but because the hardware and data of the time weren't up to realizing them, and because more was promised than could be delivered. Curiously, in the middle of the second winter (1986), Rumelhart, Hinton and Williams popularized backpropagation, the technique that decades later would make deep learning possible. The seeds of the future were planted in winter.

The statistical boom of the 1990s

In the 90s, the field took a pragmatic turn that defines it to this day: less ambition to "create minds" and more focus on solving concrete prediction problems on solid statistical foundations. Machine Learning split from symbolic AI and moved closer to statistics (which is why module 2 of this course covers statistics and probability).

Milestones of the decade:

  • 1995: Cortes and Vapnik introduce Support Vector Machines (SVM), which would dominate classification for 15 years (we'll see them in lesson 04-04).
  • 1995-2001: tree-based methods and their combinations mature, culminating in Random Forest (Breiman, 2001), a forerunner of the ensemble learning of module 7.
  • 1997: IBM's Deep Blue defeats world chess champion Garry Kasparov. (A caveat: Deep Blue was mostly brute force and rules, not pure ML, but it changed the public perception of what machines could do.)
  • 1998: Yann LeCun applies convolutional neural networks to read handwritten digits on bank checks: one of the first massive industrial applications of neural networks.

For a business like MercaFresh, this decade marks the moment when techniques such as demand forecasting or fraud detection with statistical models started to become viable in real companies (banks and large retailers were the pioneers).

Big data and the industrialization of ML (2000-2012)

With the explosion of the Internet, something decisive happened: for the first time in history, massive data became available. Every online purchase, every click, every search was being recorded. An online supermarket like MercaFresh generates more customer behavior data in a single day today than a brick-and-mortar chain of the 1980s did in a year.

Milestones of the period:

  • 2006: Netflix launches the Netflix Prize (a million dollars for whoever improved its recommendation system by 10%), popularizing ML competitions.
  • 2006: Geoffrey Hinton publishes the first work on trainable "deep" networks, and the term deep learning is coined.
  • 2009: ImageNet is released, a dataset of millions of labeled images that would become the proving ground for computer vision.
  • 2007-2010: the use of GPUs (graphics cards) for training models becomes widespread: hardware designed for video games that turned out to be perfect for the matrix operations of ML.
  • 2010: Kaggle, the data science competition platform, is founded.

The deep learning revolution (2012-2017)

The turning point has an exact date: 2012. A deep neural network called AlexNet (Krizhevsky, Sutskever and Hinton) crushed the ImageNet competition, cutting the image classification error from ~26% to ~16% in a single year, when previous annual improvements were measured in fractions of a percent. It was the incontrovertible demonstration that deep networks + lots of data + GPUs worked.

From there, everything accelerated:

  • 2014: GANs (generative adversarial networks) appear, capable of generating realistic images.
  • 2016: AlphaGo (DeepMind) defeats Lee Sedol at Go, a game considered beyond the reach of machines for decades, combining deep learning with reinforcement learning (a concept we will introduce in the next lesson).
  • The big tech companies reorganize their products around ML: machine translation, voice assistants, recommenders.

The current era: transformers and generative AI

In 2017, the paper Attention Is All You Need (Google) introduced the transformer architecture, which revolutionized first language processing and then almost everything else. On top of it were built the large language models (LLMs) such as GPT or Claude, and in 2022 the launch of ChatGPT brought generative AI (systems that generate text, images or code) to the general public.

We will not go deep into these architectures in this course (they belong to the advanced territory of deep learning, which we only touch on in lesson 07-04), but it's worth placing them in context: they are the natural evolution of the very same idea Arthur Samuel formulated in 1959 — learning from data — taken to gigantic scales of data and compute. And a key point for your career: most real business problems (forecasting demand, detecting churn, segmenting customers, as at MercaFresh) are still solved with the "classic" ML you will learn in this course, not with generative AI.

Timeline of milestones

Year Milestone Key figure(s) Era
1950 Turing test and the idea of "educating" machines Alan Turing Origins
1956 Dartmouth Conference: the term "AI" is born John McCarthy Origins
1958 The perceptron Frank Rosenblatt First enthusiasm
1959 "Machine Learning" is coined (checkers program) Arthur Samuel First enthusiasm
1969 Perceptrons: limits of the simple perceptron Minsky and Papert Decline
1974-1980 First AI winter — Winter
1986 Backpropagation is popularized Rumelhart, Hinton, Williams Seed in winter
1987-1993 Second winter (fall of expert systems) — Winter
1995 Support Vector Machines (SVM) Cortes and Vapnik Statistical boom
1997 Deep Blue beats Kasparov IBM Statistical boom
2001 Random Forest Leo Breiman Statistical boom
2006 The modern term "deep learning" is born Geoffrey Hinton Big data
2009 ImageNet Fei-Fei Li Big data
2012 AlexNet wins ImageNet: deep learning takes off Krizhevsky, Sutskever, Hinton Revolution
2016 AlphaGo beats Lee Sedol DeepMind Revolution
2017 Transformer architecture Vaswani et al. (Google) Current era
2022 ChatGPT popularizes generative AI OpenAI Current era
timeline
    title Evolution of Machine Learning
    1950-1970 : Origins and first enthusiasm : Turing, Dartmouth, perceptron
    1974-1993 : AI winters : funding cuts and disillusion, with seeds like backpropagation
    1990s : Statistical boom : SVM, trees, industrial applications
    2000-2012 : Big data : Internet, GPUs, ImageNet
    2012-2017 : Deep learning : AlexNet, GANs, AlphaGo
    2017-today : Current era : transformers and generative AI

Why is it exploding now? The three ingredients

Many of ML's central ideas are 40 or 60 years old. Why has the explosion come in the last decade? Because three ingredients that used to be missing finally came together:

  1. Data: the digitization of everyday life generates massive volumes of examples. MercaFresh illustrates it well: every online order automatically records dozens of variables (what, when, how much, from which device) that in a physical supermarket in 1990 simply did not exist.
  2. Compute: GPUs (and later purpose-built chips such as TPUs) multiplied computing power by orders of magnitude, and the cloud put it within anyone's reach by the hour. Training in 2012 what Rosenblatt could only dream of in 1958 cost a few hundred dollars.
  3. Algorithms and tools: algorithmic advances (practical backpropagation, new architectures) and, above all, open-source libraries such as scikit-learn (2007), the one we will use in this course, which democratized techniques once reserved for research labs.

You can think of it as a chemical reaction: the reagents (the ideas) had been sitting in the flask for decades; data and compute were the missing heat.

A simple way to feel the third ingredient: what in the 90s required implementing an algorithm from scratch in C over several weeks is now three lines of Python:

# What used to take a lab months of work is now an open-source library
from sklearn.linear_model import LinearRegression

model = LinearRegression()   # an algorithm with nearly 220 years of history (Gauss/Legendre)
# model.fit(X, y)            # plus all the modern machinery to train it, for free

This snippet is just a nod: the linear regression it instantiates (LinearRegression) has its roots in the 19th century, and we will study it properly in lesson 04-01. The moral is historical: the ideas are old; what's new is that now anyone can apply them.

Common Mistakes and Tips

  • Believing ML is a fad born with ChatGPT. As you have seen, there are 70 years of history behind it; generative AI is the most recent chapter, not the first.
  • Assuming the latest thing is always the best. "Veteran" techniques like regression or trees still win on a multitude of tabular business problems (like MercaFresh's), at lower cost and with more interpretability than deep learning.
  • Ignoring the lesson of the winters. Promising miraculous results to a client or your management is the historical recipe for disillusionment. Define measurable expectations (remember Mitchell's P from the previous lesson).
  • Confusing media milestones with ML advances. Deep Blue (1997) was mostly brute-force search; AlphaGo (2016) genuinely was learning. Telling them apart will sharpen your judgment.
  • Tip: when you read news about AI, mentally place it on this lesson's timeline and ask yourself which of the three ingredients (data, compute, algorithms) has made the novelty possible.

Exercises

Exercise 1

Put the following milestones in chronological order and assign each one to its era (origins, winter, statistical boom, big data, deep learning revolution, current era): AlexNet, perceptron, SVM, transformer, Lighthill report, ImageNet.

Exercise 2

Explain in your own words (4-6 lines) why the 1958 perceptron did not trigger back then the revolution that neural networks led from 2012 onwards. Lean on the "three ingredients".

Exercise 3

MercaFresh's management has read in the press that "generative AI will change everything" and proposes abandoning the classic-ML demand forecasting project to "wait for the new wave". Write a brief argument (3-5 lines) using what you learned in this lesson to respond to them.

Solutions

Solution 1

  1. Perceptron (1958) — origins / first enthusiasm.
  2. Lighthill report (1973) — trigger of the first AI winter.
  3. SVM (1995) — statistical boom of the 90s.
  4. ImageNet (2009) — big data era.
  5. AlexNet (2012) — deep learning revolution.
  6. Transformer (2017) — current era.

Solution 2

The perceptron contained the right idea (learning from examples by adjusting connections), but in 1958 two of the three ingredients were missing: there was barely any digitized data to train on, and the available compute was millions of times weaker than a modern GPU. On top of that, the algorithms for training multi-layer networks (such as backpropagation) were not popularized until 1986. When, in 2012, massive data (ImageNet), suitable hardware (GPUs) and mature algorithms finally came together, the very same idea worked spectacularly.

Solution 3

A possible argument: "Tabular business problems like demand forecasting are solved today excellently, cheaply and interpretably with classic ML, the same kind of techniques consolidated since the 90s-2000s. The history of the field teaches that waiting for the promised 'next wave' is the recipe for AI winters: inflated promises and stalled projects. The sensible thing is to capture value now with mature techniques and evaluate the new ones when they show a measurable advantage on our specific problem."

Conclusion

We have covered seventy years of history: Turing's vision, Rosenblatt's perceptron, the two winters that chilled the field, the statistical turn of the 90s, the big data era, the deep learning explosion of 2012 and the current wave of transformers and generative AI. The big takeaway is that the fundamental ideas are old, and that today's revolution is explained by the confluence of three ingredients: massive data, cheap compute and algorithms accessible to everyone. With this historical perspective in hand, in the next lesson we will organize the field from the inside: we will study the types of Machine Learning (supervised, unsupervised, semi-supervised and reinforcement) and see which one each MercaFresh problem belongs to.

Machine Learning Course

Module 1: Introduction to Machine Learning

Module 2: Foundations of Statistics and Probability

Module 3: Data Preprocessing

Module 4: Supervised Machine Learning Algorithms

Module 5: Unsupervised Machine Learning Algorithms

Module 6: Model Evaluation and Validation

Module 7: Advanced Techniques and Optimization

Module 8: Model Implementation and Deployment

Module 9: Hands-On Projects

Module 10: Additional Resources

© Copyright 2026. All rights reserved