In the previous lesson we defined Machine Learning as the discipline that enables computers to learn from data. But this idea wasn't born yesterday: it has more than seventy years of history behind it, with periods of euphoria, decades of disappointment (the famous "AI winters") and a spectacular renaissance in recent years. Knowing this evolution matters for two reasons: you will understand why ML is exploding precisely now (and not in 1960, when many of the ideas already existed), and you will gain the perspective to be swayed neither by hype nor by skepticism. In this lesson we will not explain the technical workings of the algorithms we mention; that comes in modules 4, 5 and 7.
Contents
- The origins: Turing and the founding question (1950)
- The first wave of enthusiasm: the perceptron and the pioneers (1950-1970)
- The AI winters (1974-1980 and 1987-1993)
- The statistical boom of the 1990s
- Big data and the industrialization of ML (2000-2012)
- The deep learning revolution (2012-2017)
- The current era: transformers and generative AI
- Timeline of milestones
- Why is it exploding now? The three ingredients
The origins: Turing and the founding question (1950)
In 1950, the British mathematician Alan Turing published the paper Computing Machinery and Intelligence, in which he posed the question "can machines think?" and proposed his famous Turing test: if, in a written conversation, you cannot tell whether your counterpart is a person or a machine, the machine is exhibiting intelligent behavior.
The most visionary idea in the paper, for our course, is a different one: Turing suggested that instead of programming an "adult mind" with all its rules, it would be more practical to build a "child mind" and educate it with examples. In other words, he described the essence of Machine Learning decades before it was technically possible.
Other founding milestones of the decade:
- 1956 – The Dartmouth Conference: John McCarthy coins the term Artificial Intelligence. It is considered the official birth certificate of AI as a research field.
- 1959 – Arthur Samuel: creates a program that plays checkers and improves by playing against itself. Samuel coins the term Machine Learning and defines it as "the field of study that gives computers the ability to learn without being explicitly programmed" — the intuitive definition we saw in the previous lesson.
The first wave of enthusiasm: the perceptron and the pioneers (1950-1970)
In 1958, the psychologist Frank Rosenblatt unveiled the perceptron, a machine inspired (very loosely) by the neurons of the brain, capable of learning to classify simple patterns from examples. It was the direct ancestor of today's neural networks (which we will study in lesson 04-07).
The press of the day went wild: it was even reported that the perceptron would be "the embryo of a computer able to walk, talk, see, write and be conscious of its own existence". This pattern — inflated expectations followed by disappointment — will repeat itself several times in our story.
In 1969, Marvin Minsky and Seymour Papert published the book Perceptrons, proving mathematically that the simple perceptron had severe limitations (it could not solve certain basic problems). The book was so influential that funding for neural networks all but vanished.
The AI winters (1974-1980 and 1987-1993)
An AI winter is a period in which, after broken promises, funding for and interest in the field collapse. There were two major winters:
- First winter (1974-1980): the US and UK governments drastically cut investment after highly critical reports (such as the 1973 Lighthill report) which found that actual results were nowhere near what had been promised.
- Second winter (1987-1993): the 80s saw a revival thanks to expert systems (programs with thousands of rules written by hand by human specialists: exactly the "traditional" approach we saw in the previous lesson). They worked in very narrow domains, but were extremely expensive to maintain and brittle outside their niche. When the market realized this, the bubble burst.
A lesson for the present: the winters did not come because the ideas were bad, but because the hardware and data of the time weren't up to realizing them, and because more was promised than could be delivered. Curiously, in the middle of the second winter (1986), Rumelhart, Hinton and Williams popularized backpropagation, the technique that decades later would make deep learning possible. The seeds of the future were planted in winter.
The statistical boom of the 1990s
In the 90s, the field took a pragmatic turn that defines it to this day: less ambition to "create minds" and more focus on solving concrete prediction problems on solid statistical foundations. Machine Learning split from symbolic AI and moved closer to statistics (which is why module 2 of this course covers statistics and probability).
Milestones of the decade:
- 1995: Cortes and Vapnik introduce Support Vector Machines (SVM), which would dominate classification for 15 years (we'll see them in lesson 04-04).
- 1995-2001: tree-based methods and their combinations mature, culminating in Random Forest (Breiman, 2001), a forerunner of the ensemble learning of module 7.
- 1997: IBM's Deep Blue defeats world chess champion Garry Kasparov. (A caveat: Deep Blue was mostly brute force and rules, not pure ML, but it changed the public perception of what machines could do.)
- 1998: Yann LeCun applies convolutional neural networks to read handwritten digits on bank checks: one of the first massive industrial applications of neural networks.
For a business like MercaFresh, this decade marks the moment when techniques such as demand forecasting or fraud detection with statistical models started to become viable in real companies (banks and large retailers were the pioneers).
Big data and the industrialization of ML (2000-2012)
With the explosion of the Internet, something decisive happened: for the first time in history, massive data became available. Every online purchase, every click, every search was being recorded. An online supermarket like MercaFresh generates more customer behavior data in a single day today than a brick-and-mortar chain of the 1980s did in a year.
Milestones of the period:
- 2006: Netflix launches the Netflix Prize (a million dollars for whoever improved its recommendation system by 10%), popularizing ML competitions.
- 2006: Geoffrey Hinton publishes the first work on trainable "deep" networks, and the term deep learning is coined.
- 2009: ImageNet is released, a dataset of millions of labeled images that would become the proving ground for computer vision.
- 2007-2010: the use of GPUs (graphics cards) for training models becomes widespread: hardware designed for video games that turned out to be perfect for the matrix operations of ML.
- 2010: Kaggle, the data science competition platform, is founded.
The deep learning revolution (2012-2017)
The turning point has an exact date: 2012. A deep neural network called AlexNet (Krizhevsky, Sutskever and Hinton) crushed the ImageNet competition, cutting the image classification error from ~26% to ~16% in a single year, when previous annual improvements were measured in fractions of a percent. It was the incontrovertible demonstration that deep networks + lots of data + GPUs worked.
From there, everything accelerated:
- 2014: GANs (generative adversarial networks) appear, capable of generating realistic images.
- 2016: AlphaGo (DeepMind) defeats Lee Sedol at Go, a game considered beyond the reach of machines for decades, combining deep learning with reinforcement learning (a concept we will introduce in the next lesson).
- The big tech companies reorganize their products around ML: machine translation, voice assistants, recommenders.
The current era: transformers and generative AI
In 2017, the paper Attention Is All You Need (Google) introduced the transformer architecture, which revolutionized first language processing and then almost everything else. On top of it were built the large language models (LLMs) such as GPT or Claude, and in 2022 the launch of ChatGPT brought generative AI (systems that generate text, images or code) to the general public.
We will not go deep into these architectures in this course (they belong to the advanced territory of deep learning, which we only touch on in lesson 07-04), but it's worth placing them in context: they are the natural evolution of the very same idea Arthur Samuel formulated in 1959 — learning from data — taken to gigantic scales of data and compute. And a key point for your career: most real business problems (forecasting demand, detecting churn, segmenting customers, as at MercaFresh) are still solved with the "classic" ML you will learn in this course, not with generative AI.
Timeline of milestones
| Year | Milestone | Key figure(s) | Era |
|---|---|---|---|
| 1950 | Turing test and the idea of "educating" machines | Alan Turing | Origins |
| 1956 | Dartmouth Conference: the term "AI" is born | John McCarthy | Origins |
| 1958 | The perceptron | Frank Rosenblatt | First enthusiasm |
| 1959 | "Machine Learning" is coined (checkers program) | Arthur Samuel | First enthusiasm |
| 1969 | Perceptrons: limits of the simple perceptron | Minsky and Papert | Decline |
| 1974-1980 | First AI winter | — | Winter |
| 1986 | Backpropagation is popularized | Rumelhart, Hinton, Williams | Seed in winter |
| 1987-1993 | Second winter (fall of expert systems) | — | Winter |
| 1995 | Support Vector Machines (SVM) | Cortes and Vapnik | Statistical boom |
| 1997 | Deep Blue beats Kasparov | IBM | Statistical boom |
| 2001 | Random Forest | Leo Breiman | Statistical boom |
| 2006 | The modern term "deep learning" is born | Geoffrey Hinton | Big data |
| 2009 | ImageNet | Fei-Fei Li | Big data |
| 2012 | AlexNet wins ImageNet: deep learning takes off | Krizhevsky, Sutskever, Hinton | Revolution |
| 2016 | AlphaGo beats Lee Sedol | DeepMind | Revolution |
| 2017 | Transformer architecture | Vaswani et al. (Google) | Current era |
| 2022 | ChatGPT popularizes generative AI | OpenAI | Current era |
timeline
title Evolution of Machine Learning
1950-1970 : Origins and first enthusiasm : Turing, Dartmouth, perceptron
1974-1993 : AI winters : funding cuts and disillusion, with seeds like backpropagation
1990s : Statistical boom : SVM, trees, industrial applications
2000-2012 : Big data : Internet, GPUs, ImageNet
2012-2017 : Deep learning : AlexNet, GANs, AlphaGo
2017-today : Current era : transformers and generative AI
Why is it exploding now? The three ingredients
Many of ML's central ideas are 40 or 60 years old. Why has the explosion come in the last decade? Because three ingredients that used to be missing finally came together:
- Data: the digitization of everyday life generates massive volumes of examples. MercaFresh illustrates it well: every online order automatically records dozens of variables (what, when, how much, from which device) that in a physical supermarket in 1990 simply did not exist.
- Compute: GPUs (and later purpose-built chips such as TPUs) multiplied computing power by orders of magnitude, and the cloud put it within anyone's reach by the hour. Training in 2012 what Rosenblatt could only dream of in 1958 cost a few hundred dollars.
- Algorithms and tools: algorithmic advances (practical backpropagation, new architectures) and, above all, open-source libraries such as scikit-learn (2007), the one we will use in this course, which democratized techniques once reserved for research labs.
You can think of it as a chemical reaction: the reagents (the ideas) had been sitting in the flask for decades; data and compute were the missing heat.
A simple way to feel the third ingredient: what in the 90s required implementing an algorithm from scratch in C over several weeks is now three lines of Python:
# What used to take a lab months of work is now an open-source library
from sklearn.linear_model import LinearRegression
model = LinearRegression() # an algorithm with nearly 220 years of history (Gauss/Legendre)
# model.fit(X, y) # plus all the modern machinery to train it, for freeThis snippet is just a nod: the linear regression it instantiates (LinearRegression) has its roots in the 19th century, and we will study it properly in lesson 04-01. The moral is historical: the ideas are old; what's new is that now anyone can apply them.
Common Mistakes and Tips
- Believing ML is a fad born with ChatGPT. As you have seen, there are 70 years of history behind it; generative AI is the most recent chapter, not the first.
- Assuming the latest thing is always the best. "Veteran" techniques like regression or trees still win on a multitude of tabular business problems (like MercaFresh's), at lower cost and with more interpretability than deep learning.
- Ignoring the lesson of the winters. Promising miraculous results to a client or your management is the historical recipe for disillusionment. Define measurable expectations (remember Mitchell's P from the previous lesson).
- Confusing media milestones with ML advances. Deep Blue (1997) was mostly brute-force search; AlphaGo (2016) genuinely was learning. Telling them apart will sharpen your judgment.
- Tip: when you read news about AI, mentally place it on this lesson's timeline and ask yourself which of the three ingredients (data, compute, algorithms) has made the novelty possible.
Exercises
Exercise 1
Put the following milestones in chronological order and assign each one to its era (origins, winter, statistical boom, big data, deep learning revolution, current era): AlexNet, perceptron, SVM, transformer, Lighthill report, ImageNet.
Exercise 2
Explain in your own words (4-6 lines) why the 1958 perceptron did not trigger back then the revolution that neural networks led from 2012 onwards. Lean on the "three ingredients".
Exercise 3
MercaFresh's management has read in the press that "generative AI will change everything" and proposes abandoning the classic-ML demand forecasting project to "wait for the new wave". Write a brief argument (3-5 lines) using what you learned in this lesson to respond to them.
Solutions
Solution 1
- Perceptron (1958) — origins / first enthusiasm.
- Lighthill report (1973) — trigger of the first AI winter.
- SVM (1995) — statistical boom of the 90s.
- ImageNet (2009) — big data era.
- AlexNet (2012) — deep learning revolution.
- Transformer (2017) — current era.
Solution 2
The perceptron contained the right idea (learning from examples by adjusting connections), but in 1958 two of the three ingredients were missing: there was barely any digitized data to train on, and the available compute was millions of times weaker than a modern GPU. On top of that, the algorithms for training multi-layer networks (such as backpropagation) were not popularized until 1986. When, in 2012, massive data (ImageNet), suitable hardware (GPUs) and mature algorithms finally came together, the very same idea worked spectacularly.
Solution 3
A possible argument: "Tabular business problems like demand forecasting are solved today excellently, cheaply and interpretably with classic ML, the same kind of techniques consolidated since the 90s-2000s. The history of the field teaches that waiting for the promised 'next wave' is the recipe for AI winters: inflated promises and stalled projects. The sensible thing is to capture value now with mature techniques and evaluate the new ones when they show a measurable advantage on our specific problem."
Conclusion
We have covered seventy years of history: Turing's vision, Rosenblatt's perceptron, the two winters that chilled the field, the statistical turn of the 90s, the big data era, the deep learning explosion of 2012 and the current wave of transformers and generative AI. The big takeaway is that the fundamental ideas are old, and that today's revolution is explained by the confluence of three ingredients: massive data, cheap compute and algorithms accessible to everyone. With this historical perspective in hand, in the next lesson we will organize the field from the inside: we will study the types of Machine Learning (supervised, unsupervised, semi-supervised and reinforcement) and see which one each MercaFresh problem belongs to.
Machine Learning Course
Module 1: Introduction to Machine Learning
- What is Machine Learning?
- History and evolution of Machine Learning
- Types of Machine Learning
- Applications of Machine Learning
- The Machine Learning project workflow
Module 2: Foundations of Statistics and Probability
- Basic statistics concepts
- Probability distributions
- Correlation and covariance
- Statistical inference
- Bayes' theorem
Module 3: Data Preprocessing
- Data cleaning
- Handling missing data
- Data transformation
- Encoding categorical variables
- Normalization and standardization
- Feature engineering
Module 4: Supervised Machine Learning Algorithms
- Linear regression
- Logistic regression
- Decision trees
- Support Vector Machines (SVM)
- K-Nearest Neighbors (K-NN)
- Naive Bayes
- Neural networks
Module 5: Unsupervised Machine Learning Algorithms
- Clustering: K-means
- Hierarchical clustering
- Principal Component Analysis (PCA)
- DBSCAN clustering
- Data visualization with t-SNE and UMAP
Module 6: Model Evaluation and Validation
- Data splitting: training, validation and test
- Evaluation metrics
- Cross-validation
- ROC curve and AUC
- Overfitting and underfitting
Module 7: Advanced Techniques and Optimization
- Regularization: Ridge, Lasso and Elastic Net
- Ensemble Learning
- Gradient Boosting
- Deep neural networks (Deep Learning)
- Hyperparameter optimization
Module 8: Model Implementation and Deployment
- Popular frameworks and libraries
- Deploying models to production
- Model maintenance and monitoring
- Ethical and privacy considerations
Module 9: Hands-On Projects
- Project 1: Housing price prediction
- Project 2: Image classification
- Project 3: Sentiment analysis on social media
- Project 4: Fraud detection
- Project 5: Customer segmentation
