In the previous lesson we defined deep learning and saw why TecnoMarket wants to adopt it. But an inevitable question arises: if neural networks have existed since the 1950s, why did the revolution only arrive in the last decade? Knowing the history is not an exercise in erudition: understanding why deep learning failed twice before succeeding will give you the judgment to evaluate the promises (and the limits) of the technology you are about to use. In this lesson we will travel from the 1958 perceptron to the current era, passing through the "AI winters" and the milestone that changed everything in 2012.
Contents
- The origins: the perceptron (1943-1958)
- The first AI winter (1969-1980)
- Backpropagation and the 1980s revival
- The second winter (1990s and 2000s)
- The resurgence: GPUs, big data and the perfect storm
- The ImageNet/AlexNet milestone (2012)
- The current era: from 2012 to today
- Timeline of milestones
The origins: the perceptron (1943-1958)
The story begins before modern computers. In 1943, the neurophysiologist Warren McCulloch and the mathematician Walter Pitts published a mathematical model of a neuron: they showed that simple units connected together could, in theory, compute logical functions. It was pure theory, but it planted the seed of a powerful idea: perhaps intelligence can be built by connecting simple units.
In 1958, the psychologist Frank Rosenblatt took the leap into practice with the perceptron: the first machine capable of learning from examples by adjusting its internal connections. Rosenblatt even implemented it in hardware (the Mark I Perceptron), and trained it to distinguish simple shapes in 20x20 "pixel" images.
The enthusiasm was overwhelming. The New York Times went so far as to report that the perceptron would be "the embryo of a computer able to walk, talk, see, write and be conscious of its existence". Remember that sentence: inflated expectations will be a recurring pattern in this story.
We will not go into how the perceptron works internally here (we will study it in detail in lesson 02-01); what matters now is its historical role: it proved that a machine could learn, and it unleashed the first wave of optimism.
The first AI winter (1969-1980)
In 1969, Marvin Minsky and Seymour Papert published the book Perceptrons, a rigorous mathematical analysis that demonstrated the limitations of the single-layer perceptron: there were seemingly trivial problems (such as the logical XOR function) that it could never solve, no matter how much data it was given.
It was known that stacking several layers would solve the problem, but nobody had a practical method for training multilayer networks. The combination was lethal:
- Inflated expectations that went unmet.
- A proven theoretical limitation of the available model.
- Computers of the era that were extremely limited.
Governments and agencies (the main funders) cut the money. Research on neural networks was nearly abandoned for a decade: this is the so-called first AI winter. A lesson for TecnoMarket and for any company: when a technology promises more than it can deliver, the ensuing disappointment sweeps away even what did work.
Backpropagation and the 1980s revival
The thaw arrived in 1986, when David Rumelhart, Geoffrey Hinton and Ronald Williams popularized the error backpropagation algorithm: at last there was a practical, general method for training networks with several layers, overcoming the limitation Minsky had exposed. (The algorithm had precedents in the 1970s, such as Paul Werbos's work, but it was in 1986 that the community adopted it en masse.)
With backpropagation came real successes:
- 1989: Yann LeCun applied convolutional networks with backpropagation to handwritten digit recognition. His system ended up reading zip codes for the US Postal Service and bank checks: one of the first commercial applications of deep learning.
- The theoretical foundations of many architectures we use today were developed.
Again, we will not look at how backpropagation works internally here (that is the subject of lesson 02-03). What matters historically is that it removed the theoretical obstacle... but a practical one appeared.
The second winter (1990s and 2000s)
If the theoretical problem was solved, why didn't things take off then? Because training deep networks in the 1990s ran into three walls:
| Wall | Description |
|---|---|
| Computing power | The CPUs of the era took weeks to train modest networks. Adding layers sent the cost soaring. |
| Lack of data | There was no mass Internet: there were no millions of labeled images or texts with which to feed large networks. |
| Technical difficulties | The deep networks of the time suffered training problems (such as vanishing gradients, which we will mention in module 4) that nobody yet knew how to mitigate well. |
Meanwhile, classical ML methods such as support vector machines (SVMs) delivered equal or better results at much lower cost. The research community migrated to them in droves, and neural networks once again became a marginal field: the second winter. During these years, a small group of researchers — Geoffrey Hinton, Yann LeCun and Yoshua Bengio, later nicknamed "the godfathers of deep learning" — kept working on neural networks against the prevailing current.
The resurgence: GPUs, big data and the perfect storm
Between 2006 and 2012, the three missing factors aligned:
- GPUs (graphics cards): designed for video games, they turned out to be perfect for the kind of math neural networks need (enormous numbers of simple operations in parallel). Training runs that took weeks on CPUs came down to days or hours.
- Big data: the Internet, smartphones and social networks generated gigantic datasets for the first time. In 2009, Fei-Fei Li and her team published ImageNet: more than 14 million images labeled across thousands of categories, and they organized an annual image classification competition on a subset of 1,000 categories.
- Algorithmic advances: in 2006, Hinton and collaborators demonstrated techniques for training deep networks more stably, and forcefully coined the term deep learning. New activation functions and regularization techniques (which we will see in modules 2 and 5) made training far more viable.
The engine analogy is useful: backpropagation was the engine, but the fuel (data) and the road (hardware) were missing. By around 2010, everything was finally in place.
The ImageNet/AlexNet milestone (2012)
The moment almost everyone points to as the "big bang" of modern deep learning came at the 2012 ImageNet competition. A convolutional network called AlexNet, created by Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton, and trained on two consumer GPUs, swept the field:
| Year | Winner | Approximate top-5 error |
|---|---|---|
| 2010 | Classical methods (hand-crafted features + SVM) | ~28% |
| 2011 | Improved classical methods | ~26% |
| 2012 | AlexNet (deep learning) | ~16% |
| 2015 | Deep networks (ResNet) | ~3.6% (surpassing estimated human level, ~5%) |
Cutting the error from 26% to 16% in a single year was an unprecedented leap: typical annual improvements were 1-2 points. The message was clear to the entire community: the automatic feature extraction we saw in lesson 01-01 worked better than decades of manual engineering. Within two or three years, practically all research in computer vision (and then in speech and text) migrated to deep learning, and the big tech companies signed up the pioneers: Hinton (Google), LeCun (Facebook/Meta)...
For the TecnoMarket context: the product photo classifier we will build in module 3 is a direct descendant of AlexNet, and in 2012 it would have been world-class cutting-edge technology. Today it is a course exercise: that is how fast the field has moved.
The current era: from 2012 to today
Since 2012, milestones have followed one another at a dizzying pace. As a chronological wrap-up (without going into technical detail — each topic has its place in the course):
- 2014: GANs (generative adversarial networks) appear, capable of generating new images (module 5).
- 2014-2016: recurrent networks and their variants dominate machine translation and language processing (module 4).
- 2016: AlphaGo (DeepMind) defeats the world Go champion, a game considered beyond the reach of machines for decades.
- 2017: the Transformer architecture is published ("Attention Is All You Need"), which will revolutionize language processing (we will introduce it in lesson 05-05).
- 2018-2020: ever larger pretrained language models (BERT, GPT-2, GPT-3) show that scaling data and parameters keeps delivering improvements.
- 2021-today: the era of LLMs (large language models) and generative AI: ChatGPT, Claude, Gemini, image generators such as Stable Diffusion or Midjourney. Deep learning leaves the labs and reaches the general public.
In 2018, Hinton, LeCun and Bengio received the Turing Award (the "Nobel Prize of computing") for a bet they kept alive through the winters when almost nobody believed in it.
Timeline of milestones
| Year | Milestone | Why it matters |
|---|---|---|
| 1943 | McCulloch and Pitts neuron | First mathematical model of a neuron |
| 1958 | Rosenblatt's perceptron | First machine that learns from examples |
| 1969 | Perceptrons book (Minsky and Papert) | Proves the limits of the simple perceptron; triggers the first winter |
| 1969-1980 | First AI winter | Funding and research all but disappear |
| 1986 | Backpropagation (Rumelhart, Hinton, Williams) | Practical method for training multilayer networks |
| 1989 | LeCun: handwritten digits with convolutional networks | First major commercial application |
| ~1995-2006 | Second winter | Data and computing power are lacking; SVMs dominate |
| 2006 | Hinton relaunches the term deep learning | Techniques for training deep networks stably |
| 2009 | ImageNet dataset | The large-scale data "fuel" |
| ~2010 | Widespread use of GPUs to train networks | The missing "hardware" |
| 2012 | AlexNet wins ImageNet | The big bang: DL clearly surpasses classical ML in vision |
| 2014 | GANs | Generative deep learning is born |
| 2016 | AlphaGo beats the world Go champion | DL conquers "impossible" problems |
| 2017 | Transformer architecture | Foundation of the language revolution |
| 2018 | Turing Award to Hinton, LeCun and Bengio | Historic recognition of the field |
| 2020-today | LLMs and mass generative AI | DL reaches the general public and businesses |
Common Mistakes and Tips
- Mistake: believing that deep learning is a "new" technology. The core ideas are more than 60 years old; what is new is the data, the hardware and some refinements. This matters: the theoretical foundations are very well established.
- Mistake: thinking that progress was linear. There were two winters of near-total abandonment. Technologies advance in leaps, shaped by expectations, funding and hardware.
- Mistake: attributing the revolution solely to "better algorithms". The key algorithm (backpropagation) dates from 1986. Without GPUs and without big data there would have been no revolution. When you evaluate projects at your company, always ask about the available data and compute, not just about the model.
- Tip: be wary of grandiose headlines — back in 1958 they were already promising conscious machines. The field's history teaches you to distinguish between real progress (measurable, like the AlexNet leap) and inflated expectations.
- Tip: when you read about a new architecture, place it on the timeline: which problem of the previous ones does it solve? That question is the best guide to understanding the field.
Exercises
Exercise 1: Order the milestones
Put the following events in chronological order and add the approximate year: (a) AlexNet wins ImageNet, (b) Rosenblatt's perceptron, (c) popularization of backpropagation, (d) publication of the book Perceptrons, (e) Transformer architecture, (f) publication of the ImageNet dataset.
Exercise 2: The causes of the winters
Explain in your own words what caused each of the two AI winters, and which specific factor ended each one. What did both winters have in common?
Exercise 3: The report for TecnoMarket
TecnoMarket's CEO asks you, skeptically: "How do I know this deep learning thing isn't another passing fad that will deflate like in the 1990s?". Write a 4-6 line answer using historical arguments from this lesson.
Solutions
Solution 1:
- (b) Rosenblatt's perceptron — 1958
- (d) Perceptrons book — 1969
- (c) Popularization of backpropagation — 1986
- (f) ImageNet dataset — 2009
- (a) AlexNet wins ImageNet — 2012
- (e) Transformer architecture — 2017
Solution 2:
- First winter (1969-1980): caused by Minsky and Papert's proof that the single-layer perceptron could not solve basic problems (such as XOR), combined with unmet, outsized expectations. It ended when backpropagation (1986) made it possible to train multilayer networks, removing the theoretical limitation.
- Second winter (1990s-2000s): caused by practical obstacles: lack of computing power, lack of large datasets and training difficulties, while methods such as SVMs delivered better results at lower cost. It ended when GPUs, big data (ImageNet) and algorithmic improvements converged around 2006-2012.
- In common: in both cases there was a gap between expectations and actual capabilities, and funding/interest collapsed once it became evident. Theory was always ahead of the means to apply it.
Solution 3 (suggested answer):
"The earlier winters happened because pieces were missing: in the 1970s there was no way to train multilayer networks, and in the 1990s there was neither enough data nor enough hardware. Today those three pieces exist and are abundant: mature algorithms with decades of theoretical grounding, massive data and cheap GPUs, even in the cloud. Moreover, we are no longer talking about laboratory promises: since 2012, deep learning has measurably outperformed earlier techniques and powers commercial products we use every day (search engines, translators, assistants). The risk of hype exists in specific overhyped applications, but the underlying technology is consolidated; that is why I propose starting with use cases with clear returns, such as automatic product photo classification."
Conclusion
We have covered more than 60 years of history: the perceptron proved that machines could learn (1958), Minsky and Papert marked out its limits (1969), backpropagation overcame them in theory (1986), and the combination of GPUs + big data overcame them in practice, with AlexNet (2012) as the turning point that opened the current era of transformers and generative models. The professional moral: deep learning is not recent magic, but an old idea that matured once the world had the data and the hardware to feed it.
You now know what deep learning is and where it comes from. In the next lesson we will take a panoramic tour of its current applications by domain — vision, language, speech, recommendation, healthcare, fraud... — and connect them with TecnoMarket's specific needs and with the course modules where you will learn to build each one.
Deep Learning Course
Module 1: Introduction to Deep Learning
- What is Deep Learning?
- History and evolution of Deep Learning
- Applications of Deep Learning
- Basic concepts of neural networks
- Setting up the work environment
Module 2: Neural Network Fundamentals
- Perceptron and Multilayer Perceptron
- Activation functions
- Forward and backward propagation
- Optimization and loss functions
- Your first complete neural network
Module 3: Convolutional Neural Networks (CNN)
- Introduction to CNNs
- Convolutional and pooling layers
- Popular CNN architectures
- CNN applications in image recognition
Module 4: Recurrent Neural Networks (RNN)
- Introduction to RNNs
- LSTM and GRU
- RNN applications in natural language processing
- Sequences and time series
Module 5: Advanced Deep Learning Techniques
- Generative Adversarial Networks (GAN)
- Autoencoders
- Transfer Learning
- Regularization and improvement techniques
- Attention mechanisms and Transformers
Module 6: Tools and Frameworks
- Introduction to TensorFlow
- Introduction to PyTorch
- Framework comparison
- Development environments and additional resources
- Saving, loading and deploying models
Module 7: Hands-On Projects
- Image classification with CNNs
- Text generation with RNNs
- Anomaly detection with Autoencoders
- Building a GAN for image generation
- Fine-tuning a pretrained model
