We now know how to audit our models (08-01) and how to understand their social and economic context (08-02). The natural next question for the TecnoMarket team is strategic: where is the field heading, and what should we adopt, watch, or ignore? This lesson offers an honest panorama of the main trends — at the conceptual level, no tutorials — always measured by the same yardstick: what does this mean for a small team with limited resources? Just as important as knowing the trends is knowing how to filter them: deep learning is a noisy field, and telling a consolidated trend from a passing fad is a professional skill in its own right, and the one we will close with.

Contents

  1. LLMs and foundation models: scale as a strategy
  2. What foundation models mean for small teams
  3. Diffusion models: the promise left pending in 07-04
  4. Multimodality: text, image and audio in a single model
  5. Efficiency as a frontier: the MobileNet arc continues
  6. AutoML and architecture search, briefly
  7. Self-supervised learning: the silent engine
  8. How to keep up without drowning
  9. Summary table: trends and relevance for TecnoMarket

LLMs and foundation models: scale as a strategy

In 05-05 we closed the historical arc we opened in 01-02: from the neuron to transformers and LLMs. The trend that has dominated the field ever since is the consolidation of the foundation model: an enormous, general-purpose model trained once on massive data, then adapted to a thousand specific tasks. It is the transfer learning pattern from 05-03 taken to the extreme: no one pretrains "a vision model for classification" anymore; they pretrain "a model of almost everything".

Three ideas define the current stage, at the level of concept:

  • Scale with predictable returns: so far, more data + more parameters + more compute has kept producing measurable improvements (the so-called scaling laws). How long that curve will last is one of the field's open questions — we will address it with the challenges in 08-04.
  • Instruction tuning: modern LLMs don't just predict text (like our modest character-by-character generator from 07-02); they are subsequently tuned to follow instructions and be useful and safe in conversation. It is the difference between an engine and a drivable car.
  • RAG and agents, as an idea: an LLM only knows what it saw in training and can invent the rest. Retrieval-augmented generation (RAG) bolts a search engine onto it: before answering, it retrieves relevant documents (TecnoMarket's catalog, its return policies) and grounds its answer in them, reducing inventions and staying up to date without retraining. Agents go one step further: the LLM does not just answer, it uses tools (querying the orders database, running a search) over several steps to complete a task. Both are today engineering patterns built on top of the model, rather than changes to the model itself.

What foundation models mean for small teams

For TecnoMarket, the question is not "how do I train an LLM?" (it never will: 08-02 explained the costs), but "how do I take advantage of one?". The two routes, with their trade-offs:

Criterion Consume via API Open model tuned in-house
Initial quality Maximum (frontier models) Good, below the frontier
Cost Pay per use; cheap to start, grows with volume Own infrastructure + talent; flatter at high volume
Data Goes out to the provider (review with legal, 08-01) Stays in-house
Dependence High (prices, versions, availability: 08-02) Low; the model is yours
Customization Limited to the prompt and light tuning Full fine-tuning on your domain (the skill from 05-03/07-05)
Maintenance The provider's Yours (and it is not little: 06-05)

The sensible strategy for a small team is usually hybrid: API for exploring and for anything non-sensitive; an open model tuned in-house when volume, privacy or specialization justify it. An encouraging note: the fine-tuning you practiced in 07-05 with MobileNetV2 is, conceptually, the same operation as tuning an open LLM on TecnoMarket's reviews. The skills from this course scale.

Diffusion models: the promise left pending in 07-04

In 07-04 we trained a DCGAN on Fashion-MNIST and were honest about expectations: training GANs is unstable (the delicate generator-discriminator balance) and quality has a ceiling. We promised then that something better was on the way. Here it is: diffusion models, which today dominate high-quality image and video generation.

The intuition, without the mathematics: instead of the GAN's adversarial duel, diffusion learns to denoise. During training, a real image is taken and noise is gradually added until it becomes pure static; the model learns to reverse each small step: "given this slightly noisy version, what did a slightly cleaner one look like?". To generate, you start from pure noise and apply many successive cleaning steps until a new image emerges. Like sculpture: the noise is the block of marble and the model chips away what doesn't belong, step by step.

Why they have displaced GANs in quality generation:

  • Stable training: it is an ordinary prediction problem (predict the added noise), without the unstable balance between two networks that we suffered in 07-04. It converges in a boring, reliable way.
  • Coverage: GANs tend toward mode collapse (generating only a few convincing variants); diffusion covers the diversity of the data better.
  • Control: step-by-step generation is easily guided with text ("a Nordic-design lamp on a white background"), which is what makes today's image generators useful.
  • The price: generating requires many steps, so it is slower and more expensive at inference than a GAN (which generates in a single pass) — although step-reduction techniques are mitigating this quickly.

Are GANs dead? No: they are still used where inference speed matters and in specific tasks, and everything you learned in 05-01/07-04 (latent spaces, training generators, evaluating samples) transfers directly. For TecnoMarket, the practical implication is clear: the next generation of its promotional image generator will not be a GAN trained in-house, but a text-guided diffusion model — consumed via API or open — with the synthetic content rules from 08-01 applying exactly the same.

Multimodality: text, image and audio in a single model

The course treated each modality separately: CNNs for images (module 3), RNNs/LSTMs for text and sequences (module 4). The clear trend is fusion: multimodal models that understand and generate text, image and audio jointly, learning a shared space where "the photo of a coffee maker" and the sentence "stainless steel espresso machine" end up close together (a natural evolution of the representation spaces we saw in the autoencoders of 05-02).

For TecnoMarket, multimodality is not science fiction but a product roadmap:

  • Visual search: a customer photographs a lamp they saw in a bar and finds the similar ones in the catalog. (Technically: images and text queries projected into the same space, nearest-neighbor search.)
  • Enriched product pages: a model that looks at the product photo and writes a description consistent with what is visible — the fusion of our classifier (07-01/07-05) and our generator (07-02) into a single system, with the same human review before publishing.
  • Customer service with context: "does this cable work with the TV I bought?" answered by understanding the text, the history and the product photos.
  • Catalog quality control: automatically detecting that a product page's photo and description contradict each other.

Efficiency as a frontier: the MobileNet arc continues

Against the race for scale, the counter-trend that matters just as much: doing more with less. We know it well: we chose MobileNetV2 in 07-05 precisely for its efficiency, and in 08-02 we saw the environmental and economic motivation. The techniques defining this frontier, at the level of the idea:

  • Quantization: storing and computing the weights with fewer bits (from 32 bits down to 8, even 4). Models several times smaller and faster, often with minimal quality loss. We already mentioned it at deployment time in 06-05.
  • Distillation: training a small model (the "student") to imitate the outputs of a large one (the "teacher"), inheriting much of its performance at a fraction of the cost. The teacher is used once; the student serves in production.
  • Pruning: removing weights and neurons that barely contribute (recall how much redundancy the networks tolerated in 05-04).
  • Edge AI: the sum of the above brings models onto the device: the customer's phone runs the visual search without sending the photo to any server — lower latency, lower server cost and better privacy (08-01). The natural showcase of the MobileNet arc: the architecture was born for this.

For small teams this trend is great news: every advance in efficiency brings frontier capabilities closer to modest budgets.

AutoML and architecture search, briefly

Part of the design we did by hand in this course (how many layers, how many filters, which learning rate) can be automated: AutoML explores configurations automatically, and its ambitious version, neural architecture search (NAS), designs the architecture itself (in fact, variants of MobileNet and EfficientNet came out of NAS processes). The honest reading for TecnoMarket: automated hyperparameter search is useful and accessible; full NAS is expensive and rarely worthwhile compared to starting from a good published architecture, which is exactly the strategy (transfer learning) the team already masters. A real trend, but of moderate relevance for small teams: human judgment about which problem to solve and with which data remains unautomated.

Self-supervised learning: the silent engine

Underneath almost everything above sits an idea you already know without knowing it. In 05-02, the autoencoder learned useful representations without labels: its task was to reconstruct the input, and the labels were... the data itself. That is self-supervised learning: inventing a pretext task whose labels come for free from the data, and using it to learn representations.

Foundation models exist thanks to this idea taken to scale:

  • LLMs are pretrained by predicting the next word: every text on the internet brings its own "labels" (the words that follow). Our generator from 07-02 did exactly this, character by character, in miniature.
  • In vision, pretext tasks like reconstructing masked regions of an image or matching different views of the same scene produce representations that rival supervised ones.

Why it matters: labeling is the bottleneck (we suffered it in every module 7 project, and 08-02 showed its labor side). Self-supervised learning turns unlabeled data — which is abundant — into representations, and saves the few available labels for the final tuning. For TecnoMarket: its thousands of unlabeled photos and reviews stop being dead storage and become raw material for pretraining, its own or someone else's.

How to keep up without drowning

The field publishes more than anyone can read, and marketing inflates results. Professional criteria for filtering the hype:

  1. Are there reproducible results? Published code and weights, evaluations on known benchmarks run by third parties. Distrust cherry-picked demos and figures that exist only in a press release.
  2. Does it solve a problem you have? The TecnoMarket question. A trend irrelevant to your use cases can wait on the watch list.
  3. Does it survive six months? Few ideas die as fast as they are born. Waiting two publication cycles before adopting filters out most of the noise without losing anything essential.
  4. Is the total cost accounted for? Inference, maintenance, provider dependence (08-02), ethical auditing (08-01). The paper never includes the invoice.

Stable places to look (types of source, rather than volatile URLs): the papers and technical blogs of the major research labs; the field's reference conferences (NeurIPS, ICML, ICLR, CVPR for vision, ACL for language) and their preprint repository (arXiv); the official documentation and examples of the module 6 frameworks, which absorb the techniques that mature; annual state-of-the-field surveys; and technical communities where practitioners compare real results. A realistic routine: a couple of fixed hours per week, following a few good sources, with your own "watch / try / adopt" list.

Summary table: trends and relevance for TecnoMarket

Trend Maturity Relevance for TecnoMarket Reasonable action
LLMs / foundation models Consolidated and evolving fast High: customer service, descriptions, search Adopt via API under the 08-02 conditions; evaluate an open model if volume grows
RAG / agents Young but already practical patterns High: answers grounded in the real catalog and policies Pilot internally before facing customers
Diffusion models Consolidated for images; video evolving Medium-high: the natural replacement for the 07-04 GAN Adopt an existing tool; do not train in-house; 08-01 rules
Multimodality Maturing rapidly High in the medium term: visual search, enriched product pages Watch, and prototype the visual search
Efficiency (quantization, distillation, edge) Consolidated and improving High: costs, latency, privacy Adopt quantization at deployment; the MobileNet arc always followed this path
AutoML / NAS Mature for hyperparameter search; NAS expensive Low-medium Use hyperparameter search; skip NAS
Self-supervised learning Consolidated as the engine of foundation models Medium directly (leveraged via pretrained models) Understand the concept; exploit unlabeled data when the time comes

Common Mistakes and Tips

  • Confusing "the latest" with "the best for your problem". The demand predictor from 04-04, with its LSTM and its baselines, may still be the right solution even though time-series transformers exist. Change techniques when a metric on your data justifies it, not when a headline says so.
  • Adopting with no exit route. Every adoption of an API or external model must carry a documented plan B (08-02). Trends get discontinued too.
  • Ignoring a trend because "that's for the big players". Foundation models are trained by a few, but consumed from anywhere: the lesson of transfer learning (05-03) is that the frontier reaches small teams with one or two years' delay and minimal cost.
  • Reading papers without a depth gradient. A sensible order: the abstract and figures of many, the introduction of quite a few, the full detail of very few. Reading everything in depth doesn't scale; reading nothing leaves you at the mercy of marketing.
  • Tip: keep the summary table alive as a team document and review it every quarter: it is your technology radar, cheap and honest.
  • Tip: when you try a trend, test it against a baseline of yours that is already measured (the discipline of 04-04 and of the frozen set from 06-05). "It seems better" is not a data point; +3 points over your baseline is.

Exercises

Exercise 1: renewing the image generator

Marketing asks for "better promotional images than our 07-04 GAN's". Options: (A) improve the in-house DCGAN with more data and epochs; (B) use a text-guided diffusion service via API; (C) deploy an open diffusion model on your own infrastructure. Evaluate all three using the API vs. open table criteria and the properties of diffusion versus GANs, and recommend one for a 3-person team, citing which 08-01 obligations apply to your choice.

Exercise 2: the RAG chatbot

Management proposes an LLM-based customer service assistant. The team debates between connecting it "as is" or building a RAG system over TecnoMarket's catalog and policies. Explain in your own words what RAG contributes in this specific case, which main risk it reduces (and why that risk exists in a bare LLM), what new maintenance work it introduces, and what role human review should play in the initial phase.

Exercise 3: filtering the hype

A vendor offers TecnoMarket a "quantum-neural NAS platform that automatically designs networks 40% better, according to our internal studies". Apply the four hype-filtering criteria one by one to this offer and draft the reply (three sentences) you would give to management.

Solutions

Solution 1. (A) Improving the DCGAN: high time cost, low quality ceiling — 07-04 taught us first-hand the instability of adversarial training and the risk of mode collapse; for genuine promotional quality it is a dead end. (B) Diffusion API: immediate frontier quality, per-image cost, zero maintenance; risks: provider dependence and data leaving the company (minor here: the prompts are product descriptions, not personal data — still, review it). (C) Open diffusion in-house: control and data ownership, but it demands a dedicated GPU and operational skills that a 3-person team should not spend on a non-differentiating use case. Recommendation: B, with a reviewed contract (08-02), and with the 08-01 obligations: labeling the content as generated, a ban on misleading photorealism of the product, and human review before publishing. The 07-04 skills are not thrown away: they serve to evaluate samples and to understand what is being bought.

Solution 2. What RAG contributes: before answering, the system retrieves TecnoMarket's real product pages and policies and the LLM writes grounded in them; without RAG, the LLM answers only from its training, which includes neither the current catalog nor the company's policies. The risk it reduces: invented but plausible answers — a bare LLM will confidently "fill in" the facts it lacks (return periods, compatibilities), because its nature is to generate plausible text, the same one we saw in miniature in the invented outputs of our 07-02 generator. New maintenance: the retrievable knowledge base becomes a living artifact — it must be kept in sync with the catalog and policies (an outdated page produces wrong answers complete with citation), and retrieval quality must be evaluated periodically. Initial human review: a pilot phase with the proposed answers going to the human agent (who approves or corrects) before exposing them directly to customers — the 03-04 review queue applied to chat; low-confidence or high-impact cases (complaints, payments) remain with humans even afterwards.

Solution 3. Criterion 1 (reproducibility): "internal studies" with no code, weights or third-party evaluation — fail; the term "quantum-neural" used as marketing is, moreover, a classic smoke signal. Criterion 2 (a problem of ours): TecnoMarket does not have the problem "designing architectures": its winning strategy is starting from published architectures with transfer learning (07-05 beat the in-house CNN exactly that way). Criterion 3 (survival): there is no six-month track record of adoption to verify. Criterion 4 (total cost): a proprietary platform = maximum dependence, NAS cost high by nature, no inference or maintenance figures. Reply to management: "The offer provides no verifiable evidence for its 40% and attacks a problem we do not have, because our proven path to improvement is fine-tuning public architectures. We propose declining and, if the vendor insists, asking for a free trial measured by us against our 91% baseline on our frozen set. Without that test, there is no case."

Conclusion

The landscape is now mapped: foundation models turn scale into strategy and reach small teams through the twin API / tuned open model route — the decision structured in our table —; diffusion settles the promise of 07-04, replacing the GANs' unstable duel with stable, controllable denoising; multimodality fuses what this course treated separately and points straight at TecnoMarket's visual search and enriched product pages; efficiency (quantization, distillation, edge) continues the MobileNet arc and works in favor of the small; self-supervised learning — the autoencoder pretext of 05-02 at planetary scale — is the engine feeding it all; and hype filtering with four criteria is the skill that keeps everything else useful. But an honest panorama cannot tell only what is advancing: the field drags serious open problems — models that fail outside their distribution, generative systems that invent, systems fragile against adversaries — and, at the same time, those cracks are exactly where the work and the professional opportunity are. Challenges and opportunities, and the closing of this whole journey, are the subject of the course's final lesson.

© Copyright 2026. All rights reserved