At the close of module 7, the TecnoMarket team delivered a portfolio of five working models: an image classifier, a description generator, a fraud detector, a promotional-image GAN, and a model tuned with transfer learning. They all work. But "it works" is not the same as "it is right to use it this way". A fraud detector with good overall recall may be systematically penalizing an entire neighborhood; an image generator can fabricate misleading content; a classifier can discriminate without anyone ever programming it to. This lesson turns those vague worries into concrete tools: how bias arises, how to detect it with code, how to mitigate it, how to explain the decisions of a black box, how to handle personal data, and what to do about synthetic content. It is also where we settle the debt outstanding since lesson 05-01: the risk of deepfakes.

Contents

  1. Algorithmic bias: where it really comes from
  2. Bias at TecnoMarket: two concrete cases
  3. Practical detection: metrics by subgroup
  4. Bias mitigation and its dilemmas
  5. Transparency and explainability: opening the black box
  6. Privacy: the personal data behind the model
  7. Deepfakes and synthetic content: the debt from 05-01
  8. TecnoMarket's ethics checklist before every deployment

Algorithmic bias: where it really comes from

A common myth: "the algorithm is mathematical, therefore it is neutral". False. A neural network learns exactly what is in its data, including the patterns we would rather it did not learn. Algorithmic bias rarely comes from a malicious programmer; it comes from three far more mundane sources:

  • Biased historical data: if the data reflects unfair human decisions from the past, the model perpetuates them. It learned from the past, and the past was not neutral.
  • Unequal representation: if a group appears rarely in the dataset, the model performs worse for that group. Not out of "hatred", but out of pure statistics: fewer examples, worse generalization. It is the same phenomenon we saw with imbalanced classes in the fraud detector (07-03), applied to people.
  • Proxy variables: the model does not need to see the sensitive attribute (gender, ethnicity, age) to discriminate by it. It is enough that another variable correlates with it: postal code correlates with income level and origin; purchase history, with age; a person's name, with gender.

Some well-known real-world cases worth keeping in mind:

Case What happened Source of the bias
CV screening tool at a major tech company (2018) Penalized résumés containing the word "women's" (e.g. women's clubs) Historical data: 10 years of predominantly male hires
COMPAS (criminal justice, US) The recidivism-risk system produced more false positives for Black defendants Historical policing data + socioeconomic proxies
Commercial facial recognition (Gender Shades study, 2018) Error <1% for light-skinned men; up to ~35% for dark-skinned women Unequal representation in the training datasets
Credit scoring with proxies Denials concentrated in certain postal codes Proxy variables for income and origin

Notice the pattern: in none of these cases did anyone program the discrimination. It emerged from the data. That is the central lesson: bias is the default behavior of a model trained on real-world data, not an exotic anomaly.

Bias at TecnoMarket: two concrete cases

Let's bring this down to our portfolio. Two of the five models carry a direct risk of bias:

Case 1: the fraud detector and postal codes

The anomaly detector from 07-03 uses, among other variables, transaction and customer data. Suppose the team included the postal code (or something that correlates with it, such as the average price bracket of the delivery neighborhood). The risk: if historically more fraud was detected in certain neighborhoods — perhaps because more investigation happened there, not because there was more actual fraud — the model will learn "postal code X ⇒ suspicious". The result: legitimate customers from those neighborhoods suffer more blocks, more friction and more of the "annoyed customers" we priced at €5 in the confusion matrix in euros from 07-03. But that €5 cost is not spread evenly: it concentrates on one group. That is no longer just a business cost; it is a fairness problem.

Case 2: the photo classifier and small sellers

TecnoMarket lets third-party sellers upload products. Large sellers upload studio photos (white background, good lighting); small sellers upload homemade photos (kitchen background, yellow light, an old phone). If the classifier from 07-01/07-05 was trained mostly on studio photos, it will have worse accuracy on homemade ones. The consequence: products from small sellers get misclassified, fall into the human review queue (03-04) more often, or are simply miscategorized and never show up in searches. The model, without knowing it, favors the large sellers. This is unequal representation in its purest form: the "homemade photo" subgroup is underrepresented.

Both cases share an uncomfortable property: global metrics do not detect them. A 91% overall accuracy can hide 96% on studio photos and 74% on homemade ones. That is why detection requires disaggregating.

Practical detection: metrics by subgroup

The basic bias-auditing tool is simple, and you already know all the pieces: compute the metrics from module 7, but separately for each relevant subgroup. An example with the fraud detector, evaluating recall and false positive rate per postal-code segment:

import numpy as np
import pandas as pd

# df has: y_true (1=fraud), y_pred (1=blocked), zone (customer segment)
def metrics_by_group(df, group_col):
    rows = []
    for group, sub in df.groupby(group_col):
        tp = ((sub.y_true == 1) & (sub.y_pred == 1)).sum()
        fn = ((sub.y_true == 1) & (sub.y_pred == 0)).sum()
        fp = ((sub.y_true == 0) & (sub.y_pred == 1)).sum()
        tn = ((sub.y_true == 0) & (sub.y_pred == 0)).sum()
        rows.append({
            "group": group,
            "n": len(sub),
            "fraud_recall": tp / (tp + fn) if (tp + fn) > 0 else np.nan,
            "false_positive_rate": fp / (fp + tn) if (fp + tn) > 0 else np.nan,
            "pct_blocked": (sub.y_pred == 1).mean(),
        })
    return pd.DataFrame(rows)

print(metrics_by_group(df, "zone"))

A result like this should set off every alarm:

zone n fraud_recall false_positive_rate pct_blocked
city_center 41,200 0.83 0.008 1.1%
north_suburbs 8,300 0.85 0.041 4.9%
south_suburbs 6,100 0.81 0.038 4.5%

Recall is similar across all zones (the model catches fraud about equally well everywhere), but the false positive rate is 5 times higher in the suburbs: a legitimate customer from north_suburbs is 5 times more likely to have their purchase blocked. That is measurable bias. Keys to the method:

  • Choose the subgroups before looking at results (zone, seller type, age bracket, review language...), so you don't only look where it doesn't hurt.
  • Compare the metric that represents the harm for that group. For legitimate customers, the false positive rate; for sellers, the classification accuracy on their photos.
  • Beware of small subgroups: with n = 50, differences can be noise. Report the sample size too.
  • Automate it: this analysis must run on every retraining, just like validation against the frozen set from 06-05 — not once and never again.

Bias mitigation and its dilemmas

Once bias is detected, there are three families of intervention, ordered from most advisable to most delicate:

Strategy What it involves TecnoMarket example Dilemma
Act on the data Collect more examples of the underrepresented group; review suspicious historical labels; remove unnecessary proxies Collect and label 5,000 homemade photos; drop the postal code if it carries no legitimate signal Cost and time; sometimes the proxy does carry useful signal mixed in
Reweight the training Give minority-group examples more weight in the loss (like class_weight for the imbalance in 07-03) 3× weight for homemade photos in the 07-05 fine-tuning May lower the global metric somewhat; you must decide how much is acceptable
Per-group thresholds Apply different decision thresholds per subgroup to equalize error rates Anomaly threshold at the 97th percentile in the city center, 99th in the suburbs, to equalize false positives The most controversial: explicitly treating each group differently to achieve equal outcomes. May be legally problematic depending on jurisdiction and sector

The third one deserves a pause, because it reveals something deep: there are several mathematical definitions of "fairness" and they are mutually incompatible in the general case. Equalizing the false positive rate across groups, equalizing recall, or using the same threshold for everyone are three reasonable criteria... and except in degenerate cases you cannot satisfy all three at once. There is no theorem that decides for you: it is a values decision the team must make explicitly, document, and be able to defend. What is unacceptable is not choosing one criterion over another; it is never having consciously chosen any.

Transparency and explainability: opening the black box

A deep network with millions of parameters offers no readable explanation of why it blocked a transaction. This matters for three reasons:

  1. User trust: "your purchase has been blocked" with no reason given drives customers away.
  2. Team debugging: without explanations, the bias from the previous section is harder to diagnose.
  3. Right to an explanation: in the EU, people have the right not to be subject to fully automated decisions with significant effects, and to obtain human intervention and meaningful information about the logic applied (an idea enshrined in the GDPR; see the privacy section and its legal disclaimer).

Tools at the conceptual level (we will not implement them here):

  • Saliency maps / Grad-CAM (vision): highlight which pixels influenced the prediction most. If the classifier from 07-05 is "looking" at the kitchen background instead of the product, you will literally see it: it is the most direct way to confirm the homemade-photo bias.
  • SHAP / attribution importance (tabular): assign each input variable a contribution to the specific prediction. "This transaction was flagged because: amount 4× above the customer's average (+0.31), shipping to a new address (+0.22)...". If the postal code keeps showing up as the main contributor, the proxy is staring you in the face.
  • Simple surrogate models: train a decision tree that imitates the large model to obtain approximate, readable rules.

And the organizational safeguard we already built: the human review queue with a confidence threshold from 03-04. When the model is unsure — or when the decision is high-impact, such as blocking a customer — a person decides, not the model. The three-way decision from 07-03 (approve / review / block) is exactly this: the "review" band is where the human guarantee lives. Technical explainability and human review do not compete; they complement each other: the former helps the human reviewer decide better and faster.

flowchart LR
    T[Transaction] --> M[Model]
    M -->|low error| A[Auto-approve]
    M -->|middle band| H[Human review queue<br/>+ SHAP explanation]
    M -->|very high error| B[Hold and verify]
    H --> D[Documented human decision]

Privacy: the personal data behind the model

TecnoMarket's five models were trained on data that, in a real deployment, includes personal data: purchase histories, reviews written by customers, addresses. Practical principles:

  • Minimization: train only on the variables you need. Every column of personal data you don't use is free risk. Does the description generator from 07-02 really need the name of each review's author? No.
  • Anonymization and its limits: removing the name and ID number is not enough. The combination of postal code + date of birth + sex re-identifies most of the population; detailed purchase histories are practically fingerprints. Pseudonymization (replacing identifiers with codes) reduces the risk but does not eliminate it.
  • Memorization in generative models: generative models can memorize and regurgitate fragments of their training data. The generator from 07-02, trained on real reviews, could reproduce verbatim a review containing personal data ("it arrived late to my house on X street..."). This reinforces the human review of every generated output before publication, which we already established in 07-02.
  • Data subject rights: access, rectification, erasure. "Erasure" is awkward in deep learning: the data may be diluted across the weights. The standard practice is to remove it from the dataset and apply that in the next retraining, documenting the process.

On the European regulatory framework, purely for orientation: the GDPR regulates the processing of personal data (legal basis, minimization, rights, automated decisions) and the European AI Act classifies AI systems by risk level, imposing increasing obligations — transparency, human oversight, risk management — on high-risk ones, plus labeling obligations for certain AI-generated content.

Important disclaimer: none of the above is legal advice. This course summarizes general ideas for educational purposes; the details, deadlines, risk classifications and specific obligations depend on the case, the sector and the jurisdiction, and they change over time. Before any real deployment that processes personal data or makes decisions about people, the project must be reviewed by a compliance/legal professional.

Deepfakes and synthetic content: the debt from 05-01

In 05-01 we learned how GANs work and deliberately postponed the conversation about their misuse. In 07-04 we trained a DCGAN capable of generating reasonably convincing images of clothing. Time to settle the debt.

The same technology scales to faces and voices indistinguishable from real ones: deepfakes. The risks are concrete: identity impersonation (fake videos of executives authorizing transfers), disinformation, non-consensual intimate imagery, and a general erosion of trust ("if everything can be fake, nothing proves anything").

TecnoMarket does not make deepfakes of people, but its generator from 07-04 (and any future evolution toward realistic promotional images) raises commercial versions of the same problem:

  • Generating a "product photo" that does not correspond to the real product ⇒ misleading advertising.
  • Generating images with synthetic people and using them as if they were real customers ("María from Seville says...") ⇒ fabricated testimonials.
  • Generating synthetic reviews with the 07-02 model and publishing them as if written by customers ⇒ outright consumer fraud.

A framework for the responsible use of synthetic content, in four rules:

  1. Never present the synthetic as real when that influences purchase or trust decisions. Generated promotional images are used as creative illustration, not as product photos.
  2. Label synthetic content: a visible label ("AI-generated image") and, where the tooling allows it, technical marking (provenance metadata, standards like C2PA, watermarking). Content labeling is also aligning with regulatory obligations (see the legal disclaimer above).
  3. Human review before publishing, the rule that already governs the text generator from 07-02: nothing generated goes public without human eyes on it.
  4. Explicitly prohibited uses: the internal policy must list what is never done (fake reviews, synthetic people presented as real, imitation of brands or persons), so that it does not depend on each employee's individual judgment on a Friday afternoon.

TecnoMarket's ethics checklist before every deployment

Everything above condenses into a checklist the team appends to the end of the deployment checklist from 06-05. No model goes to production without answering it in writing:

# Question Related lesson
1 What decision does this model make or influence, and about whom? —
2 Have we computed the error metrics by relevant subgroups? Any unjustifiable differences? This lesson
3 Are there proxy variables for sensitive attributes? Are they justified and documented? This lesson
4 Do high-impact decisions go through human review? Is the threshold set correctly? 03-04, 07-03
5 Can we give the affected person a meaningful explanation? This lesson
6 Is the training data minimized, and do generative outputs avoid leaking personal data? 07-02
7 If it generates content, is it labeled as synthetic and reviewed before publishing? 05-01, 07-04
8 Who owns this model, and how often is it re-audited (alongside drift monitoring)? 06-05
9 Has the compliance/legal lead reviewed the deployment? —

Note that the checklist does not demand perfection: it demands having looked, having decided, and having documented it. Ethical negligence is almost never a bad decision; it is the absence of a decision.

Common Mistakes and Tips

  • "My model doesn't discriminate: I don't feed it the sensitive variable." Classic mistake. Proxies (postal code, history, name) carry the information all the same. The only way to know is to measure outcomes by subgroup, not to inspect the column list.
  • Auditing once. Bias reappears with every retraining and with data drift (06-05). The subgroup audit must be part of the pipeline, not a one-off event.
  • Confusing explainability with an excuse. "The model decided it" is not a valid explanation to a customer or to a regulator. If you cannot explain the decision, the high-impact decision must not be fully automated.
  • Equalizing metrics without deciding which. Before applying per-group thresholds, decide explicitly which notion of fairness you are pursuing and document why; remember you cannot satisfy them all at once.
  • Treating privacy as a database problem. The model's weights also "contain" data: generative models can memorize. Actively test whether your generator regurgitates training text.
  • Tip: turn the checklist into a living document in the repository, versioned alongside the model. When something goes wrong (it will), having the reasoning in writing is what distinguishes an honest mistake from negligence.
  • Tip: for small subgroups, always report the metric together with its n. A huge difference at n = 30 warrants collecting more data before redesigning the model.

Exercises

Exercise 1: auditing the photo classifier

The team disaggregates the accuracy of the 07-05 classifier by seller type and gets: large sellers (studio photos, n = 18,000): 94%; small sellers (homemade photos, n = 2,400): 76%. The sales director says: "the global figure is 92%, we beat the 90% target, we deploy". Analyze: (a) what is wrong with that reasoning?; (b) propose two concrete mitigations from different families, with their cost and their risk; (c) what role can the human review queue from 03-04 play while the mitigation is on its way?

Exercise 2: the per-zone threshold dilemma

In the fraud detector, the false positive rate is 0.8% in the city center and 4.1% in the northern suburbs (similar recall). An engineer proposes raising the anomaly threshold only in the suburbs until false positives equalize at 0.8%. Simulate the team discussion: give one solid argument in favor, one solid argument against, identify what additional information you would request before deciding, and what alternative option exists that does not touch the thresholds.

Exercise 3: the campaign with generated images

Marketing wants to use the (evolved) GAN to generate: (a) abstract decorative backgrounds for banners; (b) photorealistic "enhanced" product images for product pages; (c) faces of "satisfied customers" next to real anonymized reviews. Apply the responsible synthetic-content framework to each case and rule: allowed as is, allowed with conditions (which?), or prohibited (why?).

Solutions

Solution 1. (a) The 92% global figure is a weighted average that hides the fact that one subgroup receives a much worse service (76%); since small sellers are a minority (2,400 out of 20,400), their poor result barely moves the average. The "90% global" target does not protect subgroups, and the harm is asymmetric: for a small seller, misclassification means invisibility in searches — the model worsens the disadvantage of the weakest. (b) Data mitigation: collect/label several thousand homemade photos (even asking the sellers themselves, or generating data augmentation with homemade backgrounds and lighting); high labeling cost, low risk. Reweighting mitigation: repeat the 07-05 fine-tuning with higher weight for homemade photos; low cost, with a risk of losing some accuracy on studio photos, which would need quantifying (is dropping from 94% to 92.5% in exchange for rising from 76% to 86% a good deal? Almost certainly yes). (c) In the meantime, lower the confidence threshold specifically for photos identified as "non-studio", sending more of them to the human review queue: small sellers get correct (human) classification at the cost of latency and review cost, instead of getting automated wrong classification.

Solution 2. In favor: the current situation already treats the groups unequally in outcomes (a legitimate suburban customer suffers 5 times more blocks); the per-zone threshold does not create the inequality, it corrects it, and the criterion "equal false positive rate for legitimate customers" is a defensible, documentable notion of fairness. Against: it introduces an explicit rule that decides differently by zone (a proxy for origin/income), which may be hard to defend legally and in the court of public opinion ("TecnoMarket has different thresholds per neighborhood"); moreover, if real fraud differs between zones, raising the threshold in the suburbs will let real fraud through (a cost of €400/case per 07-03). Additional information: does the false-positive gap come from genuinely different fraud or from historical label bias (more investigation happened in the suburbs)? Which variables push the score up in the suburbs (look at SHAP-style attributions)? What is the n for each zone? Alternative without touching thresholds: attack the cause — remove or review the zone proxies among the input variables, relabel the questionable history, retrain, and in the meantime route the doubtful suburban band to the human review queue instead of to automatic blocking. Whatever the decision, it must be documented and pass through point 9 of the checklist (legal review).

Solution 3. (a) Abstract backgrounds: allowed with no special conditions; they represent nothing real and do not influence the perception of the product; the general review-before-publishing policy suffices. (b) Photorealistic "enhanced" product pages: allowed only under strict conditions — the image cannot show features the real product lacks (that is misleading advertising); it must be labeled as generated/retouched; mandatory human review comparing it against the real product; when in doubt, use the generated image only as ambient imagery and keep real product photos. (c) Faces of "satisfied customers": prohibited. Even if the reviews are real, attaching them to synthetic people presented as customers fabricates testimonials: it deceives about the identity of who is endorsing the product, violates rule 1 of the framework (do not present the synthetic as real when it influences trust) and edges into the deepfake territory that motivated this section. The legitimate alternative: show the reviews without a face, or with clearly illustrative, labeled avatars.

Conclusion

We have turned ethics from a statement of intent into a set of measurable practices: bias is detected by disaggregating metrics by subgroup (TecnoMarket's two cases proved it: the global figure hides the subgroup), it is mitigated by acting on data, weights or thresholds — knowing that the notions of fairness compete with each other and one must be chosen consciously —, opacity is fought with technical explainability and with the human review queue we have carried since 03-04, privacy demands minimization and vigilance over generative memorization, and synthetic content — the debt from 05-01, finally settled — is governed by four rules: never pass it off as real, label it, review it, and explicitly prohibit unacceptable uses. All of it crystallizes in the checklist that now accompanies every TecnoMarket deployment. But these team decisions happen inside a much larger context: deep learning is transforming employment, the economy and access to technology at the scale of entire societies. That is the zoom out of the next lesson: the social and economic impact of deep learning.

© Copyright 2026. All rights reserved