In 05-01 you assembled the skeleton of a DCGAN — a generator with Conv2DTranspose, a convolutional discriminator, the counterfeiter-and-police metaphor — and we left the most delicate part pending: the complete training loop. With the GradientTape you mastered in 06-01, you now have exactly the tool to write it. In this fourth project you will train a GAN end to end on Fashion-MNIST (clothing items as a stand-in for product photos) as a prototype of an image generator for TecnoMarket's promotional creatives, learning to read the dynamics of adversarial training: what is normal, what is mode collapse, and what to adjust when something goes wrong.

Contents

  1. Project brief and realistic expectations
  2. Phase 1: data — Fashion-MNIST with tf.data
  3. Phase 2: generator and discriminator (picking up 05-01)
  4. Phase 3: the two adversarial losses
  5. Phase 4: one training step with GradientTape, line by line
  6. Phase 5: the epoch loop with sample grids
  7. Phase 6: reading the dynamics — the normal, mode collapse and what to adjust
  8. Phase 7: evaluation, saving the generator and downstream use

Project brief and realistic expectations

Business context. TecnoMarket's marketing team wants to explore image generation for promotional creatives (backgrounds, product variations, campaign illustrations). Before investing, they ask for an educational prototype proving the team has mastered the generative mechanics.

Stand-in. Fashion-MNIST: 60,000 28×28 grayscale images of clothing items (t-shirts, sneakers, bags...), a reasonable stand-in for simple product photos. Small, quick to train and varied enough to expose the phenomena that matter.

Expectations, in writing before we start: a DCGAN of this size produces garments that are recognizable but blurry at 28×28. It is not StyleGAN nor a commercial product; today's production systems use diffusion models (you'll see them mentioned in 08-03), and all commercial image generation carries ethical implications we will address in 08-01. The goal here is to master the adversarial mechanism — which is also the conceptual foundation for understanding modern systems.

Phase 1: data — Fashion-MNIST with tf.data

import tensorflow as tf
import numpy as np
import matplotlib.pyplot as plt

tf.random.set_seed(42)

(x_train, _), _ = tf.keras.datasets.fashion_mnist.load_data()

# Normalize to [-1, 1]: the generator will end in tanh (05-01)
x_train = (x_train.astype("float32") - 127.5) / 127.5
x_train = x_train[..., np.newaxis]           # (60000, 28, 28, 1)

BATCH = 128
train_ds = (tf.data.Dataset.from_tensor_slices(x_train)
            .shuffle(60_000)
            .batch(BATCH, drop_remainder=True)
            .prefetch(tf.data.AUTOTUNE))

An important detail we already anticipated in 05-01: the range is [-1, 1] (not [0, 1]) because the generator's output will be tanh. Generator and real data must speak the same numeric language, or the discriminator will tell them apart by the range, not the content. There are no labels and no test set: in a GAN, the "data" is just the reference for reality.

Phase 2: generator and discriminator (picking up 05-01)

We instantiate the DCGAN skeleton from 05-01 adapted to 28×28. We won't re-explain the pieces (review them there); we annotate the decisions:

from tensorflow.keras import layers, models

NOISE_DIM = 100

def build_generator():
    return models.Sequential([
        layers.Input(shape=(NOISE_DIM,)),
        layers.Dense(7 * 7 * 256, use_bias=False),
        layers.BatchNormalization(), layers.LeakyReLU(),
        layers.Reshape((7, 7, 256)),
        layers.Conv2DTranspose(128, 5, strides=1, padding="same", use_bias=False),
        layers.BatchNormalization(), layers.LeakyReLU(),        # 7x7
        layers.Conv2DTranspose(64, 5, strides=2, padding="same", use_bias=False),
        layers.BatchNormalization(), layers.LeakyReLU(),        # 14x14
        layers.Conv2DTranspose(1, 5, strides=2, padding="same",
                               activation="tanh"),              # 28x28, [-1,1]
    ], name="generator")

def build_discriminator():
    return models.Sequential([
        layers.Input(shape=(28, 28, 1)),
        layers.Conv2D(64, 5, strides=2, padding="same"),
        layers.LeakyReLU(), layers.Dropout(0.3),                # 14x14
        layers.Conv2D(128, 5, strides=2, padding="same"),
        layers.LeakyReLU(), layers.Dropout(0.3),                # 7x7
        layers.Flatten(),
        layers.Dense(1),                                        # logit: real or fake
    ], name="discriminator")

generator = build_generator()
discriminator = build_discriminator()

Reminders from 05-01, applied: Conv2DTranspose with strides=2 doubles the resolution (7→14→28), BN and LeakyReLU stabilize the generator, the discriminator uses dropout (not BN) and returns a logit with no sigmoid — the loss will handle it.

Phase 3: the two adversarial losses

The game from 05-01, now in code. Both losses start from the same binary cross-entropy over logits:

bce = tf.keras.losses.BinaryCrossentropy(from_logits=True)

def discriminator_loss(real_logits, fake_logits):
    # The police wants: real -> 1, fake -> 0
    real_loss = bce(tf.ones_like(real_logits), real_logits)
    fake_loss = bce(tf.zeros_like(fake_logits), fake_logits)
    return real_loss + fake_loss

def generator_loss(fake_logits):
    # The counterfeiter wants ITS fakes to look real (label 1)
    return bce(tf.ones_like(fake_logits), fake_logits)

# Two separate optimizers: each network learns on its own (05-01)
opt_gen = tf.keras.optimizers.Adam(2e-4, beta_1=0.5)
opt_disc = tf.keras.optimizers.Adam(2e-4, beta_1=0.5)

The asymmetry is the essence: the discriminator scores the same fake images with a target of 0 and the generator with a target of 1. The hyperparameters lr=2e-4, beta_1=0.5 are the classic DCGAN values — they work; don't touch them on the first pass.

Phase 4: one training step with GradientTape, line by line

Here everything comes together: the GradientTape from 06-01 lets us compute two sets of gradients from a single graph and apply them to different networks. Annotated line by line:

@tf.function                      # compiles the step to a graph: ~5-10x faster (06-01)
def train_step(real_images):
    noise = tf.random.normal([BATCH, NOISE_DIM])          # 1. batch of noise

    with tf.GradientTape() as tape_g, tf.GradientTape() as tape_d:
        fake_images = generator(noise, training=True)      # 2. counterfeit

        real_logits = discriminator(real_images, training=True)  # 3. judge the real
        fake_logits = discriminator(fake_images, training=True)  # 4. judge the fake

        g_loss = generator_loss(fake_logits)               # 5. how well it deceives
        d_loss = discriminator_loss(real_logits, fake_logits)  # 6. how well it detects

    # 7. Gradients of EACH network's loss with respect to ITS variables
    grads_g = tape_g.gradient(g_loss, generator.trainable_variables)
    grads_d = tape_d.gradient(d_loss, discriminator.trainable_variables)

    # 8. Each optimizer updates only its own network
    opt_gen.apply_gradients(zip(grads_g, generator.trainable_variables))
    opt_disc.apply_gradients(zip(grads_d, discriminator.trainable_variables))
    return g_loss, d_loss

Fine points you must understand, not just copy:

  • Two tapes, one pass: both record the same operations, but step 7 asks each tape for gradients only of its loss with respect to its variables. Updating the generator doesn't touch the discriminator, and vice versa — if you mixed variables, each network would sabotage the other.
  • training=True on both networks, always: the generator's BN must use batch statistics even when its images feed the discriminator's loss.
  • @tf.function: the same decorator we saw in 06-01; in a custom loop like this, the speed difference is enormous.
  • This is the pattern that Keras fit() doesn't hand you ready-made: two networks, two opposing losses, one simultaneous step. That's why the promise from 05-01 had to wait for 06-01.

Phase 5: the epoch loop with sample grids

In a GAN, the loss doesn't tell the whole truth (phase 6), so the loop periodically generates a grid of samples from the same fixed noise, to compare epochs on equal terms:

fixed_noise = tf.random.normal([16, NOISE_DIM], seed=42)   # ALWAYS the same

def save_grid(epoch):
    samples = generator(fixed_noise, training=False)
    samples = (samples + 1) / 2                    # [-1,1] -> [0,1] for plotting
    plt.figure(figsize=(4, 4))
    for i in range(16):
        plt.subplot(4, 4, i + 1)
        plt.imshow(samples[i, :, :, 0], cmap="gray")
        plt.axis("off")
    plt.savefig(f"logs/gan/grid_epoch_{epoch:03d}.png")
    plt.close()

EPOCHS = 50
for epoch in range(1, EPOCHS + 1):
    g_losses, d_losses = [], []
    for image_batch in train_ds:
        gl, dl = train_step(image_batch)
        g_losses.append(float(gl)); d_losses.append(float(dl))
    print(f"Epoch {epoch:3d} | G: {np.mean(g_losses):.3f} "
          f"| D: {np.mean(d_losses):.3f}")
    if epoch % 5 == 0 or epoch == 1:
        save_grid(epoch)
        generator.save(f"models/generator_epoch_{epoch:03d}.keras")

On a GPU, each epoch takes ~20-40 s (50 epochs ≈ half an hour); on CPU it is slow — cut down to 15-20 epochs or use Colab (01-05). We save generator checkpoints every 5 epochs: in GANs, the "best epoch" is chosen by looking at grids, not losses, so being able to roll back matters.

Phase 6: reading the dynamics — the normal, mode collapse and what to adjust

Honest, typical visual evolution (your exact epochs will vary):

Epochs What you'll see in the grid
1-3 Gray noise with blotches: nothing recognizable
5-10 Light blobs on a dark background: blurry "proto-garments"
15-25 Clear silhouettes: t-shirts, trousers, shoes become distinguishable
30-50 Recognizable garments with basic texture; edges still soft

Losses: what is normal. Unlike a classifier, here the losses must not converge to zero: it is an equilibrium, not a descent. Healthy: G oscillating around ~0.7-1.5 and D around ~1.0-1.3, both moving with no clear trend. Warning signs:

Symptom Diagnosis What to adjust (from 05-01)
D → 0 and G keeps climbing Discriminator too strong: the generator gets no useful signal Lower the discriminator's lr (e.g. 1e-4), or add noise/label smoothing (0.9 instead of 1) to the reals
The grid's 16 samples are nearly identical Mode collapse: the generator found one image that deceives and repeats it More entropy: raise the generator's lr slightly, revisit label smoothing, restart from an earlier checkpoint
Everything oscillates violently and the samples get worse Learning rates too high Halve both lrs
Grids stagnant for 15+ epochs Dead equilibrium Try more generator capacity or more epochs; sometimes it just needs time

Diagnosis works with the grids first and the curves second — exactly the reverse of the previous projects.

Phase 7: evaluation, saving the generator and downstream use

Qualitative evaluation. The minimal protocol: (1) fixed-noise grids epoch by epoch — are they improving?; (2) diversity — generate 64 fresh samples and check that several garment types appear; (3) a nearest-neighbor glance — for some generated sample, find the most similar real image in the dataset and verify it is not a memorized copy.

Quantitative metrics, concept only. Research uses the FID (Fréchet Inception Distance): it compares statistics of real and generated images in the feature space of a pretrained network — lower is better. For this prototype it's enough to know it exists and that disciplined visual inspection is the practical standard at this scale.

Delivery. Only the generator is deployed — the discriminator was the personal trainer and stays home:

generator.save("models/creatives_generator_v1.keras")

# Downstream use: new images from noise
gen = tf.keras.models.load_model("models/creatives_generator_v1.keras")
new_images = gen(tf.random.normal([8, NOISE_DIM]), training=False)
new_images = ((new_images + 1) / 2).numpy()   # to [0,1]: the "un-preprocessing" travels documented (06-05)

An honest report for marketing: "we have mastered the generative mechanics and can produce low-resolution synthetic images; production-quality creatives require diffusion models (08-03) and a prior analysis of ethical and rights implications (08-01) — every published synthetic image must be identified as such".

Common Mistakes and Tips

  • Training the discriminator to perfection "first": a perfect discriminator gives the generator near-zero gradient. They must grow together: one step each, as in our loop.
  • Changing the grid's noise every epoch: without fixed noise you can't tell whether the generator improved or the sample merely changed. The fixed noise is your own private frozen set (06-05).
  • Reading the losses as in a classifier: the generator's loss going up doesn't mean it got worse — maybe the discriminator improved. Grids first.
  • Forgetting drop_remainder=True: a final batch of a different size can break shapes inside @tf.function.
  • Despairing at epoch 10: GANs are slow to get going and non-linear in their improvement. Judge every 5-10 epochs, with checkpoints to return to the best point.

Exercises

  1. Add label smoothing to the discriminator (reals = 0.9 instead of 1.0) and compare the stability of the curves and grids against the original version over 20 epochs.
  2. Train with NOISE_DIM = 2 and visualize what the model generates by sweeping a mesh of noise values (for example, from -2 to 2 in each dimension). What do you observe about the latent space? Does mode collapse appear?
  3. Interpolate in the latent space: take two noise vectors z1, z2, generate images for z = (1-t)*z1 + t*z2 with t from 0 to 1 in 8 steps, and display the transition. Relate it to the embeddings of 03-04.

Solutions

  1. Change in discriminator_loss: bce(tf.ones_like(real_logits) * 0.9, real_logits). Typical effect: D's loss stops sinking toward 0, G's oscillates less and the grids progress more steadily — one of the 05-01 remedies in action.
  2. With only 2 latent dimensions, the mesh [(x, y) for x in np.linspace(-2,2,8) for y in np.linspace(-2,2,8)] produces a grid where neighboring regions yield similar garments: you are visualizing the entire latent space. Overall diversity drops (2 dimensions leave little room) and the risk of collapse grows: you'll see few garment classes represented. Conclusion: the noise dimension bounds how much variety the generator can encode.
  3. for t in np.linspace(0, 1, 8): imgs.append(gen((1-t)*z1 + t*z2, training=False)). The transition is smooth: a t-shirt gradually morphs into trousers passing through plausible intermediate shapes. As with the embeddings of 03-04, proximity in the latent space encodes semantic similarity — the generator has organized the noise into a continuous map of garments.

Conclusion

Fourth project delivered and the oldest promise of the course settled: the complete adversarial loop, written by hand with two GradientTapes, two losses and two optimizers, with fixed-noise grids as the evaluation instrument and a practical diagnosis of GAN dynamics (equilibrium, not convergence; mode collapse and its remedies). The result is an honest prototype — recognizable garments, not product photography — and a clear report on what production would require. One project remains, and it is the one that closes the circle: in 07-05 we return to the image classifier from 07-01 and beat its 85% with transfer learning, comparing figures head to head to make TecnoMarket's final decision.

© Copyright 2026. All rights reserved