In 04-04 you learned to turn sequences into sliding windows to predict "the next value", and we promised that the same trick would serve to generate text. The time has come. In this second project of the module you will build a draft generator for product descriptions aimed at TecnoMarket sellers: a model that, character by character, learns the style of the product listings and proposes new text. Beyond the usual project phases, this lesson introduces the only new technical piece you need: temperature sampling, which controls the balance between safe-but-repetitive text and creative-but-chaotic text.

Contents

  1. Project brief and the character-by-character approach
  2. Phase 1: the description corpus
  3. Phase 2: character-level vectorization and sequence windows
  4. Phase 3: Embedding + LSTM model
  5. Phase 4: training
  6. Phase 5: temperature sampling
  7. Phase 6: iterative generation and typical results
  8. Limits of the approach and responsible use

Project brief and the character-by-character approach

Business context. TecnoMarket sellers take a long time writing descriptions for their items and many end up empty or telegraphic. We want an assistant that generates a draft in the house style, which the seller edits and approves. A human always signs off on the final text.

Why character by character? It is the simplest, most instructive formulation of generation: the model sees a sequence of characters and predicts which one comes next. It is exactly the problem from 04-04 (predicting the next step of a series) with characters in place of numbers: the same sliding windows, the same classification head as in 04-03. With vocabularies of ~50 symbols there are no unknown words to manage and the model fits on any machine. The price — coherence limited to short sentences — is something we will analyze at the end.

Phase 1: the description corpus

For the prototype we use a synthetic corpus of fictional TecnoMarket descriptions. In a real case you would use the history of approved listings; here we generate material with a homogeneous style (you can also practice with any public English text, but our own corpus keeps the project's thread):

import numpy as np
import tensorflow as tf

tf.random.set_seed(42)  # 06-04: reproducibility

products = ["wireless earbuds", "mechanical keyboard", "curved monitor",
            "robot vacuum", "air fryer", "security camera",
            "bluetooth speaker", "espresso machine", "gaming mouse",
            "smart lamp", "digital scale", "tower fan"]
adjectives = ["compact", "powerful", "quiet", "elegant", "versatile",
              "ergonomic", "durable", "lightweight"]
benefits = ["long lasting battery", "stable connection", "modern design",
            "easy setup", "low power consumption", "large capacity",
            "control from your phone", "premium quality materials"]
closers = ["ideal for the home.", "perfect for your daily routine.",
           "the best value for money.", "ships within 24 hours."]

rng = np.random.default_rng(42)
sentences = []
for _ in range(3000):
    p, a = rng.choice(products), rng.choice(adjectives)
    b1, b2 = rng.choice(benefits, size=2, replace=False)
    c = rng.choice(closers)
    sentences.append(f"{a} {p} with {b1} and {b2}. {c}\n")

corpus = "".join(sentences)
print(len(corpus))            # ~300,000 characters
print(corpus[:200])

Notice two decisions: the corpus is all lowercase plain ASCII (a small, clean vocabulary) and each description ends in \n, which the model will learn as the "end of description" signal. Around 300,000 characters is little for high-quality generation, but enough for the model to learn the style — calibrate your expectations accordingly.

Phase 2: character-level vectorization and sequence windows

In 04-03 we used TextVectorization at the word level; here we go down to the character with StringLookup, and apply the sliding windows from 04-04 with tf.data (06-01):

vocab = sorted(set(corpus))
print(len(vocab))   # ~45 symbols: letters, digits, space, punctuation, \n

char_to_id = tf.keras.layers.StringLookup(vocabulary=vocab, mask_token=None)
id_to_char = tf.keras.layers.StringLookup(vocabulary=vocab, invert=True,
                                          mask_token=None)

ids = char_to_id(tf.strings.unicode_split(corpus, "UTF-8"))

WINDOW = 80   # window length: ~one full description
ds = tf.data.Dataset.from_tensor_slices(ids)
ds = ds.batch(WINDOW + 1, drop_remainder=True)   # chunks of 81 characters

def split_input_target(chunk):
    return chunk[:-1], chunk[1:]   # input: 80 chars; target: the same shifted by 1

BATCH = 64
train_ds = (ds.map(split_input_target)
              .shuffle(10_000)
              .batch(BATCH, drop_remainder=True)
              .prefetch(tf.data.AUTOTUNE))

The key is in split_input_target: for the input "echanical keyboar", the target is "chanical keyboard" — at every time step the model learns to predict the following character. It is the same window→next-step idea from 04-04, but training all 80 positions of the window at once, which multiplies the learning signal per example.

Phase 3: Embedding + LSTM model

The architecture reuses the pieces from 04-02 and 04-03: an Embedding (of characters here, not words) and an LSTM with return_sequences=True, because we need a prediction at every time step, not only at the end:

from tensorflow.keras import layers, models

VOCAB = len(char_to_id.get_vocabulary())   # includes the [UNK] token

model = models.Sequential([
    layers.Embedding(VOCAB, 64),                     # each char -> 64-dim vector
    layers.LSTM(256, return_sequences=True),         # a state for every position
    layers.Dropout(0.2),                             # regularization (05-04)
    layers.Dense(VOCAB),                             # logits: a score per char
])

model.compile(
    optimizer="adam",
    loss=tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True),
)
model.build(input_shape=(None, None))
model.summary()   # ~0.4M parameters

Two important details:

  • The final layer has no softmax (from_logits=True in the loss): in the sampling phase we will manipulate the logits directly to apply the temperature.
  • The loss is the usual CCE (02-04) applied per time step: Keras averages the cross-entropy of the 80 predictions in each window.

Phase 4: training

callbacks = [
    tf.keras.callbacks.ModelCheckpoint("models/best_generator.keras",
                                       monitor="loss", save_best_only=True),
    tf.keras.callbacks.EarlyStopping(monitor="loss", patience=3,
                                     restore_best_weights=True),
]
history = model.fit(train_ds, epochs=30, callbacks=callbacks)

A guide to reading the loss (with a ~45-symbol vocabulary, the initial loss hovers around ln(45) ≈ 3.8 — random prediction):

Loss What the model generates
~3.8 Random character soup
~2.0 Pronounceable pseudo-words, spaces in the right places
~1.2 Real words from the corpus, approximate syntax
< 0.8 Sentences with the structure of the descriptions

With this synthetic corpus (very regular) the loss drops fast, down to 0.5-0.7 in 15-25 epochs and a few minutes of CPU. With a real, more varied corpus you would expect higher values and more time. We don't monitor validation here: in prototype-stage generation what interests us is capturing the corpus style, and the decisive evaluation will be qualitative (phase 6).

Phase 5: temperature sampling

With the model trained, how do we pick the next character from its logits? Here comes the central concept of the lesson.

  • Greedy: always pick the most probable character. Deterministic and "safe", but it falls into loops: with long lasting battery and long lasting battery and ....
  • Temperature sampling: divide the logits by a temperature T before the softmax and sample from the resulting distribution:
def next_char(logits, temperature):
    logits = logits / temperature
    sampled_id = tf.random.categorical(logits[tf.newaxis, :], num_samples=1)
    return int(sampled_id[0, 0])

The effect of T on the distribution:

Temperature Effect Typical output
T → 0 Almost greedy: the most probable always wins Correct but repetitive, sentences cloned from the corpus
T = 0.5 Conservative: spreads little probability around A good default compromise
T = 1.0 The distribution exactly as the model learned it Varied, with the occasional stumble
T = 1.5 Flattens the distribution: unlikely characters get their chance Invents words: versatle air fryer with stble connecion

Temperature doesn't change the model: it changes how much risk you accept when sampling. For commercial drafts, low-to-medium temperatures (0.4-0.8) are the sensible range.

Phase 6: iterative generation and typical results

Generating means iterating: predict a character, append it to the sequence, predict again. A didactic version (it re-processes the whole sequence at every step; good enough for a prototype):

def generate(seed, n_chars=150, temperature=0.5):
    gen_ids = char_to_id(tf.strings.unicode_split(seed, "UTF-8"))
    gen_ids = list(gen_ids.numpy())
    for _ in range(n_chars):
        window = tf.constant([gen_ids[-WINDOW:]])      # last window
        logits = model(window)[0, -1, :]               # logits of the last step
        gen_ids.append(next_char(logits, temperature))
    chars = id_to_char(tf.constant(gen_ids))
    return tf.strings.reduce_join(chars).numpy().decode("utf-8")

print(generate("robot vacuum ", temperature=0.5))

Real typical outputs from a model trained like this (yours will vary):

  • T=0.2: robot vacuum with long lasting battery and easy setup. ideal for the home. — impeccable, but almost traced from the corpus.
  • T=0.7: robot vacuum with control from your phone and low power consumption. the best value for money. — combines pieces in new, correct ways: the sweet spot.
  • T=1.4: robot vacuum ergonamic with grate capasity and stabel connction. ships within 24 huors. — creativity run wild: spelling mistakes and broken syntax.

Delivery. Just as in 07-01, the generator is packaged with its preprocessing (the StringLookup travels with the model if you integrate it into an inference model, 06-05) and would be served behind a FastAPI endpoint (POST /draft with product and temperature) that returns 3 drafts for the seller to choose from and edit.

Limits of the approach and responsible use

Be honest about what you have built:

  • Short coherence: a character LSTM keeps the thread for a few dozen characters; in long texts it contradicts itself or rambles. Our corpus of short sentences fits exactly within that limit — that's why it works.
  • It knows no facts: it can write "long lasting battery" for a product with no battery. It generates style, not truth.
  • Commercial generation systems use transformers/LLMs built on the attention you saw in 05-05: context of thousands of tokens plus general knowledge. The temperature sampling mechanism you learned here is exactly the same one they use — this project has taught you the engine in miniature.
  • Responsible use at TecnoMarket: the model produces drafts; a human reviews, corrects the technical specs and approves before publishing. It is the same philosophy as the review queue from 03-04: automate the mechanical, supervise what reaches the customer.

Common Mistakes and Tips

  • Forgetting return_sequences=True: the LSTM would return only the last state and the shapes wouldn't match the 80-position target. Revisit 04-02 if the shape error throws you.
  • Putting a softmax in the last layer and also from_logits=True: a silent double softmax that degrades training. Pick one of the two; here, logits.
  • Evaluating generation by the loss alone: a low loss on a repetitive corpus can mean pure memorization. Generate and read: qualitative inspection is part of the evaluation.
  • Using a high temperature "for more creativity" in production: in commercial copy, invented words destroy customer trust. Start at 0.5 and adjust with examples in front of you.
  • Windows longer than the descriptions for no reason: they stretch training without improving a corpus of short sentences. Fit WINDOW to the real text.

Exercises

  1. Train the same model with LSTM(128) and with LSTM(512) and compare final loss, time per epoch and the quality of three samples at T=0.7. Where do the diminishing returns set in?
  2. Add to the corpus 10% of descriptions with a new format (for example, starting with "deal: "). Retrain and check how often and how faithfully the model generates that format.
  3. Implement generate_batch(seed, k, temperature) returning k distinct drafts and discarding any that contain words outside a dictionary built from the corpus (automatic quality control ahead of the human review).

Solutions

  1. With 128 units the loss typically stays 0.1-0.2 higher and more spelling stumbles appear; with 512 the loss barely improves (the corpus is simple) and the time per epoch multiplies by ~3. On this corpus, 256 is the sweet spot: extra capacity adds nothing because there is no more complexity to learn.
  2. It's enough to add sentences.append(f"deal: {a} {p} with {b1}. {c}\n") inside an extra loop (~300 sentences). After retraining, seeds beginning with "deal: " complete the format consistently; without that seed, the pattern shows up spontaneously in around 10% of the samples — the model reproduces the corpus frequencies.
  3. Build the dictionary with valid_words = set(corpus.split()); generate in a loop with generate, split each draft with .split() and accept it only if all(w.strip('.') in valid_words for w in draft.split()). Return the first k accepted (with a cap on attempts). This cheap filter removes most high-temperature artifacts before the seller ever sees them.

Conclusion

Second project delivered: a draft generator that learns TecnoMarket's style character by character, with the sliding windows from 04-04, the Embedding+LSTM tandem from 04-03 and a temperature sampling you now know how to read and tune. Just as valuable is what you know it does not do: long-range coherence and factual accuracy are left to the transformers of 05-05, and every generated text goes through human review. In the next lesson we change problems again and keep the promise from 05-02: using an autoencoder's reconstruction error to detect fraudulent transactions at TecnoMarket — our first project with genuinely imbalanced data.

© Copyright 2026. All rights reserved