We closed module 1 with the list of nine candidate use cases for NovaMarket and with the promise of building the conceptual foundations. We start with the most important of all: the notion of an intelligent agent. Almost everything you will see in the rest of the course (a recommender, a route planner, a customer-service assistant, even a language model) can be described as an agent that perceives its environment, decides and acts. Learning to describe a system in those terms, and to characterise the environment in which it operates, is what later allows you to choose the right technique: a problem in a deterministic, fully observable environment is not solved the same way as one in a stochastic, partially observable environment. In this lesson we will define what an agent is and its perceive-decide-act cycle, what it means for it to act rationally, how to describe any system with the PEAS scheme, what properties environments have, what types of agent exist and, as a bridge to module 3, how a problem is formulated so that an agent can solve it. All of it applied to three NovaMarket systems: the customer-service assistant, the recommender and the route planner.

Contents

  1. What an intelligent agent is
  2. The perceive-decide-act cycle
  3. Rationality and performance measure
  4. The PEAS description applied to NovaMarket
  5. Properties of environments
  6. Types of agent
  7. Problem formulation: a bridge to module 3
  8. Example in Python: from the reflex agent to the utility-based agent

  1. What an intelligent agent is

In lesson 01-02 we adopted the view of AI as the construction of systems that act rationally. The concept that makes that definition operational is the agent. An agent is anything that:

  • perceives its environment through sensors, and
  • acts on that environment through actuators.

The definition is deliberately broad. A thermostat is an agent (sensor: thermometer; actuator: boiler switch). A warehouse robot is an agent (sensors: cameras, barcode readers, wheel encoders; actuators: motors, gripper). And a program that answers emails is one too, even though it has no body: its sensors are the data inputs it receives (the text of the email, the orders database) and its actuators are the outputs it produces (a reply email, a call to the warehouse API). Purely software agents are sometimes called software agents or softbots.

Basic vocabulary we will use throughout the course:

Term Definition Example in NovaMarket's customer-service assistant
Percept What the agent receives from the environment at a given instant The customer's message: "Where is my order 48213?"
Percept sequence Everything the agent has perceived so far The previous messages of the same conversation and the data looked up
Action What the agent does to the environment Reply with the shipping status; open an incident; hand the case over to a person
Sensor Mechanism through which the percept arrives The chat channel, the query to the orders database
Actuator Mechanism through which the action is executed Sending the message to the chat, the API that creates the incident
Agent function The mapping from percept sequences to actions "If the customer asks about an order and gives the number, look up the status and reply with it"
Agent program The concrete implementation of that function The Python code (or the model) that computes it

Note the distinction between the agent's function and program. The function is an abstract description ("what it should do for each percept history"); the program is how we achieve it. The rule-based chatbot NovaMarket abandoned in 2015 and a modern assistant based on a language model can aspire to the same function; what changes radically is the program. And here the central idea of the course appears: in many cases the function is too complex to write by hand rule by rule (which is what failed in 2015), so it is learned from data, exactly the paradigm shift we saw in 01-02.

  1. The perceive-decide-act cycle

Every agent, however simple or sophisticated, repeats the same loop:

flowchart LR
    E[Environment] -- percept --> S[Sensors]
    S --> D{Decide<br/>agent function}
    D --> A[Actuators]
    A -- action --> E
  1. Perceive: the sensors capture the state (or part of the state) of the environment.
  2. Decide: the agent program, from the current percept and, depending on the type of agent, from what it remembers, chooses an action.
  3. Act: the actuators execute the action, which modifies the environment.
  4. The environment changes (because of the agent's action and, often, on its own), and the cycle starts again.

Let us see it with the recommender we built at the end of 01-03. On each visit by a customer to a product page:

  • Percept: the product they are viewing (and, in more advanced versions, their history and their cart).
  • Decision: the recommend function, which consults the co-purchase table and picks the products with the most matches.
  • Action: showing those products in the "also bought" box.
  • Change in the environment: the customer clicks (or not), adds to the cart (or not); if we record that, the next turn of the loop will have more data.

This last point matters: when the agent's actions influence future percepts (what you recommend today changes what is bought tomorrow), the agent and the environment form a closed system, and that has consequences we will see when we talk about sequential environments.

  1. Rationality and performance measure

An agent that perceives and acts is not necessarily intelligent: the ELIZA of 01-01 perceived and acted. What distinguishes a rational agent is that it chooses, at each moment, the action that maximises the expected value of its performance measure, given the evidence provided by its percepts and whatever built-in knowledge it has.

Let us unpack the sentence, because every piece matters:

  • Performance measure: the criterion by which the agent's success is evaluated. It is set by the designer, not the agent, and it must always express what we want to achieve in the environment, not how we think the agent should behave. It is the most delicate part of the design.
  • Expected value: since the agent rarely has certainty about the consequences of its actions, they are evaluated on average, weighted by their probability. A recommender does not know whether the customer will buy, but it can estimate that the probability is higher with one product than with another.
  • Given the evidence and knowledge: rationality is judged with what the agent knows, not with what is known after the fact. A route planner that picks the motorway and runs into an unforeseen accident has not been irrational; it would have been irrational to ignore a traffic warning it already had.

Rationality is not the same as omniscience (knowing the actual outcome of every action) nor perfection (always getting it right). It is "doing the best possible with what is known". And a rational agent must also gather information (ask the customer for the order number before replying) and learn from experience when the environment allows it. An agent whose behaviour depends only on what its designers programmed has little autonomy; one that corrects its decisions with what it perceives has more.

The performance measure matters more than it seems

Choosing the performance measure badly produces agents that "cheat": they optimise exactly what they were asked for, which was not what we wanted. Diego describes it with a real NovaMarket anecdote: for a while the customer-service team was paid an incentive for "conversations closed per hour"; the result was that many conversations were closed without solving the problem. An artificial agent does exactly the same, only faster and without feeling embarrassed. Compare these possible measures for the assistant:

Performance measure Behaviour it rewards Problem
Number of conversations closed Closing fast Closes without resolving
Average response time Replying fast Replies with anything
Percentage of queries resolved without human intervention Avoiding escalation Never hands over cases that do require it
Customer satisfaction (survey) + cost per conversation + percentage of correct escalations Resolving well, cheaply and knowing when to ask for help Harder to measure, but it is what we want

The last row is more complex, but it is the one that reflects the real goal. Practical rule: define the performance measure in terms of the desired outcome in the environment, combine several criteria if needed, and always ask yourself "what would an agent do if it wanted to maximise this and cared about nothing else?". We will return to this idea in 02-04 (ethics) and in 04-05 (model evaluation).

  1. The PEAS description applied to NovaMarket

To design an agent it helps to start by describing its task environment with four elements, remembered by the acronym PEAS:

  • Performance (performance measure): how is success measured?
  • Environment: in what world does it act? What is out there?
  • Actuators: what can it do?
  • Sensors: what can it perceive?

Let us apply it to three of NovaMarket's use cases. Marta and Diego filled in this table in a meeting, and it is a good exercise in how PEAS forces you to ask the right questions before writing a line of code:

Element Customer-service assistant (case 7) Recommender (case 1) Delivery route planner (case 5)
Performance Customer satisfaction, percentage of queries resolved correctly, cost per conversation, well-judged escalations to a human (neither too many nor too few) Additional revenue from recommendations, click-through and purchase rate, diversity of what is recommended, not annoying (irrelevant recommendations are penalised) Total kilometres and hours, deliveries within the promised slot, fuel cost, compliance with drivers' working hours
Environment Customers with questions of all kinds (shipping, returns, warranties, invoices), orders and incidents database, company policies, other channels (email, phone) Catalogue of ~12,000 products, ~3,000 orders/day, browsing and purchase history, stock, active campaigns Street network of each city, traffic, ~3,000 orders/day split between own delivery and couriers, vans and drivers, promised time slots, Zaragoza and Getafe warehouses
Actuators Send messages to the chat, look up and create incidents, start a return, hand over to a person Show N products on the product page, in the cart or in an email; sort the listing Assign orders to vans, set the order of stops, send the route to the driver, replan during the day
Sensors Customer's text, customer/order identifier, order status in the database, conversation history Product on screen, customer identifier, cart, history, time, device List of the day's orders with addresses and slots, GPS position of the vans, real-time traffic, driver's incidents

Observations that come out of the table:

  • In all three cases the performance measure is composite: it is never a single figure. Diego insisted on including cost in all three.
  • The route planner's environment is physical and changes on its own (traffic); the recommender's is purely digital. That will affect the technique.
  • The assistant's sensors include "conversation history": that is the clue that the assistant will need memory (we will see it in the types of agent).

When in module 8 we tackle how to develop an AI project, the PEAS description will be one of the first deliverables.

  1. Properties of environments

Not all environments are equally hard. They are characterised along six dimensions, and the combination determines which techniques are viable:

Dimension "Easy" option "Hard" option Key question
Observability Fully observable: the sensors give the complete relevant state Partially observable: there is hidden or noisy information Do I see everything I need in order to decide?
Determinism Deterministic: the next state is fixed by the current state and the action Stochastic: there is chance or factors I do not control If I repeat the same action in the same situation, does the same thing always happen?
Episodic / sequential Episodic: each decision is independent of the previous ones Sequential: today's decisions condition tomorrow's Does my current action affect my future decisions?
Static / dynamic Static: the environment does not change while the agent deliberates Dynamic: it changes even if the agent does nothing Can I think calmly, or does the world keep moving?
Discrete / continuous Discrete: a finite number of clearly separated states, actions and percepts Continuous: values that vary smoothly (positions, times, quantities) Can I enumerate the options?
Agents Single-agent: only my agent acts Multi-agent: there are other agents (cooperative or competitive) whose decisions matter Is anyone else deciding in this world?

Some classic examples to calibrate intuition: a crossword is fully observable, deterministic, sequential, static, discrete and single-agent (the easiest case). Chess is the same except that it is multi-agent (competitive). An autonomous taxi is partially observable, stochastic, sequential, dynamic, continuous and multi-agent (the hardest case, which is why it took so long). Poker adds to chess partial observability (you cannot see your opponent's cards) and the randomness of the deal.

Let us now classify NovaMarket's three systems:

Dimension Customer-service assistant Recommender Route planner
Observability Partial: the customer does not always say what is wrong; internal data may be out of date Partial: we do not know the customer's intentions or budget; we only see their behaviour Partial: future traffic and incidents are unknown; GPS has error
Determinism Stochastic: the same reply produces different reactions in different customers Stochastic: the same recommendation sometimes converts and sometimes does not Stochastic: variable travel times, recipient not at home
Episodic / sequential Sequential within a conversation (each message depends on the previous ones) Almost episodic on each visit, but sequential in the long run (what you recommend changes what is bought and, therefore, future data) Sequential: each stop conditions the next ones
Static / dynamic Dynamic (the customer types while the agent "thinks"; in practice, semi-dynamic) Static on each request (the catalogue does not change in the milliseconds it takes to respond) Dynamic: traffic and new orders change while planning
Discrete / continuous Discrete in actions (replies, hand-overs), although text is an enormous space Discrete: finite catalogue Continuous in positions and times; discrete in the choice of stop order
Agents Multi-agent: the customer is another agent with their own goals Single-agent as a first approximation (although it competes with the customer's own browsing and with other stores) Multi-agent: other vehicles, other drivers, couriers

Practical conclusions of the classification:

  • None of the three is the "easy case". In the real world almost no problem is; fully observable, deterministic environments appear mostly in games and puzzles, which is where module 3 will introduce search algorithms before complicating them.
  • Partial observability and stochasticity are what push towards machine learning (module 4) and reasoning under uncertainty (06-03): if you cannot see everything or predict with certainty, you need to estimate.
  • The dynamic nature of the route planner requires the agent to be fast and able to replan: a perfect algorithm that takes three hours is no use.
  • The recommender looks the simplest (static, discrete, almost episodic), which is why NovaMarket was right to choose it as its first project.

  1. Types of agent

Agents are organised on a scale of increasing sophistication according to what information they use to decide. Each level includes the previous one:

flowchart TB
    A[Simple reflex agent<br/>current percept → rule → action] --> B[Model-based reflex agent<br/>adds internal state: what I remember about the world]
    B --> C[Goal-based agent<br/>adds goals: what I want to achieve]
    C --> D[Utility-based agent<br/>adds preferences: how much I value each outcome]
    D --> E[Learning agent<br/>adds improvement with experience]

6.1 Simple reflex agent

It decides by looking only at the current percept, through condition-action rules: "if the message contains 'where is my order', reply with the shipping status". It is fast and easy to understand, but:

  • It only works well if the environment is fully observable: if the current percept is not enough to decide, it gets it wrong.
  • It has no memory: if the customer gave the order number in the previous message, this agent no longer knows it.
  • It can get stuck in loops (repeating the same question over and over).

NovaMarket's 2015 chatbot was exactly this, and its failure was largely due to operating in a partially observable, sequential environment with an architecture designed for fully observable, episodic environments.

6.2 Model-based reflex agent

It adds an internal state: a representation of what the agent believes is going on in the world, maintained with two kinds of knowledge: how the world evolves on its own and what effect the agent's own actions have. In the assistant, the internal state is the context of the conversation (order number already identified, whether a return has already been offered, how many times the customer has asked the same thing). It still decides with rules, but now the rules can consult memory. This solves partial observability within a conversation.

6.3 Goal-based agent

Rules say "what to do", but not "what for". A goal-based agent has an explicit description of the goal (deliver all of the day's orders within their slot) and chooses actions by reasoning about their consequences: "if I go first to Delicias and then to Actur, do I reach the goal?". This is more flexible: if the goal changes (for example, prioritising an urgent order), the rules do not need rewriting; changing the goal is enough. The trade-off is that deciding requires searching or planning among sequences of actions, which is what module 3 studies.

6.4 Utility-based agent

A goal says whether a state is good or bad (binary). But often there are many ways of reaching the goal and some are better than others: getting to every delivery is the goal, but doing it in 6 hours is better than in 8, and bothering a customer with a phone call has a cost. A utility function assigns to each state (or outcome) a number expressing how much the agent values it, and the rational agent chooses the action that maximises expected utility. It is the direct translation of the definition of rationality in section 3: the agent's internal utility should coincide with the external performance measure. In the recommender, the utility of recommending a product could combine the probability of purchase, the margin and a penalty for always repeating the same thing.

6.5 Learning agent

Any of the above can additionally learn: improve its decision function from experience. Conceptually it has four components: the performance element (the agent as we have described it), the learning element (which modifies it), the critic (which compares the results with the performance measure and informs the learning element) and the problem generator (which proposes exploratory actions to discover new things, such as occasionally recommending a rarely seen product to see whether it works). Module 4 devotes six lessons to how learning happens; here it is enough to place it on the map. Our learn_threshold from 01-02, which replaced Diego's €300 with €120, was a rudimentary learning element.

Summary table:

Type of agent What it uses to decide Requires of the environment NovaMarket example
Simple reflex Current percept + rules Fully observable, episodic Rule "amount > €300 → review"
Model-based reflex Percept + internal state + rules Tolerates partial observability Assistant that remembers the order number
Goal-based State + goals + reasoning about consequences Needs to be able to predict effects Route planner: "deliver everything within its slot"
Utility-based State + graded preferences Same, and allows comparing alternatives Recommender that weighs purchase probability and margin
Learning Any of the above + feedback Needs data or experience Recommender that retrains with new orders

  1. Problem formulation: a bridge to module 3

When a goal-based agent has to decide which sequence of actions leads it to the goal, the first step is to formulate the problem precisely. The standard formulation has five components:

  1. Initial state: where the agent starts. Example: the van is at the Getafe warehouse with 40 orders loaded.
  2. Actions: what it can do in each state. Example: travel to any of the pending stops.
  3. Transition model: which state results from applying an action in a state. Example: "go to stop 7" leads to a state with the van at stop 7 and 39 pending orders.
  4. Goal test: how we recognise that the goal has been reached. Example: no pending orders remain and the van has returned to the warehouse.
  5. Path cost: a number measuring the cost of a sequence of actions. Example: accumulated kilometres or minutes.

With these five elements, solving the problem consists of finding a sequence of actions (a path in the state space) that leads from the initial state to a goal state, ideally at minimum cost. We are not going to solve anything here: the algorithms that perform that search (breadth-first, depth-first, uniform-cost, A*) are the content of lesson 03-02, and the techniques for when the space is too large to explore in full are the content of 03-04. What is worth retaining now are two ideas:

  • The formulation is an abstraction: we remove all the irrelevant details (the colour of the van, the music the driver listens to) and keep what affects the decision. Choosing the level of abstraction well is half the solution.
  • The cost is the "search" version of the performance measure: if you define the cost badly (kilometres only, forgetting time slots), the agent will find optimal paths for the wrong problem.

A second, more discrete example: diagnosing an incident (case 9). Initial state: open incident with known symptoms. Actions: perform a check (did the parcel arrive? is it damaged? does the reference match?). Transition: each check reveals information and narrows down the possible causes. Goal: cause identified. Cost: number of checks or operator time. We will return to this case in module 6 with expert systems.

  1. Example in Python: from the reflex agent to the utility-based agent

We are going to build a miniature customer-service assistant and make it evolve through three of the levels in section 6. We will use only standard Python. The goal is not a useful assistant, but to see in code what distinguishes each type of agent.

8.1 Simple reflex agent

class ReflexAgent:
    """Simple reflex agent: decides by looking only at the current percept."""

    def __init__(self, rules):
        # rules: list of (condition, action) tuples.
        # condition is a function that receives the percept and returns True/False.
        self.rules = rules

    def perceive(self, message):
        # The "sensor" normalises the text: lowercase and no surplus whitespace.
        return message.lower().strip()

    def decide(self, percept):
        # Walk through the rules in order and execute the first one that holds.
        for condition, action in self.rules:
            if condition(percept):
                return action
        return "hand_over_to_human"        # default action

    def act(self, action):
        # The "actuator": here it only prints; in production it would send the message.
        responses = {
            "report_shipping_status": "Let me check your shipment status. Could you give me the order number?",
            "explain_return":         "You can return it within 30 days from your customer area.",
            "hand_over_to_human":     "I'll pass you on to a member of the team.",
        }
        print(f"[{action}] {responses[action]}")

    def step(self, message):
        # Full cycle: perceive -> decide -> act
        percept = self.perceive(message)
        action = self.decide(percept)
        self.act(action)


rules = [
    (lambda p: "order" in p or "shipping" in p,   "report_shipping_status"),
    (lambda p: "return" in p or "refund" in p,    "explain_return"),
]

assistant = ReflexAgent(rules)
assistant.step("Hello, where is my order?")
assistant.step("It's 48213")
assistant.step("I want to return a product")

Output:

[report_shipping_status] Let me check your shipment status. Could you give me the order number?
[hand_over_to_human] I'll pass you on to a member of the team.
[explain_return] You can return it within 30 days from your customer area.

Line-by-line explanation:

  • rules is a list of (condition, action) pairs. Each condition is an anonymous function (lambda) that receives the percept and returns True if the rule applies. This is the essence of the reflex agent: a percept → action table.
  • perceive is the sensor: it turns the raw message into something comparable (lowercase, no surplus whitespace). Notice how fragile it is: if the first rule looked for the full phrase "where is my order", the real message "Hello, where's my order?" (with a contraction) would not trigger it and the agent would hand over to a human for no reason. That is why the rule looks only for the word "order". Reflex agents based on text strings fail on trivial variations (contractions, punctuation, synonyms, typos), and extending the rules to cover them all is an endless race: it is one of the reasons natural language processing ended up turning to learning (modules 4 and 5).
  • The second message, "It's 48213", is the reply to the agent's question. But the reflex agent does not remember that it has just asked for an order number, so it does not recognise the reply and hands over to a human. This is the structural failure of the reflex agent in a sequential environment, and exactly the kind of conversation that broke the 2015 chatbot.

8.2 Model-based reflex agent: adding internal state

import re

class ModelBasedAgent(ReflexAgent):
    """Model-based reflex agent: keeps an internal state of the conversation."""

    def __init__(self, rules):
        super().__init__(rules)
        self.state = {"awaiting_order": False, "order": None}

    def update_state(self, percept):
        # World model: if we were waiting for a number and one arrives, store it.
        number = re.search(r"\b\d{4,6}\b", percept)
        if self.state["awaiting_order"] and number:
            self.state["order"] = number.group()
            self.state["awaiting_order"] = False

    def decide(self, percept):
        self.update_state(percept)
        if self.state["order"] and "order_just_identified" not in self.state:
            self.state["order_just_identified"] = True
            return "report_order_status"
        action = super().decide(percept)
        if action == "report_shipping_status":
            self.state["awaiting_order"] = True   # effect of our own action
        return action

    def act(self, action):
        if action == "report_order_status":
            print(f"[{action}] Your order {self.state['order']} leaves the Zaragoza warehouse today.")
        else:
            super().act(action)


rules = [
    (lambda p: "order" in p or "shipping" in p,   "report_shipping_status"),
    (lambda p: "return" in p or "refund" in p,    "explain_return"),
]

assistant = ModelBasedAgent(rules)
assistant.step("Hello, where is my order?")
assistant.step("It's 48213")

Output:

[report_shipping_status] Let me check your shipment status. Could you give me the order number?
[report_order_status] Your order 48213 leaves the Zaragoza warehouse today.

Explanation:

  • self.state is the internal state: what the agent believes about the world. Here it only stores whether it is waiting for an order number and which one it is.
  • update_state implements the "model": it uses a regular expression (\b\d{4,6}\b, an isolated number of 4 to 6 digits) to detect the order number, but it only interprets it as such if it was waiting for it. It is the combination of current percept and memory.
  • When the agent asks for the number (report_shipping_status), it records the effect of its own action (awaiting_order = True). This is the second kind of knowledge of the model-based agent: what effect my actions have on the world (in this case, on the conversation).
  • The result: the same conversation that broke the reflex agent now flows. With very little code we have gone from an agent unsuited to sequential environments to one that tolerates them. In a real assistant the state would be much richer (detected intent, products mentioned, customer's tone), and with language models that "state" is largely the conversation history itself (we will see it in 05-05).

8.3 Utility-based agent: choosing between alternatives

Suppose now that, faced with a complaint about a delay, the assistant has three possible actions and wants to pick the one with the highest expected utility. For each action we estimate the probability that the customer ends up satisfied and the cost to NovaMarket:

def expected_utility(prob_satisfaction, cost_euros, satisfied_customer_value=15.0):
    """Utility = expected benefit of satisfaction - cost of the action."""
    return prob_satisfaction * satisfied_customer_value - cost_euros

# Estimates for a customer with a 2-day delay on an 80 EUR order
options = {
    "apology_and_new_date":       {"prob": 0.55, "cost": 0.0},
    "coupon_5_euros":             {"prob": 0.80, "cost": 5.0},
    "urgent_reshipment":          {"prob": 0.90, "cost": 12.0},
}

def choose_action(options):
    best_action, best_utility = None, float("-inf")
    for action, data in options.items():
        u = expected_utility(data["prob"], data["cost"])
        print(f"{action:28s} expected utility = {u:6.2f}")
        if u > best_utility:
            best_action, best_utility = action, u
    return best_action

print("Chosen action:", choose_action(options))

Output:

apology_and_new_date         expected utility =   8.25
coupon_5_euros               expected utility =   7.00
urgent_reshipment            expected utility =   1.50
Chosen action: apology_and_new_date

Explanation:

  • expected_utility is the utility function: it translates into a number what NovaMarket values (a satisfied customer is "worth" €15 in future loyalty, a figure Marta estimated from the average repeat purchase) minus what the action costs.
  • Each action has an uncertain outcome (the customer may or may not end up satisfied), which is why we multiply by the probability: it is the "expectation" in the definition of rationality.
  • choose_action walks through the options and keeps the one with the highest expected utility. Here the apology wins: the coupon raises satisfaction a lot but its cost penalises it. Try changing satisfied_customer_value to 40 (a very valuable customer) and you will see the decision switch to coupon_5_euros. That sensitivity is a virtue: the policy adapts by changing preferences, not by rewriting rules.
  • Where do the probabilities 0.55, 0.80 and 0.90 come from? Here they are written by hand. In a real system they would be learned from historical incident data and surveys: that is the step to the learning agent, and the content of module 4.

With these three snippets you have seen the whole ladder: rules on the current percept, rules plus memory, and choice by expected utility. A modern assistant combines all three levels and adds learning.

Common Mistakes and Tips

  • Confusing the performance measure with the desired behaviour. "I want the assistant to reply fast" is a behaviour; "I want satisfied customers at low cost" is an outcome. Measure outcomes, and always ask yourself how an agent blindly optimising your metric would cheat.
  • Judging rationality after the fact. A decision turning out badly does not mean it was irrational. Evaluate with the information the agent had at the moment of deciding.
  • Choosing the architecture without classifying the environment. The 2015 chatbot failed because it used a reflex agent in a partially observable, sequential environment. Fill in the environment-properties table before deciding on the technique.
  • Filling in PEAS vaguely. "Environment: the internet" is no use. Be concrete: which data, which users, which constraints. The quality of the PEAS anticipates the quality of the project.
  • Over-engineering the agent. Not everything needs utility and learning. If the environment is genuinely simple (a clear business rule), a well-made reflex agent is cheaper, faster and more explainable. Sophistication is justified by the environment, not by fashion.
  • Forgetting the cost of actions. Diego is right to insist: an action with a high probability of success but a high cost can have worse utility than a modest, free one.
  • Tip: whenever you are shown any "intelligent" system, try to fill in its PEAS mentally and classify its environment. Within a minute you will know whether the system suits the problem and what can go wrong.

Exercises

Exercise 1: PEAS and environment of the demand-forecasting system

NovaMarket has decided to also start with demand forecasting (case 2): estimating how many units of each product will sell next week in order to place orders with suppliers and move stock between Zaragoza and Getafe. Write its PEAS description and classify the environment along the six dimensions, justifying each one in a sentence.

Exercise 2: Performance measures with a catch

For the route planner, state what undesirable behaviour each of these performance measures would produce if it were the only one: (a) minimise total kilometres; (b) maximise deliveries completed per hour; (c) minimise customer complaints. Then propose a reasonable composite measure.

Exercise 3: Extending the model-based agent

Modify ModelBasedAgent so that, if the customer repeats three times a query the agent does not know how to resolve (action hand_over_to_human), the agent detects it and adds the text "I can see I'm not managing to help you" to the hand-over message. You will need a counter in self.state. Explain which property of the environment makes this memory necessary.

Solutions

Solution 1.

Element Demand forecasting
Performance Forecast error (difference between units forecast and sold), stock-outs avoided, overstock cost, cost of transfers between warehouses
Environment Catalogue of ~12,000 products, history of ~3,000 orders/day, seasonality, campaigns and promotions, own and competitors' prices, supplier lead times
Actuators Issue a forecast per product and warehouse; suggest supplier orders and transfers between warehouses
Sensors Historical orders.csv, products.csv (category, price, stock), campaign calendar, public holidays

Environment: partially observable (we do not see customers' intentions or competitors' actions); stochastic (demand has randomness); sequential (a wrong forecast today leaves stock that conditions the following month; moreover, recommendations and promotions themselves change demand); static at the moment of computing (it is done at night with closed data; the world does not change during the computation); discrete in units, although sales time series are usually treated as continuous; single-agent as a first approximation (although competitors and suppliers are agents that have an influence). It is a typical prediction problem for supervised learning on time series, which will be picked up again in module 4.

Solution 2.

  • (a) Minimise kilometres: the agent will group nearby stops while ignoring the promised time slots; it will deliver late to whoever is "on the way" later on. It could also leave orders undelivered if that reduces kilometres (the minimum is not going out at all!), unless it is forced to deliver everything.
  • (b) Maximise deliveries per hour: it will reward easy deliveries (accessible entrances, customer at home) and postpone the difficult ones indefinitely; it incentivises "marking as delivered" without checking.
  • (c) Minimise complaints: it can lead to never promising slots (there is no breach if there is no promise) or to allocating disproportionate resources to the customers who complain most.
  • Composite measure: percentage of deliveries within the promised slot (delivering everything assigned is mandatory) + total cost (kilometres × cost per km + hours × cost per hour) + penalty for breaching working hours + customer satisfaction in delivery surveys. With weights agreed with Diego.

Solution 3.

class PatientAgent(ModelBasedAgent):
    def __init__(self, rules):
        super().__init__(rules)
        self.state["consecutive_failures"] = 0

    def decide(self, percept):
        action = super().decide(percept)
        if action == "hand_over_to_human":
            self.state["consecutive_failures"] += 1
        else:
            self.state["consecutive_failures"] = 0
        return action

    def act(self, action):
        if action == "hand_over_to_human" and self.state["consecutive_failures"] >= 3:
            print("[hand_over_to_human] I can see I'm not managing to help you. I'll pass you on to a member of the team.")
        else:
            super().act(action)

assistant = PatientAgent(rules)
for _ in range(3):
    assistant.step("I need the invoice as a PDF")

The third message produces the extended text. The memory is necessary because the environment is sequential: the right decision in the face of "I need the invoice" depends on how many times it has already been said, something the current percept on its own does not contain (partial observability of the conversation if the history is not stored). Note: in this simple example the agent hands over to a human from the very first time; in a real assistant the hand-over is an expensive action (it occupies a person) and the counter would serve to decide when to escalate instead of continuing to try.

Conclusion

In this lesson we have defined the concept that runs through the whole course: the intelligent agent, which perceives its environment through sensors and acts through actuators in a continuous perceive-decide-act cycle. We have pinned down what it means for an agent to be rational (choosing the action that maximises the expected performance measure with the available information) and we have seen that defining that measure well is the most delicate design decision. We have learned to describe any system with the PEAS scheme and to classify its environment along six dimensions (observability, determinism, episodic/sequential, static/dynamic, discrete/continuous, single/multi-agent), applying it to NovaMarket's assistant, recommender and route planner. We have climbed the ladder of types of agent (simple reflex, model-based, goal-based, utility-based, learning) and seen it in code with the evolution from ReflexAgent to ModelBasedAgent and on to choice by expected utility. Finally, we have presented problem formulation (initial state, actions, transition, goal, cost), which module 3 will use to solve problems with search algorithms.

In the next lesson, Types of Artificial Intelligence, we will change scale: instead of looking inside an agent, we will classify AI systems as a whole according to their capability (narrow, general, superintelligence), their functionality (reactive, limited memory, theory of mind, self-awareness) and other useful dichotomies, and we will place NovaMarket's nine use cases on that map. You will see that the "reactive" and "memory-based" agents we have just programmed have their direct counterpart in that classification.

Fundamentals of Artificial Intelligence (AI)

Module 1: Introduction to Artificial Intelligence

Module 2: Basic Principles of AI

Module 3: Algorithms in AI

Module 4: Machine Learning

Module 5: Neural Networks and Deep Learning

Module 6: Logic and Expert Systems

Module 7: Tools and Programming Languages in AI

Module 8: Projects and Case Studies

Module 9: Exercises and Practice

Module 10: Additional Resources

© Copyright 2026. All rights reserved