In the previous lesson, 06-01, we closed the technical write-up of the Aroma Store API with a list of certainties: offset pagination, exact counters, authorisation based on ownership of the resource, collections wrapped in data with their total, and real time as an accessory of the internal panel. All those decisions were right for that domain. Now we change domain without changing architectural style: we design CafeSocial, the tasters' network Aroma Store wants to launch, and one by one we are going to watch those certainties fall. Not because they were wrong, but because they were tied to a context that no longer exists. That contrast —the same REST discipline producing opposite designs— is the real content of this lesson.

Contents

  1. The scenario: CafeSocial, requirements and consumers
  2. Modelling the relationship graph with REST
  3. The timeline: fan-out, derived resource and opaque cursor
  4. Volume and scale: approximate counters, caching and queues
  5. User-generated content: upload, moderation and reports
  6. Privacy and resource-level authorisation
  7. Notifications and real time
  8. Aroma Store versus CafeSocial, decision by decision
  9. Why GraphQL is defensible here
  10. Common mistakes and tips
  11. Exercises and solutions

  1. The scenario: CafeSocial, requirements and consumers

CafeSocial is a vertical social network for coffee tasters. A user publishes a tasting: a photo of the coffee, the variety, the origin, the brewing method and a score from 0 to 100. Other users comment, give "likes", follow the tasters that interest them and receive a timeline with the posts of the people they follow. There are hashtags (#geisha, #v60), notifications and direct messages.

The identifiers keep the course's convention: usr_10, usr_77, pst_2100, cmt_310, ntf_55. The JSON is still camelCase, the errors still follow the {"error": {"code","message","details":[]}} catalogue and the version is still in the path (/v1). The domain changes, the conventions do not: that is precisely what makes comparison possible.

Requirements that shape the design

Requirement Target figure (fictional) Design consequence
Read/write ratio ~500:1 Everything is optimised for reads; writes can be more expensive
Timeline latency p95 < 150 ms Precompute, do not compute during the request
Size of the graph 2 M users, 180 M follow edges The relationship is a first-class resource, not a field
Follower distribution Long tail: 0.01 % exceed 100,000 followers A single distribution algorithm will not do
User content 40,000 posts with a photo per day Moderation and storage outside the API
Consistency Eventual is acceptable for the timeline and counters It can be decoupled with queues

Consumers

  • Mobile app (iOS/Android): the main consumer, with limited bandwidth and battery. Every byte and every round trip hurts.
  • Public web: profiles and posts indexable by search engines when they are public.
  • Internal moderation panel: few users, broad permissions, needs work queues.
  • Integration with Aroma Store: when a post mentions a coffee from the catalogue, the card links to /v1/coffees/cof_001 of the store's API. They are two separate APIs linked by hypermedia, not one.

Unlike the store, here there is no money in the request. Nothing demands the transactional exactness that in 06-01 forced us to solve overselling inside the UPDATE. That freedom is what makes almost everything that follows possible.


  1. Modelling the relationship graph with REST

At Aroma Store almost everything was a simple containment relationship: an order belongs to a customer, an item belongs to an order. Here the relationship is the interesting entity.

2.1 The two collections of the graph

GET /v1/users/usr_10/followers?limit=20 HTTP/1.1
GET /v1/users/usr_10/following?limit=20 HTTP/1.1

They are two views of the same edge, from its two ends. Neither is "the right one": the app needs both, and with very different cardinalities (a user follows 300 people but may have 300,000 followers).

2.2 The relationship as a resource: PUT versus POST /follow

The temptation is to invent a verb: POST /v1/follow with {"followedId": "usr_77"}. It works, and it is exactly what in 02-02 we called turning an action into a resource without need. The alternative is to treat the edge as an addressable resource:

PUT /v1/users/usr_10/following/usr_77 HTTP/1.1
Authorization: Bearer <usr_10's token>

HTTP/1.1 204 No Content
Aroma-Graph-State: following
DELETE /v1/users/usr_10/following/usr_77 HTTP/1.1

HTTP/1.1 204 No Content

The advantages, going back to the methods table of 02-03:

Aspect PUT /following/{id} POST /follow
Idempotency Yes: hitting "follow" five times leaves the same state Not guaranteed; you have to deduplicate
Retry after a timeout Safe by definition Requires an Idempotency-Key
Querying the state A GET on the same URI (204/404) You have to invent another endpoint
Undoing A DELETE on the same URI Another verb: POST /unfollow
Cacheable / linkable Yes, it has its own URI No

In a mobile app on an unstable network, idempotency is not theoretical elegance: it is the difference between a button that gets stuck halfway and one that does not. The client can retry the PUT without thinking.

GET /v1/users/usr_10/following/usr_77 returns 204 if the relationship exists and 404 if it does not, which gives the client a cheap check for rendering the button.

2.3 "Likes": user in the path or in the token?

Both forms are legitimate and it is worth understanding the trade-off:

PUT /v1/posts/pst_2100/likes/usr_10     # A: explicit subject
PUT /v1/posts/pst_2100/likes            # B: subject implicit in the token
Criterion A (explicit) B (implicit)
Self-descriptive URI Yes No: the same URI means different things depending on who calls
Impersonation risk You have to validate that the path matches the token Impossible by construction
Acting on someone else's behalf (admin, import) Straightforward Needs a separate header or endpoint
Listing who liked it GET /posts/pst_2100/likes is natural Just as natural
Caching Cacheable by URI Needs Vary: Authorization

At CafeSocial we chose A for the follow graph and for "likes", for consistency and because the moderation panel needs to be able to remove another user's "like". The rule we applied: if any legitimate consumer can act on a third party's relationship, the subject goes in the path. When the path and the token do not match and the caller lacks the appropriate scope, we respond 403 with insufficient_permissions.

2.4 When the relationship needs a body of its own

An edge with attributes stops being a simple "it exists or it does not". Following someone can carry preferences:

PUT /v1/users/usr_10/following/usr_77 HTTP/1.1
Content-Type: application/json

{"notifications": "muted", "showReposts": false}
HTTP/1.1 200 OK
Content-Type: application/json
ETag: "w/rel-usr10-usr77-3"

{
  "userId": "usr_10",
  "followedId": "usr_77",
  "createdAt": "2026-03-04T10:22:11Z",
  "notifications": "muted",
  "showReposts": false,
  "_links": {
    "self":     {"href": "/v1/users/usr_10/following/usr_77"},
    "followed": {"href": "/v1/users/usr_77"}
  }
}

A practical rule: with no attributes, 204 and an empty body; with attributes, 200 and a full representation with an ETag, and then partial modifications (PATCH) and the optimistic concurrency of 03-06 apply again. Do not invent attributes "just in case": an edge with a body costs one wider row multiplied by 180 million.


  1. The timeline: fan-out, derived resource and opaque cursor

3.1 The fan-out problem

When usr_77 posts, who pays the cost of it appearing in their followers' timelines?

graph LR
  A[usr_77 publishes pst_2100] --> B{Strategy}
  B -->|Fan-out on write| C[Distribution queue]
  C --> D[Insert into 300,000 inboxes]
  D --> E[GET /timeline reads one inbox: fast]
  B -->|Fan-out on read| F[Cheap write: 1 row]
  F --> G[GET /timeline queries 300 followees and merges]
  G --> H[High and variable latency]
Criterion Fan-out on write (push) Fan-out on read (pull)
Cost of publishing High: O(followers) Minimal: O(1)
Cost of reading the timeline Minimal: a sequential read of one inbox High: O(followees), ordered merge
p95 read latency Low and stable High and dependent on the user
Storage Enormous (duplicated per follower) Minimal
Posting with 1 M followers A million writes per post No extra cost
Deleting a post Every inbox has to be cleaned up It disappears by itself
Fits The majority of normal users Heavily followed accounts

With 500 reads for every write, fan-out on write wins almost every time. The problem is the long tail: if usr_77 is a famous taster with a million followers, every post fires a million inserts and the queue backs up for minutes.

A hybrid solution, which is the one CafeSocial adopts:

  1. Users with fewer than 10,000 followers use push: when they post, an asynchronous task inserts the reference into every follower's inbox.
  2. Users above that threshold are marked as reach accounts and use pull: they distribute nothing.
  3. GET /v1/timeline reads the user's inbox and merges on the fly the recent posts of the few reach accounts they follow (usually fewer than 20), ordering by timestamp.

It is more code and more complexity, but it is the only design that survives both ends of the distribution.

3.2 /timeline is not a collection

At Aroma Store, /coffees was a collection: you could create (POST), count (total), filter and jump to page 7. The timeline is none of that: it is a derived resource, a computed view that only exists for one user and at one instant.

Property of a normal collection On /timeline
POST to create an element Does not exist: you post to /posts, not to the timeline
A stable total Impossible: there is no closed set to count
Jumping to page N Meaningless: page 7 from 3 s ago is no longer the same page
Arbitrary filters Very limited (?withPhotoOnly=true), it is not a search engine
DELETE of an element No: you hide it or you unfollow the author

Put another way: /timeline answers "what is new for me?", not "what elements does this set contain?". Just as in 02-06 we distinguished search from listing, here we distinguish a feed from a collection.

3.3 Why offset pagination does not work

Recall the mechanism from 02-06: ?limit=20&offset=20 translates to LIMIT 20 OFFSET 20. That assumes the ordered set does not change between requests. In a timeline ordered by date descending, it changes every second.

A worked example. usr_10's timeline contains, ordered newest to oldest, the posts pst_2100 (position 1) down to pst_2001 (position 100).

The duplicates case. The client asks for the first page:

GET /v1/timeline?limit=20&offset=0
→ positions 1..20  = pst_2100 ... pst_2081

While the user is reading, 3 new posts arrive. Everything has now shifted by 3 positions. The client asks for the second page:

GET /v1/timeline?limit=20&offset=20
→ positions 21..40 of the NEW set = pst_2084 ... pst_2065

pst_2084, pst_2083 and pst_2082 had already been shown on the first page. The user sees three repeated posts.

The gaps case. With the same initial state, between the first and the second request 3 posts are deleted from among the first 20. When offset=20 is requested, positions 21..40 of the reduced set correspond to what used to be positions 24..43: pst_2078 and the two before it are never shown. The user will never see them, and there is no visible error to give it away.

Add to that the fact that OFFSET 200000 forces the engine to walk and discard 200,000 rows: the cost grows with depth. Duplicates, silent gaps and growing cost: three independent reasons, each one sufficient on its own.

3.4 The opaque cursor

At Aroma Store the cursor was an option we applied to /orders. Here it is mandatory. The cursor implements keyset pagination: instead of "skip 20 rows", it says "give me what comes before this exact point".

Contents of the cursor: timestamp + tie-breaking id. The timestamp alone is not enough because two posts can share a millisecond; the id breaks the tie and guarantees a total ordering.

// utils/cursor.js — encoding and decoding of the opaque cursor
const VERSION_KEY = 'v1';

function encodeCursor({ createdAt, id }) {
  const payload = JSON.stringify({ v: VERSION_KEY, t: createdAt, i: id });
  return Buffer.from(payload, 'utf8').toString('base64url');
}

function decodeCursor(cursor) {
  let data;
  try {
    data = JSON.parse(Buffer.from(cursor, 'base64url').toString('utf8'));
  } catch {
    throw new ApiError(400, 'invalid_cursor', 'The cursor is not valid.');
  }
  if (data.v !== VERSION_KEY || !data.t || !data.i) {
    throw new ApiError(400, 'invalid_cursor', 'The cursor is not valid.');
  }
  // Age limit: an old cursor points at an inbox that has already been trimmed.
  const ageInDays = (Date.now() - Date.parse(data.t)) / 86_400_000;
  if (ageInDays > 30) {
    throw new ApiError(410, 'expired_cursor', 'Start again from the beginning.');
  }
  return { createdAt: data.t, id: data.i };
}

The corresponding query uses tuple comparison, which takes advantage of the composite index (user_id, created_at DESC, post_id DESC):

-- We ask for limit+1 rows so we know whether there is a next page without counting the total
SELECT p.post_id, p.author_id, p.created_at
FROM inbox b
JOIN posts p ON p.post_id = b.post_id
WHERE b.user_id = :userId
  AND (p.created_at, p.post_id) < (:cursorDate, :cursorId)
ORDER BY p.created_at DESC, p.post_id DESC
LIMIT :limit + 1;

The response carries no total —there is none— and exposes the next link both in the body and in the RFC 8288 Link header, just as in the store:

HTTP/1.1 200 OK
Content-Type: application/json
Cache-Control: private, max-age=0, must-revalidate
Vary: Authorization
Link: </v1/timeline?limit=20&cursor=eyJ2IjoidjEiLCJ0IjoiMjAyNi0wOC0xNFQwOToxMjozMy40MTJaIiwiaSI6InBzdF8yMDgxIn0>; rel="next"

{
  "data": [
    {
      "id": "pst_2100",
      "author": {"id": "usr_77", "handle": "geisha_taster"},
      "score": 92,
      "method": "v60",
      "likes": 1240,
      "approximateLikes": true,
      "_links": {"self": {"href": "/v1/posts/pst_2100"}}
    }
  ],
  "pagination": {
    "next": "eyJ2IjoidjEiLCJ0IjoiMjAyNi0wOC0xNFQwOToxMjozMy40MTJaIiwiaSI6InBzdF8yMDgxIn0",
    "hasMore": true
  }
}

Why opaque. The cursor is base64url of a JSON object, not encrypted: anybody can read it. Opacity is a contract, not a security measure: by not documenting its innards, tomorrow we can move from (date, id) to a position identifier in a distributed index without breaking any client. What we do guarantee is validation: if the client tampers with it, decoding fails (400 invalid_cursor) or the values do not pass the type check. And because the cursor does not contain the user's identifier —that always comes from the token— tampering with it does not let you read someone else's inbox. That point is essential: never put authorisation information inside a cursor.

Maximum age. Inboxes are trimmed to the last 800 entries and to 30 days. An older cursor points at nothing, so we respond 410 Gone with expired_cursor and the client goes back to the beginning, instead of returning an empty list that the app would read as "end of content".


  1. Volume and scale

4.1 Deliberately approximate counters

At Aroma Store, stock had to be exact: that is where the condition inside the UPDATE that solved overselling came from. Here, SELECT COUNT(*) FROM likes WHERE post_id = 'pst_2100' on every read of the timeline means 20 counting queries per screen, multiplied by millions of screens.

The decision: a denormalised counter, updated asynchronously and declared as approximate.

// services/likes.js
async function addLike(postId, userId) {
  const created = await likeRepository.insertIfMissing(postId, userId);
  if (created) {
    // The counter is not updated here: it is aggregated in batches every 5 seconds.
    await queue.publish('counters.likes', { postId, delta: 1 });
  }
  return created; // lets us answer 201 the first time and 204 on retries
}

The consequences have to be accepted explicitly, not hidden away:

  • The response sets "approximateLikes": true when the counter goes above 1,000. Below that it is recomputed on the fly and is exact: users notice a discrepancy in 12 but not in 12,480.
  • A nightly job reconciles the counters against the real table.
  • The user always sees their own "like" reflected immediately (read-your-own-writes), even if the global number lags behind: that is what stops the interface looking broken.

4.2 Caching profiles and the "it depends who is asking" problem

Picking up 04-06, the profile is the ideal candidate for caching: it is read constantly and changes little.

GET /v1/users/usr_77 HTTP/1.1
If-None-Match: "prof-usr77-v18"

HTTP/1.1 304 Not Modified
ETag: "prof-usr77-v18"
Cache-Control: private, max-age=60
Vary: Authorization

The delicate part is what can be cached in a shared way:

Resource Directive Reason
Public profile, no session public, max-age=300 The same for everybody
Profile seen by an authenticated user private, max-age=60 + Vary: Authorization Includes "do I follow them?", "have they blocked me?"
Timeline private, max-age=0, must-revalidate Unique per user and per instant
Post image (CDN) public, max-age=31536000, immutable URL with a content fingerprint

Vary: Authorization is correct but, in practice, it destroys the shared cache: every token produces a different entry. That is why the real strategy is to split the response in two: the objective profile data (handle, biography, photo) is served cacheable and public, and whatever depends on the observer (followingThisUser, hasBlockedMe) is requested separately or marked private. Caching the mixture is the classic mistake, and its most serious version is a proxy handing the profile "as seen by another user" to somebody who should not have it.

4.3 Queues, 202 and task resources

Fan-out distribution and other long operations do not fit inside the request cycle. Just as in 04-06 with the store's reports:

POST /v1/users/usr_10/export HTTP/1.1

HTTP/1.1 202 Accepted
Location: /v1/tasks/tsk_9001
Retry-After: 10

{"id": "tsk_9001", "status": "in_progress",
 "_links": {"self": {"href": "/v1/tasks/tsk_9001"}}}

Posting is the same thing in reverse: POST /v1/posts responds 201 immediately with the created post —the author already sees it on their profile— while the distribution to the inboxes happens behind the scenes. It is eventual consistency turned into a contract: a follower may take a few seconds to see it, and that is documented.

4.4 The cost of denormalisation

Denormalising is not free. At CafeSocial we pay for it: multiplied storage (a post appears in hundreds of thousands of inboxes), a more fragile write path (if the queue fails, there are incomplete inboxes and a repair process is needed), and two sources of truth that can diverge. The rule is that the canonical source is still /posts/{id}; the inboxes are a rebuildable cache. Anything that can be regenerated from the canonical source is acceptable to denormalise; anything that cannot, is not.


  1. User-generated content

5.1 Image upload with a pre-signed URL

The API does not receive the megabytes of the photo. An authorisation is issued for uploading directly to object storage:

sequenceDiagram
  participant App as Mobile app
  participant API as CafeSocial API
  participant St as Storage
  App->>API: POST /v1/posts/pst_2100/images {type, size}
  API-->>App: 201 {uploadUrl, fields, expiresAt, imageId}
  App->>St: PUT uploadUrl (photo bytes)
  St-->>App: 200 OK
  App->>API: POST /v1/posts/pst_2100/images/img_44/confirmation
  API->>St: Verify type, size and file header
  API-->>App: 200 {status: pending_moderation}
POST /v1/posts/pst_2100/images HTTP/1.1
Content-Type: application/json

{"contentType": "image/jpeg", "sizeBytes": 2411520}
HTTP/1.1 201 Created
Location: /v1/posts/pst_2100/images/img_44

{
  "id": "img_44",
  "uploadUrl": "https://storage.cafesocial.example/uploads/img_44?signature=...",
  "method": "PUT",
  "expiresAt": "2026-08-14T09:27:00Z",
  "maxSizeBytes": 5242880,
  "status": "pending_upload"
}

The reasons for not passing the bytes through the API are concrete:

  • A Node process tied up for 8 seconds with a 5 MB upload is a process that is serving nobody else.
  • Scaling the upload traffic is decoupled from scaling the business logic.
  • Storage and the CDN already solve resumption, multipart uploads and geographical distribution.
  • Proxy and load-balancer timeouts stop being a problem.

Validation. The pre-signed URL limits type and size, but that is not enough: the client declares image/jpeg and uploads whatever it likes. That is why confirmation is mandatory and checks the real bytes on the server (magic numbers in the file header, dimensions, size), reprocesses the image into several sizes and strips the EXIF metadata —which includes GPS coordinates: publishing the exact location of a user's home would be a personal data leak. An image left unconfirmed for 15 minutes is deleted and the post stays a draft.

5.2 Moderation

The status mirrors that of the Aroma Store reviews: pending_moderation | published | rejected. What changes is the scale, which forces two stages:

  1. Automatic, immediately: an image and text classifier. Low score, published straight away; medium score, human queue; high score, immediate rejection with the possibility of appeal.
  2. Human, over a work queue:
GET /v1/moderation/pending?limit=50&cursor=... HTTP/1.1
Authorization: Bearer <token with the moderation:read scope>

POST /v1/moderation/posts/pst_2100/approval HTTP/1.1
POST /v1/moderation/posts/pst_2100/rejection HTTP/1.1
Content-Type: application/json

{"reason": "unrelated_content", "notifyAuthor": true}

Design details that matter:

  • The moderation queue lives under /moderation/*, a namespace of its own with its own scopes (moderation:read, moderation:write). Mixing it into /posts would force the same endpoint to behave radically differently depending on the scope, which is exactly what makes tests ungovernable.
  • Approval and rejection are action resources (POST to a sub-resource) because they are not idempotent in their effects: they fire notifications and they are audited. Every decision is stored with the moderator, a timestamp and a reason.
  • The queue is paginated by cursor: it grows and is consumed at the same time.

5.3 Reports

POST /v1/posts/pst_2100/reports HTTP/1.1

{"reason": "spam", "comment": "Posts the same link on every tasting"}
HTTP/1.1 202 Accepted
{"status": "received", "_links": {"self": {"href": "/v1/reports/rep_77"}}}

We respond 202 and not 201 with a verdict: the report is accepted, not resolved. The reporter must not be able to see the detailed status or the moderator's identity, and the reported user must not learn who reported them. Several reports about the same post are aggregated into a single case.

An essential warning. Everything above is the engineering part, which is the easy part. A real service with user content has legal obligations that are not technical decisions: lawful bases and GDPR rights, deadlines for taking down illegal content, protection of minors and age verification, data retention and disclosure to authorities, moderation transparency and routes of appeal. They vary by country and over time. Design it with legal and compliance advice from the start; a deletion endpoint that misses the legal deadline is a legal problem, not a TODO in the backlog.


  1. Privacy and resource-level authorisation

At Aroma Store the question was "who owns this?": order ord_5001 belongs to cus_842, so only cus_842 and the administrators see it. Binary and simple.

At CafeSocial the question is "who is asking and what do we let them see?", and the answer is no longer yes or no, but a different representation:

Who asks for usr_77 (private profile) What they get
usr_77 themselves Everything, including drafts and email address
An accepted follower Full profile and posts
An authenticated user who does not follow them Handle, photo, bio, counters; no posts
A user blocked by usr_77 404 user_not_found
Unauthenticated 404 if the profile is private; a reduced card if it is public

6.1 Projection by visibility

The pattern: the service layer returns the full entity and a projection decides which fields survive, according to the relationship between observer and observed.

// presentation/projections/user.js
const FIELDS = {
  owner:    ['id','handle','name','bio','photo','email','followers','following','private'],
  follower: ['id','handle','name','bio','photo','followers','following','private'],
  public:   ['id','handle','photo','bio','followers','private'],
};

function projectUser(user, context) {
  const level = visibilityLevelFor(user, context); // owner|follower|public|hidden
  if (level === 'hidden') throw new ApiError(404, 'user_not_found', 'Not found.');
  const output = {};
  for (const field of FIELDS[level]) output[field] = user[field];
  output._links = linksForLevel(user, level);
  return output;
}

Three rules that avoid the usual failures:

  1. An allowlist, never a blocklist. A new field on the entity must not leak just because it has been added.
  2. The projection applies inside lists too. GET /users/usr_77/followers trims out blocked and private users: the list returned may be shorter than the counter shown on the profile, and that is correct.
  3. Visibility is decided in the query, not afterwards. Filtering in memory after fetching 50 rows breaks cursor pagination: you would ask for 20 and return 14.

6.2 404 instead of 403

If usr_77 blocks usr_10 and the latter requests GET /v1/users/usr_77, a 403 insufficient_permissions would be technically honest and it would leak information: it would confirm that the account exists, and by the difference between responses you could enumerate who has blocked whom or discover which handles are registered. When the mere existence of the resource is sensitive information, you respond 404.

The general criterion: 403 when the resource is known to the caller or its existence reveals nothing (the moderator without the right scope); 404 when revealing the existence is itself a leak. And you have to be consistent in response times: a 404 that takes 5 ms when the resource does not exist and 40 ms when it exists but is blocked leaks the information again through a side channel.

6.3 The effect on caching and tests

  • Caching: any response whose shape depends on the observer is private and carries Vary: Authorization. A shared CDN can only serve what is objectively public. This reduces performance and it is a cost accepted consciously.
  • Tests: the matrix grows all at once. Every sensitive endpoint needs cases for owner, follower, non-follower, blocked and anonymous. At CafeSocial it is handled with a parameterised table of cases:
describe('GET /v1/users/usr_77 by observer', () => {
  const cases = [
    { observer: 'owner',    status: 200, includes: ['email'],  excludes: [] },
    { observer: 'follower', status: 200, includes: ['name'],   excludes: ['email'] },
    { observer: 'stranger', status: 200, includes: ['handle'], excludes: ['name','email'] },
    { observer: 'blocked',  status: 404, includes: [],         excludes: [] },
  ];
  for (const c of cases) it(`observer ${c.observer} → ${c.status}`, async () => { /* ... */ });
});

  1. Notifications and real time

At Aroma Store, SSE was an extra for the internal panel. Here, immediacy is the product: a social network where the notification arrives two minutes late is perceived as broken.

Pure REST would force polling: the app asks GET /v1/notifications every 10 seconds. With 500,000 active users that is 50,000 requests per second, 99 % of which return the same thing. Not even with ETag and 304 (which save bandwidth, not requests) do the numbers add up.

The design combines three pieces:

GET /v1/notifications?limit=30&cursor=... HTTP/1.1

{"data": [
   {"id": "ntf_55", "type": "like", "read": false,
    "actor": {"id": "usr_77", "handle": "geisha_taster"},
    "resource": {"href": "/v1/posts/pst_2100"},
    "createdAt": "2026-08-14T09:12:33Z"}
 ],
 "unread": 4,
 "pagination": {"next": "eyJ2IjoidjEi...", "hasMore": true}}
POST /v1/notifications/read HTTP/1.1
{"upTo": "ntf_55"}          # marks as read up to a point, idempotent

The history is paginated by cursor (the same mechanism as the timeline) and the live channel only carries notices that there is something new, not the full content: the client receives the signal and reloads over REST. That way the real-time channel does not turn into a second API to maintain in parallel.

Mechanism Direction When to choose it at CafeSocial Cost
Polling with ETag/304 Client→server A fallback when everything else fails; non-urgent data Constant requests
Long polling Client→server Compatibility with hostile networks Held-open connections
SSE Server→client Unread counter, new posts: it is what the web uses Unidirectional; automatic reconnection and Last-Event-ID out of the box
WebSockets Bidirectional Direct messages with "typing…" and read receipts Its own infrastructure, per-connection state, harder to scale
Webhooks (01-07) Server→server Integrations: telling Aroma Store about a tasting with a high score Retries, HMAC signature, "at least once" delivery
Push notifications (APNs/FCM) Server→device App closed or in the background A dependency on external platforms

CafeSocial uses SSE for the web and the notification counter, WebSockets only on the direct messages screen, push when the app is not active and webhooks signed with HMAC-SHA256 towards Aroma Store, reusing exactly the mechanism from 06-01.


  1. Aroma Store versus CafeSocial, decision by decision

This is the table that makes sense of the two lessons together.

Decision Aroma Store CafeSocial What changes it
Read/write ratio ~10:1 ~500:1 Justifies precomputing and denormalising
Pagination limit/offset on coffees; cursor on orders Cursor mandatory on every feed The set changes between requests
total on collections Yes, useful and cheap Does not exist on feeds There is no closed set to count
Consistency Strong and transactional (stock, payments) Eventual and documented There is no money in the request
Counters Exact by definition Approximate by design The cost of counting against the value of exactness
Authorisation "who owns it?" → 200/403 "who is asking?" → projections and 404 The existence of the resource is information
Shape of the response Stable for everybody Variable by observer Privacy and blocks
Caching public with ETag on the catalogue Almost everything private + Vary It depends on the observer
Writes Synchronous, with optimistic concurrency Asynchronous, queue and 202 Fan-out does not fit inside the request
Cost of fan-out Non-existent Dominant: hybrid push/pull The follower distribution has a long tail
Files No significant uploads Pre-signed URL, outside the API Volume and occupation time
Real time SSE as an extra for the panel SSE + WebSockets as the product Immediacy is the perceived value
Moderation Reviews, low volume, manual review Two stages, dedicated queue, reports 40,000 items a day
Relationships Simple containment (order→customer) A graph: the edge is a resource with PUT/DELETE 180 M edges with semantics of their own
Idempotency Idempotency-Key on payments PUT, idempotent by design across the graph Retries on mobile networks

The conclusion is not that one design is better. It is that there is no universal REST design: there are domain-dependent decisions, and professional competence consists in knowing which one applies and being able to justify it. If somebody offers you "the right way to paginate" without asking about the domain, they are selling you an answer before they have heard the question.


  1. Why GraphQL is defensible here

In 01-07 we concluded that for Aroma Store GraphQL would have been complexity with no return: few screen types, high value from HTTP caching on the catalogue, one very stable main consumer. At CafeSocial the arguments flip sign:

  • Screens with heterogeneous data. The detail view of a post needs the post, the author, whether I follow them, the first five comments with their authors, the "like" counter, whether I liked it, the hashtags and the linked catalogue coffee. In REST that is 5-7 requests, or a bespoke composite endpoint that ages badly.
  • Mobile clients with limited bandwidth. The list needs 6 fields per post; the detail view, 30. With REST you end up inventing ?fields= or ?view=summary, which is GraphQL done badly.
  • Fast client evolution. A redesign of the timeline every quarter means, in REST, negotiating contract changes with the backend every quarter.
  • A graph is queried like a graph. "The latest comments from the people I follow on posts that my followers also liked" is a natural query in GraphQL and a contorted endpoint in REST.

And what you lose, which has to go on the same scales:

What you gain What you lose
One request per screen Intermediate HTTP caching: everything is POST /graphql
The client chooses the fields ETag/304 and CDNs stop helping
Evolution without versioning routes Status codes: almost everything is 200 with errors
A typed, self-documenting schema The client can build extremely expensive queries: you need depth limits, complexity limits and persisted queries
A single entry point Per-endpoint observability and rate limiting stop working as they are
— The N+1 problem in the resolvers: it forces DataLoader from day one

The realistic decision at CafeSocial —and the one you will see in many companies— is hybrid: GraphQL for the mobile app's screens, REST for what is public and indexable (profiles and posts cacheable in a CDN), for third-party integrations and for webhooks. Choosing GraphQL is not abandoning what you have learned: the resources, the states, idempotency, cursor pagination and observer-based authorisation remain exactly the same problems; only the transport layer changes.


Common Mistakes and Tips

  1. Using POST /follow because "it is an action". You lose idempotency, and querying and deleting for free. If the action creates or destroys a relationship between two identifiable entities, that relationship has a URI: PUT/DELETE (02-03).
  2. Putting authorisation information inside the cursor. A cursor with a userId inside is a privilege escalation waiting to happen. The subject always comes from the token.
  3. Confusing opaque with secure. Base64url encrypts nothing. Opacity is freedom to change the format, not protection: always validate the decoded content.
  4. Returning a total on a feed. It forces an expensive COUNT over a set that changes, and the resulting number is false the moment it is sent. Use hasMore.
  5. Filtering by visibility after paginating. You ask for 20, hide 6 and return 14: the client thinks it is reaching the end. Visibility goes in the query.
  6. A blocklist of fields in the projections. The day you add recoveryEmail to the entity, it will publish itself. Always an allowlist.
  7. 403 where existence already leaks information. And watch out for response times and error messages too, which leak through side channels.
  8. Caching as public a response that depends on the observer. It is the fastest route to a proxy serving one user's private profile to another. When in doubt, private.
  9. Uploading large files through the API. It ties up processes, collides with proxy timeouts and does not scale. Pre-signed URL and later confirmation.
  10. Trusting the Content-Type declared by the client. Verify the real bytes, reprocess the image and strip the EXIF metadata before publishing it.
  11. Designing fan-out with a single algorithm. Pure push dies with heavily followed accounts; pure pull dies on latency. The hybrid is ugly and it is the one that works.
  12. A process tip: write the comparison table from section 8 before you start coding. It forces you to justify every decision against a concrete alternative and it is the best defence against copying the last project's design out of inertia.

Exercises

Exercise 1 — The graph as a resource. Design the endpoints for "muting" a user you follow (you stop seeing their posts in your timeline, but you remain their follower). State the method, the URI, the status codes and whether it needs a body. Justify why you would not model it as POST /v1/mute.

Exercise 2 — A tamper-proof cursor. A client sends ?cursor=eyJ2IjoidjEiLCJ0IjoiMjAzMC0wMS0wMVQwMDowMDowMFoiLCJpIjoicHN0Xzk5OTk5In0 (a date in the future). Explain what the system returns, why that is not a security flaw, and add the missing validation to decodeCursor.

Exercise 3 — Projection and pagination together. GET /v1/users/usr_77/followers?limit=20 must hide the users who have blocked the observer. Explain why filtering in memory breaks pagination and sketch the correct SQL query with a cursor.

Solutions

Solution 1. Muting is an attribute of an existing relationship, not a new relationship:

PATCH /v1/users/usr_10/following/usr_77
Content-Type: application/json
If-Match: "w/rel-usr10-usr77-3"

{"muted": true}

HTTP/1.1 200 OK
ETag: "w/rel-usr10-usr77-4"
{"userId":"usr_10","followedId":"usr_77","muted":true,"createdAt":"2026-03-04T10:22:11Z"}

Codes: 200 on success; 404 if you do not follow that user (there is no relationship to modify); 412 if the If-Match does not match; 422/400 with invalid_data if the body does not validate. POST /v1/mute is not used because the relationship already has a URI: creating a parallel verb would duplicate the resource, lose the idempotency of a PATCH against a specific state and force you to invent POST /v1/unmute. An equally valid alternative: a PUT over the complete relationship with all its attributes, if you prefer to avoid PATCH.

Solution 2. The system decodes the JSON correctly (a valid v, t and i present), the age check does not trigger because the date is in the future, and the query WHERE (created_at, id) < ('2030-01-01', 'pst_99999') simply returns the first page, since everything is earlier than that date. It is not a security flaw because the userId in the WHERE comes from the token, not from the cursor: tampering with it only lets you reposition yourself within your own inbox, something the user can already do by paginating. Even so, it is worth rejecting it in order to detect broken clients:

  const t = Date.parse(data.t);
  if (Number.isNaN(t)) {
    throw new ApiError(400, 'invalid_cursor', 'The cursor is not valid.');
  }
  if (t > Date.now() + 60_000) {           // 1 min of slack for clock skew
    throw new ApiError(400, 'invalid_cursor', 'The cursor is not valid.');
  }
  if (typeof data.i !== 'string' || !/^pst_[0-9]+$/.test(data.i)) {
    throw new ApiError(400, 'invalid_cursor', 'The cursor is not valid.');
  }

Solution 3. Filtering in memory breaks pagination because the LIMIT is applied before the filter: you ask for 20 rows, discard the 6 belonging to users who have blocked you and return 14, while the cursor advances as if you had delivered 20. The client sees irregular pages and, if a whole page comes out empty, reads it as the content having run out. The filter has to go in the query:

SELECT f.follower_id, f.created_at
FROM follows f
WHERE f.followed_id = :profileId
  AND NOT EXISTS (
        SELECT 1 FROM blocks b
        WHERE b.blocker_id = f.follower_id
          AND b.blocked_id = :observerId )
  AND (f.created_at, f.follower_id) < (:cursorDate, :cursorId)
ORDER BY f.created_at DESC, f.follower_id DESC
LIMIT :limit + 1;

A consequence that has to be documented: the number of items walked may not match the profile's followers counter, because that counter is global and the list is relative to the observer. And because the response depends on the observer, it is private with Vary: Authorization.


Conclusion

We have designed CafeSocial with the same tools as Aroma Store —resources, methods, status codes, headers, hypermedia— and we have ended up with an almost opposite design. The follow relationship stopped being a field and became a resource with an idempotent PUT and a DELETE. The timeline stopped being a collection and became a derived resource with no POST, no total and no page N, held up by a hybrid fan-out and reachable only through an opaque cursor of timestamp plus id. The counters stopped being exact because counting stopped being worth it. The images left the API. Authorisation stopped asking who owns the resource and started asking who is observing, with projections by visibility and 404 wherever existence is itself information. And real time stopped being an ornament and became the product.

If this lesson leaves one single idea, let it be the one in the table in section 8: there is no universal REST design, there are domain-dependent decisions, and a professional is distinguished by being able to name the alternative they rejected and why. That is also the reason GraphQL, indefensible for the store's catalogue, is a serious option here, as long as the price is accepted: losing HTTP caching, status codes and control over the cost of queries.

One last matter remains, and it is the one that separates a well-designed API from an API that is still alive three years later: what happens after you publish it. In 06-03, Evolving and Maintaining an API in Production, we will look at what to do when an apparently harmless change breaks a consumer in production, how to write a post-mortem that is actually useful, how contract debt is accumulated and paid off, and how a migration to v2 is planned with its deprecation period and its governance. Because designing an API well is hard, but changing it without breaking whoever already uses it is harder still.

REST API Course: Principles of Designing and Developing RESTful APIs

Module 1: Introduction to RESTful APIs

Module 2: Designing RESTful APIs

Module 3: Building RESTful APIs

Module 4: Best Practices and Security

Module 5: Tools and Frameworks

Module 6: Case Studies and Projects

© Copyright 2026. All rights reserved