In the previous lesson, 06-01, we closed the technical write-up of the Aroma Store API with a list of certainties: offset pagination, exact counters, authorisation based on ownership of the resource, collections wrapped in data with their total, and real time as an accessory of the internal panel. All those decisions were right for that domain. Now we change domain without changing architectural style: we design CafeSocial, the tasters' network Aroma Store wants to launch, and one by one we are going to watch those certainties fall. Not because they were wrong, but because they were tied to a context that no longer exists. That contrast —the same REST discipline producing opposite designs— is the real content of this lesson.
Contents
- The scenario: CafeSocial, requirements and consumers
- Modelling the relationship graph with REST
- The timeline: fan-out, derived resource and opaque cursor
- Volume and scale: approximate counters, caching and queues
- User-generated content: upload, moderation and reports
- Privacy and resource-level authorisation
- Notifications and real time
- Aroma Store versus CafeSocial, decision by decision
- Why GraphQL is defensible here
- Common mistakes and tips
- Exercises and solutions
- The scenario: CafeSocial, requirements and consumers
CafeSocial is a vertical social network for coffee tasters. A user publishes a tasting: a photo of the coffee, the variety, the origin, the brewing method and a score from 0 to 100. Other users comment, give "likes", follow the tasters that interest them and receive a timeline with the posts of the people they follow. There are hashtags (#geisha, #v60), notifications and direct messages.
The identifiers keep the course's convention: usr_10, usr_77, pst_2100, cmt_310, ntf_55. The JSON is still camelCase, the errors still follow the {"error": {"code","message","details":[]}} catalogue and the version is still in the path (/v1). The domain changes, the conventions do not: that is precisely what makes comparison possible.
Requirements that shape the design
| Requirement | Target figure (fictional) | Design consequence |
|---|---|---|
| Read/write ratio | ~500:1 | Everything is optimised for reads; writes can be more expensive |
| Timeline latency | p95 < 150 ms | Precompute, do not compute during the request |
| Size of the graph | 2 M users, 180 M follow edges | The relationship is a first-class resource, not a field |
| Follower distribution | Long tail: 0.01 % exceed 100,000 followers | A single distribution algorithm will not do |
| User content | 40,000 posts with a photo per day | Moderation and storage outside the API |
| Consistency | Eventual is acceptable for the timeline and counters | It can be decoupled with queues |
Consumers
- Mobile app (iOS/Android): the main consumer, with limited bandwidth and battery. Every byte and every round trip hurts.
- Public web: profiles and posts indexable by search engines when they are public.
- Internal moderation panel: few users, broad permissions, needs work queues.
- Integration with Aroma Store: when a post mentions a coffee from the catalogue, the card links to
/v1/coffees/cof_001of the store's API. They are two separate APIs linked by hypermedia, not one.
Unlike the store, here there is no money in the request. Nothing demands the transactional exactness that in 06-01 forced us to solve overselling inside the
UPDATE. That freedom is what makes almost everything that follows possible.
- Modelling the relationship graph with REST
At Aroma Store almost everything was a simple containment relationship: an order belongs to a customer, an item belongs to an order. Here the relationship is the interesting entity.
2.1 The two collections of the graph
They are two views of the same edge, from its two ends. Neither is "the right one": the app needs both, and with very different cardinalities (a user follows 300 people but may have 300,000 followers).
2.2 The relationship as a resource: PUT versus POST /follow
The temptation is to invent a verb: POST /v1/follow with {"followedId": "usr_77"}. It works, and it is exactly what in 02-02 we called turning an action into a resource without need. The alternative is to treat the edge as an addressable resource:
PUT /v1/users/usr_10/following/usr_77 HTTP/1.1
Authorization: Bearer <usr_10's token>
HTTP/1.1 204 No Content
Aroma-Graph-State: followingThe advantages, going back to the methods table of 02-03:
| Aspect | PUT /following/{id} |
POST /follow |
|---|---|---|
| Idempotency | Yes: hitting "follow" five times leaves the same state | Not guaranteed; you have to deduplicate |
| Retry after a timeout | Safe by definition | Requires an Idempotency-Key |
| Querying the state | A GET on the same URI (204/404) |
You have to invent another endpoint |
| Undoing | A DELETE on the same URI |
Another verb: POST /unfollow |
| Cacheable / linkable | Yes, it has its own URI | No |
In a mobile app on an unstable network, idempotency is not theoretical elegance: it is the difference between a button that gets stuck halfway and one that does not. The client can retry the PUT without thinking.
GET /v1/users/usr_10/following/usr_77 returns 204 if the relationship exists and 404 if it does not, which gives the client a cheap check for rendering the button.
2.3 "Likes": user in the path or in the token?
Both forms are legitimate and it is worth understanding the trade-off:
PUT /v1/posts/pst_2100/likes/usr_10 # A: explicit subject
PUT /v1/posts/pst_2100/likes # B: subject implicit in the token| Criterion | A (explicit) | B (implicit) |
|---|---|---|
| Self-descriptive URI | Yes | No: the same URI means different things depending on who calls |
| Impersonation risk | You have to validate that the path matches the token | Impossible by construction |
| Acting on someone else's behalf (admin, import) | Straightforward | Needs a separate header or endpoint |
| Listing who liked it | GET /posts/pst_2100/likes is natural |
Just as natural |
| Caching | Cacheable by URI | Needs Vary: Authorization |
At CafeSocial we chose A for the follow graph and for "likes", for consistency and because the moderation panel needs to be able to remove another user's "like". The rule we applied: if any legitimate consumer can act on a third party's relationship, the subject goes in the path. When the path and the token do not match and the caller lacks the appropriate scope, we respond 403 with insufficient_permissions.
2.4 When the relationship needs a body of its own
An edge with attributes stops being a simple "it exists or it does not". Following someone can carry preferences:
PUT /v1/users/usr_10/following/usr_77 HTTP/1.1
Content-Type: application/json
{"notifications": "muted", "showReposts": false}HTTP/1.1 200 OK
Content-Type: application/json
ETag: "w/rel-usr10-usr77-3"
{
"userId": "usr_10",
"followedId": "usr_77",
"createdAt": "2026-03-04T10:22:11Z",
"notifications": "muted",
"showReposts": false,
"_links": {
"self": {"href": "/v1/users/usr_10/following/usr_77"},
"followed": {"href": "/v1/users/usr_77"}
}
}A practical rule: with no attributes, 204 and an empty body; with attributes, 200 and a full representation with an ETag, and then partial modifications (PATCH) and the optimistic concurrency of 03-06 apply again. Do not invent attributes "just in case": an edge with a body costs one wider row multiplied by 180 million.
- The timeline: fan-out, derived resource and opaque cursor
3.1 The fan-out problem
When usr_77 posts, who pays the cost of it appearing in their followers' timelines?
graph LR
A[usr_77 publishes pst_2100] --> B{Strategy}
B -->|Fan-out on write| C[Distribution queue]
C --> D[Insert into 300,000 inboxes]
D --> E[GET /timeline reads one inbox: fast]
B -->|Fan-out on read| F[Cheap write: 1 row]
F --> G[GET /timeline queries 300 followees and merges]
G --> H[High and variable latency]
| Criterion | Fan-out on write (push) | Fan-out on read (pull) |
|---|---|---|
| Cost of publishing | High: O(followers) | Minimal: O(1) |
| Cost of reading the timeline | Minimal: a sequential read of one inbox | High: O(followees), ordered merge |
| p95 read latency | Low and stable | High and dependent on the user |
| Storage | Enormous (duplicated per follower) | Minimal |
| Posting with 1 M followers | A million writes per post | No extra cost |
| Deleting a post | Every inbox has to be cleaned up | It disappears by itself |
| Fits | The majority of normal users | Heavily followed accounts |
With 500 reads for every write, fan-out on write wins almost every time. The problem is the long tail: if usr_77 is a famous taster with a million followers, every post fires a million inserts and the queue backs up for minutes.
A hybrid solution, which is the one CafeSocial adopts:
- Users with fewer than 10,000 followers use push: when they post, an asynchronous task inserts the reference into every follower's inbox.
- Users above that threshold are marked as reach accounts and use pull: they distribute nothing.
GET /v1/timelinereads the user's inbox and merges on the fly the recent posts of the few reach accounts they follow (usually fewer than 20), ordering by timestamp.
It is more code and more complexity, but it is the only design that survives both ends of the distribution.
3.2 /timeline is not a collection
At Aroma Store, /coffees was a collection: you could create (POST), count (total), filter and jump to page 7. The timeline is none of that: it is a derived resource, a computed view that only exists for one user and at one instant.
| Property of a normal collection | On /timeline |
|---|---|
POST to create an element |
Does not exist: you post to /posts, not to the timeline |
A stable total |
Impossible: there is no closed set to count |
| Jumping to page N | Meaningless: page 7 from 3 s ago is no longer the same page |
| Arbitrary filters | Very limited (?withPhotoOnly=true), it is not a search engine |
DELETE of an element |
No: you hide it or you unfollow the author |
Put another way: /timeline answers "what is new for me?", not "what elements does this set contain?". Just as in 02-06 we distinguished search from listing, here we distinguish a feed from a collection.
3.3 Why offset pagination does not work
Recall the mechanism from 02-06: ?limit=20&offset=20 translates to LIMIT 20 OFFSET 20. That assumes the ordered set does not change between requests. In a timeline ordered by date descending, it changes every second.
A worked example. usr_10's timeline contains, ordered newest to oldest, the posts pst_2100 (position 1) down to pst_2001 (position 100).
The duplicates case. The client asks for the first page:
While the user is reading, 3 new posts arrive. Everything has now shifted by 3 positions. The client asks for the second page:
pst_2084, pst_2083 and pst_2082 had already been shown on the first page. The user sees three repeated posts.
The gaps case. With the same initial state, between the first and the second request 3 posts are deleted from among the first 20. When offset=20 is requested, positions 21..40 of the reduced set correspond to what used to be positions 24..43: pst_2078 and the two before it are never shown. The user will never see them, and there is no visible error to give it away.
Add to that the fact that OFFSET 200000 forces the engine to walk and discard 200,000 rows: the cost grows with depth. Duplicates, silent gaps and growing cost: three independent reasons, each one sufficient on its own.
3.4 The opaque cursor
At Aroma Store the cursor was an option we applied to /orders. Here it is mandatory. The cursor implements keyset pagination: instead of "skip 20 rows", it says "give me what comes before this exact point".
Contents of the cursor: timestamp + tie-breaking id. The timestamp alone is not enough because two posts can share a millisecond; the id breaks the tie and guarantees a total ordering.
// utils/cursor.js — encoding and decoding of the opaque cursor
const VERSION_KEY = 'v1';
function encodeCursor({ createdAt, id }) {
const payload = JSON.stringify({ v: VERSION_KEY, t: createdAt, i: id });
return Buffer.from(payload, 'utf8').toString('base64url');
}
function decodeCursor(cursor) {
let data;
try {
data = JSON.parse(Buffer.from(cursor, 'base64url').toString('utf8'));
} catch {
throw new ApiError(400, 'invalid_cursor', 'The cursor is not valid.');
}
if (data.v !== VERSION_KEY || !data.t || !data.i) {
throw new ApiError(400, 'invalid_cursor', 'The cursor is not valid.');
}
// Age limit: an old cursor points at an inbox that has already been trimmed.
const ageInDays = (Date.now() - Date.parse(data.t)) / 86_400_000;
if (ageInDays > 30) {
throw new ApiError(410, 'expired_cursor', 'Start again from the beginning.');
}
return { createdAt: data.t, id: data.i };
}The corresponding query uses tuple comparison, which takes advantage of the composite index (user_id, created_at DESC, post_id DESC):
-- We ask for limit+1 rows so we know whether there is a next page without counting the total
SELECT p.post_id, p.author_id, p.created_at
FROM inbox b
JOIN posts p ON p.post_id = b.post_id
WHERE b.user_id = :userId
AND (p.created_at, p.post_id) < (:cursorDate, :cursorId)
ORDER BY p.created_at DESC, p.post_id DESC
LIMIT :limit + 1;The response carries no total —there is none— and exposes the next link both in the body and in the RFC 8288 Link header, just as in the store:
HTTP/1.1 200 OK
Content-Type: application/json
Cache-Control: private, max-age=0, must-revalidate
Vary: Authorization
Link: </v1/timeline?limit=20&cursor=eyJ2IjoidjEiLCJ0IjoiMjAyNi0wOC0xNFQwOToxMjozMy40MTJaIiwiaSI6InBzdF8yMDgxIn0>; rel="next"
{
"data": [
{
"id": "pst_2100",
"author": {"id": "usr_77", "handle": "geisha_taster"},
"score": 92,
"method": "v60",
"likes": 1240,
"approximateLikes": true,
"_links": {"self": {"href": "/v1/posts/pst_2100"}}
}
],
"pagination": {
"next": "eyJ2IjoidjEiLCJ0IjoiMjAyNi0wOC0xNFQwOToxMjozMy40MTJaIiwiaSI6InBzdF8yMDgxIn0",
"hasMore": true
}
}Why opaque. The cursor is base64url of a JSON object, not encrypted: anybody can read it. Opacity is a contract, not a security measure: by not documenting its innards, tomorrow we can move from (date, id) to a position identifier in a distributed index without breaking any client. What we do guarantee is validation: if the client tampers with it, decoding fails (400 invalid_cursor) or the values do not pass the type check. And because the cursor does not contain the user's identifier —that always comes from the token— tampering with it does not let you read someone else's inbox. That point is essential: never put authorisation information inside a cursor.
Maximum age. Inboxes are trimmed to the last 800 entries and to 30 days. An older cursor points at nothing, so we respond 410 Gone with expired_cursor and the client goes back to the beginning, instead of returning an empty list that the app would read as "end of content".
- Volume and scale
4.1 Deliberately approximate counters
At Aroma Store, stock had to be exact: that is where the condition inside the UPDATE that solved overselling came from. Here, SELECT COUNT(*) FROM likes WHERE post_id = 'pst_2100' on every read of the timeline means 20 counting queries per screen, multiplied by millions of screens.
The decision: a denormalised counter, updated asynchronously and declared as approximate.
// services/likes.js
async function addLike(postId, userId) {
const created = await likeRepository.insertIfMissing(postId, userId);
if (created) {
// The counter is not updated here: it is aggregated in batches every 5 seconds.
await queue.publish('counters.likes', { postId, delta: 1 });
}
return created; // lets us answer 201 the first time and 204 on retries
}The consequences have to be accepted explicitly, not hidden away:
- The response sets
"approximateLikes": truewhen the counter goes above 1,000. Below that it is recomputed on the fly and is exact: users notice a discrepancy in 12 but not in 12,480. - A nightly job reconciles the counters against the real table.
- The user always sees their own "like" reflected immediately (read-your-own-writes), even if the global number lags behind: that is what stops the interface looking broken.
4.2 Caching profiles and the "it depends who is asking" problem
Picking up 04-06, the profile is the ideal candidate for caching: it is read constantly and changes little.
GET /v1/users/usr_77 HTTP/1.1
If-None-Match: "prof-usr77-v18"
HTTP/1.1 304 Not Modified
ETag: "prof-usr77-v18"
Cache-Control: private, max-age=60
Vary: AuthorizationThe delicate part is what can be cached in a shared way:
| Resource | Directive | Reason |
|---|---|---|
| Public profile, no session | public, max-age=300 |
The same for everybody |
| Profile seen by an authenticated user | private, max-age=60 + Vary: Authorization |
Includes "do I follow them?", "have they blocked me?" |
| Timeline | private, max-age=0, must-revalidate |
Unique per user and per instant |
| Post image (CDN) | public, max-age=31536000, immutable |
URL with a content fingerprint |
Vary: Authorization is correct but, in practice, it destroys the shared cache: every token produces a different entry. That is why the real strategy is to split the response in two: the objective profile data (handle, biography, photo) is served cacheable and public, and whatever depends on the observer (followingThisUser, hasBlockedMe) is requested separately or marked private. Caching the mixture is the classic mistake, and its most serious version is a proxy handing the profile "as seen by another user" to somebody who should not have it.
4.3 Queues, 202 and task resources
Fan-out distribution and other long operations do not fit inside the request cycle. Just as in 04-06 with the store's reports:
POST /v1/users/usr_10/export HTTP/1.1
HTTP/1.1 202 Accepted
Location: /v1/tasks/tsk_9001
Retry-After: 10
{"id": "tsk_9001", "status": "in_progress",
"_links": {"self": {"href": "/v1/tasks/tsk_9001"}}}Posting is the same thing in reverse: POST /v1/posts responds 201 immediately with the created post —the author already sees it on their profile— while the distribution to the inboxes happens behind the scenes. It is eventual consistency turned into a contract: a follower may take a few seconds to see it, and that is documented.
4.4 The cost of denormalisation
Denormalising is not free. At CafeSocial we pay for it: multiplied storage (a post appears in hundreds of thousands of inboxes), a more fragile write path (if the queue fails, there are incomplete inboxes and a repair process is needed), and two sources of truth that can diverge. The rule is that the canonical source is still /posts/{id}; the inboxes are a rebuildable cache. Anything that can be regenerated from the canonical source is acceptable to denormalise; anything that cannot, is not.
- User-generated content
5.1 Image upload with a pre-signed URL
The API does not receive the megabytes of the photo. An authorisation is issued for uploading directly to object storage:
sequenceDiagram
participant App as Mobile app
participant API as CafeSocial API
participant St as Storage
App->>API: POST /v1/posts/pst_2100/images {type, size}
API-->>App: 201 {uploadUrl, fields, expiresAt, imageId}
App->>St: PUT uploadUrl (photo bytes)
St-->>App: 200 OK
App->>API: POST /v1/posts/pst_2100/images/img_44/confirmation
API->>St: Verify type, size and file header
API-->>App: 200 {status: pending_moderation}
POST /v1/posts/pst_2100/images HTTP/1.1
Content-Type: application/json
{"contentType": "image/jpeg", "sizeBytes": 2411520}HTTP/1.1 201 Created
Location: /v1/posts/pst_2100/images/img_44
{
"id": "img_44",
"uploadUrl": "https://storage.cafesocial.example/uploads/img_44?signature=...",
"method": "PUT",
"expiresAt": "2026-08-14T09:27:00Z",
"maxSizeBytes": 5242880,
"status": "pending_upload"
}The reasons for not passing the bytes through the API are concrete:
- A Node process tied up for 8 seconds with a 5 MB upload is a process that is serving nobody else.
- Scaling the upload traffic is decoupled from scaling the business logic.
- Storage and the CDN already solve resumption, multipart uploads and geographical distribution.
- Proxy and load-balancer timeouts stop being a problem.
Validation. The pre-signed URL limits type and size, but that is not enough: the client declares image/jpeg and uploads whatever it likes. That is why confirmation is mandatory and checks the real bytes on the server (magic numbers in the file header, dimensions, size), reprocesses the image into several sizes and strips the EXIF metadata —which includes GPS coordinates: publishing the exact location of a user's home would be a personal data leak. An image left unconfirmed for 15 minutes is deleted and the post stays a draft.
5.2 Moderation
The status mirrors that of the Aroma Store reviews: pending_moderation | published | rejected. What changes is the scale, which forces two stages:
- Automatic, immediately: an image and text classifier. Low score, published straight away; medium score, human queue; high score, immediate rejection with the possibility of appeal.
- Human, over a work queue:
GET /v1/moderation/pending?limit=50&cursor=... HTTP/1.1
Authorization: Bearer <token with the moderation:read scope>
POST /v1/moderation/posts/pst_2100/approval HTTP/1.1
POST /v1/moderation/posts/pst_2100/rejection HTTP/1.1
Content-Type: application/json
{"reason": "unrelated_content", "notifyAuthor": true}Design details that matter:
- The moderation queue lives under
/moderation/*, a namespace of its own with its own scopes (moderation:read,moderation:write). Mixing it into/postswould force the same endpoint to behave radically differently depending on the scope, which is exactly what makes tests ungovernable. - Approval and rejection are action resources (
POSTto a sub-resource) because they are not idempotent in their effects: they fire notifications and they are audited. Every decision is stored with the moderator, a timestamp and a reason. - The queue is paginated by cursor: it grows and is consumed at the same time.
5.3 Reports
POST /v1/posts/pst_2100/reports HTTP/1.1
{"reason": "spam", "comment": "Posts the same link on every tasting"}We respond 202 and not 201 with a verdict: the report is accepted, not resolved. The reporter must not be able to see the detailed status or the moderator's identity, and the reported user must not learn who reported them. Several reports about the same post are aggregated into a single case.
An essential warning. Everything above is the engineering part, which is the easy part. A real service with user content has legal obligations that are not technical decisions: lawful bases and GDPR rights, deadlines for taking down illegal content, protection of minors and age verification, data retention and disclosure to authorities, moderation transparency and routes of appeal. They vary by country and over time. Design it with legal and compliance advice from the start; a deletion endpoint that misses the legal deadline is a legal problem, not a
TODOin the backlog.
- Privacy and resource-level authorisation
At Aroma Store the question was "who owns this?": order ord_5001 belongs to cus_842, so only cus_842 and the administrators see it. Binary and simple.
At CafeSocial the question is "who is asking and what do we let them see?", and the answer is no longer yes or no, but a different representation:
Who asks for usr_77 (private profile) |
What they get |
|---|---|
usr_77 themselves |
Everything, including drafts and email address |
| An accepted follower | Full profile and posts |
| An authenticated user who does not follow them | Handle, photo, bio, counters; no posts |
A user blocked by usr_77 |
404 user_not_found |
| Unauthenticated | 404 if the profile is private; a reduced card if it is public |
6.1 Projection by visibility
The pattern: the service layer returns the full entity and a projection decides which fields survive, according to the relationship between observer and observed.
// presentation/projections/user.js
const FIELDS = {
owner: ['id','handle','name','bio','photo','email','followers','following','private'],
follower: ['id','handle','name','bio','photo','followers','following','private'],
public: ['id','handle','photo','bio','followers','private'],
};
function projectUser(user, context) {
const level = visibilityLevelFor(user, context); // owner|follower|public|hidden
if (level === 'hidden') throw new ApiError(404, 'user_not_found', 'Not found.');
const output = {};
for (const field of FIELDS[level]) output[field] = user[field];
output._links = linksForLevel(user, level);
return output;
}Three rules that avoid the usual failures:
- An allowlist, never a blocklist. A new field on the entity must not leak just because it has been added.
- The projection applies inside lists too.
GET /users/usr_77/followerstrims out blocked and private users: the list returned may be shorter than the counter shown on the profile, and that is correct. - Visibility is decided in the query, not afterwards. Filtering in memory after fetching 50 rows breaks cursor pagination: you would ask for 20 and return 14.
6.2 404 instead of 403
If usr_77 blocks usr_10 and the latter requests GET /v1/users/usr_77, a 403 insufficient_permissions would be technically honest and it would leak information: it would confirm that the account exists, and by the difference between responses you could enumerate who has blocked whom or discover which handles are registered. When the mere existence of the resource is sensitive information, you respond 404.
The general criterion: 403 when the resource is known to the caller or its existence reveals nothing (the moderator without the right scope); 404 when revealing the existence is itself a leak. And you have to be consistent in response times: a 404 that takes 5 ms when the resource does not exist and 40 ms when it exists but is blocked leaks the information again through a side channel.
6.3 The effect on caching and tests
- Caching: any response whose shape depends on the observer is
privateand carriesVary: Authorization. A shared CDN can only serve what is objectively public. This reduces performance and it is a cost accepted consciously. - Tests: the matrix grows all at once. Every sensitive endpoint needs cases for owner, follower, non-follower, blocked and anonymous. At CafeSocial it is handled with a parameterised table of cases:
describe('GET /v1/users/usr_77 by observer', () => {
const cases = [
{ observer: 'owner', status: 200, includes: ['email'], excludes: [] },
{ observer: 'follower', status: 200, includes: ['name'], excludes: ['email'] },
{ observer: 'stranger', status: 200, includes: ['handle'], excludes: ['name','email'] },
{ observer: 'blocked', status: 404, includes: [], excludes: [] },
];
for (const c of cases) it(`observer ${c.observer} → ${c.status}`, async () => { /* ... */ });
});
- Notifications and real time
At Aroma Store, SSE was an extra for the internal panel. Here, immediacy is the product: a social network where the notification arrives two minutes late is perceived as broken.
Pure REST would force polling: the app asks GET /v1/notifications every 10 seconds. With 500,000 active users that is 50,000 requests per second, 99 % of which return the same thing. Not even with ETag and 304 (which save bandwidth, not requests) do the numbers add up.
The design combines three pieces:
GET /v1/notifications?limit=30&cursor=... HTTP/1.1
{"data": [
{"id": "ntf_55", "type": "like", "read": false,
"actor": {"id": "usr_77", "handle": "geisha_taster"},
"resource": {"href": "/v1/posts/pst_2100"},
"createdAt": "2026-08-14T09:12:33Z"}
],
"unread": 4,
"pagination": {"next": "eyJ2IjoidjEi...", "hasMore": true}}The history is paginated by cursor (the same mechanism as the timeline) and the live channel only carries notices that there is something new, not the full content: the client receives the signal and reloads over REST. That way the real-time channel does not turn into a second API to maintain in parallel.
| Mechanism | Direction | When to choose it at CafeSocial | Cost |
|---|---|---|---|
Polling with ETag/304 |
Client→server | A fallback when everything else fails; non-urgent data | Constant requests |
| Long polling | Client→server | Compatibility with hostile networks | Held-open connections |
| SSE | Server→client | Unread counter, new posts: it is what the web uses | Unidirectional; automatic reconnection and Last-Event-ID out of the box |
| WebSockets | Bidirectional | Direct messages with "typing…" and read receipts | Its own infrastructure, per-connection state, harder to scale |
| Webhooks (01-07) | Server→server | Integrations: telling Aroma Store about a tasting with a high score | Retries, HMAC signature, "at least once" delivery |
| Push notifications (APNs/FCM) | Server→device | App closed or in the background | A dependency on external platforms |
CafeSocial uses SSE for the web and the notification counter, WebSockets only on the direct messages screen, push when the app is not active and webhooks signed with HMAC-SHA256 towards Aroma Store, reusing exactly the mechanism from 06-01.
- Aroma Store versus CafeSocial, decision by decision
This is the table that makes sense of the two lessons together.
| Decision | Aroma Store | CafeSocial | What changes it |
|---|---|---|---|
| Read/write ratio | ~10:1 | ~500:1 | Justifies precomputing and denormalising |
| Pagination | limit/offset on coffees; cursor on orders |
Cursor mandatory on every feed | The set changes between requests |
total on collections |
Yes, useful and cheap | Does not exist on feeds | There is no closed set to count |
| Consistency | Strong and transactional (stock, payments) | Eventual and documented | There is no money in the request |
| Counters | Exact by definition | Approximate by design | The cost of counting against the value of exactness |
| Authorisation | "who owns it?" → 200/403 |
"who is asking?" → projections and 404 |
The existence of the resource is information |
| Shape of the response | Stable for everybody | Variable by observer | Privacy and blocks |
| Caching | public with ETag on the catalogue |
Almost everything private + Vary |
It depends on the observer |
| Writes | Synchronous, with optimistic concurrency | Asynchronous, queue and 202 |
Fan-out does not fit inside the request |
| Cost of fan-out | Non-existent | Dominant: hybrid push/pull | The follower distribution has a long tail |
| Files | No significant uploads | Pre-signed URL, outside the API | Volume and occupation time |
| Real time | SSE as an extra for the panel | SSE + WebSockets as the product | Immediacy is the perceived value |
| Moderation | Reviews, low volume, manual review | Two stages, dedicated queue, reports | 40,000 items a day |
| Relationships | Simple containment (order→customer) | A graph: the edge is a resource with PUT/DELETE |
180 M edges with semantics of their own |
| Idempotency | Idempotency-Key on payments |
PUT, idempotent by design across the graph |
Retries on mobile networks |
The conclusion is not that one design is better. It is that there is no universal REST design: there are domain-dependent decisions, and professional competence consists in knowing which one applies and being able to justify it. If somebody offers you "the right way to paginate" without asking about the domain, they are selling you an answer before they have heard the question.
- Why GraphQL is defensible here
In 01-07 we concluded that for Aroma Store GraphQL would have been complexity with no return: few screen types, high value from HTTP caching on the catalogue, one very stable main consumer. At CafeSocial the arguments flip sign:
- Screens with heterogeneous data. The detail view of a post needs the post, the author, whether I follow them, the first five comments with their authors, the "like" counter, whether I liked it, the hashtags and the linked catalogue coffee. In REST that is 5-7 requests, or a bespoke composite endpoint that ages badly.
- Mobile clients with limited bandwidth. The list needs 6 fields per post; the detail view, 30. With REST you end up inventing
?fields=or?view=summary, which is GraphQL done badly. - Fast client evolution. A redesign of the timeline every quarter means, in REST, negotiating contract changes with the backend every quarter.
- A graph is queried like a graph. "The latest comments from the people I follow on posts that my followers also liked" is a natural query in GraphQL and a contorted endpoint in REST.
And what you lose, which has to go on the same scales:
| What you gain | What you lose |
|---|---|
| One request per screen | Intermediate HTTP caching: everything is POST /graphql |
| The client chooses the fields | ETag/304 and CDNs stop helping |
| Evolution without versioning routes | Status codes: almost everything is 200 with errors |
| A typed, self-documenting schema | The client can build extremely expensive queries: you need depth limits, complexity limits and persisted queries |
| A single entry point | Per-endpoint observability and rate limiting stop working as they are |
| — | The N+1 problem in the resolvers: it forces DataLoader from day one |
The realistic decision at CafeSocial —and the one you will see in many companies— is hybrid: GraphQL for the mobile app's screens, REST for what is public and indexable (profiles and posts cacheable in a CDN), for third-party integrations and for webhooks. Choosing GraphQL is not abandoning what you have learned: the resources, the states, idempotency, cursor pagination and observer-based authorisation remain exactly the same problems; only the transport layer changes.
Common Mistakes and Tips
- Using
POST /followbecause "it is an action". You lose idempotency, and querying and deleting for free. If the action creates or destroys a relationship between two identifiable entities, that relationship has a URI:PUT/DELETE(02-03). - Putting authorisation information inside the cursor. A cursor with a
userIdinside is a privilege escalation waiting to happen. The subject always comes from the token. - Confusing opaque with secure. Base64url encrypts nothing. Opacity is freedom to change the format, not protection: always validate the decoded content.
- Returning a
totalon a feed. It forces an expensiveCOUNTover a set that changes, and the resulting number is false the moment it is sent. UsehasMore. - Filtering by visibility after paginating. You ask for 20, hide 6 and return 14: the client thinks it is reaching the end. Visibility goes in the query.
- A blocklist of fields in the projections. The day you add
recoveryEmailto the entity, it will publish itself. Always an allowlist. 403where existence already leaks information. And watch out for response times and error messages too, which leak through side channels.- Caching as
publica response that depends on the observer. It is the fastest route to a proxy serving one user's private profile to another. When in doubt,private. - Uploading large files through the API. It ties up processes, collides with proxy timeouts and does not scale. Pre-signed URL and later confirmation.
- Trusting the
Content-Typedeclared by the client. Verify the real bytes, reprocess the image and strip the EXIF metadata before publishing it. - Designing fan-out with a single algorithm. Pure push dies with heavily followed accounts; pure pull dies on latency. The hybrid is ugly and it is the one that works.
- A process tip: write the comparison table from section 8 before you start coding. It forces you to justify every decision against a concrete alternative and it is the best defence against copying the last project's design out of inertia.
Exercises
Exercise 1 — The graph as a resource. Design the endpoints for "muting" a user you follow (you stop seeing their posts in your timeline, but you remain their follower). State the method, the URI, the status codes and whether it needs a body. Justify why you would not model it as POST /v1/mute.
Exercise 2 — A tamper-proof cursor. A client sends ?cursor=eyJ2IjoidjEiLCJ0IjoiMjAzMC0wMS0wMVQwMDowMDowMFoiLCJpIjoicHN0Xzk5OTk5In0 (a date in the future). Explain what the system returns, why that is not a security flaw, and add the missing validation to decodeCursor.
Exercise 3 — Projection and pagination together. GET /v1/users/usr_77/followers?limit=20 must hide the users who have blocked the observer. Explain why filtering in memory breaks pagination and sketch the correct SQL query with a cursor.
Solutions
Solution 1. Muting is an attribute of an existing relationship, not a new relationship:
PATCH /v1/users/usr_10/following/usr_77
Content-Type: application/json
If-Match: "w/rel-usr10-usr77-3"
{"muted": true}
HTTP/1.1 200 OK
ETag: "w/rel-usr10-usr77-4"
{"userId":"usr_10","followedId":"usr_77","muted":true,"createdAt":"2026-03-04T10:22:11Z"}Codes: 200 on success; 404 if you do not follow that user (there is no relationship to modify); 412 if the If-Match does not match; 422/400 with invalid_data if the body does not validate. POST /v1/mute is not used because the relationship already has a URI: creating a parallel verb would duplicate the resource, lose the idempotency of a PATCH against a specific state and force you to invent POST /v1/unmute. An equally valid alternative: a PUT over the complete relationship with all its attributes, if you prefer to avoid PATCH.
Solution 2. The system decodes the JSON correctly (a valid v, t and i present), the age check does not trigger because the date is in the future, and the query WHERE (created_at, id) < ('2030-01-01', 'pst_99999') simply returns the first page, since everything is earlier than that date. It is not a security flaw because the userId in the WHERE comes from the token, not from the cursor: tampering with it only lets you reposition yourself within your own inbox, something the user can already do by paginating. Even so, it is worth rejecting it in order to detect broken clients:
const t = Date.parse(data.t);
if (Number.isNaN(t)) {
throw new ApiError(400, 'invalid_cursor', 'The cursor is not valid.');
}
if (t > Date.now() + 60_000) { // 1 min of slack for clock skew
throw new ApiError(400, 'invalid_cursor', 'The cursor is not valid.');
}
if (typeof data.i !== 'string' || !/^pst_[0-9]+$/.test(data.i)) {
throw new ApiError(400, 'invalid_cursor', 'The cursor is not valid.');
}Solution 3. Filtering in memory breaks pagination because the LIMIT is applied before the filter: you ask for 20 rows, discard the 6 belonging to users who have blocked you and return 14, while the cursor advances as if you had delivered 20. The client sees irregular pages and, if a whole page comes out empty, reads it as the content having run out. The filter has to go in the query:
SELECT f.follower_id, f.created_at
FROM follows f
WHERE f.followed_id = :profileId
AND NOT EXISTS (
SELECT 1 FROM blocks b
WHERE b.blocker_id = f.follower_id
AND b.blocked_id = :observerId )
AND (f.created_at, f.follower_id) < (:cursorDate, :cursorId)
ORDER BY f.created_at DESC, f.follower_id DESC
LIMIT :limit + 1;A consequence that has to be documented: the number of items walked may not match the profile's followers counter, because that counter is global and the list is relative to the observer. And because the response depends on the observer, it is private with Vary: Authorization.
Conclusion
We have designed CafeSocial with the same tools as Aroma Store —resources, methods, status codes, headers, hypermedia— and we have ended up with an almost opposite design. The follow relationship stopped being a field and became a resource with an idempotent PUT and a DELETE. The timeline stopped being a collection and became a derived resource with no POST, no total and no page N, held up by a hybrid fan-out and reachable only through an opaque cursor of timestamp plus id. The counters stopped being exact because counting stopped being worth it. The images left the API. Authorisation stopped asking who owns the resource and started asking who is observing, with projections by visibility and 404 wherever existence is itself information. And real time stopped being an ornament and became the product.
If this lesson leaves one single idea, let it be the one in the table in section 8: there is no universal REST design, there are domain-dependent decisions, and a professional is distinguished by being able to name the alternative they rejected and why. That is also the reason GraphQL, indefensible for the store's catalogue, is a serious option here, as long as the price is accepted: losing HTTP caching, status codes and control over the cost of queries.
One last matter remains, and it is the one that separates a well-designed API from an API that is still alive three years later: what happens after you publish it. In 06-03, Evolving and Maintaining an API in Production, we will look at what to do when an apparently harmless change breaks a consumer in production, how to write a post-mortem that is actually useful, how contract debt is accumulated and paid off, and how a migration to v2 is planned with its deprecation period and its governance. Because designing an API well is hard, but changing it without breaking whoever already uses it is harder still.
REST API Course: Principles of Designing and Developing RESTful APIs
Module 1: Introduction to RESTful APIs
- What Is an API?
- History and Evolution of APIs
- HTTP Fundamentals for APIs
- Basic Principles of REST
- The Richardson Maturity Model and HATEOAS
- REST vs. SOAP
- REST Compared with GraphQL, gRPC and Webhooks
Module 2: Designing RESTful APIs
- RESTful API Design Principles
- Resources and URIs
- HTTP Methods
- HTTP Status Codes
- Representations, Headers and Content Negotiation
- Filtering, Sorting, Pagination and Search
- API Versioning
- API Documentation
Module 3: Building RESTful APIs
- Setting Up the Development Environment
- Building a Basic Server
- Handling Requests and Responses
- Input Data Validation
- Persistence and the Data Access Layer
- Authentication and Authorisation
- Error Handling
- Testing and Validation
Module 4: Best Practices and Security
- API Design Best Practices
- Security in RESTful APIs
- OAuth 2.0 and OpenID Connect in Practice
- Rate Limiting and Throttling
- CORS and Security Policies
- HTTP Caching and Performance
- Observability: Logs, Metrics and Traces
Module 5: Tools and Frameworks
- Postman for API Testing
- Swagger and OpenAPI for Documentation
- Popular Frameworks for RESTful APIs
- Contracts, Mocks and Automated API Testing
- Continuous Integration and Deployment
- API Gateways and Developer Portals
