So far you have worked at the transport level. You have opened sockets, you have designed BTCP/1 and BTDP/1 from scratch, you have decided where each message ends and what codes the server returns. That is excellent for understanding how things work, but it has an obvious problem: nobody else speaks your protocols. BTCP/1 is understood only by the client you wrote yourself.
This lesson moves up a level. We are going to stop inventing protocols and use the one that already rules the planet: HTTP. Every external service BiblioTech might want to talk to —a metadata catalogue by ISBN, a cover service, a mail gateway, any API— speaks HTTP. And the good news is that HTTP is not magic: it is exactly what travels over the socket of 09-02, plain text with an agreed format, delimited by lines. You are going to see it written byte by byte.
Java has two APIs for speaking HTTP. The old one, HttpURLConnection, from 1996, is the one we study today. The modern one, Java 11's HttpClient, is that of 09-06. And it is worth saying honestly up front: HttpURLConnection is an old, verbose API full of traps. You would not use it to write new code. But it is in the whole standard library, it appears in mountains of legacy code, and its concepts —methods, headers, status codes, input and error streams— are the same ones you will need in the modern API. Learning it means understanding HTTP hands-on.
By the end, BiblioTech will query an external metadata service by ISBN and download a book's cover to disk.
Contents
- Anatomy of a URL
URLversusURI- Encoding parameters with
URLEncoder - HTTP in depth: the request
- HTTP in depth: the response
- A real exchange, byte by byte
- HTTP methods
- Status codes
- Key headers
- The simple case:
URL.openStream() HttpURLConnection: real control- Time limits: never without them
getInputStreamversusgetErrorStream- Sending a body:
POSTwithsetDoOutput - Redirects
- Compression and HTTPS
- BiblioTech: metadata client and cover downloads
- An honest assessment of this API
- Common Mistakes and Tips
- Exercises
- Anatomy of a URL
A URL (uniform resource locator) says where something is and how to reach it. It has six parts, and it is worth being able to name them all.
https://api.nexussoftware.com:8443/v1/books/search?isbn=978-0000000001&format=json#summary
\___/ \____________________/\__/\_______________/\__________________________/\_______/
| | | | | |
scheme host port path query fragment| Part | Example | Notes |
|---|---|---|
| Scheme | https |
The protocol. It determines the default port |
| Authority | api.nexussoftware.com:8443 |
Host and port; it may include user:password@ (obsolete and insecure) |
| Host | api.nexussoftware.com |
Name or IP. Resolved by DNS (09-01) |
| Port | 8443 |
If omitted, the scheme's: 80 for http, 443 for https |
| Path | /v1/books/search |
Which resource is being asked for |
| Query | isbn=978-...&format=json |
Parameters, key=value separated by & |
| Fragment | summary |
It is never sent to the server. It is for the client |
That last point surprises a lot of people: the fragment (#something) is purely local. The browser uses it to scroll to a section of the page; the server never sees it. If you try to use the fragment to pass information to a service, it will not get there.
In Java:
import java.net.URL;
URL url = new URL("https://api.nexussoftware.com:8443/v1/books/search"
+ "?isbn=978-0000000001&format=json#summary");
System.out.println("Protocol : " + url.getProtocol()); // https
System.out.println("Host : " + url.getHost()); // api.nexussoftware.com
System.out.println("Port : " + url.getPort()); // 8443
System.out.println("Default : " + url.getDefaultPort()); // 443
System.out.println("Path : " + url.getPath()); // /v1/books/search
System.out.println("Query : " + url.getQuery()); // isbn=978-...&format=json
System.out.println("Fragment : " + url.getRef()); // summary
System.out.println("File : " + url.getFile()); // path + ? + queryCareful with
getPort(). It returns-1if the port does not appear explicitly in the URL, not the default port. To get the effective port you have to combine it withgetDefaultPort():int port = url.getPort() != -1 ? url.getPort() : url.getDefaultPort();Forgetting it produces the classic attempt to connect to port -1.
URL versus URI
URL versus URIJava has two classes for this and the difference matters.
java.net.URL |
java.net.URI |
|
|---|---|---|
| What it represents | A resource that can be accessed | An identifier, nothing more |
| Can open connections | Yes (openConnection, openStream) |
No |
| Validates the syntax | Barely | Yes, strictly |
| Needs to know the scheme | Yes: new URL("foo://x") throws |
No: any scheme is fine |
equals() |
It performs DNS resolution (!) | Text comparison |
Normalises paths (.., .) |
No | Yes, with normalize() |
| Since which version | 1.0 | 1.4 |
The point that bites hardest is URL's equals():
URL a = new URL("http://example.com/page");
URL b = new URL("http://93.184.216.34/page");
// This comparison PERFORMS A DNS LOOKUP and may take seconds.
// And it returns true if both names resolve to the same address.
boolean same = a.equals(b);URL.equals() and URL.hashCode() resolve the host by DNS. That means that:
- Comparing two URLs may block for seconds.
- Putting a
URLin aHashSetor as aHashMapkey causes DNS lookups on insertion and on lookup. - With no network, the behaviour changes.
It is an acknowledged design mistake, and the practical rule is simple: use URI to represent, manipulate and compare; convert to URL only at the moment of opening the connection.
import java.net.URI;
import java.net.URL;
// 1. Build and manipulate with URI: strict validation, no DNS.
URI uri = new URI("https", "api.nexussoftware.com", "/v1/books/search",
"isbn=978-0000000001", null);
// scheme host path query fragment
System.out.println(uri); // https://api.nexussoftware.com/v1/books/search?isbn=978-0000000001
// 2. Convert to URL only in order to connect.
URL url = uri.toURL();URI also has utilities that URL does not have:
URI base = URI.create("https://api.nexussoftware.com/v1/");
URI full = base.resolve("books/978-0000000001");
// -> https://api.nexussoftware.com/v1/books/978-0000000001
URI messy = URI.create("https://api.nexussoftware.com/v1/../v2/./books");
System.out.println(messy.normalize());
// -> https://api.nexussoftware.com/v2/booksNote. In Java 20 the
URLconstructors were deprecated, precisely to push towardsURI.create(...).toURL(). If you compile with a recent version and see deprecation warnings aboutnew URL(...), that is the reason, and the solution is the one we have just recommended.
- Encoding parameters with
URLEncoder
URLEncoderA URL only accepts a restricted set of characters. Spaces, accented letters, the signs &, =, ? and #, and any non-ASCII character must be encoded in the form %XX, where XX is the hexadecimal value of the byte.
Without encoding, things fail in creative ways:
Searching for the title "Naïve Set Theory":
WRONG: /search?title=Naïve Set Theory
-> The space breaks the HTTP request (the server thinks the path
ends at "Naïve" and that "Set" is the protocol version).
-> The ï, unencoded, arrives as bytes the server may
interpret in another encoding.
RIGHT: /search?title=Na%C3%AFve+Set+Theory
-> The space is +, and the ï is its two UTF-8 bytes: C3 AF.And the genuinely dangerous case: a value containing & or = injects parameters.
Searching for author = "Bloch & Gamma"
WRONG: /search?author=Bloch & Gamma&admin=true
-> The server sees TWO parameters: author="Bloch " and " Gamma"
-> And if the attacker writes "&admin=true", they inject a parameter.
RIGHT: /search?author=Bloch+%26+GammaIn Java:
import java.net.URLEncoder;
import java.net.URLDecoder;
import java.nio.charset.StandardCharsets;
String title = "Naïve Set Theory";
String encoded = URLEncoder.encode(title, StandardCharsets.UTF_8);
System.out.println(encoded); // Na%C3%AFve+Set+Theory
String back = URLDecoder.decode(encoded, StandardCharsets.UTF_8);
System.out.println(back); // Naïve Set TheoryThe charset is mandatory. There is an overload
URLEncoder.encode(String)with no charset, deprecated for decades, that uses the platform encoding. With it, the same URL comes out differently on Linux and on Windows. Always use the two-argument version withStandardCharsets.UTF_8, which is also what any modern server expects.
An important trap
URLEncoder is designed for form values (application/x-www-form-urlencoded), not for paths. Its visible difference: it encodes the space as +, not as %20.
| Context | Space | Correct tool |
|---|---|---|
| Value of a query parameter | + (or %20, both valid) |
URLEncoder |
| Segment of the path | %20. A + in the path is a literal + |
URI with the multi-argument constructor |
// WRONG: in the path, + does not mean space.
String path = "/books/" + URLEncoder.encode("Effective Java", UTF_8);
// -> /books/Effective+Java (the server will look for a book with a + in the title)
// RIGHT: let URI encode the path correctly.
URI uri = new URI("https", "api.nexussoftware.com",
"/books/Effective Java", null, null);
// -> https://api.nexussoftware.com/books/Effective%20JavaURI's multi-argument constructor encodes each component according to its own rules. It is the correct form and the one we will use in BiblioTech.
A helper for building query strings:
package com.nexussoftware.bibliotech.network;
import java.net.URLEncoder;
import java.nio.charset.StandardCharsets;
import java.util.LinkedHashMap;
import java.util.Map;
/**
* Builder of query strings with correct encoding.
* LinkedHashMap so the order is predictable (useful when debugging and caching).
*/
public class UrlQuery {
private final Map<String, String> parameters = new LinkedHashMap<>();
public UrlQuery with(String key, String value) {
if (value != null) {
parameters.put(key, value);
}
return this; // chainable
}
public UrlQuery with(String key, int value) {
parameters.put(key, String.valueOf(value));
return this;
}
/** Returns "a=1&b=2" with everything encoded, or "" if there are no parameters. */
public String build() {
StringBuilder sb = new StringBuilder();
for (Map.Entry<String, String> e : parameters.entrySet()) {
if (sb.length() > 0) {
sb.append('&');
}
// BOTH the key AND the value are encoded: a key with
// odd characters breaks the request just as a value does.
sb.append(URLEncoder.encode(e.getKey(), StandardCharsets.UTF_8));
sb.append('=');
sb.append(URLEncoder.encode(e.getValue(), StandardCharsets.UTF_8));
}
return sb.toString();
}
}String query = new UrlQuery()
.with("title", "Naïve Set Theory")
.with("author", "Bloch & Gamma")
.with("max", 10)
.build();
// title=Na%C3%AFve+Set+Theory&author=Bloch+%26+Gamma&max=10
- HTTP in depth: the request
Here comes the moment when everything fits together. An HTTP request is plain text sent over a TCP socket, with an agreed format. Exactly what you have known how to do since 09-02.
GET /v1/books/978-0000000001 HTTP/1.1\r\n <- request line
Host: api.nexussoftware.com\r\n <- headers
Accept: application/json\r\n
User-Agent: BiblioTech/1.0\r\n
Connection: close\r\n
\r\n <- EMPTY LINE: end of headers
<- (the body would go here, if there were one)Four elements:
- Request line:
METHOD PATH VERSION. Separated by a space. - Headers:
Name: value, one per line. The names are case-insensitive. - Empty line: marks the end of the headers. Mandatory.
- Body (optional): the data, in
POST,PUT,PATCH.
Details that matter:
- The delimiter is
\r\n, not\n. HTTP demands it strictly. If you write this by hand with aPrintWriterandprintln(), on Linux you will send only\nand some servers will reject it. It is exactly the warning of 09-02 about the platform line separator. - The
Hostheader is mandatory in HTTP/1.1. It is what lets a single IP serve hundreds of different sites (virtual hosting): the server decides which site to serve by looking at that header. - The empty line is essential. Without it, the server goes on waiting for headers and your request is never processed: the same deadlock as the forgotten
flushof 09-02, from a different cause.
You can write it by hand right now
That is a complete HTTP client written with nc. And with what you know from 09-02, you could write it in Java in twenty lines. HTTP is nothing more than a text protocol over TCP, like the BTCP/1 you designed; the difference is that this one is understood by half the planet.
- HTTP in depth: the response
The response has the same structure, with a different first line.
HTTP/1.1 200 OK\r\n <- status line
Content-Type: application/json; charset=utf-8\r\n <- headers
Content-Length: 79\r\n
Date: Wed, 05 Aug 2026 09:14:22 GMT\r\n
Server: nginx/1.24.0\r\n
\r\n <- empty line
{"isbn":"978-0000000001","title":"Effective Java","author":"Bloch","pages":416}- Status line:
VERSION CODE TEXT. The code is what matters; the text is informative. - Headers: as in the request.
- Empty line.
- Body: the content.
How does the client know where the body ends? It is the same question as the delimiter problem of 09-01, and HTTP solves it in three ways:
| Mechanism | How | When it is used |
|---|---|---|
Content-Length: 79 |
Length prefix: exactly 79 bytes | The usual one, when the size is known |
Transfer-Encoding: chunked |
Chunks, each preceded by its size in hexadecimal, and a chunk of size 0 at the end | When the size is not known in advance (generated content) |
| Connection close | The body ends when the server closes | HTTP/1.0, or Connection: close with no length |
The three techniques of the delimiter section of 09-01, all together. HTTP is a case study in protocol design, and you now have the context to appreciate it.
- A real exchange, byte by byte
Let us see it for real, with no Java in between. Start a local server that shows what reaches it:
Terminal 1:
Terminal 2:
In terminal 1 exactly what curl sent appears:
Now type the reply yourself in terminal 1 (remember the empty line before the body) and press Ctrl+D:
HTTP/1.1 200 OK
Content-Type: application/json; charset=utf-8
Content-Length: 50
{"isbn":"978-0000000001","title":"Effective Java"}And in terminal 2, curl shows the complete exchange:
* Connected to localhost (127.0.0.1) port 8080
> GET /v1/books/978-0000000001 HTTP/1.1
> Host: localhost:8080
> User-Agent: curl/8.5.0
> Accept: */*
>
< HTTP/1.1 200 OK
< Content-Type: application/json; charset=utf-8
< Content-Length: 50
<
{"isbn":"978-0000000001","title":"Effective Java"}You have just played HTTP server with nc, just as in 09-02 you played BTCP server. It is the same exercise, with another protocol. And it makes the central idea of this lesson crystal clear: HTTP is text over a socket, and everything HttpURLConnection does is format that text for you and parse the response.
sequenceDiagram
participant C as Client (BiblioTech)
participant S as Metadata server
Note over C,S: TCP connection: three-way handshake (09-01)
C->>S: GET /v1/books/978-0000000001 HTTP/1.1
C->>S: Host: api.nexussoftware.com
C->>S: Accept: application/json
C->>S: (empty line)
Note over S: Processes the request
S-->>C: HTTP/1.1 200 OK
S-->>C: Content-Type: application/json
S-->>C: Content-Length: 79
S-->>C: (empty line)
S-->>C: {"isbn":"978-...","title":"Effective Java"}
Note over C,S: Close, or reuse if keep-alive
- HTTP methods
| Method | What for | Safe? | Idempotent? | Body? |
|---|---|---|---|---|
| GET | Retrieve a resource | Yes | Yes | No |
| HEAD | Like GET but headers only | Yes | Yes | No |
| POST | Create, or send data to be processed | No | No | Yes |
| PUT | Replace a whole resource | No | Yes | Yes |
| PATCH | Modify partially | No | No | Yes |
| DELETE | Delete a resource | No | Yes | Rare |
| OPTIONS | What can be done with this resource | Yes | Yes | No |
The two properties in the table have direct practical consequences:
Safe means it modifies nothing on the server. A GET can be repeated, cached and prefetched with no consequences. That is why it is a serious mistake to use GET for actions that change something: a web crawler, or the browser's prefetching, would run the action without anybody asking for it.
Idempotent means that repeating it produces the same result as doing it once. And this decides whether you can retry after a timeout:
GET,PUT,DELETE,HEAD: retrying is safe.POSTandPATCH: retrying may duplicate the operation.
If you send a POST that registers a loan and the wait times out, you do not know whether the server processed it or not. Retrying may create two loans. It is exactly the same reasoning that in 09-04 led to leaving LEND on TCP instead of UDP, applied here. The professional solution is the idempotency key: the client generates a unique identifier, sends it in a header, and the server rejects the second request with the same key. We mention it in 09-06 when we talk about retries.
- Status codes
The first digit indicates the family, and that lets you decide without reading the text — the same idea you applied in BTCP/1.
| Family | Meaning | What to do |
|---|---|---|
| 1xx | Informational | Rare; you will almost never see it in practice |
| 2xx | Success | Process the response |
| 3xx | Redirect | Go somewhere else (often automatic) |
| 4xx | Client error | You got it wrong. Do not retry without changing something |
| 5xx | Server error | They got it wrong. Retrying may make sense |
The ones you will really see:
| Code | Name | When |
|---|---|---|
| 200 | OK | All good |
| 201 | Created | Resource created (typical reply to a POST) |
| 204 | No Content | Fine, but there is no body (typical of DELETE) |
| 301 | Moved Permanently | It moved for good. Update your links |
| 302 | Found | It moved temporarily |
| 304 | Not Modified | It has not changed since your last query; use your cache |
| 400 | Bad Request | Your request is malformed |
| 401 | Unauthorized | You have not authenticated. The name is misleading |
| 403 | Forbidden | You have authenticated, but you have no permission |
| 404 | Not Found | The resource does not exist |
| 405 | Method Not Allowed | That resource does not accept that method |
| 409 | Conflict | State conflict (the same 409 as in BTCP/1) |
| 429 | Too Many Requests | Rate limit exceeded. Look at Retry-After |
| 500 | Internal Server Error | Generic server failure |
| 502 | Bad Gateway | An intermediary got no response from the real server |
| 503 | Service Unavailable | Overloaded or under maintenance. Usually temporary |
| 504 | Gateway Timeout | An intermediary ran out of patience |
Which ones deserve a retry (you will apply this in 09-06):
| Code | Retry? | Reason |
|---|---|---|
| 429 | Yes, waiting whatever Retry-After says |
It is what the server asks of you |
| 502, 503, 504 | Yes, with increasing backoff | Transient infrastructure failures |
| 500 | With caution | It may be a deterministic failure that will repeat |
| 4xx in general | No | Retrying the same thing gives the same thing |
And the number-one classic mistake with this API, which deserves its own box:
A 4xx or 5xx code does NOT throw an exception. A 404 or 500 response is a perfectly valid HTTP response, correctly delivered. From the network's point of view, everything went fine. You have to check
getResponseCode()by hand, always. Assuming "if there was no exception, it went well" is the cause of an enormous number of bugs.
- Key headers
| Header | Direction | What for |
|---|---|---|
Host |
Request | Mandatory in HTTP/1.1. Which site is being asked for |
Accept |
Request | Which formats the client accepts: application/json |
Accept-Encoding |
Request | Which compressions it accepts: gzip, deflate |
Accept-Language |
Request | Preferred languages: en-GB, en;q=0.9 |
User-Agent |
Request | Who you are. Set an identifying one, not Java's default |
Authorization |
Request | Credentials: Bearer <token> or Basic <base64> |
Content-Type |
Both | What format the body has |
Content-Length |
Both | How many bytes the body has |
Content-Encoding |
Response | How the body is compressed |
Location |
Response | Where to go on a redirect (3xx) |
Retry-After |
Response | How many seconds to wait before retrying (429, 503) |
Cache-Control |
Both | Caching policy |
ETag |
Response | Version identifier, for conditional requests |
Set-Cookie / Cookie |
Response / Request | Session state |
Two concrete tips:
Set an identifying User-Agent. By default Java sends something like Java/17.0.9, which many services block because they associate it with badly behaved crawlers. A BiblioTech/1.0 (+https://nexussoftware.com/bibliotech) identifies you and gives whoever administers the service a way of contacting you if something goes wrong.
Read the response's Content-Type to find the charset. It is the only correct way of decoding the body:
If you ignore it and assume UTF-8, it will work 95 % of the time and give you corrupt characters the remaining 5 %. We will write a helper that extracts it.
- The simple case:
URL.openStream()
URL.openStream()To download something with no ceremony:
import java.io.BufferedReader;
import java.io.InputStream;
import java.io.InputStreamReader;
import java.net.URI;
import java.net.URL;
import java.nio.charset.StandardCharsets;
URL url = URI.create("http://localhost:8080/v1/books/978-0000000001").toURL();
// openStream() = openConnection().getInputStream(), abbreviated.
try (BufferedReader reader = new BufferedReader(
new InputStreamReader(url.openStream(), StandardCharsets.UTF_8))) {
String line;
while ((line = reader.readLine()) != null) {
System.out.println(line);
}
}You will recognise the stack: InputStreamReader with an explicit charset over an InputStream, and a BufferedReader on top. It is the same one as in module 7 and the same one as in 09-02. Only the source of the InputStream changes.
Its limitations rule it out for almost everything:
| Limitation | Consequence |
|---|---|
| No time limit | It can block indefinitely. An absolute deal-breaker |
GET only |
No use for sending anything |
| Headers cannot be set | No Accept, no Authorization, no User-Agent |
| The status code cannot be seen | A 404 throws FileNotFoundException; a 500, IOException. Indistinguishable |
| Response headers cannot be seen | You do not know the Content-Type or the charset |
It is fine for a test main or for reading a local resource. For anything else, HttpURLConnection.
HttpURLConnection: real control
HttpURLConnection: real controlThe usage flow has a strict order that must be respected:
graph TD
A["URI.create(...).toURL()"] --> B["url.openConnection()<br/>does NOT connect yet"]
B --> C["cast to HttpURLConnection"]
C --> D["setRequestMethod<br/>setRequestProperty<br/>setConnectTimeout<br/>setReadTimeout<br/>setDoOutput"]
D --> E["write the body if it is a POST"]
E --> F["getResponseCode()<br/>HERE it really connects"]
F -->|2xx or 3xx| G["getInputStream()"]
F -->|4xx or 5xx| H["getErrorStream()"]
G --> I["disconnect()"]
H --> I
The two points that confuse everybody:
openConnection() does not connect. It returns a configuration object. The real connection happens on the first call that needs the response: getResponseCode(), getInputStream() or getHeaderFields(). That is why all the configuration must be done before those calls; afterwards, it is ignored silently or throws IllegalStateException.
You have to cast. openConnection() is declared to return URLConnection, and all the HTTP methods are on the HttpURLConnection subclass.
A complete and correct example:
package com.nexussoftware.bibliotech.network;
import java.io.BufferedReader;
import java.io.IOException;
import java.io.InputStream;
import java.io.InputStreamReader;
import java.net.HttpURLConnection;
import java.net.URI;
import java.net.URL;
import java.nio.charset.Charset;
import java.nio.charset.StandardCharsets;
import java.util.List;
import java.util.Map;
import java.util.logging.Logger;
/** An HTTP GET request with HttpURLConnection, done correctly. */
public class GetExample {
private static final Logger LOG = Logger.getLogger(GetExample.class.getName());
public static void main(String[] args) throws IOException {
URL url = URI.create("http://localhost:8080/v1/books/978-0000000001").toURL();
// 1. openConnection() does NOT connect: it returns a configuration object.
HttpURLConnection connection = (HttpURLConnection) url.openConnection();
try {
// 2. Configuration. ALL of this must go BEFORE getResponseCode().
connection.setRequestMethod("GET");
connection.setRequestProperty("Accept", "application/json");
connection.setRequestProperty("User-Agent", "BiblioTech/1.0");
connection.setRequestProperty("Accept-Charset", "UTF-8");
// NEVER without time limits. Both of them, and they differ (09-02).
connection.setConnectTimeout(5_000); // establishing the connection
connection.setReadTimeout(10_000); // waiting for data
// 3. HERE it really connects and reads the response.
int code = connection.getResponseCode();
String message = connection.getResponseMessage();
System.out.println("Status: " + code + " " + message);
// 4. Response headers.
System.out.println("Content-Type : " + connection.getContentType());
System.out.println("Content-Length: " + connection.getContentLengthLong());
for (Map.Entry<String, List<String>> e : connection.getHeaderFields().entrySet()) {
// CAREFUL: the null key holds the STATUS LINE.
// It is an API oddity that surprises you the first time.
String name = e.getKey() == null ? "(status line)" : e.getKey();
System.out.println(" " + name + ": " + e.getValue());
}
// 5. The body: getInputStream if it went well, getErrorStream if not.
// THIS distinction is classic mistake number two.
InputStream body = (code >= 200 && code < 400)
? connection.getInputStream()
: connection.getErrorStream();
if (body == null) {
System.out.println("(no body)");
return;
}
Charset charset = charsetOf(connection.getContentType());
try (BufferedReader reader = new BufferedReader(
new InputStreamReader(body, charset))) {
String line;
while ((line = reader.readLine()) != null) {
System.out.println(line);
}
}
} finally {
// 6. disconnect() releases the connection (or returns it to the internal pool).
connection.disconnect();
}
}
/**
* Extracts the charset from the Content-Type. Assuming UTF-8 works
* almost always and fails exactly when it hurts most.
*/
static Charset charsetOf(String contentType) {
if (contentType != null) {
for (String part : contentType.split(";")) {
String p = part.strip();
if (p.toLowerCase().startsWith("charset=")) {
String name = p.substring("charset=".length())
.replace("\"", "").strip();
try {
return Charset.forName(name);
} catch (Exception e) {
LOG.warning("Unknown charset: " + name + "; using UTF-8");
}
}
}
}
return StandardCharsets.UTF_8; // sensible fallback
}
}Notice the oddity of getHeaderFields(): the entry with the null key holds the status line. It is an API detail that shows up as soon as you iterate over it and that baffles you if you are not expecting it.
About disconnect(): its name is misleading. It does not necessarily close the socket, because Java keeps a pool of persistent connections internally. What it does is state that you have finished with that connection, and if you have consumed the whole body, the connection can be reused. If you walk away without reading the body, it really closes and you lose the reuse. That is why it is worth always reading the complete response, even when you do not care about it.
- Time limits: never without them
You have already seen this in 09-02 and 09-03, but with HTTP there is an aggravating factor: you are talking to a third-party service you do not control, over a network you do not control.
connection.setConnectTimeout(5_000); // establishing the TCP connection
connection.setReadTimeout(10_000); // waiting for data, per operation| Time limit | Covers | Without it |
|---|---|---|
setConnectTimeout |
The three-way handshake | It may take over a minute to give up |
setReadTimeout |
Each read operation | It may wait indefinitely |
The default value of both is 0, which means infinite. An external service that accepts the connection and never replies leaves your thread blocked forever. If that happens on a thread of a bounded pool, a few cases exhaust the pool and your application stops working without a single error in the log. It is exactly the scenario you know from 09-03, with the difference that here it depends on a third party.
And a limitation to be aware of: setReadTimeout is per operation, not total. A server that sends one byte every nine seconds keeps your read alive indefinitely without ever exhausting a ten-second deadline. HttpURLConnection has no total request time limit; the HttpClient of 09-06 does, with HttpRequest.timeout(), and it is one of its improvements.
Reasonable values:
| Type of service | Connection | Read |
|---|---|---|
| Internal, same network | 1-2 s | 3-5 s |
| External, fast API | 3-5 s | 10 s |
| External, heavy operation | 5 s | 30-60 s |
| Large file download | 5 s | 30 s (per operation, not total) |
getInputStream versus getErrorStream
getInputStream versus getErrorStreamThe second classic mistake, and one of the greatest time-wasters.
// BROKEN CODE.
int code = connection.getResponseCode();
InputStream input = connection.getInputStream(); // <-- THROWS on a 404With a 4xx or 5xx code, getInputStream() throws IOException (FileNotFoundException for the 404). And what is serious is what is lost: the error body, which almost always contains the explanation of what you did wrong.
HTTP/1.1 400 Bad Request
Content-Type: application/json
{"error":"invalid_isbn","message":"The ISBN must have 13 digits","field":"isbn"}That JSON is exactly what you need in order to debug, and getInputStream() stops you reading it. The correct form:
int code = connection.getResponseCode();
InputStream body = (code >= 200 && code < 400)
? connection.getInputStream()
: connection.getErrorStream();With two nuances:
getErrorStream()may returnnullif the server sent no error body. It has to be checked.getErrorStream()throws no exception, not even when there is nothing. It returnsnulland that is that.
A helper that settles the pattern once and for all:
/**
* Reads the response body, whether it comes through the normal stream or the error one.
* Returns an empty string if there is no body.
*/
static String readBody(HttpURLConnection connection, int code) throws IOException {
InputStream input = (code >= 200 && code < 400)
? connection.getInputStream()
: connection.getErrorStream();
if (input == null) {
return "";
}
Charset charset = charsetOf(connection.getContentType());
// readAllBytes is convenient but UNBOUNDED: a giant response
// exhausts memory. We bound it, as in 09-03.
try (InputStream stream = input) {
byte[] bytes = stream.readNBytes(MAX_BODY);
if (bytes.length == MAX_BODY) {
LOG.warning("Response truncated to " + MAX_BODY + " bytes");
}
return new String(bytes, charset);
}
}
- Sending a body:
POST with setDoOutput
POST with setDoOutputTo send data you have to enable output explicitly:
package com.nexussoftware.bibliotech.network;
import java.io.IOException;
import java.io.OutputStream;
import java.net.HttpURLConnection;
import java.net.URI;
import java.net.URL;
import java.nio.charset.StandardCharsets;
/** POST with a JSON body using HttpURLConnection. */
public class PostExample {
public static void main(String[] args) throws IOException {
URL url = URI.create("http://localhost:8080/v1/loans").toURL();
HttpURLConnection connection = (HttpURLConnection) url.openConnection();
try {
connection.setRequestMethod("POST");
// setDoOutput(true) is what enables getOutputStream().
// CAREFUL: it also changes the default method to POST, so
// if you call setDoOutput(true) after a setRequestMethod("GET"),
// the request goes out as POST. A classic trap of this API.
connection.setDoOutput(true);
connection.setRequestProperty("Content-Type", "application/json; charset=utf-8");
connection.setRequestProperty("Accept", "application/json");
connection.setRequestProperty("User-Agent", "BiblioTech/1.0");
connection.setConnectTimeout(5_000);
connection.setReadTimeout(10_000);
// The body, built by hand. Escaping the quotes and the
// backslashes is essential; doing it PROPERLY requires a JSON
// library, and that is 11-07.
String json = "{\"isbn\":\"978-0000000001\",\"employee\":\"Marta Ruiz\"}";
byte[] body = json.getBytes(StandardCharsets.UTF_8);
// Content-Length in BYTES, not in characters: with accents
// they differ, and a wrong value corrupts the request.
connection.setFixedLengthStreamingMode(body.length);
try (OutputStream output = connection.getOutputStream()) {
output.write(body);
output.flush(); // the usual flush
}
int code = connection.getResponseCode();
System.out.println("Status: " + code);
System.out.println("Body: " + readBody(connection, code));
} finally {
connection.disconnect();
}
}
}Four points that deserve attention:
setDoOutput(true) changes the method to POST implicitly. If you write setRequestMethod("GET") and then setDoOutput(true), the request goes out as POST. It is one of the most quoted traps of this API.
Content-Length is measured in bytes. "Naïve".length() is 5 characters but 6 bytes in UTF-8. Putting the length in characters corrupts the request.
The body-sending modes:
| Mode | When | Effect |
|---|---|---|
setFixedLengthStreamingMode(n) |
The size is known | Sends with Content-Length. The preferable one |
setChunkedStreamingMode(n) |
It is not known (generated content) | Sends with Transfer-Encoding: chunked |
| Neither | By default | Keeps the whole body in memory before sending. Bad for large files |
Escaping the JSON by hand is a stopgap. The example works because the values are simple, but a title with quotation marks or a backslash would break the JSON. Building and parsing JSON properly requires a library —Jackson— and that is 11-07. Here we do it by hand and are aware of the limitation.
A form instead of JSON
connection.setRequestProperty("Content-Type",
"application/x-www-form-urlencoded; charset=utf-8");
// Here it IS correct to use URLEncoder: this is exactly its format.
String body = new UrlQuery()
.with("isbn", "978-0000000001")
.with("employee", "Marta Ruiz")
.build();
// isbn=978-0000000001&employee=Marta+Ruiz
- Redirects
When the server replies 301, 302, 303, 307 or 308, the Location header states where to go.
HttpURLConnection follows redirects automatically by default, which is convenient and sometimes undesirable.
// Global, for the whole JVM. Avoid it: it affects code that is not yours.
HttpURLConnection.setFollowRedirects(false);
// Per instance. This is the one you should use.
connection.setInstanceFollowRedirects(false);Cases where it is worth disabling them:
- You want to know there was a redirect, for example to update a stored URL after a 301.
- You need to control the number of hops, so as not to fall into an infinite loop.
- You send credentials: by following automatically, the
Authorizationheader could be sent to a host other than the intended one. It is a real leakage risk.
And an important limitation: HttpURLConnection does not follow redirects between http and https. If you request http://example.com and it replies 301 towards https://example.com, the library does not follow the hop and returns the 301 to you with no useful body. It is an endless source of confusion, because curl and browsers do follow it.
/** Follows redirects by hand, with a limit and scheme control. */
static HttpURLConnection followRedirects(URL url, int maxHops)
throws IOException {
URL current = url;
for (int hop = 0; hop <= maxHops; hop++) {
HttpURLConnection connection = (HttpURLConnection) current.openConnection();
connection.setInstanceFollowRedirects(false); // we manage them ourselves
connection.setConnectTimeout(5_000);
connection.setReadTimeout(10_000);
int code = connection.getResponseCode();
if (code < 300 || code >= 400) {
return connection; // not a redirect: we have arrived
}
String target = connection.getHeaderField("Location");
connection.disconnect();
if (target == null) {
throw new IOException("Redirect " + code + " with no Location header");
}
// resolve() handles relative Location values ("/new/path"),
// which are perfectly legal and surprising if unexpected.
current = current.toURI().resolve(target).toURL();
LOG.info("Redirect " + code + " -> " + current);
}
throw new IOException("Too many redirects (more than " + maxHops + ")");
}
- Compression and HTTPS
Compression with gzip
A compressed JSON body can take up a fifth of the space. HttpURLConnection advertises gzip by default and decompresses it by itself... but only if you do not touch the Accept-Encoding header. If you set it by hand, automatic decompression is disabled and you receive compressed bytes.
// If you set this BY HAND, YOU have to decompress.
connection.setRequestProperty("Accept-Encoding", "gzip");
int code = connection.getResponseCode();
InputStream input = connection.getInputStream();
if ("gzip".equalsIgnoreCase(connection.getContentEncoding())) {
input = new GZIPInputStream(input); // a module 7 decorator
}GZIPInputStream is just another decorator, like BufferedInputStream. Module 7 keeps paying off.
Recommendation: do not touch Accept-Encoding and let the library handle it. Only do so if you need explicit control.
HTTPS
If the URL begins with https, openConnection() returns an HttpsURLConnection, a subclass of HttpURLConnection. Nothing else has to be done: the TLS encryption, the certificate validation and the host name check happen transparently.
URL url = URI.create("https://api.nexussoftware.com/v1/books").toURL();
HttpsURLConnection connection = (HttpsURLConnection) url.openConnection();
// ... the same as always ...
// Additional methods, if you need to inspect the certificate:
System.out.println("Cipher : " + connection.getCipherSuite());
System.out.println("Certificate: " + connection.getServerCertificates()[0]);Errors you will see and what they mean:
| Exception | Cause | Correct solution |
|---|---|---|
SSLHandshakeException: PKIX path building failed |
The certificate is not signed by an authority Java recognises (self-signed, or an internal CA) | Import the certificate into the trust store, do not disable validation |
SSLHandshakeException: No name matching X found |
The certificate is for another host name | Use the correct name |
SSLException: Received fatal alert: protocol_version |
Incompatible TLS versions | Update the JDK or the server |
Never disable certificate validation. You will find snippets on the internet with a
TrustManagerthat accepts everything and aHostnameVerifierthat always returnstrue. That removes all of TLS's security: it turns HTTPS into HTTP with extra steps, and leaves the connection open to a man-in-the-middle attack. If you have an internal certificate, the solution is to import it into the trust store (keytool -importcert) or to use anSSLContextwith a store of your own. Network security is covered thoroughly in 12-07.
- BiblioTech: metadata client and cover downloads
Now everything together. Nexus Software has an internal bibliographic metadata service, and BiblioTech is going to query it.
The service contract
GET /v1/books/{isbn}
200 -> {"isbn":"...","title":"...","author":"...","pages":416,
"cover":"https://metadata.nexussoftware.local/covers/978-0000000001.jpg"}
404 -> {"error":"not_found","message":"Unknown ISBN"}
429 -> Retry-After header with the seconds to wait
GET /covers/{isbn}.jpg
200 -> binary JPEG imageThe client
package com.nexussoftware.bibliotech.network;
import com.nexussoftware.bibliotech.exception.BiblioTechException;
import java.io.IOException;
import java.io.InputStream;
import java.io.OutputStream;
import java.net.HttpURLConnection;
import java.net.SocketTimeoutException;
import java.net.URI;
import java.net.URL;
import java.net.UnknownHostException;
import java.nio.charset.Charset;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
import java.util.logging.Level;
import java.util.logging.Logger;
/**
* Client of Nexus Software's external metadata service,
* with HttpURLConnection.
*
* It shows the correct use of the old API: time limits, explicit checking
* of the status code, getErrorStream for errors, bounded reading
* and translation into domain exceptions.
*/
public class MetadataClient {
private static final Logger LOG = Logger.getLogger(MetadataClient.class.getName());
private static final int CONNECT_TIMEOUT_MS = 5_000;
private static final int READ_TIMEOUT_MS = 10_000;
/** No legitimate metadata card exceeds this. A memory defence. */
private static final int MAX_BODY = 256 * 1024;
/** No legitimate cover exceeds this. */
private static final long MAX_COVER = 5L * 1024 * 1024;
private static final String USER_AGENT = "BiblioTech/1.0 (+https://nexussoftware.com)";
private final String base;
public MetadataClient(String base) {
// It is normalised so paths can be concatenated without duplicating slashes.
this.base = base.endsWith("/") ? base.substring(0, base.length() - 1) : base;
}
/** The metadata of a book exactly as the service returns it. */
public record Metadata(String isbn, String title, String author,
int pages, String coverUrl) {
}
// =================================================================
// Metadata query
// =================================================================
/** Queries an ISBN. Returns null if the service does not know it (404). */
public Metadata query(String isbn) throws BiblioTechException {
validateIsbn(isbn);
HttpURLConnection connection = null;
try {
// The ISBN goes in the PATH, so it is encoded with URI's
// multi-argument constructor, not with URLEncoder (which would use + for space).
URI uri = new URI("http", null, hostOf(base), portOf(base),
pathOf(base) + "/v1/books/" + isbn, null, null);
URL url = uri.toURL();
connection = (HttpURLConnection) url.openConnection();
connection.setRequestMethod("GET");
connection.setRequestProperty("Accept", "application/json");
connection.setRequestProperty("User-Agent", USER_AGENT);
connection.setConnectTimeout(CONNECT_TIMEOUT_MS);
connection.setReadTimeout(READ_TIMEOUT_MS);
connection.setInstanceFollowRedirects(true);
// HERE it really connects.
int code = connection.getResponseCode();
LOG.fine(() -> "GET " + url + " -> " + code);
// A 404 is not a program error: it is "I do not have it".
if (code == HttpURLConnection.HTTP_NOT_FOUND) {
consumeAndClose(connection, code);
return null;
}
// 429: the service is asking us to slow down.
if (code == 429) {
String wait = connection.getHeaderField("Retry-After");
consumeAndClose(connection, code);
throw new BiblioTechException(
"The metadata service has rate-limited the requests."
+ (wait == null ? "" : " Retry in " + wait + " s."));
}
if (code != HttpURLConnection.HTTP_OK) {
// The error body almost always explains what happened:
// reading it is the difference between debugging in two minutes
// or in two hours.
String detail = readBody(connection, code);
LOG.warning("Response " + code + " from the service: " + trim(detail));
throw new BiblioTechException(
"The metadata service replied " + code);
}
String body = readBody(connection, code);
return parse(body, isbn);
} catch (SocketTimeoutException e) {
// TRANSIENT: worth a retry with increasing backoff.
LOG.warning("Timed out querying the metadata of " + isbn);
throw new BiblioTechException(
"The metadata service is not replying in time.", e);
} catch (UnknownHostException e) {
// PERMANENT: misconfiguration.
LOG.severe("Unresolvable metadata host: " + base);
throw new BiblioTechException(
"The metadata service cannot be found.", e);
} catch (IOException | java.net.URISyntaxException e) {
LOG.log(Level.WARNING, "Failure querying the metadata of " + isbn, e);
throw new BiblioTechException(
"Error querying the metadata service.", e);
} finally {
if (connection != null) {
connection.disconnect();
}
}
}
// =================================================================
// Cover download (binary to disk)
// =================================================================
/**
* Downloads the cover to a file. It combines HTTP with module 7's NIO.2.
* ATOMIC write: first to a temporary file, then moved. That way an
* interrupted download never leaves a half-written JPEG in the catalogue.
*/
public Path downloadCover(String coverUrl, Path target)
throws BiblioTechException {
HttpURLConnection connection = null;
Path temporary = null;
try {
URL url = URI.create(coverUrl).toURL();
// http and https only: without this, a "file:///etc/passwd" URL
// received from the service would make us read local files.
String scheme = url.getProtocol();
if (!scheme.equals("http") && !scheme.equals("https")) {
throw new BiblioTechException("Scheme not allowed: " + scheme);
}
connection = (HttpURLConnection) url.openConnection();
connection.setRequestMethod("GET");
connection.setRequestProperty("Accept", "image/jpeg, image/png, image/*");
connection.setRequestProperty("User-Agent", USER_AGENT);
connection.setConnectTimeout(CONNECT_TIMEOUT_MS);
connection.setReadTimeout(30_000); // an image takes longer
int code = connection.getResponseCode();
if (code != HttpURLConnection.HTTP_OK) {
consumeAndClose(connection, code);
throw new BiblioTechException(
"The cover could not be downloaded: HTTP " + code);
}
// Content-Length is INDICATIVE: it may be missing (-1) or lie.
// It is checked before AND during the download.
long announced = connection.getContentLengthLong();
if (announced > MAX_COVER) {
consumeAndClose(connection, code);
throw new BiblioTechException(
"Cover too large: " + announced + " bytes");
}
String type = connection.getContentType();
if (type != null && !type.startsWith("image/")) {
consumeAndClose(connection, code);
throw new BiblioTechException("The resource is not an image: " + type);
}
Files.createDirectories(target.getParent());
temporary = Files.createTempFile(target.getParent(), "cover-", ".tmp");
long downloaded = 0;
try (InputStream input = connection.getInputStream();
OutputStream output = Files.newOutputStream(temporary)) {
byte[] buffer = new byte[8192];
int read;
while ((read = input.read(buffer)) != -1) {
downloaded += read;
// The limit is checked DURING the download as well:
// the Content-Length may lie or be absent.
if (downloaded > MAX_COVER) {
throw new BiblioTechException(
"The cover exceeds the limit during the download");
}
// write(buffer, 0, read): never buffer.length (module 7).
output.write(buffer, 0, read);
}
}
// Atomic move: the final file appears complete or it does not appear.
Files.move(temporary, target,
StandardCopyOption.REPLACE_EXISTING,
StandardCopyOption.ATOMIC_MOVE);
temporary = null; // no longer needs cleaning up
long total = downloaded;
LOG.info(() -> "Cover downloaded: " + target + " (" + total + " bytes)");
return target;
} catch (SocketTimeoutException e) {
throw new BiblioTechException("Timed out downloading the cover.", e);
} catch (IOException e) {
LOG.log(Level.WARNING, "Failure downloading " + coverUrl, e);
throw new BiblioTechException("The cover could not be downloaded.", e);
} finally {
if (connection != null) {
connection.disconnect();
}
// Cleaning up the temporary file if something failed halfway.
if (temporary != null) {
try {
Files.deleteIfExists(temporary);
} catch (IOException e) {
LOG.fine("Could not delete the temporary file: " + temporary);
}
}
}
}
// =================================================================
// Utilities
// =================================================================
private String readBody(HttpURLConnection connection, int code) throws IOException {
// The distinction everybody forgets: with 4xx/5xx you have to
// use getErrorStream, because getInputStream THROWS.
InputStream input = (code >= 200 && code < 400)
? connection.getInputStream()
: connection.getErrorStream();
if (input == null) {
return "";
}
Charset charset = charsetOf(connection.getContentType());
try (InputStream stream = input) {
// Bounded readNBytes, not readAllBytes: a huge response
// must not exhaust our memory.
byte[] bytes = stream.readNBytes(MAX_BODY);
if (bytes.length == MAX_BODY) {
LOG.warning("Response truncated to " + MAX_BODY + " bytes");
}
return new String(bytes, charset);
}
}
/**
* Consumes and closes the body even when we do not care about it.
* Without this, the connection does not return to the internal pool and
* the reuse is lost, which in HTTP is worth a full network round trip.
*/
private void consumeAndClose(HttpURLConnection connection, int code) {
try {
InputStream input = (code >= 200 && code < 400)
? connection.getInputStream()
: connection.getErrorStream();
if (input != null) {
try (InputStream stream = input) {
stream.readNBytes(MAX_BODY);
}
}
} catch (IOException e) {
LOG.fine("Failure consuming the body: " + e.getMessage());
}
}
static Charset charsetOf(String contentType) {
if (contentType != null) {
for (String part : contentType.split(";")) {
String p = part.strip();
if (p.toLowerCase().startsWith("charset=")) {
String name = p.substring(8).replace("\"", "").strip();
try {
return Charset.forName(name);
} catch (Exception e) {
LOG.warning("Unknown charset: " + name);
}
}
}
}
return StandardCharsets.UTF_8;
}
/**
* Extraction of JSON fields BY SUBSTRING SEARCH.
*
* THIS IS A TEACHING STOPGAP, and it has to be said plainly.
* It works with this service's particular, simple response, and it
* breaks with: values containing the searched substring, escapes (\"),
* nesting, arrays, different spacing, fields in another order or
* null values. Parsing JSON properly requires a library, and
* that is done with JACKSON IN 11-07. Do not take this to production.
*/
private Metadata parse(String json, String requestedIsbn)
throws BiblioTechException {
try {
String title = textField(json, "title");
String author = textField(json, "author");
String cover = textField(json, "cover");
int pages = intField(json, "pages");
if (title == null) {
throw new BiblioTechException(
"Service response without the 'title' field");
}
return new Metadata(requestedIsbn, title,
author == null ? "(unknown)" : author,
pages, cover);
} catch (RuntimeException e) {
throw new BiblioTechException(
"The service response could not be interpreted", e);
}
}
/** Looks for "field":"value" and returns the value. A stopgap, see the comment. */
private String textField(String json, String field) {
String mark = "\"" + field + "\"";
int i = json.indexOf(mark);
if (i < 0) {
return null;
}
int colon = json.indexOf(':', i + mark.length());
if (colon < 0) {
return null;
}
int open = json.indexOf('"', colon);
if (open < 0) {
return null;
}
int close = json.indexOf('"', open + 1);
if (close < 0) {
return null;
}
return json.substring(open + 1, close);
}
/** Looks for "field":123 and returns the number, or 0 if it is not found. */
private int intField(String json, String field) {
String mark = "\"" + field + "\"";
int i = json.indexOf(mark);
if (i < 0) {
return 0;
}
int colon = json.indexOf(':', i + mark.length());
if (colon < 0) {
return 0;
}
int j = colon + 1;
while (j < json.length() && !Character.isDigit(json.charAt(j))) {
if (json.charAt(j) == ',' || json.charAt(j) == '}') {
return 0;
}
j++;
}
int start = j;
while (j < json.length() && Character.isDigit(json.charAt(j))) {
j++;
}
return start == j ? 0 : Integer.parseInt(json.substring(start, j));
}
private void validateIsbn(String isbn) throws BiblioTechException {
if (isbn == null || isbn.isBlank() || isbn.length() > 20) {
throw new BiblioTechException("Invalid ISBN");
}
// Allow list: digits and hyphens only. It prevents injecting paths
// ("../admin") or parameters ("?x=1") into the URL.
for (int i = 0; i < isbn.length(); i++) {
char c = isbn.charAt(i);
if (!Character.isDigit(c) && c != '-') {
throw new BiblioTechException("ISBN with characters that are not allowed");
}
}
}
private String trim(String text) {
return text.length() > 200 ? text.substring(0, 200) + "..." : text;
}
// Simple decomposition of the base URL for the URI constructor.
private String hostOf(String base) throws java.net.URISyntaxException {
return new URI(base).getHost();
}
private int portOf(String base) throws java.net.URISyntaxException {
return new URI(base).getPort();
}
private String pathOf(String base) throws java.net.URISyntaxException {
String p = new URI(base).getPath();
return p == null ? "" : p;
}
}Testing it with no external service
Nexus Software does not exist, so let us play the service with nc, as in 09-02:
Terminal 1:
printf 'HTTP/1.1 200 OK\r\nContent-Type: application/json; charset=utf-8\r\nContent-Length: 144\r\nConnection: close\r\n\r\n{"isbn":"978-0000000001","title":"Effective Java","author":"Joshua Bloch","pages":416,"cover":"http://localhost:8080/covers/978-0000000001.jpg"}' | nc -l 8080Terminal 2:
package com.nexussoftware.bibliotech.presentation;
import com.nexussoftware.bibliotech.exception.BiblioTechException;
import com.nexussoftware.bibliotech.network.MetadataClient;
import com.nexussoftware.bibliotech.network.MetadataClient.Metadata;
import java.nio.file.Path;
public class MetadataTest {
public static void main(String[] args) {
MetadataClient client = new MetadataClient("http://localhost:8080");
try {
Metadata m = client.query("978-0000000001");
if (m == null) {
System.out.println("The service does not know that ISBN");
return;
}
System.out.println("Title : " + m.title());
System.out.println("Author : " + m.author());
System.out.println("Pages : " + m.pages());
System.out.println("Cover : " + m.coverUrl());
if (m.coverUrl() != null) {
Path target = Path.of("covers", m.isbn() + ".jpg");
client.downloadCover(m.coverUrl(), target);
System.out.println("Saved in " + target.toAbsolutePath());
}
} catch (BiblioTechException e) {
System.err.println("ERROR: " + e.getMessage());
if (e.getCause() != null) {
System.err.println("Cause: " + e.getCause());
}
}
}
}Title : Effective Java
Author : Joshua Bloch
Pages : 416
Cover : http://localhost:8080/covers/978-0000000001.jpgTests worth running
- Reply with a 404 from
ncand check thatqueryreturnsnullwith no exception. - Reply with a 500 and a JSON error body. You will see the error body in the log thanks to
getErrorStream(); withgetInputStream()you would have had only anIOExceptionwith no information. - Reply with nothing and wait. After ten seconds the
setReadTimeoutfires. Remove it and check that it waits indefinitely. - Compare the same request with
curl -v. It is the way to know whether the problem is yours or the server's. - Use an ISBN with
../:client.query("../admin")is rejected in the validation, before touching the network.
- An honest assessment of this API
You have learned to use it properly. Now the honest assessment of why you would not use it for new code.
| Problem | Detail |
|---|---|
| Verbosity | A simple request is 30 lines with its resource handling |
| Configuration by side effect | setDoOutput(true) changes the method to POST without saying so |
getInputStream versus getErrorStream |
A distinction that should not exist and that everybody forgets |
| No total time limit | Per operation only: a slow server can hold you indefinitely |
| It does not follow redirects between schemes | http→https fails, unlike curl and browsers |
| A mutable object with states | Configuring after connecting fails silently or throws IllegalStateException |
| No asynchrony | Every request blocks the thread |
| HTTP/1.1 only | No HTTP/2, no multiplexing |
| No WebSocket | Out of its scope |
URL.equals() does DNS |
A consequence of the 1996 design |
| Hard to test | There is no interface to substitute; you have to intercept with URLStreamHandler |
All of this is solved by Java 11's HttpClient, which you will see in 09-06.
So why learn it? For three solid reasons:
- It is everywhere. Millions of lines of Java code in production use it. You will read it and maintain it.
- It teaches HTTP hands-on. Being verbose, it forces you to know the methods, the codes, the headers and the streams. With
HttpClienteverything works so well that it can be used without understanding what happens underneath. - The concepts transfer intact. Methods, status codes, headers, time limits, redirects, charsets: all of that reappears in the modern API with better packaging. The hard part of HTTP is not the API, it is HTTP.
Common Mistakes and Tips
Assuming that a 404 or a 500 throws an exception. Mistake number one. An error response is a valid, correctly delivered response. Check getResponseCode() always.
Using getInputStream() with an error code. It throws IOException and takes away the error body, which is exactly what explains the problem. With 4xx and 5xx, getErrorStream().
Not setting time limits. By default they are infinite. An external service that accepts and does not reply blocks your thread forever, and with a bounded pool, exhausts the pool without a single error in the log.
Configuring after connecting. openConnection() does not connect, but getResponseCode() does. All configuration goes before.
Forgetting that setDoOutput(true) changes the method to POST. A classic, silent trap.
Computing Content-Length with String.length(). Those are characters, not bytes. With accents they differ and the request goes out corrupt. Use getBytes(UTF_8).length.
Using URLEncoder for path segments. It encodes the space as +, which in a path is a literal +. For paths, URI's multi-argument constructor.
Using URLEncoder.encode(String) with no charset. It is deprecated and uses the platform encoding: the same URL comes out differently on each system.
Comparing URL objects or putting them in a HashSet. equals() and hashCode() perform DNS resolution: they block and depend on the network. Use URI.
Assuming UTF-8 without reading the Content-Type. It works almost always and gives you corrupt text exactly when you least expect it.
Using readAllBytes() on a remote response. With no size limit, a huge —or malicious— response exhausts your memory. readNBytes with a cap.
Not consuming the body when you do not care about it. The connection does not return to the internal pool and you lose the reuse, which in HTTP costs a full network round trip.
Setting Accept-Encoding: gzip by hand and not decompressing. Touching that header disables automatic decompression and you receive compressed bytes. Either do not touch it, or decompress yourself.
Disabling TLS certificate validation. It removes all of HTTPS's security and opens the door to a man-in-the-middle attack. If the certificate is internal, import it (12-07).
Downloading straight to the final file. An interrupted download leaves a corrupt file that looks valid. Download to a temporary file and move it at the end.
Trusting the Content-Length. It may be missing (-1) or lie. Check the limit during the download too.
Golden tip for debugging. When something does not work, make the same request with curl -v and compare. If curl works and your code does not, the difference is in your headers or in the method. And if you need to see what your Java code sends, nc -l 8080 shows it to you byte by byte — the same trick as in 09-02, still the most effective tool there is.
Exercises
Exercise 1: HTTP inspector
Write a class HttpInspector with a method inspect(String url) that shows a complete report of a URL, in the style of curl -v but in Java.
Requirements:
- Break down and show every part of the URL (scheme, host, effective port, path, query, fragment).
- First make a
HEADrequest —which does not download the body— and show the status code, the message and all the headers in order. - If the
HEADreturns 405 (method not allowed, which does happen), retry withGET. - Disable automatic redirect following and show the complete chain of hops with their codes and their
Locationvalues, with a limit of 5. - Show the size of the body, the content type, the detected charset and whether it is compressed.
- Measure and show the connection time and the total time.
- Mandatory time limits and differentiated handling of the exceptions.
Test it against nc -l 8080 with replies you make up, including a chain of two redirects.
Exercise 2: Client with retries that respects Retry-After
Write ResilientHttpClient, a wrapper over HttpURLConnection that applies a correct retry policy.
Requirements:
- A method
get(String url)returning arecord Response(int code, String body, Map<String,List<String>> headers, int attempts). - Retry only on:
SocketTimeoutException,ConnectException, and codes 429, 502, 503, 504. - Never retry on 4xx (except 429) or on
UnknownHostException. - Increasing backoff: 200 ms, 400, 800, 1600, with a maximum of 4 attempts.
- Add to the wait a random component of up to 20 % (jitter) and explain in a comment what problem it avoids.
- If the response carries
Retry-After, respect it instead of the computed wait, with a cap of 30 s. Accept the format in seconds (the date format may be ignored, stating so). - A method
post(String url, String body, String contentType)that does not retry by default, with an explicit parameter to allow it, and a comment explaining whyPOSTis different. - Log every attempt with the logger.
Exercise 3: BiblioTech cover synchroniser
Write CoverSynchroniser, which walks BiblioTech's catalogue and downloads the covers that are missing.
Requirements:
- For each catalogue material with no local cover, query
MetadataClientand download the image. - Before downloading, do a
HEADto check the type and the size, and skip those that are not images or exceed 5 MB. - Conditional download: if the local file already exists, send the
If-Modified-Sinceheader with its modification date in HTTP format and skip the download if the server replies 304 Not Modified. Look up the HTTP date format and generate it without usingjava.time(that is 10-05): withSimpleDateFormatin the GMT zone andLocale.US, noting in a comment that in 10-05 it is done better. - Download to a temporary file and move atomically.
- At most 2 requests per second to the service, so as not to swamp it. Implement the rate limiter yourself.
- Final report: downloaded, already up to date (304), skipped, failed, total bytes and time.
- All sequential: doing it in parallel is 09-06, and you will mention it in a comment.
Solutions
Solution 1
package com.nexussoftware.bibliotech.network;
import java.io.IOException;
import java.io.InputStream;
import java.net.ConnectException;
import java.net.HttpURLConnection;
import java.net.SocketTimeoutException;
import java.net.URI;
import java.net.URL;
import java.net.UnknownHostException;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
import java.util.TreeMap;
/**
* HTTP inspector: a complete report of a URL, in the style of curl -v.
* A diagnostic tool for Nexus Software's systems team.
*/
public class HttpInspector {
private static final int MAX_HOPS = 5;
private static final int CONNECT_TIMEOUT_MS = 5_000;
private static final int READ_TIMEOUT_MS = 10_000;
private static final int MAX_BODY = 1024 * 1024;
/** One hop of the redirect chain. */
private record Hop(String url, int code, String message, String target) {
}
public void inspect(String urlText) {
System.out.println("=".repeat(78));
System.out.println("INSPECTION OF: " + urlText);
System.out.println("=".repeat(78));
URL url;
try {
// URI validates the syntax strictly; URL does not.
url = URI.create(urlText).toURL();
} catch (Exception e) {
System.out.println("Invalid URL: " + e.getMessage());
return;
}
showParts(url);
List<Hop> chain = new ArrayList<>();
URL current = url;
HttpURLConnection connection = null;
try {
// --- Follow the redirect chain by hand ---
for (int hop = 0; hop <= MAX_HOPS; hop++) {
connection = open(current, "HEAD");
long t0 = System.nanoTime();
int code = connection.getResponseCode();
long msConnect = (System.nanoTime() - t0) / 1_000_000;
// Many servers do not accept HEAD and return 405.
// We retry with GET, which is always supported.
if (code == HttpURLConnection.HTTP_BAD_METHOD) {
System.out.println("\n(HEAD returned 405; retrying with GET)");
connection.disconnect();
connection = open(current, "GET");
t0 = System.nanoTime();
code = connection.getResponseCode();
msConnect = (System.nanoTime() - t0) / 1_000_000;
}
String target = connection.getHeaderField("Location");
chain.add(new Hop(current.toString(), code,
connection.getResponseMessage(), target));
if (code < 300 || code >= 400 || target == null) {
showResponse(connection, code, msConnect);
break;
}
// Location may be relative: resolve() handles it.
current = current.toURI().resolve(target).toURL();
connection.disconnect();
connection = null;
}
showChain(chain);
} catch (UnknownHostException e) {
System.out.println("\nHost DOES NOT RESOLVE: " + e.getMessage());
} catch (ConnectException e) {
System.out.println("\nCONNECTION REFUSED: there is no server on that port");
} catch (SocketTimeoutException e) {
System.out.println("\nTIMED OUT: the server is not replying");
} catch (Exception e) {
System.out.println("\nFAILURE: " + e);
} finally {
if (connection != null) {
connection.disconnect();
}
}
}
private HttpURLConnection open(URL url, String method) throws IOException {
HttpURLConnection c = (HttpURLConnection) url.openConnection();
c.setRequestMethod(method);
c.setRequestProperty("User-Agent", "BiblioTech-Inspector/1.0");
c.setRequestProperty("Accept", "*/*");
c.setConnectTimeout(CONNECT_TIMEOUT_MS);
c.setReadTimeout(READ_TIMEOUT_MS);
// We manage them ourselves, so we can show them.
c.setInstanceFollowRedirects(false);
return c;
}
private void showParts(URL url) {
// getPort() returns -1 if it is not explicit: it has to be combined
// with getDefaultPort() to know the effective port.
int port = url.getPort() != -1 ? url.getPort() : url.getDefaultPort();
System.out.println("\n--- PARTS OF THE URL ---");
System.out.printf(" %-16s %s%n", "Scheme", url.getProtocol());
System.out.printf(" %-16s %s%n", "Host", url.getHost());
System.out.printf(" %-16s %d%s%n", "Port", port,
url.getPort() == -1 ? " (default for the scheme)" : " (explicit)");
System.out.printf(" %-16s %s%n", "Path",
url.getPath().isEmpty() ? "/" : url.getPath());
System.out.printf(" %-16s %s%n", "Query",
url.getQuery() == null ? "(none)" : url.getQuery());
System.out.printf(" %-16s %s%n", "Fragment",
url.getRef() == null ? "(none)"
: url.getRef() + " <- NOT sent to the server");
}
private void showResponse(HttpURLConnection c, int code, long ms)
throws IOException {
System.out.println("\n--- RESPONSE ---");
System.out.printf(" %-16s %d %s (%s)%n", "Status", code,
c.getResponseMessage(), family(code));
System.out.printf(" %-16s %d ms%n", "Time", ms);
System.out.println("\n--- HEADERS ---");
// TreeMap for alphabetical order. The null key carries the status line.
Map<String, List<String>> headers = new TreeMap<>((a, b) -> {
if (a == null) return -1;
if (b == null) return 1;
return a.compareToIgnoreCase(b);
});
headers.putAll(c.getHeaderFields());
for (Map.Entry<String, List<String>> e : headers.entrySet()) {
String name = e.getKey() == null ? "(status line)" : e.getKey();
for (String value : e.getValue()) {
System.out.printf(" %-24s %s%n", name + ":", value);
}
}
System.out.println("\n--- CONTENT ---");
System.out.printf(" %-16s %s%n", "Type",
c.getContentType() == null ? "(not stated)" : c.getContentType());
System.out.printf(" %-16s %s%n", "Charset",
MetadataClient.charsetOf(c.getContentType()));
long length = c.getContentLengthLong();
System.out.printf(" %-16s %s%n", "Length",
length < 0 ? "(not stated: chunked or close)" : length + " bytes");
System.out.printf(" %-16s %s%n", "Compression",
c.getContentEncoding() == null ? "(none)" : c.getContentEncoding());
// With HEAD there is no body, but with the fallback GET there is.
if ("GET".equals(c.getRequestMethod())) {
InputStream input = (code >= 200 && code < 400)
? c.getInputStream() : c.getErrorStream();
if (input != null) {
try (InputStream stream = input) {
byte[] bytes = stream.readNBytes(MAX_BODY);
System.out.printf(" %-16s %d bytes read%n",
"Actual body", bytes.length);
String text = new String(bytes,
MetadataClient.charsetOf(c.getContentType()));
System.out.println("\n--- FIRST LINES OF THE BODY ---");
String[] lines = text.split("\n", 6);
for (int i = 0; i < Math.min(5, lines.length); i++) {
// We truncate: never dump network data without a limit.
String l = lines[i];
System.out.println(" " + (l.length() > 100
? l.substring(0, 100) + "..." : l));
}
}
}
}
}
private void showChain(List<Hop> chain) {
if (chain.size() <= 1) {
return;
}
System.out.println("\n--- REDIRECT CHAIN (" + chain.size() + ") ---");
for (int i = 0; i < chain.size(); i++) {
Hop h = chain.get(i);
System.out.printf(" %d. %d %s%n %s%n", i + 1, h.code(),
h.message(), h.url());
if (h.target() != null) {
System.out.println(" -> Location: " + h.target());
}
}
}
private String family(int code) {
return switch (code / 100) {
case 1 -> "informational";
case 2 -> "SUCCESS";
case 3 -> "redirect";
case 4 -> "CLIENT ERROR";
case 5 -> "SERVER ERROR";
default -> "unknown";
};
}
public static void main(String[] args) {
HttpInspector inspector = new HttpInspector();
inspector.inspect(args.length > 0 ? args[0]
: "http://localhost:8080/v1/books/978-0000000001?format=json#summary");
}
}Output against an nc that replies with a redirect and then a 200:
==============================================================================
INSPECTION OF: http://localhost:8080/v1/books/978-0000000001?format=json#summary
==============================================================================
--- PARTS OF THE URL ---
Scheme http
Host localhost
Port 8080 (explicit)
Path /v1/books/978-0000000001
Query format=json
Fragment summary <- NOT sent to the server
--- RESPONSE ---
Status 200 OK (SUCCESS)
Time 4 ms
--- HEADERS ---
(status line): HTTP/1.1 200 OK
Content-Length: 50
Content-Type: application/json; charset=utf-8
--- CONTENT ---
Type application/json; charset=utf-8
Charset UTF-8
Length 50 bytes
Compression (none)
--- REDIRECT CHAIN (2) ---
1. 302 Found
http://localhost:8080/v1/books/978-0000000001?format=json
-> Location: /v2/books/978-0000000001
2. 200 OK
http://localhost:8080/v2/books/978-0000000001Comments. Four details this exercise teaches. The fragment appears in no sent header: it is shown in the parts of the URL and then disappears, which is exactly the correct behaviour. getPort() returns -1 when the port is not explicit and has to be combined with getDefaultPort() — trying to connect to port -1 is a common mistake. The Location may be relative (/v2/books/...), which is perfectly legal according to the standard and breaks any code that treats it as an absolute URL; URI.resolve() handles it. And the 405 with HEAD happens in practice more than you would think: quite a few servers implement only GET, and a diagnostic tool has to allow for it.
Solution 2
package com.nexussoftware.bibliotech.network;
import java.io.IOException;
import java.io.InputStream;
import java.io.OutputStream;
import java.net.ConnectException;
import java.net.HttpURLConnection;
import java.net.SocketTimeoutException;
import java.net.URI;
import java.net.URL;
import java.net.UnknownHostException;
import java.nio.charset.StandardCharsets;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ThreadLocalRandom;
import java.util.logging.Logger;
/**
* HTTP client with a correct retry policy.
*
* The rule that governs everything: retry what is TRANSIENT
* (timeouts, 429, infrastructure 5xx) and never what is
* PERMANENT (4xx, non-existent host).
*/
public class ResilientHttpClient {
private static final Logger LOG =
Logger.getLogger(ResilientHttpClient.class.getName());
private static final int MAX_ATTEMPTS = 4;
private static final int INITIAL_WAIT_MS = 200;
private static final int MAX_RETRY_AFTER_S = 30;
private static final int MAX_BODY = 1024 * 1024;
private static final int CONNECT_TIMEOUT_MS = 5_000;
private static final int READ_TIMEOUT_MS = 10_000;
public record Response(int code, String body,
Map<String, List<String>> headers, int attempts) {
public boolean success() {
return code >= 200 && code < 300;
}
}
// =================================================================
// GET: idempotent, retried
// =================================================================
public Response get(String url) throws IOException {
return execute(url, "GET", null, null, true);
}
// =================================================================
// POST: NOT idempotent, NOT retried by default
// =================================================================
/**
* POST with no retries.
*
* WHY IT IS DIFFERENT: if a POST times out, WE DO NOT KNOW
* whether the server processed it. The request may have arrived, been
* executed and only the reply lost. Retrying would duplicate the operation:
* two loans, two charges, two orders.
*
* The professional solution is the IDEMPOTENCY KEY: the client generates
* a unique identifier per operation, sends it in a header
* (Idempotency-Key) and the server rejects the second request with the
* same key. With that, retrying IS safe. Without it, it is not.
*/
public Response post(String url, String body, String contentType)
throws IOException {
return execute(url, "POST", body, contentType, false);
}
/** POST with retries: only if the service guarantees idempotence. */
public Response postIdempotent(String url, String body, String contentType)
throws IOException {
return execute(url, "POST", body, contentType, true);
}
// =================================================================
// Core
// =================================================================
private Response execute(String urlText, String method, String body,
String contentType, boolean retry)
throws IOException {
URL url = URI.create(urlText).toURL();
int wait = INITIAL_WAIT_MS;
IOException lastFailure = null;
int maximum = retry ? MAX_ATTEMPTS : 1;
for (int attempt = 1; attempt <= maximum; attempt++) {
HttpURLConnection connection = null;
try {
connection = (HttpURLConnection) url.openConnection();
connection.setRequestMethod(method);
connection.setRequestProperty("User-Agent", "BiblioTech/1.0");
connection.setRequestProperty("Accept", "application/json, */*");
connection.setConnectTimeout(CONNECT_TIMEOUT_MS);
connection.setReadTimeout(READ_TIMEOUT_MS);
if (body != null) {
connection.setDoOutput(true);
connection.setRequestProperty("Content-Type",
contentType == null ? "application/json; charset=utf-8"
: contentType);
byte[] bytes = body.getBytes(StandardCharsets.UTF_8);
// Length in BYTES, not in characters.
connection.setFixedLengthStreamingMode(bytes.length);
try (OutputStream output = connection.getOutputStream()) {
output.write(bytes);
output.flush();
}
}
int code = connection.getResponseCode();
Map<String, List<String>> headers = connection.getHeaderFields();
final int n = attempt;
LOG.fine(() -> method + " " + url + " -> " + code
+ " (attempt " + n + ")");
// A transient code: worth a retry if there is room left.
if (isTransient(code) && attempt < maximum) {
long waitMs = waitAfterCode(connection, code, wait);
consume(connection, code);
LOG.warning(method + " " + url + " -> " + code
+ "; retry " + (attempt + 1) + " in " + waitMs + " ms");
sleep(waitMs);
wait *= 2;
continue;
}
String text = readBody(connection, code);
return new Response(code, text, headers, attempt);
} catch (UnknownHostException e) {
// PERMANENT: retrying will not make the name exist.
LOG.severe("Unresolvable host: " + url.getHost());
throw e;
} catch (SocketTimeoutException | ConnectException e) {
// TRANSIENT: they may resolve themselves.
lastFailure = e;
if (attempt >= maximum) {
break;
}
long waitMs = withJitter(wait);
LOG.warning(e.getClass().getSimpleName() + " on " + url
+ "; retry " + (attempt + 1) + " in " + waitMs + " ms");
sleep(waitMs);
wait *= 2;
} finally {
if (connection != null) {
connection.disconnect();
}
}
}
throw lastFailure != null ? lastFailure
: new IOException("No response after " + maximum + " attempts");
}
// =================================================================
// Retry policy
// =================================================================
private boolean isTransient(int code) {
// 429: we are being asked to slow down. 502/503/504: infrastructure failures.
// The 500 is NOT included: it is usually a deterministic failure that will repeat.
// Nor are the 4xx: the request is wrong and retrying it gives the same result.
return code == 429 || code == 502 || code == 503 || code == 504;
}
/** If the server says how long to wait, we do as it says. */
private long waitAfterCode(HttpURLConnection connection, int code, int computed) {
String retryAfter = connection.getHeaderField("Retry-After");
if (retryAfter != null) {
try {
// Retry-After accepts two formats: seconds, or an HTTP date.
// Here we only handle the seconds one; the date one requires
// date parsing, and that is done properly in 10-05.
long seconds = Long.parseLong(retryAfter.strip());
long capped = Math.min(seconds, MAX_RETRY_AFTER_S);
LOG.info("The server asks us to wait " + seconds
+ " s (we apply " + capped + " s)");
return capped * 1000;
} catch (NumberFormatException e) {
LOG.fine("Retry-After in date format; ignored: " + retryAfter);
}
}
return withJitter(computed);
}
/**
* Adds a random component of up to 20 %.
*
* WHAT PROBLEM IT AVOIDS: the "thundering herd". If a hundred clients fail at
* the same time because the server went down, and they all retry exactly at
* 200 ms, the hundred requests arrive together again and take it down once
* more, in a cycle that repeats indefinitely. Spreading the
* retries out breaks the synchronisation and distributes the load over time.
*/
private long withJitter(long base) {
long variation = (long) (base * 0.2);
return base + ThreadLocalRandom.current().nextLong(-variation, variation + 1);
}
private void sleep(long ms) throws IOException {
try {
Thread.sleep(Math.max(0, ms));
} catch (InterruptedException e) {
Thread.currentThread().interrupt(); // 08-02: restore the flag
throw new IOException("Retry interrupted", e);
}
}
// =================================================================
// Utilities
// =================================================================
private String readBody(HttpURLConnection connection, int code) throws IOException {
InputStream input = (code >= 200 && code < 400)
? connection.getInputStream() : connection.getErrorStream();
if (input == null) {
return "";
}
try (InputStream stream = input) {
byte[] bytes = stream.readNBytes(MAX_BODY);
return new String(bytes, MetadataClient.charsetOf(connection.getContentType()));
}
}
/** Consumes the body so the connection returns to the internal pool. */
private void consume(HttpURLConnection connection, int code) {
try {
InputStream input = (code >= 200 && code < 400)
? connection.getInputStream() : connection.getErrorStream();
if (input != null) {
try (InputStream stream = input) {
stream.readNBytes(MAX_BODY);
}
}
} catch (IOException e) {
LOG.fine("Failure consuming the body: " + e.getMessage());
}
}
}A test with an nc replying 503 with Retry-After:
WARNING: GET http://localhost:8080/v1/books -> 503; retry 2 in 2000 ms
INFO: The server asks us to wait 2 s (we apply 2 s)
WARNING: GET http://localhost:8080/v1/books -> 503; retry 3 in 431 ms
Response[code=200, attempts=3]Comments. Three central ideas.
The table of what gets retried is the heart of the exercise, and every exclusion has its reason. The 4xx are not retried because the request is malformed and sending the same thing again gives the same thing. The 500 is excluded even though it is a 5xx because it usually indicates a deterministic failure —a programming error on the server— that will repeat identically. UnknownHostException is not retried because a name that does not exist is not going to start existing in 400 ms.
The jitter looks like a minor detail and it is not. Without it, a hundred clients failing simultaneously retry simultaneously, and the synchronised burst takes down the service that was recovering. It repeats indefinitely. With 20 % randomness the retries are spread out and the server receives a gradual load it can actually absorb. It is a standard technique in any serious distributed system.
Respecting Retry-After with a cap combines courtesy with prudence: the server is obeyed, since it knows better than you when it will be ready, but with a limit, because a Retry-After: 3600 cannot leave your thread blocked for an hour. And notice that the date format is explicitly ignored with a reason given: parsing HTTP dates correctly is java.time's job, and that is 10-05.
Solution 3
package com.nexussoftware.bibliotech.network;
import com.nexussoftware.bibliotech.domain.Material;
import com.nexussoftware.bibliotech.exception.BiblioTechException;
import com.nexussoftware.bibliotech.service.ConcurrentCatalog;
import com.nexussoftware.bibliotech.network.MetadataClient.Metadata;
import java.io.IOException;
import java.io.InputStream;
import java.io.OutputStream;
import java.net.HttpURLConnection;
import java.net.URI;
import java.net.URL;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
import java.text.SimpleDateFormat;
import java.util.Date;
import java.util.List;
import java.util.Locale;
import java.util.TimeZone;
import java.util.logging.Level;
import java.util.logging.Logger;
/**
* Synchronises BiblioTech's catalogue covers with the external
* metadata service. Sequential and respectful of the service.
*/
public class CoverSynchroniser {
private static final Logger LOG =
Logger.getLogger(CoverSynchroniser.class.getName());
private static final long MAX_COVER = 5L * 1024 * 1024;
private static final int REQUESTS_PER_SECOND = 2;
private static final int CONNECT_TIMEOUT_MS = 5_000;
private static final int READ_TIMEOUT_MS = 30_000;
private static final String USER_AGENT = "BiblioTech/1.0";
private final ConcurrentCatalog catalog;
private final MetadataClient metadata;
private final Path directory;
/** Instant of the last network access, for the rate limiter. */
private long lastRequest = 0;
private int downloaded = 0;
private int upToDate = 0;
private int skipped = 0;
private int failed = 0;
private long totalBytes = 0;
public CoverSynchroniser(ConcurrentCatalog catalog,
MetadataClient metadata, Path directory) {
this.catalog = catalog;
this.metadata = metadata;
this.directory = directory;
}
public void synchronise() throws IOException {
Files.createDirectories(directory);
long start = System.currentTimeMillis();
List<Material> materials = catalog.all();
System.out.println("Synchronising the covers of " + materials.size()
+ " materials...\n");
// SEQUENTIAL on purpose. Doing it in PARALLEL -which would cut the
// total time drastically, because almost all of it is network waiting-
// requires the asynchronous HttpClient with sendAsync and allOf, and that
// is exactly what is done in 09-06.
for (Material material : materials) {
process(material);
}
report(System.currentTimeMillis() - start);
}
private void process(Material material) {
String isbn = material.getIsbn();
Path target = directory.resolve(isbn + ".jpg");
try {
rateLimit();
Metadata m = metadata.query(isbn);
if (m == null || m.coverUrl() == null || m.coverUrl().isBlank()) {
System.out.printf(" %-18s SKIPPED (no cover in the service)%n", isbn);
skipped++;
return;
}
rateLimit();
long bytes = downloadIfChanged(m.coverUrl(), target, isbn);
if (bytes < 0) {
upToDate++; // 304: we already had it up to date
} else if (bytes == 0) {
skipped++; // rejected by type or size
} else {
downloaded++;
totalBytes += bytes;
}
} catch (BiblioTechException | IOException e) {
failed++;
System.out.printf(" %-18s FAILED: %s%n", isbn, e.getMessage());
LOG.log(Level.FINE, "Failure synchronising " + isbn, e);
}
}
/**
* Downloads the cover if it has changed.
* @return bytes downloaded; -1 if the server returned 304; 0 if it was rejected.
*/
private long downloadIfChanged(String urlText, Path target, String isbn)
throws IOException, BiblioTechException {
URL url = URI.create(urlText).toURL();
if (!url.getProtocol().startsWith("http")) {
throw new BiblioTechException("Scheme not allowed: " + url.getProtocol());
}
// --- Phase 1: HEAD to check the type and the size ---
HttpURLConnection head = (HttpURLConnection) url.openConnection();
try {
head.setRequestMethod("HEAD");
head.setRequestProperty("User-Agent", USER_AGENT);
head.setConnectTimeout(CONNECT_TIMEOUT_MS);
head.setReadTimeout(READ_TIMEOUT_MS);
int code = head.getResponseCode();
if (code == HttpURLConnection.HTTP_OK) {
String type = head.getContentType();
long length = head.getContentLengthLong();
if (type != null && !type.startsWith("image/")) {
System.out.printf(" %-18s SKIPPED (not an image: %s)%n", isbn, type);
return 0;
}
if (length > MAX_COVER) {
System.out.printf(" %-18s SKIPPED (%d bytes, over the limit)%n",
isbn, length);
return 0;
}
}
// A failing HEAD does not stop us trying the GET: many
// servers do not implement it properly.
} finally {
head.disconnect();
}
// --- Phase 2: conditional GET ---
HttpURLConnection connection = (HttpURLConnection) url.openConnection();
Path temporary = null;
try {
connection.setRequestMethod("GET");
connection.setRequestProperty("User-Agent", USER_AGENT);
connection.setRequestProperty("Accept", "image/jpeg, image/png, image/*");
connection.setConnectTimeout(CONNECT_TIMEOUT_MS);
connection.setReadTimeout(READ_TIMEOUT_MS);
// CONDITIONAL DOWNLOAD: if we already have the file, we ask the
// server to send it only if it has changed since then.
// A 304 saves the entire transfer.
if (Files.exists(target)) {
long modified = Files.getLastModifiedTime(target).toMillis();
connection.setRequestProperty("If-Modified-Since",
httpDate(new Date(modified)));
}
int code = connection.getResponseCode();
if (code == HttpURLConnection.HTTP_NOT_MODIFIED) { // 304
System.out.printf(" %-18s UP TO DATE (304, not changed)%n", isbn);
return -1;
}
if (code != HttpURLConnection.HTTP_OK) {
throw new IOException("HTTP " + code + " downloading the cover");
}
// Download to a TEMPORARY file: an interruption does not leave a
// half-written JPEG that looks valid until somebody tries to open it.
temporary = Files.createTempFile(directory, "cover-", ".tmp");
long received = 0;
try (InputStream input = connection.getInputStream();
OutputStream output = Files.newOutputStream(temporary)) {
byte[] buffer = new byte[8192];
int read;
while ((read = input.read(buffer)) != -1) {
received += read;
// The limit is checked HERE too: the HEAD's Content-Length
// may be missing or lying.
if (received > MAX_COVER) {
throw new IOException("The cover exceeds the limit while downloading");
}
output.write(buffer, 0, read); // never buffer.length
}
}
Files.move(temporary, target,
StandardCopyOption.REPLACE_EXISTING,
StandardCopyOption.ATOMIC_MOVE);
temporary = null;
System.out.printf(" %-18s DOWNLOADED (%d bytes)%n", isbn, received);
return received;
} finally {
connection.disconnect();
if (temporary != null) {
try {
Files.deleteIfExists(temporary);
} catch (IOException e) {
LOG.fine("Could not delete the temporary file " + temporary);
}
}
}
}
/**
* HTTP date format (RFC 7231): "Wed, 05 Aug 2026 09:14:22 GMT".
*
* SimpleDateFormat is used because java.time is 10-05. Two details
* that are essential and that almost everybody forgets:
* - Locale.US: without it, the day and month names come out in the
* system language ("mié", "ago") and the server does not understand them.
* - The GMT zone: the format demands it explicitly.
* In 10-05 this is done with DateTimeFormatter.RFC_1123_DATE_TIME,
* which is immutable and thread-safe; SimpleDateFormat is NOT.
*/
private String httpDate(Date date) {
SimpleDateFormat format = new SimpleDateFormat(
"EEE, dd MMM yyyy HH:mm:ss zzz", Locale.US);
format.setTimeZone(TimeZone.getTimeZone("GMT"));
return format.format(date);
}
/**
* Rate limiter: no more than REQUESTS_PER_SECOND to the service.
* Being a good citizen stops your IP being blocked, and avoids
* provoking the 429s that would then have to be handled.
*/
private void rateLimit() {
long interval = 1000 / REQUESTS_PER_SECOND;
long since = System.currentTimeMillis() - lastRequest;
if (since < interval) {
try {
Thread.sleep(interval - since);
} catch (InterruptedException e) {
Thread.currentThread().interrupt(); // 08-02
}
}
lastRequest = System.currentTimeMillis();
}
private void report(long ms) {
System.out.println();
System.out.println("=== COVER SYNCHRONISATION ===");
System.out.printf("%-22s %d%n", "Downloaded", downloaded);
System.out.printf("%-22s %d%n", "Up to date (304)", upToDate);
System.out.printf("%-22s %d%n", "Skipped", skipped);
System.out.printf("%-22s %d%n", "Failed", failed);
System.out.printf("%-22s %.1f KB%n", "Bytes downloaded", totalBytes / 1024.0);
System.out.printf("%-22s %.1f s%n", "Total time", ms / 1000.0);
System.out.printf("%-22s %s%n", "Directory", directory.toAbsolutePath());
}
}Typical output:
Synchronising the covers of 5 materials...
978-0000000001 DOWNLOADED (48213 bytes)
978-0000000002 UP TO DATE (304, not changed)
978-0000000003 DOWNLOADED (39104 bytes)
978-0000000004 SKIPPED (no cover in the service)
978-0000000005 FAILED: HTTP 404 downloading the cover
=== COVER SYNCHRONISATION ===
Downloaded 2
Up to date (304) 1
Skipped 1
Failed 1
Bytes downloaded 85.3 KB
Total time 5.4 sComments. Four points.
The conditional download with If-Modified-Since is HTTP's most profitable optimisation and almost nobody uses it. A 304 has an empty body: it saves the whole transfer in exchange for one network round trip. In a catalogue of a thousand books with 50 KB covers, synchronising without the conditional moves 50 MB every time; with it, a few kilobytes unless something has genuinely changed.
The HTTP date format has two traps and both are in the code. Without Locale.US, a system in Spanish generates mié, 05 ago 2026 and the server ignores it silently, so the conditional stops working without anybody noticing. And without setTimeZone("GMT"), the date comes out in local time and the server interprets it as GMT, with an offset of an hour or two. It is exactly the kind of problem java.time solves at the root in 10-05.
The rate limiter is courtesy and prudence at once. A client that fires a thousand requests in two seconds ends up with its IP blocked, and in the meantime provokes the 429s that would then have to be handled with retries. Two per second is slow but sustainable.
And the implicit TODO that points at the next lesson. Synchronising five covers takes 5.4 seconds, and practically all of that time is network waiting: the CPU is idle. A thousand covers would take almost twenty minutes. Since the downloads are independent of each other, doing them in parallel would cut the time by nearly the parallelism factor — and that is exactly what you will do in 09-06 with sendAsync and allOf. With HttpURLConnection you would have to set up the pool and the tasks by hand; with the modern API, it is a chain of three calls.
Conclusion
You have moved up a level: you have stopped inventing protocols and learned to speak the one everybody understands.
You know how to break a URL into its six parts —scheme, host, port, path, query and fragment— with the details that bite: the fragment is never sent to the server and getPort() returns -1 when it is not explicit. And you know the difference between URL and URI and the rule that follows from it: use URI to represent, manipulate and compare, and convert to URL only in order to connect, because URL.equals() performs DNS resolution and turns a simple comparison into a blocking network operation.
You can handle parameter encoding and you know why it is essential: without it a space breaks the request, an accented letter arrives corrupt and an & in a value injects parameters that the server treats as its own. With the distinction almost nobody knows: URLEncoder encodes the space as +, which is correct in a query value and wrong in a path segment, where URI's multi-argument constructor has to be used.
And above all you understand HTTP, you do not use it blind. You know that a request is a line of method, path and version, some headers, a mandatory empty line and an optional body; that the delimiter is \r\n and not \n; that the Host header is mandatory because it is what makes virtual hosting possible. You know that the response has the same shape with a status line, and how the end of the body is marked: Content-Length, Transfer-Encoding: chunked or the connection close — the three techniques of the delimiter problem of 09-01, all together in the same protocol. And you have seen it byte by byte by playing HTTP server with nc, exactly as you played BTCP server in 09-02: HTTP is text over a TCP socket, and that is all.
You know the methods and the two properties that govern their use: safe —it modifies nothing, it can be cached and prefetched— and idempotent —repeating it gives the same result—, which is the property that decides whether you can retry after a timeout. GET, PUT and DELETE yes; POST and PATCH no, because a retry may duplicate the operation. And you know the status codes by family, with the rule you applied in your own BTCP/1: the first digit decides without reading the text. Along with the list of which deserve a retry —429 respecting Retry-After, and 502, 503, 504— and which do not.
You know how to use HttpURLConnection correctly, which is no small thing: openConnection() does not connect —getResponseCode() does, which is why all the configuration goes first—; mandatory time limits for connection and for reading, both defaulting to infinite; getErrorStream() with 4xx and 5xx, because getInputStream() throws and takes away the error body that explains the problem; setDoOutput(true) to send a body, with the trap that it changes the method to POST without warning and with the Content-Length measured in bytes, not characters; automatic redirects that do not jump between http and https; and transparent gzip unless you touch Accept-Encoding, in which case decompressing is up to you.
And you know what needs to be known about HTTPS: that it works by itself, that certificate errors have identifiable causes, and the rule that is not negotiable: never disable certificate validation, because it turns HTTPS into HTTP with extra steps and opens the door to a man in the middle (12-07).
BiblioTech has started talking to the outside world. MetadataClient queries Nexus Software's service by ISBN and downloads covers to disk combining HTTP with module 7's NIO.2: scheme validation so that a file:// URL cannot make it read local files, checking the type and the size before and during the download because the Content-Length may lie, writing to a temporary file and an atomic move so an interruption never leaves a half-written JPEG, bounded reading with readNBytes instead of readAllBytes, and translation of every failure into BiblioTechException telling transient from permanent. Plus UrlQuery, HttpInspector, ResilientHttpClient with its retry policy and its jitter against the thundering herd, and CoverSynchroniser with its conditional download and its rate limiter.
And you have seen, honestly, the JSON stopgap: extracting fields by searching for substrings works with this service's particular response and breaks with escapes, nesting, arrays or a change in the order of the fields. It is flagged as what it is —a teaching stopgap— because doing it properly requires Jackson, and that is 11-07.
With the final assessment it deserves: HttpURLConnection is verbose, configured by side effects, without a total time limit, without asynchrony, HTTP/1.1 only and hard to test. You would not use it for new code. But it is in the whole standard library, you will find it in legacy code, and —the important part— it has taught you HTTP hands-on. The hard part of HTTP was never the API.
In the next lesson, The Modern HTTP Client, comes the reward. Java 11's java.net.http with its three pieces —HttpClient, HttpRequest, HttpResponse—, immutable and with fluent builders, a client that is created once and reused with its internal connection pool, genuine total time limits, HTTP/2 with multiplexing, and HttpResponse<T> with body handlers that give you text, lines, a stream or a file directly. And above all, the moment when two modules meet: sendAsync returns a CompletableFuture<HttpResponse<String>>, and everything you learned in 08-07 —thenApply, thenCompose, exceptionally, orTimeout, allOf— applies as it is to query the metadata of several ISBNs in parallel and compose a report without blocking a single thread. With the performance problem the cover synchroniser left open solved in three calls. It is module 9's closing lesson.
Java Programming Course
Module 1: Introduction to Java
- Introduction to Java
- Setting Up the Development Environment
- Basic Syntax and Structure
- Variables and Data Types
- Operators
- Console Input and Output
- Your First Complete Program: BiblioTech
Module 2: Control Flow
- Conditional Statements
- Loops
- Switch Statements
- Break and Continue
- Debugging and Execution Traces
- Project: The BiblioTech Interactive Menu
Module 3: Object-Oriented Programming
- Introduction to OOP
- Classes and Objects
- Methods
- Constructors
- Inheritance
- Polymorphism
- Encapsulation
- Abstraction
- The Object Class: equals, hashCode and toString
Module 4: Advanced Object-Oriented Programming
- Interfaces
- Abstract Classes
- Inner Classes
- Anonymous Classes
- Lambda Expressions
- Functional Interfaces and Method References
- Enums and Records
Module 5: Data Structures and Collections
- Arrays
- The Collections Framework
- ArrayList
- LinkedList
- HashMap
- HashSet
- Queue and Deque
- Stack
- Sorting and Searching Collections
Module 6: Exception Handling
- Introduction to Exceptions
- The Try-Catch Block
- Throw and Throws
- Custom Exceptions
- The Finally Block
- Try-with-resources and AutoCloseable
- Error Handling Strategies and Logging
Module 7: File Input/Output
- Reading Files
- Writing Files
- File Streams
- BufferedReader and BufferedWriter
- Serialization
- The NIO.2 API: Path and Files
- Interchange Formats: CSV and Properties
Module 8: Multithreading and Concurrency
- Introduction to Multithreading
- Creating Threads
- Thread Lifecycle
- Synchronization
- Concurrency Utilities
- Concurrent Collections and Atomic Variables
- Asynchronous Tasks with CompletableFuture
Module 9: Networking
- Introduction to Networking
- Sockets
- ServerSocket
- DatagramSocket and DatagramPacket
- URL and HttpURLConnection
- The Modern HTTP Client
Module 10: Advanced Topics
- Generics
- Annotations
- Reflection
- Java 8 Features: Streams and Optional
- Dates and Times with java.time
- Java 9 and Beyond
- Memory, Garbage Collection and Performance
Module 11: Java Frameworks and Libraries
- Introduction to Java Frameworks
- Spring Framework
- Hibernate
- JUnit
- Maven
- Advanced Testing with Mockito
- Essential Ecosystem Libraries
