You have been hearing about the buffer for three lessons. In 07-01 it explained why reading character by character is a hundred times slower than necessary. In 07-02 it explained why a just-written file has zero bytes. In 07-03 it appeared in every decorator chain and in the performance table with a factor of sixty.

This lesson is the buffer's lesson. You will see what it does internally, why that improvement is so large, and the two classes that bring it to the world of text: BufferedReader, with the readLine() that is Java's most used I/O method, and BufferedWriter, with its portable newLine().

And you will see something more important than performance: stream processing. A BufferedReader lets you traverse a one-gigabyte file with constant memory consumption, line by line, holding nothing more than the current line. That capability is what turns BiblioTech into a system able to import a complete inventory, and it is what you will build in the final case study: CatalogImporter, which reads thousands of lines, validates each one, counts the good and the bad, logs the reasons and returns a report.

try-with-resources everywhere, as always. With write buffers there is an extra reason not to forget it: close() is what flushes the buffer to the disk. Without it, you have written nothing.

Contents

  1. What exactly a buffer does
  2. Without a buffer and with a buffer: the diagram
  3. The figures that justify the difference
  4. Construction by decoration and buffer size
  5. BufferedReader.readLine() and its contract
  6. The canonical reading loop
  7. ready() and why it is no good as an end condition
  8. mark and reset: looking without consuming
  9. BufferedWriter: write, newLine and flush
  10. The PrintWriter + BufferedWriter + FileWriter trio
  11. Files.newBufferedReader and newBufferedWriter
  12. Processing a large file with constant memory
  13. Reading from the console with BufferedReader versus Scanner
  14. BiblioTech: CatalogImporter with a complete report
  15. Common Mistakes and Tips
  16. Exercises

  1. What exactly a buffer does

A buffer is an intermediate array in memory that groups operations together in order to reduce the number of calls to the operating system.

Remember the diagram from 07-01: between your code and the disk there is an expensive frontier, the system call. Every time you cross it, the processor switches context, jumps into the kernel, does its job and comes back. That journey costs orders of magnitude more than executing a few instructions inside your own process.

What a BufferedReader does:

  1. When you ask it for a character, it looks at its internal array.
  2. If it has data, it gives you one without leaving the process memory.
  3. If it is empty, it asks the wrapped stream for 8192 characters in one go, stores them in its array and gives you the first one.

The result is that one read in every 8192 reaches the operating system; the other 8191 are array accesses.

The same principle, inverted, in writing:

  1. When you ask it to write, it stores in its array.
  2. When the array fills up —or when somebody calls flush() or close()— it dumps it in a single go.

From that comes the complete explanation of two things you have already seen and which now fit together:

  • Why FileReader.read() is so slow: it has no buffer. Every call is a system call.
  • Why a just-written file has 0 bytes (07-02): the data is in the buffer array, not on the disk.

  1. Without a buffer and with a buffer: the diagram

flowchart TD
    subgraph WITHOUT["WITHOUT a buffer: FileReader.read()"]
        A1["read() number 1"] --> S1["system call"]
        A2["read() number 2"] --> S2["system call"]
        A3["read() number 3"] --> S3["system call"]
        A4["... 8192 times"] --> S4["... 8192 calls"]
        S1 --> D1["DISK"]
        S2 --> D1
        S3 --> D1
        S4 --> D1
    end

    style D1 fill:#ffcdd2
flowchart TD
    subgraph WITH["WITH a buffer: BufferedReader.read()"]
        B1["read() number 1"] --> BUF["internal array<br/>of 8192 characters"]
        B2["read() number 2"] --> BUF
        B3["read() number 3"] --> BUF
        B4["... 8192 times"] --> BUF
        BUF -->|"ONCE only,<br/>when it is empty"| SYS["system call"]
        SYS --> D2["DISK"]
    end

    style BUF fill:#c8e6c9
    style D2 fill:#c8e6c9

The numerical comparison of the complete case, with a 5 MB file:

Without a buffer With an 8 KB buffer
Calls to read() from your code 5,242,880 5,242,880
Calls to the operating system 5,242,880 640
Additional memory used 0 16 KB
Indicative time ~4,500 ms ~60 ms

The code makes the same number of calls; what changes is how many cross the frontier. Eight thousand times fewer, in exchange for sixteen kilobytes of memory. It is probably the best memory-for-time trade that exists in programming.

  1. The figures that justify the difference

An experiment you can run:

import java.io.BufferedReader;
import java.io.FileReader;
import java.io.IOException;
import java.nio.charset.StandardCharsets;

public class MeasureBuffer {

    /** Without a buffer: one system call per character. */
    static long withoutBuffer(String path) throws IOException {
        long n = 0;
        try (FileReader reader = new FileReader(path, StandardCharsets.UTF_8)) {
            while (reader.read() != -1) {
                n++;
            }
        }
        return n;
    }

    /** With a buffer, same loop: the buffer absorbs the calls. */
    static long withBuffer(String path) throws IOException {
        long n = 0;
        try (BufferedReader reader = new BufferedReader(
                new FileReader(path, StandardCharsets.UTF_8))) {
            while (reader.read() != -1) {
                n++;
            }
        }
        return n;
    }

    /** With a buffer and by lines: the idiomatic form. */
    static long byLines(String path) throws IOException {
        long n = 0;
        try (BufferedReader reader = new BufferedReader(
                new FileReader(path, StandardCharsets.UTF_8))) {
            String line;
            while ((line = reader.readLine()) != null) {
                n += line.length();
            }
        }
        return n;
    }

    public static void main(String[] args) throws IOException {
        String path = "data/catalog-large.txt";        // about 5 MB

        measure("Unbuffered, character by character", () -> withoutBuffer(path));
        measure("Buffered, character by character",   () -> withBuffer(path));
        measure("Buffered, by lines",                 () -> byLines(path));
    }

    interface Measurable { long run() throws IOException; }

    static void measure(String name, Measurable m) throws IOException {
        long start = System.nanoTime();
        long result = m.run();
        long ms = (System.nanoTime() - start) / 1_000_000;
        System.out.printf("%-35s %6d ms  (%d characters)%n", name, ms, result);
    }
}

Indicative result:

Unbuffered, character by character     4512 ms  (5242880 characters)
Buffered, character by character         68 ms  (5242880 characters)
Buffered, by lines                       41 ms  (5158400 characters)

Three observations:

  1. The difference between the first and the second is a single word in the code: new BufferedReader(...). A factor of 66.
  2. The third is better still because readLine() works directly on the buffer array, looking for the line break, without calling read() once per character.
  3. The count in the third is lower because readLine() does not include the line break. That detail is part of its contract and is explained in section 5.

  1. Construction by decoration and buffer size

BufferedReader and BufferedWriter are filters in the sense of 07-03: they wrap another Reader or Writer.

import java.io.*;
import java.nio.charset.StandardCharsets;

// Reading, default size (8192 characters)
BufferedReader reader = new BufferedReader(
        new FileReader("data/catalog.txt", StandardCharsets.UTF_8));

// Reading, explicit size
BufferedReader large = new BufferedReader(
        new FileReader("data/catalog.txt", StandardCharsets.UTF_8), 65536);

// Writing
BufferedWriter writer = new BufferedWriter(
        new FileWriter("data/report.txt", StandardCharsets.UTF_8));

// Over a bridge, when the source is not a file (07-03)
BufferedReader fromConsole = new BufferedReader(
        new InputStreamReader(System.in, StandardCharsets.UTF_8));

About the size:

Size When
8192 (default) Almost always. It matches the file system block
32768 – 65536 Very large files read sequentially, if you have measured an improvement
1024 or less Never, except with very restricted memory
More than 1 MB Counterproductive: it stops fitting in cache and presses the collector

And the warning from 07-03, which applies just the same here: if you open a hundred streams at once with 1 MB buffers, that is 100 MB of heap. The default size is very well chosen; change it only with a measurement in front of you.

A detail worth pointing out: wrapping something that already has a buffer contributes nothing. new BufferedReader(new BufferedReader(...)) is redundant, and so is new BufferedInputStream(new ByteArrayInputStream(...)), because an array in memory makes no system calls. The buffer serves where there is an expensive frontier to cross.

  1. BufferedReader.readLine() and its contract

readLine() is, by a wide margin, Java's most used I/O method. Its contract has four points, and all four matter:

public String readLine() throws IOException

1. It returns the line without the terminator. If the file contains Effective Java\n, readLine() returns "Effective Java", 14 characters long. The \n is consumed but not included. That is why the count in the section 3 experiment came out lower.

2. It recognises the three conventions. \n, \r\n and \r are all treated as end of line, regardless of the system. A file written on Windows is read correctly on Linux with no effort. This is the reason section 8 of 07-02 concluded that, for your own program, the separator does not matter.

3. It returns null at the end of the file. Not "", not an exception: null. And that distinction is essential, because an empty line returns "", which is not the same thing:

Remaining content readLine() returns
Effective Java\n "Effective Java"
\n (empty line) "" (empty string, length 0)
Nothing: end of file null
No final break (last line without \n) "No final break"

4. It blocks until it has a complete line. With a file it is instantaneous. With the console or a socket, it waits until a line break arrives or the stream is closed. It is what makes the console wait for you to press Enter.

A borderline case worth knowing: a line with no terminator at the end of the file is returned anyway. The next call returns null. It is correct and avoids losing the last line of files generated by tools that do not add a final break.

And a sizing warning: readLine() loads the whole line into memory. With a file whose content is a single 2 GB line —they exist: database dumps, JSON on one line—, readLine() tries to build a 2 GB String and causes OutOfMemoryError. It is rare, but when it happens it is baffling, because the code "processes line by line" and still runs out of memory.

  1. The canonical reading loop

This is the idiomatic form, and it must always be written the same way:

try (BufferedReader reader = new BufferedReader(
        new FileReader(path, StandardCharsets.UTF_8))) {

    String line;
    while ((line = reader.readLine()) != null) {
        process(line);
    }
}

Broken down:

  1. String line; is declared outside the loop, because the condition needs to see it.
  2. line = reader.readLine() reads and assigns. The assignment is an expression, as in 07-01.
  3. The inner brackets are mandatory: without them, line = reader.readLine() != null would try to assign a boolean to a String and would not compile.
  4. != null is the end. Not !line.isEmpty(), which would stop at the first blank line.

The three mistakes that replace this loop, all of which fail:

// WRONG 1: it stops at the first empty line
while (!(line = reader.readLine()).isEmpty()) { }
// Also: NullPointerException on reaching the end, because readLine returns null

// WRONG 2: it reads every line TWICE and skips every other one
while (reader.readLine() != null) {
    process(reader.readLine());       // this is the NEXT line
}

// WRONG 3: ready() does not mean "there is data left". Section 7
while (reader.ready()) {
    process(reader.readLine());
}

The second is especially treacherous because it half works: it processes the even lines and skips the odd ones, and with a small test file it can look correct.

With a line counter, which is the norm in real code:

int lineNumber = 0;
String line;
while ((line = reader.readLine()) != null) {
    lineNumber++;

    String clean = line.trim();
    if (clean.isEmpty() || clean.startsWith("#")) {
        continue;                        // blanks and comments: 02-04
    }
    process(clean, lineNumber);          // the number, so the error can be reported
}

That lineNumber is not decorative: it is what lets you say "line 4,217: the year is not a number" instead of "import error". It is the direct application of 06-04 —an exception must carry the context whoever reads it needs— to file importing.

  1. ready() and why it is no good as an end condition

ready() exists and its name invites misunderstanding:

public boolean ready() throws IOException

It returns true if a read would not block, that is, if there is data available in the buffer or ready at the source. And that is not the same as "there is something left to read".

// WRONG: incorrect use of ready()
try (BufferedReader reader = new BufferedReader(new FileReader(path, UTF_8))) {
    while (reader.ready()) {              // <-- BUG
        System.out.println(reader.readLine());
    }
}

Why it fails, in three real scenarios:

Scenario What happens
Small local file It usually works. That is why the bug survives testing
Large file, buffer empty at that instant ready() returns false and the loop ends halfway
Console ready() is false while the user thinks: the loop never reads anything
Network socket (module 9) ready() is false between packets: data is lost
File on a slow network drive It fails intermittently and irreproducibly

It is the worst kind of bug: it works in development and fails in production with large files or slow sources, non-deterministically.

The only correct end condition is readLine() != null. ready() is for something else: checking whether you can read without blocking, in a program that has other work to do while it waits. That case belongs to module 8.

  1. mark and reset: looking without consuming

A stream does not allow you to go back... unless the buffer remembers. BufferedReader offers that limited capability:

public void mark(int readAheadLimit) throws IOException
public void reset() throws IOException
public boolean markSupported()

mark(n) marks the current position and promises to be able to return to it as long as no more than n characters are read. reset() goes back to the mark.

Typical use: inspecting the beginning of a file to decide how to process it, without losing that first line.

import java.io.BufferedReader;
import java.io.FileReader;
import java.io.IOException;
import java.nio.charset.StandardCharsets;

public class HeaderDetection {

    /**
     * Detects whether the file has a header; if it does not, the first
     * line is data and must NOT be lost.
     */
    public static void process(String path) throws IOException {
        try (BufferedReader reader = new BufferedReader(
                new FileReader(path, StandardCharsets.UTF_8))) {

            reader.mark(8192);                     // a generous margin
            String first = reader.readLine();

            boolean hasHeader = first != null
                    && (first.startsWith("#") || first.toLowerCase().contains("title"));

            if (!hasHeader) {
                reader.reset();                    // put the line back into the stream
                System.out.println("No header: the first line is data");
            } else {
                System.out.println("Header detected: " + first);
            }

            String line;
            while ((line = reader.readLine()) != null) {
                System.out.println("  data: " + line);
            }
        }
    }
}

Its limits, which must be respected:

  • mark can fail if the readAheadLimit is too large. The buffer has to grow to that size; asking for mark(Integer.MAX_VALUE) tries to reserve 2 GB.
  • Reading more than n characters invalidates the mark and reset() throws IOException.
  • Not every Reader supports it. markSupported() tells you. BufferedReader does; a plain FileReader does not.

In practice it is little used: it is nearly always simpler to read the first line and decide what to do with it in a variable. It is worth knowing because it appears in format-analysis code, and because it explains why the buffer size parameter exists.

  1. BufferedWriter: write, newLine and flush

The writing counterpart. Its API is short:

import java.io.BufferedWriter;
import java.io.FileWriter;
import java.io.IOException;
import java.nio.charset.StandardCharsets;

public class BufferedWriting {

    public static void write(String path) throws IOException {
        try (BufferedWriter writer = new BufferedWriter(
                new FileWriter(path, StandardCharsets.UTF_8))) {

            writer.write("# BiblioTech catalogue");
            writer.newLine();                          // the SYSTEM separator

            writer.write("BOOK;978-0000000001;Effective Java");
            writer.newLine();

            writer.write("BOOK;978-0000000002;Design Patterns");
            writer.newLine();
        }
        // close() -> flush() -> the data reaches the disk (07-02)
    }
}
Method What it does
write(String) Writes the string into the buffer
write(String, int, int) Writes a substring
write(char[], int, int) Writes part of an array
write(int) Writes one character
newLine() Writes System.lineSeparator()
flush() Dumps the buffer into the wrapped stream
close() flush() + cascading close

Two things to highlight:

newLine() versus write("\n"). newLine() uses the system separator; write("\n") always writes LF. It is the decision from section 8 of 07-02: newLine() for reports, a fixed "\n" for data files that are compared or versioned.

Unlike PrintWriter, BufferedWriter does throw IOException. All its methods declare it. That is a serious argument in its favour in code where the file content matters: there is no checkError() you can forget.

Comparison of the two ways of writing buffered text:

BufferedWriter PrintWriter over BufferedWriter
I/O errors Throws IOException It stores them: you must call checkError()
Formatting It has none printf, format
Line break Explicit newLine() println() adds it
Writing objects Only String and char[] print(Object) uses toString()
When to use it Data: the failure must be detected Reports and readable output

  1. The PrintWriter + BufferedWriter + FileWriter trio

This three-layer chain is the most usual one for writing text in Java, and you can now explain exactly what each layer contributes:

try (PrintWriter output = new PrintWriter(          // 3. formatting
        new BufferedWriter(                          // 2. buffer
            new FileWriter(path,                     // 1. file + charset
                    StandardCharsets.UTF_8)))) {

    output.printf("%-25s %8.2f%n", "Effective Java", 3.75);
}
Layer Type (07-03) What it contributes What happens if you remove it
FileWriter Node Connection to the file and encoding There is no file. It is essential
BufferedWriter Filter Groups the writes; newLine() It works, but much more slowly
PrintWriter Filter println, printf, print(Object) It works, but you format by hand

An honest nuance worth knowing: PrintWriter already has its own internal buffer, so the intermediate layer contributes less than people think. In many cases, new PrintWriter(new FileWriter(path, UTF_8)) performs practically the same. The three-layer chain is written out of habit and because it is explicit, not because the improvement is large.

What does change depending on the outer layer is error handling, and that is the real decision:

// Option A: PrintWriter. Convenient, but errors have to be asked for.
try (PrintWriter output = new PrintWriter(
        new BufferedWriter(new FileWriter(path, UTF_8)))) {
    output.printf("%-25s %8.2f%n", title, fine);
    if (output.checkError()) {
        throw new IOException("Failed to write " + path);
    }
}

// Option B: BufferedWriter. More verbose, but errors propagate on their own.
try (BufferedWriter output = new BufferedWriter(new FileWriter(path, UTF_8))) {
    output.write(String.format("%-25s %8.2f", title, fine));
    output.newLine();
}

Option B uses String.format to keep the formatting and BufferedWriter to keep the exceptions. It is the one we will use in BiblioTech for the data files; A is left for the query reports.

  1. Files.newBufferedReader and newBufferedWriter

Since Java 7 there has been a shorter way of building these chains, and it is the one you will see in modern code:

import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;

// Reading: equivalent to new BufferedReader(new InputStreamReader(
//          new FileInputStream(...), UTF_8))
try (BufferedReader reader = Files.newBufferedReader(
        Path.of("data/catalog.txt"), StandardCharsets.UTF_8)) {

    String line;
    while ((line = reader.readLine()) != null) {
        process(line);
    }
}

// Writing, with explicit options
try (BufferedWriter writer = Files.newBufferedWriter(
        Path.of("data/report.txt"), StandardCharsets.UTF_8)) {
    writer.write("...");
    writer.newLine();
}

Advantages over building it manually:

Manual Files.newBufferedReader
Length Three nested constructors One call
Charset It can be forgotten A mandatory parameter in practice
Exception if it does not exist FileNotFoundException NoSuchFileException, more specific
Opening options Only append The complete StandardOpenOption set

This is the recommended form in new code. Path, Files and StandardOpenOption are the NIO.2 API, developed in full in 07-06. It is introduced here because it returns exactly the BufferedReader of this lesson: the object is the same, only the way it is built changes.

  1. Processing a large file with constant memory

This is the capability that makes BufferedReader important, beyond performance.

The problem from 07-01: loading a whole file into memory consumes as much memory as the file, and with large files it causes OutOfMemoryError. The solution: process in a stream.

import java.io.BufferedReader;
import java.io.FileReader;
import java.io.IOException;
import java.nio.charset.StandardCharsets;

public class StreamProcessing {

    /**
     * Traverses a file of ANY size with constant memory.
     *
     * At any instant there is in memory: the buffer (8 KB), the current line
     * and the accumulators. It makes no difference if the file is 1 MB or 50 GB.
     */
    public static Summary analyse(String path) throws IOException {
        long lines = 0;
        long characters = 0;
        long blankLines = 0;
        int  maxLength = 0;
        String longestLine = "";

        try (BufferedReader reader = new BufferedReader(
                new FileReader(path, StandardCharsets.UTF_8))) {

            String line;
            while ((line = reader.readLine()) != null) {
                lines++;
                characters += line.length();

                if (line.isBlank()) {
                    blankLines++;
                }
                if (line.length() > maxLength) {
                    maxLength = line.length();
                    longestLine = line;         // ONE is kept, not all of them
                }
                // 'line' becomes available to the collector on the next pass
            }
        }
        return new Summary(lines, characters, blankLines, maxLength, longestLine);
    }

    public record Summary(long lines, long characters, long blankLines,
                          int maxLength, String longestLine) { }
}

Memory comparison with a 2 GB file:

Approach Peak memory Result
Files.readString() > 4 GB OutOfMemoryError
Files.readAllLines() > 5 GB (a list of millions of Strings) OutOfMemoryError
BufferedReader line by line ~50 KB It works

The rule, which 07-01 already hinted at and which now has its tool:

If the file can grow with use, process it in a stream. Configuration, templates and files of a few KB can be loaded whole. Data, logs, exports and imports, never.

And the important nuance: processing in a stream forces you to design the algorithm differently. You can only make one pass and you only see one line at a time. Counting, summing, finding the maximum, or filtering and writing to another file work perfectly. Sorting the whole file does not: that requires an external sort, which is a different problem.

  1. Reading from the console with BufferedReader versus Scanner

Since module 1 you have used Scanner for the console. BufferedReader is the alternative:

import java.io.BufferedReader;
import java.io.IOException;
import java.io.InputStreamReader;
import java.nio.charset.StandardCharsets;

public class ConsoleWithBufferedReader {

    public static void main(String[] args) throws IOException {
        // System.in is an InputStream (bytes): the BRIDGE is needed (07-03)
        BufferedReader console = new BufferedReader(
                new InputStreamReader(System.in, StandardCharsets.UTF_8));

        System.out.print("Material reference: ");
        String reference = console.readLine();

        System.out.print("Loan days: ");
        String text = console.readLine();

        // BufferedReader does NOT parse: you have to convert by hand
        int days;
        try {
            days = Integer.parseInt(text.trim());
        } catch (NumberFormatException e) {
            System.out.println("'" + text + "' is not a valid number");
            return;
        }

        System.out.printf("Loan of %s for %d days%n", reference, days);

        // NOTE: the BufferedReader is NOT closed. Closing it would close System.in
        // for the whole application, with no way of reopening it (06-06).
    }
}

The complete comparison:

Scanner BufferedReader
Type parsing nextInt, nextDouble, nextBoolean None: Integer.parseInt by hand
Reading a line nextLine() readLine()
Buffer 1024 characters 8192 characters
Speed Slower: it uses regular expressions Noticeably faster
Exceptions Unchecked (InputMismatchException) Checked (IOException)
End of input hasNextLine() returns false readLine() returns null
Sensitive to Locale Yes: decimal comma or dot No: it does not parse
Line-break trap Yes: nextInt() does not consume it No: readLine() consumes the whole line
Regular expressions useDelimiter, hasNext(pattern) It has none
Closing System.in Dangerous Dangerous

When to use each:

  • Scanner for interactive menus and console input, which is what BiblioTechMenu has been doing since module 2. Its parsing pays off and performance is irrelevant when you are waiting for somebody to type.
  • BufferedReader for files, always. In a hundred-thousand-line file, the speed difference is several seconds.
  • BufferedReader for bulk piped input, when the program receives data on stdin from another process.

And the warning you already know from 06-07, applicable to both: never close System.in. It is a global application resource; closing it leaves the program with no input forever. It is the explicit exception to the close-everything rule that 06-06 documented.

  1. BiblioTech: CatalogImporter with a complete report

Time for the lesson's big case study, and the second debt to be settled: in 06-06 you declared CatalogImporter as a sketch; here it is written in full.

The requirements, which are those of a real import:

  1. Read a file of thousands of lines with constant memory.
  2. Validate each line without aborting the import on the first bad one.
  3. Count the correct and the discarded ones, with the reason for each discard.
  4. Log the errors with the logger, not on the console.
  5. Return a report the caller can show or save.
  6. Be transactional within reason: do not leave the catalogue half done if the whole file fails.
package com.nexussoftware.bibliotech.service;

import java.io.BufferedReader;
import java.io.File;
import java.io.FileReader;
import java.io.IOException;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.logging.Level;
import java.util.logging.Logger;

import com.nexussoftware.bibliotech.domain.Book;
import com.nexussoftware.bibliotech.domain.Material;
import com.nexussoftware.bibliotech.domain.DuplicateReferenceException;
import com.nexussoftware.bibliotech.infrastructure.FileFormat;

/**
 * Bulk import of the BiblioTech catalogue from a text file.
 *
 * Expected format (one line per material):
 *   BOOK;isbn;title;author;year
 *
 * Policy (06-07):
 *   - A bad line does NOT abort the import: it is discarded and counted.
 *   - A missing or unreadable file DOES abort: there is nothing to import.
 *   - Materials are accumulated and dumped into the catalogue AT THE END, so
 *     as not to leave it half done if the file turns out to be entirely corrupt.
 *
 * It processes in a STREAM: constant memory whatever the file size.
 */
public class CatalogImporter {

    private static final Logger LOG = Logger.getLogger(CatalogImporter.class.getName());

    private static final int EXPECTED_FIELDS = 5;
    private static final int MIN_YEAR = 1450;
    private static final int MAX_YEAR = 2100;

    /** If more than this fraction of lines fails, the file is suspicious. */
    private static final double CORRUPT_FILE_THRESHOLD = 0.5;

    /** Detailed errors that are kept; the rest are only counted. */
    private static final int MAX_DETAILED_ERRORS = 100;

    private final File file;

    public CatalogImporter(String path) {
        this.file = new File(Objects.requireNonNull(path, "The path cannot be null"));
    }

    // ------------------------- THE REPORT -------------------------

    /** Reason a line is discarded. An enum (04-07): no magic strings. */
    public enum DiscardReason {
        FIELD_COUNT("Incorrect number of fields"),
        UNKNOWN_TYPE("Material type not recognised"),
        EMPTY_FIELD("Mandatory field empty"),
        YEAR_NOT_NUMERIC("The year is not a number"),
        YEAR_OUT_OF_RANGE("Year outside the accepted range"),
        INVALID_ISBN("The ISBN does not have the expected format"),
        DUPLICATE("Reference or ISBN already present");

        private final String description;
        DiscardReason(String description) { this.description = description; }
        public String getDescription() { return description; }
    }

    /** A discarded line, with everything needed to correct it. */
    public record DiscardedLine(int number, DiscardReason reason, String detail,
                                String content) {
        @Override
        public String toString() {
            return String.format("Line %d [%s]: %s | %s",
                    number, reason.name(), detail, truncate(content));
        }
        private static String truncate(String s) {
            return (s.length() <= 60) ? s : s.substring(0, 57) + "...";
        }
    }

    /** Complete result of the import. */
    public static class Report {
        private final List<Material> imported = new ArrayList<>();
        private final List<DiscardedLine> discarded = new ArrayList<>();
        private final Map<DiscardReason, Integer> byReason = new LinkedHashMap<>();

        private int linesRead = 0;
        private int linesIgnored = 0;          // blanks and comments
        private int totalDiscards = 0;         // includes the undetailed ones
        private long milliseconds = 0;
        private boolean aborted = false;
        private String abortReason = null;

        void countLine()        { linesRead++; }
        void countIgnored()     { linesIgnored++; }
        void add(Material m)    { imported.add(m); }

        void discard(DiscardedLine d) {
            totalDiscards++;
            byReason.merge(d.reason(), 1, Integer::sum);        // 05-05
            if (discarded.size() < MAX_DETAILED_ERRORS) {
                discarded.add(d);
            }
        }

        void abort(String reason) {
            aborted = true;
            abortReason = reason;
        }

        public List<Material> getImported()             { return List.copyOf(imported); }
        public List<DiscardedLine> getDiscardedLines()  { return List.copyOf(discarded); }
        public Map<DiscardReason, Integer> getByReason() { return Map.copyOf(byReason); }
        public int getLinesRead()      { return linesRead; }
        public int getLinesIgnored()   { return linesIgnored; }
        public int getImportedTotal()  { return imported.size(); }
        public int getTotalDiscards()  { return totalDiscards; }
        public long getMilliseconds()  { return milliseconds; }
        public boolean isAborted()     { return aborted; }
        public String getAbortReason() { return abortReason; }

        public double successRate() {
            int processed = imported.size() + totalDiscards;
            return (processed == 0) ? 0.0 : (imported.size() * 100.0) / processed;
        }

        /** Readable report for the presentation layer. */
        public String summary() {
            StringBuilder sb = new StringBuilder();
            sb.append("=== IMPORT REPORT ===\n");

            if (aborted) {
                sb.append("  IMPORT ABORTED: ").append(abortReason).append('\n');
                return sb.toString();
            }

            sb.append(String.format("  Lines read       : %d%n", linesRead));
            sb.append(String.format("  Ignored          : %d (blank lines and comments)%n",
                    linesIgnored));
            sb.append(String.format("  Imported         : %d%n", imported.size()));
            sb.append(String.format("  Discarded        : %d%n", totalDiscards));
            sb.append(String.format("  Success rate     : %.1f%%%n", successRate()));
            sb.append(String.format("  Time             : %d ms%n", milliseconds));

            if (!byReason.isEmpty()) {
                sb.append("  --- Discards by reason ---\n");
                for (Map.Entry<DiscardReason, Integer> e : byReason.entrySet()) {
                    sb.append(String.format("    %-22s %4d  (%s)%n",
                            e.getKey().name(), e.getValue(),
                            e.getKey().getDescription()));
                }
            }
            if (!discarded.isEmpty()) {
                sb.append("  --- First discarded lines ---\n");
                int shown = Math.min(10, discarded.size());
                for (int i = 0; i < shown; i++) {
                    sb.append("    ").append(discarded.get(i)).append('\n');
                }
                if (totalDiscards > shown) {
                    sb.append(String.format("    ... and %d more%n",
                            totalDiscards - shown));
                }
            }
            return sb.toString();
        }
    }

    // ------------------------- THE IMPORT -------------------------

    /**
     * Imports the file into the given catalogue.
     *
     * @return the report. Never null, not even if it aborts.
     */
    public Report importInto(Catalog catalog) {
        Objects.requireNonNull(catalog, "The catalogue cannot be null");

        Report report = new Report();
        long start = System.nanoTime();

        LOG.info(() -> "Starting import from " + file.getAbsolutePath());

        // BufferedReader: constant memory even if the file has millions of
        // lines. Explicit charset, the same one the exporter uses.
        try (BufferedReader reader = new BufferedReader(
                new FileReader(file, FileFormat.CHARSET))) {

            String line;
            int number = 0;

            while ((line = reader.readLine()) != null) {    // the canonical loop
                number++;
                report.countLine();
                processLine(line, number, report);
            }

        } catch (java.io.FileNotFoundException e) {
            report.abort("The file " + file.getAbsolutePath() + " does not exist");
            LOG.log(Level.WARNING, "Import aborted: file not found", e);
            return finish(report, start);

        } catch (IOException e) {
            report.abort("I/O failure reading " + file.getAbsolutePath()
                    + ": " + e.getMessage());
            LOG.log(Level.SEVERE, "Import aborted by an I/O failure", e);
            return finish(report, start);
        }

        // Sanity check: if more than half fails, something is wrong with the
        // whole file (wrong charset, different format, columns swapped).
        // Better to reject it than to put rubbish into the catalogue.
        int processed = report.getImportedTotal() + report.getTotalDiscards();
        if (processed > 10
                && report.getTotalDiscards() > processed * CORRUPT_FILE_THRESHOLD) {

            report.abort(String.format(
                    "%.0f%% of the lines failed (%d of %d). The file does not seem "
                            + "to have the expected format; nothing is imported.",
                    100 - report.successRate(), report.getTotalDiscards(), processed));

            LOG.severe(() -> "Import rejected: " + report.getAbortReason());
            return finish(report, start);
        }

        // Dumping into the catalogue AT THE END: nothing has been touched until here
        flushToCatalog(catalog, report);
        return finish(report, start);
    }

    /** Validates and builds the material of a line. Never throws upwards. */
    private void processLine(String line, int number, Report report) {
        String clean = line.trim();

        if (clean.isEmpty() || clean.startsWith(FileFormat.COMMENT)) {
            report.countIgnored();
            return;
        }

        // The -1 keeps the trailing empty fields (07-01, solution 3)
        String[] fields = clean.split(FileFormat.FIELD_SEPARATOR, -1);

        if (fields.length != EXPECTED_FIELDS) {
            report.discard(new DiscardedLine(number, DiscardReason.FIELD_COUNT,
                    "there are " + fields.length + " and " + EXPECTED_FIELDS
                            + " were expected", clean));
            return;
        }

        String type   = fields[0].trim().toUpperCase();
        String isbn   = fields[1].trim();
        String title  = fields[2].trim();
        String author = fields[3].trim();
        String yearTx = fields[4].trim();

        if (!"BOOK".equals(type)) {
            report.discard(new DiscardedLine(number, DiscardReason.UNKNOWN_TYPE,
                    "'" + type + "'", clean));
            return;
        }
        if (title.isEmpty() || isbn.isEmpty()) {
            report.discard(new DiscardedLine(number, DiscardReason.EMPTY_FIELD,
                    title.isEmpty() ? "title" : "isbn", clean));
            return;
        }
        if (!isbn.matches("\\d{3}-\\d{10}")) {
            report.discard(new DiscardedLine(number, DiscardReason.INVALID_ISBN,
                    "'" + isbn + "' does not match NNN-NNNNNNNNNN", clean));
            return;
        }

        int year;
        try {
            year = Integer.parseInt(yearTx);
        } catch (NumberFormatException e) {
            report.discard(new DiscardedLine(number, DiscardReason.YEAR_NOT_NUMERIC,
                    "'" + yearTx + "'", clean));
            return;
        }
        if (year < MIN_YEAR || year > MAX_YEAR) {
            report.discard(new DiscardedLine(number, DiscardReason.YEAR_OUT_OF_RANGE,
                    year + " outside [" + MIN_YEAR + ", " + MAX_YEAR + "]", clean));
            return;
        }

        report.add(new Book(title, author.isEmpty() ? "Unknown" : author, isbn, year));
    }

    /**
     * Dumps what was imported into the catalogue.
     *
     * Duplicates are detected HERE, because only the catalogue knows about them.
     * A duplicate does not invalidate the whole import: it is discarded and we carry on.
     */
    private void flushToCatalog(Catalog catalog, Report report) {
        List<Material> pending = report.getImported();
        List<Material> accepted = new ArrayList<>(pending.size());

        for (Material m : pending) {
            try {
                catalog.register(m);
                accepted.add(m);
            } catch (DuplicateReferenceException e) {
                report.discard(new DiscardedLine(0, DiscardReason.DUPLICATE,
                        e.getMessage(), m.getReference()));
            }
        }
        report.imported.clear();
        report.imported.addAll(accepted);          // only what was really registered
    }

    private Report finish(Report report, long start) {
        report.milliseconds = (System.nanoTime() - start) / 1_000_000;

        if (report.isAborted()) {
            LOG.warning(() -> "Import aborted: " + report.getAbortReason());
        } else {
            LOG.info(() -> String.format(
                    "Import finished: %d imported, %d discarded, %d ms",
                    report.getImportedTotal(), report.getTotalDiscards(),
                    report.getMilliseconds()));
        }
        return report;
    }
}

And its use from the presentation layer:

package com.nexussoftware.bibliotech.presentation;

import com.nexussoftware.bibliotech.service.Catalog;
import com.nexussoftware.bibliotech.service.CatalogImporter;

public class ImportDemo {

    public static void main(String[] args) {
        Catalog catalog = new Catalog();

        CatalogImporter importer =
                new CatalogImporter("data/catalog-full.txt");

        CatalogImporter.Report report = importer.importInto(catalog);

        System.out.println(report.summary());
        System.out.println("Materials in the catalogue: " + catalog.size());
    }
}

Output with a 5,000-line file containing some errors:

=== IMPORT REPORT ===
  Lines read       : 5003
  Ignored          : 3 (blank lines and comments)
  Imported         : 4962
  Discarded        : 38
  Success rate     : 99.2%
  Time             : 87 ms
  --- Discards by reason ---
    YEAR_NOT_NUMERIC         14  (The year is not a number)
    INVALID_ISBN             11  (The ISBN does not have the expected format)
    FIELD_COUNT               7  (Incorrect number of fields)
    DUPLICATE                 4  (Reference or ISBN already present)
    EMPTY_FIELD               2  (Mandatory field empty)
  --- First discarded lines ---
    Line 47 [YEAR_NOT_NUMERIC]: 'nineteen hundred' | BOOK;978-0000000047;A book about something;Author Name;ni...
    Line 112 [FIELD_COUNT]: there are 4 and 5 were expected | BOOK;978-0000000112;Another;Author
    Line 340 [INVALID_ISBN]: '97800000340' does not match NNN-NNNNNNNNNN | BOOK;97800000340;A long enough title to be truncated here...
    ... and 35 more

The seven design decisions to understand in this class:

  1. BufferedReader and the canonical loop. Constant memory: the file could have five million lines and the consumption would be the same.
  2. A bad line does not abort. It is discarded with its reason and its line number. Whoever imports wants to see all the problems at once.
  3. A mostly bad file does abort. If more than half fails, the file is not the one expected —wrong charset, columns swapped, a different format— and putting that rubbish into the catalogue would be worse than not importing. It is the recoverable/unrecoverable distinction of 06-07 applied to importing.
  4. The dump into the catalogue happens at the end. Until that moment the catalogue has not been touched, so an abort leaves it exactly as it was. It is the state consistency of 06-05, in import form.
  5. Details are capped at 100. A corrupt file with a million bad lines would generate a million DiscardedLine objects and exhaust the memory. All are counted and the first hundred are detailed: enough information to fix things, and bounded.
  6. DiscardReason is an enum, not a string. It allows grouping, counting and translating, and it eliminates magic strings. It is 04-07 applied.
  7. All the diagnostics go to the logger; the report is returned. The service class does not print: it returns an object and the caller decides whether to show it on the console, save it or email it. It is the layer separation of 06-07.

Common Mistakes and Tips

  • Using ready() as an end condition. The bug that works in development and fails in production with large files or slow sources. The correct condition is always readLine() != null.
  • Calling readLine() twice per pass. It processes the even lines and skips the odd ones. With a small test file it can look correct.
  • Confusing null with "". null is end of file; "" is an empty line. Using isEmpty() as an end condition stops the reading at the first gap and then throws NullPointerException.
  • Forgetting the brackets in while ((line = reader.readLine()) != null). It does not compile, and it is the idiomatic form to memorise.
  • Expecting readLine() to include the line break. It does not include it. If you are rewriting the file, you have to put it back with newLine().
  • Not closing the BufferedWriter. The file is left empty: the data is in the buffer. When writing, not closing is losing data (07-02).
  • Believing BufferedWriter writes immediately. It does not, and by design. If you need another process to see the data now, flush().
  • Closing System.in. It leaves the application with no input forever. It is the documented exception to the close-everything rule.
  • Wrapping something that already has a buffer. new BufferedReader(new BufferedReader(...)) contributes nothing; over an array in memory, neither does it.
  • Setting an enormous buffer "just in case". A hundred streams with 1 MB buffers are 100 MB of heap, and beyond 64 KB the improvement is negligible.
  • Using PrintWriter without checkError() for data that matters. It swallows IOExceptions. For data files, BufferedWriter, whose methods do declare them.
  • Using Scanner to read a hundred-thousand-line file. It works and takes several extra seconds. BufferedReader for files, Scanner for menus.
  • Accumulating all the lines in a list "to process them later". It cancels the advantage of stream processing and reintroduces the OutOfMemoryError.
  • Storing an error object for every bad line with no limit. A corrupt file of a million lines exhausts the memory while reporting that there is a problem.
  • Tip: always carry the line number. "Line 4,217: the year is not a number" is actionable; "import error" is not. It is 06-04 applied to files.
  • Tip: put a sanity threshold in every import. If more than half fails, the problem is the file, not the lines. Rejecting it wholesale is more correct than importing rubbish.
  • Tip: let the service return the report, not print it. The caller decides whether to show it, save it or send it. The service layer does not know who is on the other side.
  • Tip: use Files.newBufferedReader in new code. It is shorter, it forces the charset and it gives more specific exceptions. You will see it in full in 07-06.

Exercises

Exercise 1: file statistics counter

Write FileStatistics to analyse a text file in a single pass and with constant memory, producing:

  1. Number of lines, of words and of characters.
  2. Blank lines and comment lines (those starting with #).
  3. Average, minimum and maximum line length.
  4. The three most frequent words (use a HashMap from module 5; do not use the Streams API, which is 10-04).
  5. The number of the longest line and its first 60 characters.

Requirements: BufferedReader, explicit charset, try-with-resources, and a record for the result. It must work the same with a 1 KB file as with a 1 GB one.

Exercise 2: line filter with rewriting

Write CatalogFilter to read a catalogue file and write another file containing only the lines that meet a criterion, without loading either of them into memory:

  1. filter(String source, String target, LinePredicate criterion) where LinePredicate is your own functional interface (04-06) with boolean matches(String line, int number).
  2. Keep the comment lines and the header.
  3. Count lines read, written and discarded, and return the result in a record.
  4. Use atomic writing (07-02): write to a temporary file and rename at the end.
  5. A main that filters by publication year after 2000 and by author containing a given text.

Exercise 3: comparator of two catalogues

Write CatalogComparator to compare two catalogue files and report the differences:

  1. Materials only in the first (removals).
  2. Materials only in the second (additions).
  3. Materials in both but with different data (changes), indicating which field changed.
  4. Identical materials (the count only).

Requirements: read each file only once with BufferedReader, index by ISBN with a HashMap (05-05), and produce a report with printf. Explain in a comment why this algorithm does need to load one of the two files into memory and which it should be, and what you would do if neither of them fitted.

Solutions

Solution 1

package com.nexussoftware.bibliotech.util;

import java.io.BufferedReader;
import java.io.FileReader;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.util.ArrayList;
import java.util.Comparator;
import java.util.HashMap;
import java.util.List;
import java.util.Map;

/**
 * Statistics of a text file in A SINGLE PASS and with memory constant
 * with respect to the file size.
 *
 * The only thing that grows is the frequency map, which depends on the
 * number of DISTINCT words, not on the file size. A 10 GB file in English
 * has of the order of a hundred thousand distinct words: perfectly
 * manageable.
 */
public class FileStatistics {

    private static final int TOP_WORDS = 3;
    private static final int SAMPLE_LENGTH = 60;

    public record Statistics(
            long lines, long words, long characters,
            long blankLines, long commentLines,
            double averageLength, int minLength, int maxLength,
            int longestLineNumber, String longestLineSample,
            List<Map.Entry<String, Integer>> frequentWords) {

        public String report() {
            StringBuilder sb = new StringBuilder();
            sb.append("=== FILE STATISTICS ===\n");
            sb.append(String.format("  Lines             : %d%n", lines));
            sb.append(String.format("    blank           : %d%n", blankLines));
            sb.append(String.format("    comments        : %d%n", commentLines));
            sb.append(String.format("  Words             : %d%n", words));
            sb.append(String.format("  Characters        : %d%n", characters));
            sb.append(String.format("  Average length    : %.1f%n", averageLength));
            sb.append(String.format("  Minimum length    : %d%n", minLength));
            sb.append(String.format("  Maximum length    : %d (line %d)%n",
                    maxLength, longestLineNumber));
            sb.append(String.format("    sample          : %s%n", longestLineSample));
            sb.append("  Most frequent words:\n");
            for (Map.Entry<String, Integer> e : frequentWords) {
                sb.append(String.format("    %-20s %6d%n", e.getKey(), e.getValue()));
            }
            return sb.toString();
        }
    }

    public static Statistics analyse(String path) throws IOException {
        long lines = 0, words = 0, characters = 0;
        long blanks = 0, comments = 0;
        int  minimum = Integer.MAX_VALUE, maximum = 0;
        int  longestNumber = 0;
        String longestSample = "";

        // Only this grows, and with the number of DISTINCT words
        Map<String, Integer> frequencies = new HashMap<>();

        try (BufferedReader reader = new BufferedReader(
                new FileReader(path, StandardCharsets.UTF_8))) {

            String line;
            while ((line = reader.readLine()) != null) {      // canonical loop
                lines++;
                characters += line.length();

                int length = line.length();
                if (length < minimum) { minimum = length; }
                if (length > maximum) {
                    maximum = length;
                    longestNumber = (int) lines;
                    longestSample = length <= SAMPLE_LENGTH
                            ? line
                            : line.substring(0, SAMPLE_LENGTH) + "...";
                }

                String clean = line.trim();
                if (clean.isEmpty()) {
                    blanks++;
                    continue;
                }
                if (clean.startsWith("#")) {
                    comments++;
                    continue;
                }

                // Splits on anything that is not a letter or a digit
                for (String word : clean.toLowerCase().split("[^\\p{L}\\p{N}]+")) {
                    if (word.length() > 2) {              // ignores short articles
                        words++;
                        frequencies.merge(word, 1, Integer::sum);      // 05-05
                    }
                }
                // 'line' is discarded here: constant memory
            }
        }

        if (minimum == Integer.MAX_VALUE) {
            minimum = 0;                                  // empty file
        }

        return new Statistics(
                lines, words, characters, blanks, comments,
                lines == 0 ? 0.0 : (double) characters / lines,
                minimum, maximum, longestNumber, longestSample,
                mostFrequent(frequencies, TOP_WORDS));
    }

    /**
     * The n entries with the highest value.
     *
     * Without the Streams API (10-04): list + Comparator + sort, which is
     * exactly what 05-09 covered.
     */
    private static List<Map.Entry<String, Integer>> mostFrequent(
            Map<String, Integer> frequencies, int n) {

        List<Map.Entry<String, Integer>> entries = new ArrayList<>(frequencies.entrySet());

        // Descending by value; on a tie, alphabetical by key
        entries.sort(Comparator
                .comparing(Map.Entry<String, Integer>::getValue).reversed()
                .thenComparing(Map.Entry::getKey));

        return List.copyOf(entries.subList(0, Math.min(n, entries.size())));
    }

    public static void main(String[] args) throws IOException {
        String path = (args.length > 0) ? args[0] : "data/catalog.txt";
        System.out.println(analyse(path).report());
    }
}

Output:

=== FILE STATISTICS ===
  Lines             : 5003
    blank           : 1
    comments        : 2
  Words             : 19842
  Characters        : 248915
  Average length    : 49.8
  Minimum length    : 0
  Maximum length    : 87 (line 1204)
    sample          : BOOK;978-0000001204;Introduction to distributed system archi...
  Most frequent words:
    book                   5000
    programming             412
    systems                 287

The key point is in the header comment: the only piece of data that grows with the content is the frequency map, and it grows with the number of distinct words, not with the file size. Everything else is fixed-size accumulators. That is why this analysis works with a 10 GB file. Notice too that the "longest line" is stored truncated: storing it whole would be a risk if that line were 100 MB long.

Solution 2

package com.nexussoftware.bibliotech.service;

import java.io.BufferedReader;
import java.io.BufferedWriter;
import java.io.File;
import java.io.FileReader;
import java.io.FileWriter;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.util.logging.Logger;

/**
 * Filters a catalogue file writing another one with the lines meeting
 * a criterion, WITHOUT loading either of them into memory.
 *
 * It reads and writes in a stream simultaneously: at any instant there is
 * one line and two 8 KB buffers in memory.
 */
public class CatalogFilter {

    private static final Logger LOG = Logger.getLogger(CatalogFilter.class.getName());

    private static final String COMMENT = "#";

    /**
     * Filtering criterion. Your own functional interface (04-06): it receives
     * the line and its number, in case the criterion depends on the position.
     */
    @FunctionalInterface
    public interface LinePredicate {
        boolean matches(String line, int number);
    }

    public record FilterResult(int read, int written, int discarded,
                               int commentsKept, long milliseconds) {

        public String summary() {
            return String.format(
                    "Filtered: %d read, %d written, %d discarded, "
                            + "%d comments kept (%d ms)",
                    read, written, discarded, commentsKept, milliseconds);
        }
    }

    /**
     * Filters source into target ATOMICALLY (07-02).
     *
     * If something fails halfway, the previous target is left untouched.
     */
    public FilterResult filter(String source, String target, LinePredicate criterion)
            throws IOException {

        long start = System.nanoTime();

        File targetFile = new File(target);
        File temp = new File(targetFile.getAbsolutePath() + ".tmp");

        int read = 0, written = 0, discarded = 0, comments = 0;
        boolean completed = false;

        try {
            // Simultaneous reading and writing, both buffered.
            // Both resources are declared in the same try-with-resources:
            // they close in reverse order, the writer first (06-06).
            try (BufferedReader reader = new BufferedReader(
                         new FileReader(source, StandardCharsets.UTF_8));
                 BufferedWriter writer = new BufferedWriter(
                         new FileWriter(temp, StandardCharsets.UTF_8))) {

                String line;
                int number = 0;

                while ((line = reader.readLine()) != null) {
                    number++;
                    read++;

                    String clean = line.trim();

                    // Comments and headers: always kept
                    if (clean.isEmpty() || clean.startsWith(COMMENT)) {
                        writer.write(line);
                        writer.newLine();
                        comments++;
                        continue;
                    }

                    if (criterion.matches(line, number)) {
                        writer.write(line);
                        writer.newLine();            // readLine does NOT return the break
                        written++;
                    } else {
                        discarded++;
                    }
                }
                // BufferedWriter throws real IOExceptions: no checkError
            }

            // Atomic replacement
            if (targetFile.exists() && !targetFile.delete()) {
                throw new IOException("Could not replace " + target);
            }
            if (!temp.renameTo(targetFile)) {
                throw new IOException("Could not rename " + temp.getName());
            }
            completed = true;

        } finally {
            if (!completed && temp.exists() && !temp.delete()) {
                LOG.warning(() -> "Temporary file left undeleted: " + temp.getAbsolutePath());
            }
        }

        FilterResult result = new FilterResult(read, written, discarded,
                comments, (System.nanoTime() - start) / 1_000_000);

        LOG.info(result::summary);
        return result;
    }

    // ------------------------- CRITERIA -------------------------

    /** Publication year after the given one. Field 5 of the format. */
    public static LinePredicate publishedAfter(int year) {
        return (line, number) -> {
            String[] fields = line.split(";", -1);
            if (fields.length < 5) {
                return false;
            }
            try {
                return Integer.parseInt(fields[4].trim()) > year;
            } catch (NumberFormatException e) {
                return false;               // unreadable year: does not match
            }
        };
    }

    /** The author contains the given text, ignoring case. */
    public static LinePredicate authorContains(String text) {
        String wanted = text.toLowerCase();
        return (line, number) -> {
            String[] fields = line.split(";", -1);
            return fields.length >= 4 && fields[3].toLowerCase().contains(wanted);
        };
    }

    /** Combination of two criteria (lambda composition, 04-05). */
    public static LinePredicate and(LinePredicate a, LinePredicate b) {
        return (line, number) -> a.matches(line, number) && b.matches(line, number);
    }

    public static void main(String[] args) throws IOException {
        CatalogFilter filter = new CatalogFilter();

        System.out.println(filter.filter(
                "data/catalog.txt",
                "data/catalog-recent.txt",
                publishedAfter(2000)).summary());

        System.out.println(filter.filter(
                "data/catalog.txt",
                "data/catalog-bloch.txt",
                authorContains("bloch")).summary());

        System.out.println(filter.filter(
                "data/catalog.txt",
                "data/catalog-recent-fowler.txt",
                and(publishedAfter(1995), authorContains("fowler"))).summary());
    }
}

Output:

Filtered: 5003 read, 3418 written, 1582 discarded, 3 comments kept (61 ms)
Filtered: 5003 read, 12 written, 4988 discarded, 3 comments kept (48 ms)
Filtered: 5003 read, 7 written, 4993 discarded, 3 comments kept (52 ms)

The four teaching points:

  1. Simultaneous stream reading and writing. Neither file is loaded: the consumption is one line and two buffers, whatever the size. It is the pattern the command-line tools have always used.
  2. Both resources go in the same try-with-resources. They close in reverse order —the writer first— and closing exceptions are recorded as suppressed (06-06). Without that close, the writer's buffer would never be flushed.
  3. writer.newLine() is compulsory. readLine() does not return the terminator, so if you do not put it back, the output file comes out with every line stuck together. It is the most frequent mistake when rewriting files.
  4. Your own functional interface allows criteria to be composed with lambdas and combined with and(...), exactly like RateRule and MaterialFilter in 04-06. In 10-04 you will see that the Streams API has this very thing with Predicate and and.

Solution 3

package com.nexussoftware.bibliotech.service;

import java.io.BufferedReader;
import java.io.FileReader;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.util.ArrayList;
import java.util.HashMap;
import java.util.List;
import java.util.Map;

/**
 * Compares two catalogue files and reports additions, removals and changes.
 *
 * WHY THIS ALGORITHM DOES NEED MEMORY:
 * To know whether a material from file B exists in A, you have to be able to
 * look it up. Since streams are SEQUENTIAL and do not allow going back
 * (07-03), the only way to consult A while traversing B is to have A
 * indexed in memory.
 *
 * WHICH ONE TO LOAD: the smaller one. The big one is traversed in a stream.
 *
 * IF NEITHER FITS: there are two classic ways out. Sort both files by key
 * and do a "merge" traversing them in parallel, which needs constant
 * memory; or load only the KEYS (the ISBNs) with a hash of each line,
 * which takes a tiny fraction. The second is the usual one.
 */
public class CatalogComparator {

    private static final String SEPARATOR = ";";
    private static final String COMMENT = "#";

    /** Data of a material exactly as it appears in the file. */
    private record Item(String isbn, String title, String author, int year) {

        /** Which fields differ from another item with the same ISBN. */
        List<String> differencesWith(Item other) {
            List<String> changes = new ArrayList<>();
            if (!title.equals(other.title)) {
                changes.add("title: '" + title + "' -> '" + other.title + "'");
            }
            if (!author.equals(other.author)) {
                changes.add("author: '" + author + "' -> '" + other.author + "'");
            }
            if (year != other.year) {
                changes.add("year: " + year + " -> " + other.year);
            }
            return changes;
        }
    }

    public record Differences(List<String> removals, List<String> additions,
                              List<String> changes, int identical) {

        public String report() {
            StringBuilder sb = new StringBuilder();
            sb.append("=== CATALOGUE COMPARISON ===\n");
            sb.append(String.format("  Identical      : %d%n", identical));
            sb.append(String.format("  Additions      : %d%n", additions.size()));
            sb.append(String.format("  Removals       : %d%n", removals.size()));
            sb.append(String.format("  Changes        : %d%n", changes.size()));

            printBlock(sb, "ADDITIONS (only in the new one)", additions);
            printBlock(sb, "REMOVALS (only in the old one)", removals);
            printBlock(sb, "CHANGES", changes);
            return sb.toString();
        }

        private void printBlock(StringBuilder sb, String title, List<String> lines) {
            if (lines.isEmpty()) {
                return;
            }
            sb.append("  --- ").append(title).append(" ---\n");
            int shown = Math.min(10, lines.size());
            for (int i = 0; i < shown; i++) {
                sb.append("    ").append(lines.get(i)).append('\n');
            }
            if (lines.size() > shown) {
                sb.append(String.format("    ... and %d more%n", lines.size() - shown));
            }
        }
    }

    /**
     * Compares two catalogues.
     *
     * @param oldPath the one loaded INTO MEMORY (the smaller one, ideally)
     * @param newPath the one traversed IN A STREAM
     */
    public Differences compare(String oldPath, String newPath) throws IOException {

        // PHASE 1: index the old one. A single pass.
        Map<String, Item> oldItems = index(oldPath);

        List<String> additions = new ArrayList<>();
        List<String> changes = new ArrayList<>();
        int identical = 0;

        // PHASE 2: traverse the new one IN A STREAM, consulting the index
        try (BufferedReader reader = new BufferedReader(
                new FileReader(newPath, StandardCharsets.UTF_8))) {

            String line;
            while ((line = reader.readLine()) != null) {
                Item fresh = parse(line);
                if (fresh == null) {
                    continue;                        // comment or bad line
                }

                // remove() rather than get(): what is left at the end are the REMOVALS.
                // It saves a second pass and an extra structure.
                Item old = oldItems.remove(fresh.isbn());

                if (old == null) {
                    additions.add(fresh.isbn() + " - " + fresh.title());
                    continue;
                }

                List<String> diff = old.differencesWith(fresh);
                if (diff.isEmpty()) {
                    identical++;
                } else {
                    changes.add(fresh.isbn() + " - " + String.join("; ", diff));
                }
            }
        }

        // PHASE 3: what is left over in the index are the removals
        List<String> removals = new ArrayList<>();
        for (Item i : oldItems.values()) {
            removals.add(i.isbn() + " - " + i.title());
        }

        removals.sort(null);                         // natural order (05-09)
        additions.sort(null);
        changes.sort(null);

        return new Differences(removals, additions, changes, identical);
    }

    /** Loads a catalogue indexed by ISBN. A single pass, buffered. */
    private Map<String, Item> index(String path) throws IOException {
        Map<String, Item> index = new HashMap<>();

        try (BufferedReader reader = new BufferedReader(
                new FileReader(path, StandardCharsets.UTF_8))) {

            String line;
            while ((line = reader.readLine()) != null) {
                Item i = parse(line);
                if (i != null) {
                    index.put(i.isbn(), i);
                }
            }
        }
        return index;
    }

    /** Returns null if the line is not a valid record. */
    private Item parse(String line) {
        String clean = line.trim();
        if (clean.isEmpty() || clean.startsWith(COMMENT)) {
            return null;
        }
        String[] c = clean.split(SEPARATOR, -1);
        if (c.length != 5) {
            return null;
        }
        try {
            return new Item(c[1].trim(), c[2].trim(), c[3].trim(),
                    Integer.parseInt(c[4].trim()));
        } catch (NumberFormatException e) {
            return null;
        }
    }

    public static void main(String[] args) throws IOException {
        Differences d = new CatalogComparator()
                .compare("data/catalog-2025.txt", "data/catalog-2026.txt");
        System.out.println(d.report());
    }
}

Output:

=== CATALOGUE COMPARISON ===
  Identical      : 4871
  Additions      : 92
  Removals       : 37
  Changes        : 14
  --- ADDITIONS (only in the new one) ---
    978-0000005001 - Hexagonal architecture
    978-0000005002 - Modern Java
    ... and 90 more
  --- REMOVALS (only in the old one) ---
    978-0000000341 - COBOL handbook
    ... and 36 more
  --- CHANGES ---
    978-0000000002 - title: 'Cafe Society' -> 'Café Society'
    978-0000000047 - year: 2018 -> 2019
    ... and 12 more

The three points to take away:

  1. The header comment answers the question in the brief. A stream is sequential and does not allow going back, so consulting one file while traversing another forces you to hold one in memory. Load the small one, traverse the big one. And if neither fits, the two classic ways out are the merge of sorted files —constant memory— or loading only the keys with a hash of each line, which is a tiny fraction of the content.
  2. remove() instead of get(). A trick that saves work: by taking out of the index whatever is found, what remains at the end is exactly the removals. Without it you would need a second pass and an extra structure.
  3. Look at the first change in the output: 'Cafe Society' -> 'Café Society'. That is an old file written without accents, or —more likely— an old file written with the wrong charset and corrected afterwards. It is exactly the kind of difference a catalogue comparator brings to light, and one more reason for always specifying the charset.

Conclusion

You now know what a buffer is and why it is everywhere.

You understand that it is an intermediate array that groups operations in order to reduce the system calls: one in every 8192 instead of all of them. You have the figures that justify it —from 4,500 ms to 68 ms just by wrapping the stream, a factor of sixty-six— and you know that the price is sixteen kilobytes of memory, probably the best memory-for-time trade there is. And you now also have the complete explanation of two things you accepted earlier: why FileReader.read() is so slow, and why a just-written file has zero bytes.

You know how to build them by decoration over FileReader/FileWriter or over a bridge, and you know how to choose the size: 8192 unless you have measured, with the warning that a hundred streams with one-megabyte buffers are a hundred megabytes of heap, and that wrapping something that already has a buffer contributes nothing.

You have mastered readLine() and its four contract points: it returns the line without the terminator, it recognises the three end-of-line conventions —which is why a Windows file reads without trouble on Linux—, it returns null at the end and not "", and it blocks until it has a complete line. You write the canonical loop without hesitating, with the compulsory brackets and != null as the only correct condition, and you instantly recognise the three mistakes that replace it: the isEmpty() that stops at the first blank line, the double readLine() that processes every other one, and the ready() that works in development and fails in production with large files or slow sources. You also know that readLine() loads the whole line into memory, and what that means with a file that is a single giant line.

You know mark/reset for inspecting the beginning of a stream without losing it, with its real limits. And you know BufferedWriter with its portable newLine() and —the argument that weighs most— its methods do declare IOException, unlike PrintWriter, which swallows them. Hence the rule: PrintWriter for reports, BufferedWriter for data whose failure must be detected. You know exactly what each layer of the PrintWriter + BufferedWriter + FileWriter trio contributes, including the honest observation that the middle layer contributes less than people believe, because PrintWriter already has its own buffer.

You know that Files.newBufferedReader and newBufferedWriter build exactly these objects in a shorter way, forcing the charset and with more specific exceptions, and that this is the recommended form in new code. It is NIO.2, and it is the next lesson.

And you have the capability that goes beyond performance: stream processing. A two-gigabyte file that blows up with readString() or readAllLines() is traversed with BufferedReader using fifty kilobytes. You know that this forces you to design the algorithm differently —a single pass, one line at a time— and what can and cannot be done that way. You also know the complete comparison between BufferedReader and Scanner, with the criterion for choosing: Scanner for interactive menus, BufferedReader for files and bulk input; and the warning, once again, never to close System.in.

BiblioTech really imports. CatalogImporter traverses thousands of lines with constant memory, validates each one against seven rules, discards the bad ones with their line number and a reason typed in an enum, counts all of them and details the first hundred so as not to exhaust the memory reporting a problem, aborts the whole import if more than half fails —because then the problem is the file, not the lines—, dumps into the catalogue only at the end so as not to leave it half done, and returns a report instead of printing anything, because a service class does not know who is on the other side. It is module 5, module 6 and this one, working together in a single class.

Until now, everything BiblioTech saves is text you have decided how to format: BOOK;isbn;title;author;year. It works, but it has an obvious limit. What if you wanted to save the complete state of the system —the loans with their references to materials and employees, the reservations, the history, the nested incidents— with all its relationships intact? Writing that object graph by hand, with its cross-references and its shared objects, would be an enormous job full of details.

In lesson 07-05, Serialisation, you will see that Java knows how to do it on its own. You will learn what it means to turn an object graph into bytes preserving the shared references; the marker interface Serializable, which picks up the marker interfaces of 04-01, with ObjectOutputStream and ObjectInputStream; what is saved and what is not, with transient and static; the serialVersionUID you already declared in your module 6 exceptions without fully knowing what it was for, and the InvalidClassException that appears when the class changes; which class changes are compatible and which break the existing files; and —the most important part of the lesson— the real security risks of deserialising data you do not control, which is one of the most exploited attack vectors in Java's history. By the end of it, BiblioTech will save and restore a complete session between runs, and you will know exactly when not to use this tool.

Java Programming Course

Module 1: Introduction to Java

Module 2: Control Flow

Module 3: Object-Oriented Programming

Module 4: Advanced Object-Oriented Programming

Module 5: Data Structures and Collections

Module 6: Exception Handling

Module 7: File Input/Output

Module 8: Multithreading and Concurrency

Module 9: Networking

Module 10: Advanced Topics

Module 11: Java Frameworks and Libraries

Module 12: Building Real-World Applications

© Copyright 2026. All rights reserved