The previous lesson ended on an irony: BiblioTech has a @Transactional on lend that does absolutely nothing, because there is no database underneath that could commit or roll back. All the persistence is still what it was in module 7: hand-written CSV files.

This lesson replaces them with a real relational database.

Hibernate is the most widely used ORM in the Java world, and the reference implementation of Jakarta Persistence (formerly JPA). Its job is to translate between two worlds that do not fit together: the world of objects, with inheritance, references and navigation; and the world of tables, with rows, columns and foreign keys. That job is tedious, repetitive and easy to get wrong, and that is why a tool exists to automate it.

But an ORM is not a transparent layer you can use without understanding it. It is probably the piece most people abuse without knowing what is going on: queries that multiply by a hundred, exceptions that appear outside the transaction, entities that save themselves without anybody calling save, and production schemas destroyed by one misplaced property. All of that has an explanation, and this lesson gives it.

The advantage you arrive with is the usual one: Hibernate reads @Entity and @Column by reflection, exactly as your AnnotatedExporter read @CsvField. And Spring Data JPA will give you working repositories without writing the implementation, using the same Proxy.newProxyInstance from 10-03.

By the end you will know what an ORM automates and what it costs; you will map complete entities and relationships; you will understand the persistence context and an entity's states; you will know why the LazyInitializationException appears and what the N+1 problem is, with its solutions; you will write JPQL queries with parameters —and know why concatenating strings is a vulnerability—; you will handle transactions, propagation and optimistic locking; and BiblioTech will store its data in H2 with Material, Employee and Loan as real entities.

Contents

  1. The object-relational mismatch
  2. What you gain and what you lose compared with JDBC
  3. JDBC in twenty lines: the baseline
  4. JPA versus Hibernate: specification and implementation
  5. Getting JPA running with H2 in BiblioTech
  6. The first entity: @Entity, @Table, @Id
  7. @GeneratedValue and its strategies
  8. @Column and type mapping
  9. @Enumerated: why STRING and never ORDINAL
  10. LocalDate, Instant and your own converters
  11. @Embeddable and @Embedded
  12. Relationships: the four types
  13. The owning side and mappedBy
  14. FetchType: lazy versus eager loading
  15. The LazyInitializationException
  16. The N+1 problem
  17. Solving it: JOIN FETCH and @EntityGraph
  18. CascadeType and orphanRemoval
  19. The persistence context and the EntityManager
  20. The four states of an entity
  21. Automatic dirty checking
  22. flush and the ordering of operations
  23. The operations: persist, merge, find, getReference, remove
  24. Queries: JPQL
  25. Named queries, the Criteria API and native SQL
  26. SQL injection: why you never concatenate
  27. Transactions: @Transactional with a real database
  28. Propagation and isolation
  29. Optimistic locking with @Version
  30. First- and second-level caches
  31. ddl-auto and why never update in production
  32. Spring Data JPA
  33. BiblioTech: from CSV to H2
  34. Common Mistakes and Tips
  35. Exercises

  1. The object-relational mismatch

The problem has a name of its own: the object-relational impedance mismatch. The object model and the relational model were invented for different things and they do not fit together naturally.

The concrete points of friction:

Concept In objects In tables
Identity == (reference) and equals (value): two concepts Primary key: just one
Inheritance Material sealed with Book, Magazine, Dvd Does not exist
Association A directed reference (loan.material()) A foreign key, bidirectional by nature
Navigation Chained: loan.material().publisher().country() With a JOIN, all at once
Granularity Small objects: Address, Money Columns in the same row
Collections List, Set, Map Related rows
Types LocalDate, Optional, enums, record DATE, VARCHAR, NUMERIC
null Absence of a reference NULL with three-valued semantics

Look at the inheritance row, because it is the most brutal: the relational model has no inheritance. BiblioTech's sealed Material hierarchy, which in 10-06 allowed exhaustive switch expressions checked by the compiler, simply cannot be expressed in SQL. Somebody has to decide how it is represented, and there are three different strategies, each with its trade-offs.

And the navigation row is the one that causes the most performance problems: in Java, loan.material().publisher() is free; in SQL, every hop can be a query. That is where section 16's N+1 problem comes from.

An ORM (Object-Relational Mapper) is a tool that translates automatically between the two worlds. It does not remove the mismatch: it manages it. And knowing it exists is the difference between using Hibernate well and suffering it.

  1. What you gain and what you lose compared with JDBC

Honestly, because there are strong opinions in both directions and both are partly right.

Aspect Plain JDBC ORM (Hibernate/JPA)
Boilerplate Masses of it: ResultSet to object, by hand, column by column Almost none
Control over the SQL Total Partial; recoverable with native SQL
Learning curve Low (if you know SQL) High
Straightforward performance Predictable Good, with caches that help
Performance when misused Bad, and your fault Catastrophic (N+1, massive eager loading)
Portability between engines Low: the SQL varies High: dialects
Complex queries / reports Natural Awkward; you end up using native SQL
Writing an object graph Very laborious Trivial: cascades
Transparency You see every statement It generates SQL you do not see unless you ask
Caching You build it First and second level included

The reasonable professional rule:

  • ORM for the domain CRUD, which is 80% of the code of a business application: saving a loan, loading an employee, updating a material.
  • Direct SQL (JDBC, JdbcTemplate or jOOQ) for reports, complex aggregations and bulk processing.

The two coexist perfectly well in the same project, and that is what BiblioTech does.

  1. JDBC in twenty lines: the baseline

To understand what Hibernate automates, you first have to see what had to be written. This is pure JDBC reading loans:

package com.nexussoftware.bibliotech.persistence;

import java.sql.*;
import java.time.LocalDate;
import java.util.*;

public class JdbcLoanRepository {

    private final DataSource dataSource;

    public JdbcLoanRepository(DataSource dataSource) {
        this.dataSource = dataSource;
    }

    public List<Loan> findByEmployee(String employeeEmail) {
        String sql = """
                SELECT id, isbn, employee_email, loan_date,
                       due_date, return_date
                FROM loans
                WHERE employee_email = ?
                """;

        List<Loan> result = new ArrayList<>();

        // try-with-resources from module 6: closes the three resources in reverse order
        try (Connection connection = dataSource.getConnection();
             PreparedStatement statement = connection.prepareStatement(sql)) {

            statement.setString(1, employeeEmail);   // a parameter, NEVER concatenation

            try (ResultSet rows = statement.executeQuery()) {
                while (rows.next()) {
                    // MANUAL column-by-column mapping: this is what the ORM automates
                    Date returned = rows.getDate("return_date");
                    result.add(new Loan(
                            rows.getLong("id"),
                            rows.getString("isbn"),
                            rows.getString("employee_email"),
                            rows.getDate("loan_date").toLocalDate(),
                            rows.getDate("due_date").toLocalDate(),
                            returned == null
                                    ? Optional.empty()
                                    : Optional.of(returned.toLocalDate())));
                }
            }
        } catch (SQLException e) {
            throw new PersistenceException("Error finding loans for " + employeeEmail, e);
        }
        return result;
    }
}

What is in there, and what Hibernate is going to eliminate:

  1. Hand-written SQL for every query.
  2. Column-by-column mapping, with getLong, getString, getDate.
  3. Type conversion: java.sql.Date to LocalDate, null to Optional.empty().
  4. Resource management: three nested try-with-resources.
  5. Exception translation: SQLException into module 6's hierarchy.
  6. Column names as strings: a misspelled "due_date" fails at runtime.

Multiply that by the forty queries of a real application and by six entities, and you get several thousand lines that add no value and have to be maintained every time a column changes.

JDBC is not bad. It is low-level. With Hibernate, that whole query becomes:

List<Loan> loans = repository.findByEmployeeEmail(employeeEmail);

And you do not even have to write that line, as you will see in section 32.

  1. JPA versus Hibernate: specification and implementation

Another classic confusion worth clearing up.

Jakarta Persistence (JPA) is a specification: a set of standardised interfaces and annotations, in the jakarta.persistence package. It defines EntityManager, @Entity, @Id, @OneToMany, JPQL... but it contains no working code.

Hibernate is an implementation of that specification. It is what actually generates the SQL, manages the persistence context and talks to the database.

graph TD
    A["Your BiblioTech code"] --> B["JPA API<br/>jakarta.persistence<br/>@Entity, EntityManager, JPQL"]
    B --> C["Hibernate 6.x<br/>(the implementation)"]
    B -.->|alternatives| D["EclipseLink"]
    B -.-> E["OpenJPA"]
    C --> F["JDBC"]
    F --> G[("H2 / PostgreSQL /<br/>MySQL / Oracle")]
Aspect JPA (jakarta.persistence) Hibernate (org.hibernate)
What it is A specification, interfaces An implementation with code
Package jakarta.persistence.* org.hibernate.*
Examples @Entity, EntityManager, JPQL Session, @BatchSize, HQL, filters
Portable to another implementation Yes No

The practical recommendation: program against JPA whenever you can and use Hibernate extensions only when you genuinely need them. It is the same facade principle you will see with SLF4J in 11-07.

Version warning, important. Jakarta EE renamed every javax.* package to jakarta.*. Spring Boot 3 and Hibernate 6 use jakarta.persistence. If you copy an example with import javax.persistence.Entity, it is from Spring Boot 2 or earlier and it will not compile. It is the most frequent copy-and-paste error today.

  1. Getting JPA running with H2 in BiblioTech

Two dependencies added to 11-02's pom.xml:

<!-- Hibernate + JPA + Spring Data + transaction management -->
<dependency>
    <groupId>org.springframework.boot</groupId>
    <artifactId>spring-boot-starter-data-jpa</artifactId>
</dependency>

<!-- H2: an in-memory database. NOTHING to install -->
<dependency>
    <groupId>com.h2database</groupId>
    <artifactId>h2</artifactId>
    <scope>runtime</scope>
</dependency>

And the configuration in application.yml:

spring:
  datasource:
    url: jdbc:h2:mem:bibliotech;DB_CLOSE_DELAY=-1
    username: sa
    password:
    driver-class-name: org.h2.Driver

  h2:
    console:
      enabled: true          # http://localhost:8080/h2-console (dev only)
      path: /h2-console

  jpa:
    hibernate:
      ddl-auto: create-drop  # ONLY in development. See section 31
    show-sql: true           # shows the generated SQL
    properties:
      hibernate:
        format_sql: true     # formats it readably
        jdbc:
          batch_size: 25
    open-in-view: false      # IMPORTANT. See section 15

That is all it takes. Run:

mvn spring-boot:run -Dspring-boot.run.profiles=dev

And at start-up you will see the conditions from section 19 of 11-02 being met: DataSourceAutoConfiguration detects H2 on the classpath, sees that you have not defined a DataSource, and creates one. HibernateJpaAutoConfiguration detects Hibernate and assembles the EntityManagerFactory and the transaction manager. Zero configuration code.

Two notes about show-sql: it is indispensable while you are learning —you are going to see exactly what SQL Hibernate generates, and that is half the learning— and it must be switched off in production, where proper logging goes through SLF4J (11-07):

logging:
  level:
    org.hibernate.SQL: DEBUG                        # the statements
    org.hibernate.orm.jdbc.bind: TRACE              # the parameters

  1. The first entity: @Entity, @Table, @Id

We turn BiblioTech's Book into an entity:

package com.nexussoftware.bibliotech.domain;

import jakarta.persistence.*;
import java.time.LocalDate;

@Entity                                  // "this class maps to a table"
@Table(name = "materials",               // explicit table name
       indexes = @Index(name = "idx_material_isbn", columnList = "isbn", unique = true))
public class Book {

    @Id                                  // primary key
    @GeneratedValue(strategy = GenerationType.IDENTITY)
    private Long id;

    @Column(name = "isbn", nullable = false, unique = true, length = 20)
    private String isbn;

    @Column(nullable = false, length = 200)
    private String title;

    @Column(name = "publication_year")
    private Integer publicationYear;

    @Column(name = "available_copies", nullable = false)
    private int availableCopies;

    // JPA REQUIRES a no-argument constructor (it may be protected)
    protected Book() { }

    public Book(String isbn, String title, int availableCopies) {
        this.isbn = isbn;
        this.title = title;
        this.availableCopies = availableCopies;
    }

    // getters and (only the necessary) setters
    public Long getId() { return id; }
    public String getIsbn() { return isbn; }
    public String getTitle() { return title; }
    public int getAvailableCopies() { return availableCopies; }

    // Domain behaviour: the entity is not a bag of data
    public void lendOneCopy() {
        if (availableCopies <= 0) {
            throw new NoCopiesAvailableException(isbn);
        }
        availableCopies--;
    }

    public void returnOneCopy() { availableCopies++; }
}

The requirements JPA imposes on an entity, worth knowing because the errors are cryptic:

Requirement Reason
Annotated with @Entity It is what makes it mappable
An @Id Without identity there is no row
A no-argument constructor Hibernate instantiates it by reflection (10-03)
A non-final class It needs to generate proxy subclasses for lazy loading
Non-final fields It fills them in by reflection
Methods cannot be final Lazy-loading interception

And there you have why a record cannot be a JPA entity: it is final, its fields are final and it has no no-argument constructor. BiblioTech's record types (Card, SessionSummary) will still be perfect as DTOs and as query projections, but entities have to be mutable classes. It is a real ORM trade-off, not a whim.

About @Table: if you leave it out, the table is named after the class. Putting it in explicitly is good practice, because Loan would map to a table LOAN, and words such as ORDER, USER or GROUP are reserved in SQL and give bewildering errors.

  1. @GeneratedValue and its strategies

Who generates the primary key. It is a decision with a real performance impact.

Strategy How it works Advantages Drawbacks
IDENTITY The engine's auto-increment column Simple; works everywhere Prevents insert batching: it forces an immediate INSERT to learn the id
SEQUENCE A database sequence The best: it can reserve ids in batches and group inserts Not every engine has sequences (older MySQL)
TABLE An auxiliary table of counters Portable everywhere Slow; contention. Avoid it
AUTO Hibernate chooses Convenient Unpredictable across engines
(none) You assign it Full control; natural keys You have to guarantee uniqueness

The recommendation with Hibernate 6 and a database that supports sequences:

@Id
@GeneratedValue(strategy = GenerationType.SEQUENCE, generator = "material_seq")
@SequenceGenerator(name = "material_seq", sequenceName = "material_seq",
                   allocationSize = 50)   // reserves 50 ids at once: 1 query per 50 inserts
private Long id;

The reason allocationSize matters: with IDENTITY, Hibernate cannot batch inserts, because it needs to execute each INSERT so the engine returns the id. With SEQUENCE and allocationSize=50, it can accumulate 50 inserts and send them as one batch. On an import of 10,000 materials, the difference is minutes.

For this lesson's H2 examples we will use IDENTITY for simplicity, but know the reason SEQUENCE is preferred in production.

  1. @Column and type mapping

@Column describes the column. Its useful attributes:

Attribute What it does Example
name The column name @Column(name = "loan_date")
nullable NOT NULL in the generated schema nullable = false
unique A uniqueness constraint unique = true
length VARCHAR length (255 by default) length = 20
precision / scale For BigDecimal precision = 10, scale = 2
insertable / updatable Exclude from INSERT or UPDATE updatable = false
columnDefinition Raw SQL. It breaks portability A last resort

The types JPA maps automatically:

Java type Typical SQL column
String VARCHAR
int, Integer, long, Long INTEGER, BIGINT
boolean, Boolean BOOLEAN
BigDecimal NUMERIC(p,s)
LocalDate DATE
LocalDateTime TIMESTAMP
Instant TIMESTAMP (UTC)
byte[] BLOB / VARBINARY
enum Depends on @Enumerated (section 9)
Anything else Needs an AttributeConverter (section 10)

Two concrete warnings about BiblioTech:

One: never double for money. It was said back in 01-04 and here it matters more than ever. Fines are BigDecimal with explicit precision and scale:

@Column(name = "fine_amount", precision = 10, scale = 2)
private BigDecimal fineAmount;

Two: Optional is not mapped. Optional<LocalDate> returnDate was an excellent decision in 10-04 for the public API, but JPA cannot persist Optional. The solution is a nullable field and a getter that returns Optional:

@Column(name = "return_date")
private LocalDate returnDate;                    // the field: it may be null

public Optional<LocalDate> getReturnDate() {     // the API: Optional
    return Optional.ofNullable(returnDate);
}

That way you keep 10-04's guarantee —callers never receive null— and JPA can persist the field. It is the standard pattern.

  1. @Enumerated: why STRING and never ORDINAL

BiblioTech has had three enums since module 4: Severity, MaterialType and LoanStatus. Mapping them has a trap with serious consequences.

public enum LoanStatus { ACTIVE, OVERDUE, RETURNED }
// BAD: by default JPA uses ORDINAL (the index: 0, 1, 2)
@Enumerated(EnumType.ORDINAL)
private LoanStatus status;

// GOOD: always STRING
@Enumerated(EnumType.STRING)
@Column(length = 20, nullable = false)
private LoanStatus status;

Why ORDINAL is dangerous. With it, the database stores 0, 1, 2. Now imagine that six months from now somebody adds a status in its logical place:

public enum LoanStatus { ACTIVE, RENEWED, OVERDUE, RETURNED }
//                          0       1        2         3

Every loan stored with 1 (OVERDUE) now reads as RENEWED, and the 2s (RETURNED) as OVERDUE. All the historical data is silently corrupted. No exception, no warning. It is discovered weeks later, when somebody is chased for a fine on a book they returned.

On top of that, ORDINAL makes the database unreadable: SELECT * FROM loans returns numbers that mean nothing.

Aspect ORDINAL STRING
What it stores The index (0, 1, 2) The name ("ACTIVE")
Space Less A little more
Reordering or inserting constants Corrupts the data No problem
Renaming a constant No problem A migration is needed (and the failure is visible)
Readable in SQL No Yes

The rule is absolute: @Enumerated(EnumType.STRING), always. ORDINAL's space saving does not remotely compensate for the risk. And the most dangerous part is that ORDINAL is the default: if you forget the annotation, you get the bad behaviour.

  1. LocalDate, Instant and your own converters

Good news: since JPA 2.2, java.time maps natively. All the work from 10-05 carries straight over:

@Column(name = "loan_date", nullable = false)
private LocalDate loanDate;               // -> DATE

@Column(name = "due_date", nullable = false)
private LocalDate dueDate;                // -> DATE

@Column(name = "recorded_at", nullable = false)
private Instant recordedAt;               // -> TIMESTAMP in UTC

The distinction from 10-05 is still the right one: LocalDate for business dates (a loan falls due "on the 20th", with no time and no zone) and Instant for audit timestamps (an absolute instant on the timeline).

And java.util.Date and Calendar are never used again. If you see @Temporal(TemporalType.DATE) in an example, it is pre-Java 8 code.

Converters for your own types

And what if you want to persist a type JPA does not know? That is where AttributeConverter comes in. BiblioTech has a record Isbn as a value object:

package com.nexussoftware.bibliotech.domain;

public record Isbn(String value) {
    public Isbn {
        if (value == null || !value.matches("\\d{3}-\\d{10}")) {
            throw new InvalidIsbnException(value);
        }
    }
}

The converter:

package com.nexussoftware.bibliotech.persistence;

import jakarta.persistence.AttributeConverter;
import jakarta.persistence.Converter;
import com.nexussoftware.bibliotech.domain.Isbn;

@Converter(autoApply = true)   // applies to ALL fields of type Isbn
public class IsbnConverter implements AttributeConverter<Isbn, String> {

    @Override
    public String convertToDatabaseColumn(Isbn isbn) {
        return isbn == null ? null : isbn.value();
    }

    @Override
    public Isbn convertToEntityAttribute(String column) {
        return column == null ? null : new Isbn(column);
    }
}

With autoApply = true, any Isbn field is converted on its own. Without it, you have to mark it:

@Convert(converter = IsbnConverter.class)
@Column(name = "isbn", nullable = false, unique = true)
private Isbn isbn;

This lets you have value objects with validation in the domain and simple columns in the database: the best of both worlds. Other common uses: encrypting a sensitive field, storing a short list as a comma-separated string, or mapping a boolean to 'Y'/'N' in a legacy database.

  1. @Embeddable and @Embedded

An embeddable is an object that has no identity of its own and whose fields are stored in the same row as the entity containing it. It gives structure to the object model without creating tables.

package com.nexussoftware.bibliotech.domain;

import jakarta.persistence.Embeddable;

@Embeddable
public class PublicationData {

    private String publisher;
    private Integer year;
    private String language;

    protected PublicationData() { }

    public PublicationData(String publisher, Integer year, String language) {
        this.publisher = publisher;
        this.year = year;
        this.language = language;
    }

    public boolean isRecent(int currentYear) {   // behaviour of its own
        return year != null && currentYear - year <= 3;
    }
}

Usage:

@Entity
public class Book {

    @Embedded
    private PublicationData publication;

    // If the same embeddable appears twice, the columns must be renamed:
    @Embedded
    @AttributeOverrides({
        @AttributeOverride(name = "publisher", column = @Column(name = "original_publisher")),
        @AttributeOverride(name = "year",      column = @Column(name = "original_year"))
    })
    private PublicationData originalPublication;
}

The resulting table has columns publisher, year, language, original_publisher, original_year... in the same row of materials. There is no publication_data table and no JOIN.

Concept @Entity @Embeddable
A table of its own Yes No: columns in the container's
Identity (@Id) Yes No
Lifecycle Its own That of the containing entity
Can be queried alone Yes No
Example Book, Loan PublicationData, Address, Money

It is the answer to section 1's granularity problem: you can have small, cohesive objects in Java without fragmenting the schema.

  1. Relationships: the four types

BiblioTech's model:

  • A Loan belongs to one Material and one Employee.
  • A Material has many Loans.
  • An Employee has many Loans and many Reservations.
  • A Material has many Category values and a Category has many Materials.

@ManyToOne: the side that holds the foreign key

@Entity
@Table(name = "loans")
public class Loan {

    @Id @GeneratedValue(strategy = GenerationType.IDENTITY)
    private Long id;

    @ManyToOne(fetch = FetchType.LAZY, optional = false)   // ALWAYS LAZY. Section 14
    @JoinColumn(name = "material_id", nullable = false)     // the FK column
    private Material material;

    @ManyToOne(fetch = FetchType.LAZY, optional = false)
    @JoinColumn(name = "employee_id", nullable = false)
    private Employee employee;

    @Column(name = "loan_date", nullable = false)
    private LocalDate loanDate;

    @Column(name = "due_date", nullable = false)
    private LocalDate dueDate;

    @Column(name = "return_date")
    private LocalDate returnDate;

    @Enumerated(EnumType.STRING)
    @Column(nullable = false, length = 20)
    private LoanStatus status;

    @Version                                // optimistic locking. Section 29
    private Long version;
}

The loans table will have material_id and employee_id columns with foreign keys.

@OneToMany: the inverse side

@Entity
@Table(name = "employees")
public class Employee {

    @Id @GeneratedValue(strategy = GenerationType.IDENTITY)
    private Long id;

    @Column(nullable = false, unique = true)
    private String email;

    @Column(nullable = false)
    private String name;

    // mappedBy = "employee" -> the Loan field that owns the relationship
    @OneToMany(mappedBy = "employee",
               cascade = CascadeType.ALL,
               orphanRemoval = true,
               fetch = FetchType.LAZY)      // LAZY is the default here
    private List<Loan> loans = new ArrayList<>();

    // Convenience method: it keeps BOTH sides in sync
    public void addLoan(Loan loan) {
        loans.add(loan);
        loan.assignEmployee(this);
    }

    public void removeLoan(Loan loan) {
        loans.remove(loan);
        loan.assignEmployee(null);
    }
}

Those convenience methods are not optional in practice: if you add to the list without setting the field on the other side, the relationship is not persisted (section 13) and the in-memory state is left inconsistent too.

@OneToOne

@Entity
public class Employee {
    @OneToOne(mappedBy = "employee", cascade = CascadeType.ALL, fetch = FetchType.LAZY)
    private LibraryCard card;
}

@Entity
public class LibraryCard {
    @Id @GeneratedValue private Long id;

    @OneToOne(fetch = FetchType.LAZY, optional = false)
    @JoinColumn(name = "employee_id", unique = true)
    private Employee employee;          // the owning side
}

A known trap: an optional @OneToOne on the inverse side cannot be lazy with proxies, because Hibernate needs to query in order to know whether there is a row. If performance matters, make it optional = false or use a shared primary key.

@ManyToMany

@Entity
public class Material {

    @ManyToMany(fetch = FetchType.LAZY)
    @JoinTable(name = "material_category",
               joinColumns = @JoinColumn(name = "material_id"),
               inverseJoinColumns = @JoinColumn(name = "category_id"))
    private Set<Category> categories = new HashSet<>();
}

@Entity
public class Category {
    @ManyToMany(mappedBy = "categories")
    private Set<Material> materials = new HashSet<>();
}

A much-repeated piece of professional advice: a pure @ManyToMany falls short almost always. As soon as you need an attribute on the relationship —the date the category was assigned, who assigned it— you have to turn it into an intermediate entity with two @ManyToOnes. Many teams model the intermediate entity from the start.

Summary:

Annotation Where the FK lives Owning side Use in BiblioTech
@ManyToOne In this table This one Loan → Material, Loan → Employee
@OneToMany In the other table The other one (mappedBy) Employee → Loan
@OneToOne In one of the two The one with @JoinColumn Employee → LibraryCard
@ManyToMany An intermediate table The one with @JoinTable Material ↔ Category

  1. The owning side and mappedBy

This is a concept that confuses nearly everybody and causes the "I saved it and nothing was saved" bug.

In a relational database, the relationship is a single column: the foreign key. In Java, a bidirectional relationship is two fields. Somebody has to decide which of the two is in charge.

The owning side is the one that holds the foreign key. It is the only one Hibernate looks at in order to persist the relationship. The other side carries mappedBy and is read-only as far as persistence is concerned.

The consequence:

// BAD: only the inverse side is touched
Employee marta = employeeRepository.findById(1L).orElseThrow();
Loan loan = new Loan(book, LocalDate.now(clock));
marta.getLoans().add(loan);              // only the INVERSE side
loanRepository.save(loan);
// -> The employee_id column stays NULL. The relationship has NOT been saved.
// GOOD: the owning side is touched
loan.assignEmployee(marta);               // the OWNING side (@ManyToOne)
loanRepository.save(loan);
// -> employee_id is saved correctly
// BETTER: a convenience method that syncs both sides
marta.addLoan(loan);                      // sets both fields
loanRepository.save(loan);

The practical rule: whenever you have a bidirectional relationship, write convenience methods that update both sides. Never manipulate the collections directly from outside.

  1. FetchType: lazy versus eager loading

When you load a Loan, should Hibernate also load its Material and its Employee? That decision is FetchType, and it is the decision with the greatest performance impact in the whole lesson.

FetchType Behaviour When the association is loaded
EAGER Always loaded, together with the entity Immediately, with a JOIN or an extra query
LAZY A proxy is placed there; it loads on access The first time you call one of its methods

JPA's defaults are, unfortunately, inconsistent:

Annotation Default Is it a good idea?
@ManyToOne EAGER No
@OneToOne EAGER No
@OneToMany LAZY Yes
@ManyToMany LAZY Yes

Why @ManyToOne's default EAGER is bad: imagine Loan has eager @ManyToOnes to Material and Employee, and Material has an eager @ManyToOne to Publisher, which has another to Country. Loading one loan drags in five tables. And you pay that cost always, even when all you wanted was the due date.

Professional rule: set fetch = FetchType.LAZY explicitly on ALL associations. And when you need the related data, ask for it explicitly in the query with JOIN FETCH or @EntityGraph.

@ManyToOne(fetch = FetchType.LAZY)   // ALWAYS explicit
@JoinColumn(name = "material_id")
private Material material;

How a lazy proxy works, and here module 10 returns: when you load a Loan, Hibernate does not put a Material in the field. It puts an object of a generated subclass of Material that has no data, only the id. The first time you call material.getTitle(), that subclass intercepts the call, fires the SELECT and returns the value.

It is your dynamic proxy from 10-03 again, and from it come two consequences you already know:

  • An entity cannot be final: the subclass could not be generated.
  • While debugging you will see objects called Material$HibernateProxy$xYz123. That is your material, unloaded.

  1. The LazyInitializationException

The flip side of lazy loading, and one of the most famous exceptions in the Java world.

@Service
public class ReportService {

    @Transactional(readOnly = true)
    public Loan get(Long id) {
        return repository.findById(id).orElseThrow();
    }   // <- THE TRANSACTION ENDS HERE. The persistence context is closed.
}

// Somewhere else:
Loan loan = service.get(1L);
System.out.println(loan.getMaterial().getTitle());   // BOOM
org.hibernate.LazyInitializationException: could not initialize proxy
[com.nexussoftware.bibliotech.domain.Material#7] - no Session

What happened, step by step:

  1. Inside the transaction, Hibernate loaded the loan and put a proxy in the material field.
  2. The transaction ended and the persistence context was closed.
  3. On calling getTitle(), the proxy tries to fire the SELECT... and there is no longer a session to do it with.

The solutions, from best to worst:

1. Fetch what you need in the query (the correct one):

@Query("SELECT l FROM Loan l JOIN FETCH l.material JOIN FETCH l.employee WHERE l.id = :id")
Optional<Loan> findWithDetails(@Param("id") Long id);

2. Return a DTO instead of the entity (the best architecturally):

@Transactional(readOnly = true)
public LoanCard get(Long id) {
    Loan l = repository.findById(id).orElseThrow();
    // Everything is accessed INSIDE the transaction and a detached object is returned
    return new LoanCard(l.getId(), l.getMaterial().getTitle(),
                        l.getEmployee().getName(), l.getDueDate());
}

That LoanCard can perfectly well be a record — where they do fit beautifully.

3. Widen the transaction to cover the usage: sometimes a valid solution, but keeping transactions open longer than necessary has its own cost.

4. spring.jpa.open-in-view=true: discouraged, even though it is Spring Boot's default. It keeps the persistence context open for the whole HTTP request, so the exception never appears... and in exchange it hides the N+1 problem, keeps connections busy for longer and causes queries from the presentation layer. That is why section 5's application.yml explicitly contains:

spring:
  jpa:
    open-in-view: false     # let the problems show up in development

A tip: keeping it at false makes the LazyInitializationException appear on your machine instead of the performance problem appearing in production.

  1. The N+1 problem

The number one performance problem of ORMs. And it is so easy to cause that almost everybody has it without knowing.

@Transactional(readOnly = true)
public void loanReport() {
    List<Loan> loans = repository.findAll();            // 1 query

    for (Loan l : loans) {
        System.out.println(l.getMaterial().getTitle());   // 1 query EVERY TIME!
    }
}

With 500 loans, the show-sql log shows:

-- 1: the main query
select l.id, l.material_id, l.employee_id, ... from loans l

-- N: one per loan, on accessing its material
select m.id, m.isbn, m.title, ... from materials m where m.id=1
select m.id, m.isbn, m.title, ... from materials m where m.id=2
select m.id, m.isbn, m.title, ... from materials m where m.id=3
-- ... 500 times

501 queries where there should be one. At 2 ms of latency per query, a whole second lost. And if you add the employee, it is 1001.

The insidious part: in development, with 5 test loans, it is instantaneous. The problem appears in production with real data, and by then the cause is a loop that looks harmless.

graph TD
    A["findAll()<br/>1 query: 500 loans"] --> B["Loop over the 500"]
    B --> C["l.getMaterial() -> proxy"]
    C --> D["getTitle() fires a SELECT"]
    D --> E["500 additional queries"]
    E --> F["TOTAL: 501 queries<br/>instead of 1"]

An important note: switching to EAGER does not fix it. With EAGER, Hibernate always loads the associations, but it very often does so with one query per entity anyway. All you achieve is having the problem always, even when you were not going to use the materials.

How to detect it:

  1. spring.jpa.show-sql=true and count the queries in the log.
  2. Hibernate statistics: spring.jpa.properties.hibernate.generate_statistics=true.
  3. In production: metrics and distributed tracing (12-07).

  1. Solving it: JOIN FETCH and @EntityGraph

Two solutions, both from the specification.

JOIN FETCH in JPQL

public interface LoanRepository extends JpaRepository<Loan, Long> {

    @Query("""
           SELECT l FROM Loan l
           JOIN FETCH l.material
           JOIN FETCH l.employee
           WHERE l.status = :status
           """)
    List<Loan> findActiveWithDetails(@Param("status") LoanStatus status);
}

Now a single query is generated:

select l.*, m.*, e.*
from loans l
join materials m on l.material_id = m.id
join employees e on l.employee_id = e.id
where l.status = 'ACTIVE'

From 501 queries to 1.

Watch out for one detail: a JOIN FETCH over a collection can duplicate rows. If you fetch an employee with their 5 loans, the JOIN returns 5 rows and you could get the employee repeated. It is solved with SELECT DISTINCT or by asking for a Set. And a known limitation: you cannot efficiently paginate a collection JOIN FETCH; Hibernate warns that it is paginating in memory. In that case, the solution is to fetch the paginated ids in one query and the full entities in another.

@EntityGraph: declarative

public interface LoanRepository extends JpaRepository<Loan, Long> {

    @EntityGraph(attributePaths = { "material", "employee" })
    List<Loan> findByStatus(LoanStatus status);

    @EntityGraph(attributePaths = { "material", "material.categories" })
    Optional<Loan> findById(Long id);
}

The advantage: the same method can have several graphs, and there is no JPQL to write.

Comparison

Solution When Advantage Drawback
JOIN FETCH A specific, bespoke query Full control Hand-written JPQL
@EntityGraph On derived methods Declarative, reusable Less flexible
@BatchSize(size = 25) When laziness is unavoidable From 501 to 21 queries Still not 1
DTO projection Read-only queries Fetches only the needed columns They are not managed entities
EAGER Never — It makes the problem worse

The DTO projection deserves a separate mention because it is the most efficient solution for reports:

@Query("""
       SELECT new com.nexussoftware.bibliotech.domain.LoanCard(
              l.id, m.title, e.name, l.dueDate)
       FROM Loan l JOIN l.material m JOIN l.employee e
       WHERE l.status = :status
       """)
List<LoanCard> activeCards(@Param("status") LoanStatus status);

One query, only four columns, no managed entities and no risk of N+1. And LoanCard is a record — which finally finds its place in the persistence layer.

  1. CascadeType and orphanRemoval

Cascades propagate operations from the parent to the children.

CascadeType What it propagates Example in BiblioTech
PERSIST Saving the parent saves the new children Saving an Employee saves their Loans
MERGE Merging the parent merges the children Updating a detached graph
REMOVE Deleting the parent deletes the children Dangerous: deleting a Material would delete its history
REFRESH Cascading refresh Rare
DETACH Cascading detach Rare
ALL All five Convenient and dangerous
@OneToMany(mappedBy = "employee",
           cascade = { CascadeType.PERSIST, CascadeType.MERGE },
           orphanRemoval = true)
private List<Loan> loans = new ArrayList<>();

orphanRemoval = true means: if you remove a child from the collection, delete it from the database.

employee.getLoans().remove(loan);
// With orphanRemoval = true -> DELETE FROM loans WHERE id = ...
// Without it                -> UPDATE loans SET employee_id = NULL (or an error if NOT NULL)

The difference from CascadeType.REMOVE:

CascadeType.REMOVE orphanRemoval = true
Triggered when Deleting the parent Removing the child from the collection
Semantics "When you delete the whole, delete the parts" "A child without a parent makes no sense"

A serious warning: CascadeType.ALL on relationships where the child has a life of its own is a source of data loss. If Material had cascade = ALL on its loans, withdrawing a book would delete its entire loan history, with its fines and its audit trail. Use cascades only when the relationship is genuine composition (the child does not exist without the parent).

  1. The persistence context and the EntityManager

Here is JPA's conceptual heart, and what separates those who understand Hibernate from those who suffer it.

The EntityManager is the gateway into JPA. And it manages a persistence context: a cache of entities where every entity loaded or saved stays managed for as long as the transaction lasts.

Three properties of the persistence context, all with consequences:

  1. It is a first-level cache. If you ask for the same entity twice in the same transaction, the second time generates no query.
  2. It guarantees identity. Within a transaction, the same row is always the same Java object: a == b is true. That solves section 1's identity problem.
  3. It detects changes automatically. That is section 21, and it is the most surprising part.
@Transactional
public void contextExample(Long id) {
    Material a = entityManager.find(Material.class, id);   // SELECT
    Material b = entityManager.find(Material.class, id);   // no query: cache
    System.out.println(a == b);                            // true: THE SAME object
}

In Spring, the EntityManager is injected like this:

@Repository
public class JpaMaterialRepository {

    @PersistenceContext
    private EntityManager entityManager;   // Spring injects a per-transaction proxy
}

That injected EntityManager is not a real instance: it is a proxy that on each call looks up the persistence context bound to the current thread's transaction. Once again, the same mechanism from module 10.

  1. The four states of an entity

Every entity is in one of four states, and knowing which one is the key to understanding what Hibernate is going to do.

stateDiagram-v2
    [*] --> New: new Material(...)
    New --> Managed: persist()
    Managed --> Detached: transaction close<br/>or detach()
    Detached --> Managed: merge()
    Managed --> Removed: remove()
    Removed --> New: persist()
    Removed --> [*]: flush / commit
    Managed --> Managed: changes are detected on their own
    [*] --> Managed: find() / query
State What it means Is Hibernate watching it? Does it have an id?
New (transient) Just created with new No No
Managed In the persistence context Yes Yes
Detached It was managed, no longer is No Yes
Removed Marked for deletion Yes Yes

The example that clears everything up:

@Transactional
public void demonstration() {

    // NEW: just a Java object. The database knows nothing about it.
    Material book = new Book("978-0000000001", "Effective Java", 3);

    // MANAGED: it enters the context. It will be inserted on commit.
    entityManager.persist(book);

    // Managed: this change will be saved ON ITS OWN, without calling anything.
    book.setTitle("Effective Java, 3rd edition");

}   // COMMIT: INSERT with the title already corrected

// Outside the method: DETACHED. Changes are no longer detected.

And the classic error it produces:

// In a controller or service, outside a transaction
Material material = service.findByIsbn("978-0000000001");   // detached
material.setTitle("Another title");                         // NOTHING happens
// The change stays in memory and is lost.

For a change to a detached entity to be saved, it has to be re-attached with merge (section 23).

  1. Automatic dirty checking

This is the most surprising thing about JPA, and at the same time the most powerful.

@Transactional
public void renewLoan(Long loanId, int days) {
    Loan loan = repository.findById(loanId).orElseThrow();

    loan.renew(days);             // changes dueDate

    // There is NO save(). There is NO update().
}   // And yet, on commit an UPDATE is executed.
update loans set due_date=?, version=? where id=? and version=?

How? The mechanism is called dirty checking and it works like this:

  1. On loading the entity, Hibernate stores a snapshot of the state of all its fields.
  2. On flush (normally on commit), it compares the current state with the snapshot.
  3. For every differing field, it generates the corresponding UPDATE.

The practical consequences, all of them important:

One: you do not need to call save on managed entities. A great deal of Spring Data code calls save() unnecessarily. It does no harm, but it is noise.

Two: any accidental modification is persisted. If inside a transaction you touch a field of a managed entity "just to calculate something", that change goes to the database. It is a real source of baffling bugs.

Three: readOnly = true switches it off and saves work.

@Transactional(readOnly = true)   // no snapshots, no comparisons: faster
public List<Loan> list() { ... }

On queries returning many entities the difference is measurable. It is a good habit to mark readOnly = true on every read-only method.

Four: the cost grows with the number of managed entities. Loading 100,000 entities in one transaction means every flush compares 100,000 snapshots. For bulk processing you have to clear the context periodically with entityManager.clear().

  1. flush and the ordering of operations

flush is the moment when Hibernate actually writes the pending SQL to the database. It is not the same as commit.

When it happens automatically:

  1. On committing the transaction (always).
  2. Before running a query that could be affected by the pending changes.
  3. When you call entityManager.flush() explicitly.

And here comes a surprising detail: Hibernate does not execute the SQL in the order in which you call the methods. It reorders the operations like this:

  1. Entity INSERTs, in persist order
  2. UPDATEs
  3. Removal of collection elements
  4. Insertion of collection elements
  5. Entity DELETEs

That is why this code can fail incomprehensibly:

@Transactional
public void replaceMaterial(String isbn, Material replacement) {
    Material old = repository.findByIsbn(isbn).orElseThrow();
    entityManager.remove(old);          // "deletes" the one with isbn = X
    entityManager.persist(replacement); // "inserts" another with isbn = X
}   // On commit: FIRST the INSERT, THEN the DELETE
    // -> violation of the isbn uniqueness constraint

The solution is to force the order:

entityManager.remove(old);
entityManager.flush();       // execute the DELETE now
entityManager.persist(replacement);

Understanding that there is "pending SQL" executed later and reordered explains half of Hibernate's odd behaviour.

  1. The operations: persist, merge, find, getReference, remove

Operation What it does Resulting state When
persist(e) Marks a new entity for insertion Managed (the same instance) New entities
merge(e) Copies a detached entity's state into the context Managed (a different instance) Detached entities
find(C, id) Looks up by primary key Managed or null You need the data
getReference(C, id) Returns an unloaded proxy Managed (proxy) You only need the reference
remove(e) Marks for deletion Removed Deletion
detach(e) Takes it out of the context Detached Bulk processing
refresh(e) Reloads from the database Managed Discarding local changes

The difference between persist and merge (which causes real bugs)

// persist: the SAME instance becomes managed
Material book = new Book("978-0000000003", "Refactoring", 2);
entityManager.persist(book);
System.out.println(book.getId());        // it already has an id: it is the managed instance

// merge: it returns ANOTHER instance. The original stays detached.
Material detached = new Book(...);  detached.setId(7L);
Material managed = entityManager.merge(detached);

managed.setTitle("A");    // IS SAVED
detached.setTitle("B");   // NOT saved: it is still detached
System.out.println(managed == detached);   // false

The rule: always use the object merge returns. Ignoring the return value is an extremely frequent mistake.

getReference: the optimisation people forget

// You need to associate a loan with a material, but you do not need its data.

// Option A: find -> executes an unnecessary SELECT
Material material = entityManager.find(Material.class, materialId);
loan.assignMaterial(material);

// Option B: getReference -> NO query; it just creates a proxy with the id
Material material = entityManager.getReference(Material.class, materialId);
loan.assignMaterial(material);   // enough to set the FK

All the INSERT needs is the material_id, and getReference has it. You save one SELECT per operation. The flip side: if you access any property of the proxy outside the session, you get a LazyInitializationException, and if the id does not exist, the error arrives late (EntityNotFoundException on access, not on requesting the reference).

  1. Queries: JPQL

JPQL (Jakarta Persistence Query Language) looks like SQL but operates on entities and their Java attributes, not on tables and columns.

// SQL: TABLE and COLUMN names
SELECT l.* FROM loans l WHERE l.due_date < ?

// JPQL: ENTITY and ATTRIBUTE names
SELECT l FROM Loan l WHERE l.dueDate < :date

Complete examples over BiblioTech:

@Repository
public class QueryRepository {

    @PersistenceContext
    private EntityManager em;

    // A simple query with a named parameter
    public List<Loan> overdue(LocalDate date) {
        return em.createQuery("""
                SELECT l FROM Loan l
                WHERE l.dueDate < :date
                  AND l.returnDate IS NULL
                ORDER BY l.dueDate
                """, Loan.class)
                .setParameter("date", date)
                .getResultList();
    }

    // JOIN FETCH to avoid the N+1
    public List<Loan> overdueWithDetails(LocalDate date) {
        return em.createQuery("""
                SELECT l FROM Loan l
                JOIN FETCH l.material
                JOIN FETCH l.employee
                WHERE l.dueDate < :date
                """, Loan.class)
                .setParameter("date", date)
                .getResultList();
    }

    // Aggregation with GROUP BY, returning a record directly
    public List<CountByEmployee> loansByEmployee() {
        return em.createQuery("""
                SELECT new com.nexussoftware.bibliotech.domain.CountByEmployee(
                       e.name, COUNT(l))
                FROM Loan l JOIN l.employee e
                GROUP BY e.id, e.name
                HAVING COUNT(l) > 2
                ORDER BY COUNT(l) DESC
                """, CountByEmployee.class)
                .getResultList();
    }

    // Pagination
    public List<Material> page(int number, int size) {
        return em.createQuery("SELECT m FROM Material m ORDER BY m.title", Material.class)
                .setFirstResult(number * size)
                .setMaxResults(size)
                .getResultList();
    }

    // Bulk modification: it does NOT go through the persistence context
    @Modifying
    public int markOverdue(LocalDate date) {
        return em.createQuery("""
                UPDATE Loan l SET l.status = :overdue
                WHERE l.dueDate < :date AND l.returnDate IS NULL
                """)
                .setParameter("overdue", LoanStatus.OVERDUE)
                .setParameter("date", date)
                .executeUpdate();
    }
}

A warning about that last one: bulk modification queries run directly on the database and do not update entities already loaded in the persistence context, which are left out of sync. After one of those, an em.clear() is advisable.

Key differences from SQL:

JPQL SQL
FROM Loan l (entity) FROM loans l (table)
l.dueDate (attribute) l.due_date (column)
JOIN l.material (navigates the relationship) JOIN materials ON ... (explicit condition)
SELECT l returns entities SELECT * returns rows
Portable between engines Depends on the dialect

  1. Named queries, the Criteria API and native SQL

Named queries

They are declared on the entity and validated at start-up, not on execution:

@Entity
@NamedQuery(name = "Loan.overdue",
            query = """
                    SELECT l FROM Loan l
                    WHERE l.dueDate < :date AND l.returnDate IS NULL
                    """)
public class Loan { ... }
List<Loan> overdue = em.createNamedQuery("Loan.overdue", Loan.class)
        .setParameter("date", LocalDate.now(clock))
        .getResultList();

A real advantage: a syntax error in the JPQL stops the application from starting, instead of blowing up the day somebody runs that report.

The Criteria API, mentioned

For queries built dynamically —a search screen with five optional filters— concatenating JPQL by hand is ugly and dangerous. JPA offers a typed API:

CriteriaBuilder cb = em.getCriteriaBuilder();
CriteriaQuery<Loan> query = cb.createQuery(Loan.class);
Root<Loan> l = query.from(Loan.class);

List<Predicate> filters = new ArrayList<>();
if (isbn != null)  filters.add(cb.equal(l.get("material").get("isbn"), isbn));
if (from != null)  filters.add(cb.greaterThanOrEqualTo(l.get("loanDate"), from));

query.where(cb.and(filters.toArray(new Predicate[0])));
List<Loan> result = em.createQuery(query).getResultList();

It is verbose, and that is why many people prefer Spring Data's Specification or QueryDSL. It is mentioned so you know it exists and what it is for; in practice it is not used much.

Native SQL

When you need an engine-specific function, a heavily optimised query or a complex report:

List<Object[]> rows = em.createNativeQuery("""
        SELECT m.title, COUNT(l.id) AS total
        FROM materials m LEFT JOIN loans l ON l.material_id = m.id
        WHERE l.loan_date >= :from
        GROUP BY m.id, m.title
        ORDER BY total DESC
        FETCH FIRST 10 ROWS ONLY
        """)
        .setParameter("from", from)
        .getResultList();

It is perfectly legitimate. An ORM does not force you to do everything with it, and pretending otherwise leads to convoluted JPQL that nobody understands. The only loss is portability between engines.

  1. SQL injection: why you never concatenate

This is not a style recommendation. It is a security vulnerability.

// NEVER. NOT EVER.
public List<Loan> findByEmployee(String email) {
    return em.createQuery(
            "SELECT l FROM Loan l WHERE l.employee.email = '" + email + "'",
            Loan.class).getResultList();
}

If email comes from a form or from an API parameter, an attacker can send:

' OR '1'='1

And the query becomes:

SELECT l FROM Loan l WHERE l.employee.email = '' OR '1'='1'

which returns every loan of every employee. With native SQL the consequences are worse: depending on the engine and the permissions, an attacker can end up reading other tables or modifying data.

The solution is always the same, and it is trivial:

// A named parameter
em.createQuery("SELECT l FROM Loan l WHERE l.employee.email = :email", Loan.class)
  .setParameter("email", email)
  .getResultList();

// Or a positional one
em.createQuery("SELECT l FROM Loan l WHERE l.employee.email = ?1", Loan.class)
  .setParameter(1, email)
  .getResultList();

With parameters, the value is never interpreted as part of the query: it travels separately and is treated as a literal piece of data, whatever it contains. It is the same reason section 3's JDBC example used a PreparedStatement with ? and setString.

The rule, without exceptions: every value coming from outside goes in as a parameter. Not even "when I know it is a number" — because tomorrow that code gets copied somewhere it is not.

The only things that cannot be parameterised are table and column names and the direction of an ORDER BY. If you need those to be dynamic, validate them against an allow-list of permitted values; never concatenate them as they come.

Application security —input validation, authorisation, secrets management, headers— is covered in 12-07. Here the rule is enough: always parameters.

  1. Transactions: @Transactional with a real database

Now that there is a database, 11-02's @Transactional actually does something.

A transaction satisfies the ACID properties:

Property What it guarantees
Atomicity Either every operation or none
Consistency Integrity constraints are respected
Isolation Concurrent transactions do not tread on each other
Durability What is committed survives a crash

BiblioTech's case:

@Service
public class LoanManager {

    @Transactional
    public Loan lend(String isbn, String employeeEmail) {

        Material material = materialRepository.findByIsbn(isbn)
                .orElseThrow(() -> new MaterialNotFoundException(isbn));

        Employee employee = employeeRepository.findByEmail(employeeEmail)
                .orElseThrow(() -> new EmployeeNotFoundException(employeeEmail));

        if (loanRepository.countActive(employee) >= properties.loan().maxPerEmployee()) {
            throw new LoanLimitExceededException(employeeEmail);
        }

        material.lendOneCopy();              // UPDATE (through dirty checking)

        Loan loan = new Loan(material, employee,
                LocalDate.now(clock), properties.loan().defaultDays());
        loanRepository.save(loan);           // INSERT

        auditLog.record(loan);               // INSERT

        return loan;
    }

    @Transactional(readOnly = true)          // optimised: no dirty checking
    public List<Loan> activeFor(String email) {
        return loanRepository.findActiveByEmployee(email);
    }
}

If auditLog.record fails, everything is rolled back: the copy becomes available again and no orphan loan is left behind. Without a transaction, one fewer copy and an unaudited loan would remain, and nobody would know why.

Remember the two rules from 11-02, which still hold here: rollback only on unchecked exceptions by default, and internal calls do not go through the proxy.

  1. Propagation and isolation

Propagation: what to do if a transaction is already open

Propagation Behaviour Use
REQUIRED (the default) Joins the existing one; if there is none, creates it 95% of cases
REQUIRES_NEW Suspends the current one and opens an independent one Auditing that must persist even if the rest fails
SUPPORTS Joins if there is one; if not, no transaction Queries
MANDATORY Requires one to exist already; error otherwise Internal methods that must never be called alone
NOT_SUPPORTED Suspends the current one and runs without a transaction Very long operations
NEVER Error if there is a transaction Rare
NESTED A savepoint inside the current one Limited support

The case where REQUIRES_NEW is the right answer:

@Service
public class AuditLog {

    // We want to record the ATTEMPT even if the loan is rolled back
    @Transactional(propagation = Propagation.REQUIRES_NEW)
    public void recordAttempt(String isbn, String employee, String result) {
        em.persist(new AuditEntry(isbn, employee, result, Instant.now(clock)));
    }
}

Careful: REQUIRES_NEW consumes a second connection from the pool while the first is still open. Overusing it exhausts the pool and causes deadlocks. Use it only when independence is a genuine requirement.

Isolation: how concurrent transactions see each other

Level Dirty read Non-repeatable read Phantom read Cost
READ_UNCOMMITTED Possible Possible Possible Minimal
READ_COMMITTED No Possible Possible Low (the default in PostgreSQL, Oracle, H2)
REPEATABLE_READ No No Possible Medium (the default in MySQL)
SERIALIZABLE No No No High

The three phenomena, with BiblioTech:

  • Dirty read: you read a copy as being on loan and that transaction rolls back. You read something that never existed.
  • Non-repeatable read: you read available_copies = 1, somebody else lends it, you read again and there is 0. Two reads, two results.
  • Phantom read: you count 3 active loans, somebody inserts one, you repeat the query and there are 4.
@Transactional(isolation = Isolation.REPEATABLE_READ)
public MonthlyReport generate(YearMonth month) { ... }

In practice, isolation is almost never touched. READ_COMMITTED with optimistic locking (the next section) solves 99% of cases at far less cost than raising the level.

  1. Optimistic locking with @Version

The concrete case, which picks up module 8 directly:

Marta Ruiz and Diego Alonso open the record for "Effective Java" at the same time; it has 1 available copy. Both click "Lend" in the same second.

sequenceDiagram
    participant M as Marta
    participant DB as Database
    participant D as Diego
    M->>DB: SELECT material 1 (available=1)
    D->>DB: SELECT material 1 (available=1)
    M->>DB: UPDATE available=0
    D->>DB: UPDATE available=0
    Note over DB: Two loans, one copy.<br/>LOST UPDATE

It is exactly the race condition from 08-04, but distributed: here synchronized is of no use whatsoever, because the two users can be on different instances of the application.

Two strategies:

Pessimistic locking Optimistic locking
Assumption There will be a conflict There will almost never be a conflict
How SELECT ... FOR UPDATE: it locks the row A version column checked on update
Cost High: locks, waits, deadlocks Very low
When Frequent and expensive conflicts Almost always
JPA @Lock(LockModeType.PESSIMISTIC_WRITE) @Version

Optimistic locking is a single annotation:

@Entity
public class Material {

    @Id @GeneratedValue private Long id;

    @Version                        // Hibernate manages it on its own
    private Long version;

    @Column(name = "available_copies", nullable = false)
    private int availableCopies;
}

With that, every UPDATE includes the version that was read in the condition:

update materials set available_copies=0, version=6 where id=1 and version=5

If somebody else has already updated, the row has version=6 and the condition matches no row. Hibernate detects it and throws:

jakarta.persistence.OptimisticLockException:
Row was updated or deleted by another transaction

Handling it properly:

@Service
public class LoanManager {

    @Transactional
    public Loan lend(String isbn, String email) {
        // ... as before
    }

    public Result<Loan> lendWithRetry(String isbn, String email) {
        for (int attempt = 1; attempt <= 3; attempt++) {
            try {
                return Result.success(self.lend(isbn, email));
            } catch (OptimisticLockingFailureException e) {
                log.warn("Concurrency conflict on {} (attempt {}/3)", isbn, attempt);
            }
        }
        return Result.error("The material is being modified. Please try again.");
    }
}

That Result<T> is the generic type you wrote in 10-01, and it fits perfectly here. And that retry loop is, conceptually, your RetryProxy from 10-03: in a real project you would declare it with Spring Retry's @Retryable.

Pessimistic locking, when the operation is expensive and conflict is likely:

@Lock(LockModeType.PESSIMISTIC_WRITE)
@Query("SELECT m FROM Material m WHERE m.isbn = :isbn")
Optional<Material> findForUpdate(@Param("isbn") String isbn);

It generates SELECT ... FOR UPDATE: the second waits for the first. Correct, but with a real concurrency cost.

  1. First- and second-level caches

Cache Scope Enabled by default What it stores
First level The persistence context (the transaction) Yes, always Managed entities
Second level The EntityManagerFactory (the whole application) No Entities between transactions
Query cache Global No Query results

You already know the first-level one: it is section 19's persistence context. It cannot be switched off and it is what guarantees object identity.

The second-level one is optional and shared by every transaction. It needs a provider (Ehcache, Caffeine, Hazelcast):

@Entity
@Cacheable
@org.hibernate.annotations.Cache(usage = CacheConcurrencyStrategy.READ_WRITE)
public class Material { ... }
spring:
  jpa:
    properties:
      hibernate:
        cache:
          use_second_level_cache: true
          region.factory_class: org.hibernate.cache.jcache.JCacheRegionFactory

When to enable it: data that is read a lot and changes little. BiblioTech's material catalogue or its categories are good candidates; the loans are not.

Warnings to bear in mind:

  1. Coherence: if another application modifies the database, your cache goes stale and serves old data.
  2. In a cluster: with several instances, each one has its own cache. Either you use a distributed cache or you will get inconsistencies.
  3. Measure first: enabling the second-level cache "just in case" adds complexity with no demonstrated benefit. It is exactly what 10-07 said: measure before optimising.

  1. ddl-auto and why never update in production

Hibernate can generate the schema from your entities. It is convenient and dangerous.

Value What it does When to use it
none Nothing Production
validate Checks that the schema matches the entities; fails if not Production (recommended)
update Tries to modify the schema to fit Never in production
create Drops and creates the schema at start-up Tests
create-drop Like create, and drops on shutdown Development with H2, tests

Why update is never used in production, with concrete reasons:

  1. It never deletes anything. If you remove a field, the column stays there forever. If you rename title to name, it adds name and leaves title — and it does not copy the data.
  2. It versions nothing. There is no record of which changes were applied or when, and no way to undo them.
  3. It is unpredictable. The SQL it generates depends on the dialect, the Hibernate version and the current state of the schema.
  4. It cannot be reviewed beforehand. Nobody has seen the ALTER TABLE that is about to run against the production database.
  5. It cannot do data migrations. Splitting full_name into first_name and surname needs logic no automatic tool can invent.

The correct configuration per environment:

# application-dev.yml — in-memory H2, schema regenerated on every start-up
spring:
  jpa:
    hibernate:
      ddl-auto: create-drop
# application-prod.yml — the schema is managed by a migration tool
spring:
  jpa:
    hibernate:
      ddl-auto: validate

validate is especially valuable: if the real schema does not match the entities, the application does not start, and that is much better than finding out through an error at three in the morning.

Flyway and Liquibase, mentioned

In production, the schema is managed with versioned migrations: numbered SQL files applied in order and recorded in a control table.

-- src/main/resources/db/migration/V1__initial_schema.sql
CREATE TABLE materials (
    id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
    isbn VARCHAR(20) NOT NULL UNIQUE,
    title VARCHAR(200) NOT NULL,
    available_copies INT NOT NULL DEFAULT 0,
    version BIGINT NOT NULL DEFAULT 0
);
-- V2__add_category.sql
ALTER TABLE materials ADD COLUMN category VARCHAR(50);
UPDATE materials SET category = 'GENERAL' WHERE category IS NULL;

The advantages: the schema change is in git, it is reviewable, it is reproducible and it can be tested beforehand. It is developed in 12-06, alongside deployment.

  1. Spring Data JPA

The last piece, and the one that removes the most code.

With plain JPA, a repository is about eighty lines. With Spring Data JPA:

package com.nexussoftware.bibliotech.persistence;

import org.springframework.data.jpa.repository.*;
import org.springframework.data.repository.query.Param;

public interface LoanRepository extends JpaRepository<Loan, Long> {
    // And that is it. No implementation.
}

That already gives you save, findById, findAll, deleteById, count, existsById, pagination and sorting.

Methods derived from the name

public interface LoanRepository extends JpaRepository<Loan, Long> {

    List<Loan> findByStatus(LoanStatus status);

    List<Loan> findByEmployeeEmail(String email);

    List<Loan> findByDueDateBeforeAndReturnDateIsNull(LocalDate date);

    long countByEmployeeAndStatus(Employee employee, LoanStatus status);

    Optional<Loan> findFirstByMaterialIsbnOrderByLoanDateDesc(String isbn);

    boolean existsByMaterialIsbnAndStatus(String isbn, LoanStatus status);
}

Spring Data analyses the method name and generates the query. The usual keywords:

Keyword Generated JPQL
findBy, readBy, getBy SELECT ...
And, Or AND, OR
Between, LessThan, GreaterThan, Before, After Comparisons
IsNull, IsNotNull IS NULL
Like, Containing, StartingWith LIKE
In, NotIn IN
OrderBy...Asc/Desc ORDER BY
countBy, existsBy, deleteBy Aggregation, existence, deletion
First, Top3 LIMIT

A practical tip: derived names are fantastic until they stop being so. findByDueDateBeforeAndReturnDateIsNullAndStatusNot is unreadable. When the name goes beyond three conditions, use @Query.

@Query

@Query("""
       SELECT l FROM Loan l
       JOIN FETCH l.material m
       JOIN FETCH l.employee e
       WHERE l.dueDate < :date AND l.returnDate IS NULL
       ORDER BY l.dueDate
       """)
List<Loan> overdueWithDetails(@Param("date") LocalDate date);

@Modifying
@Transactional
@Query("UPDATE Loan l SET l.status = :status WHERE l.id = :id")
int updateStatus(@Param("id") Long id, @Param("status") LoanStatus status);

@Query(value = "SELECT * FROM loans WHERE loan_date > ?1", nativeQuery = true)
List<Loan> nativeSince(LocalDate from);

Pagination

Page<Loan> findByStatus(LoanStatus status, Pageable pageable);
Pageable page = PageRequest.of(0, 20, Sort.by("dueDate").descending());
Page<Loan> result = repository.findByStatus(LoanStatus.ACTIVE, page);

result.getContent();        // the 20 on this page
result.getTotalElements();  // the total (it runs an additional COUNT)
result.getTotalPages();
result.hasNext();

If you do not need the total, Slice<T> avoids the COUNT and is faster.

How it works: dynamic proxies again

LoanRepository is an interface with no implementation. Who runs findByStatus?

At start-up, Spring Data:

  1. Scans the interfaces that extend Repository.
  2. For each one, it creates an implementation on the fly with Proxy.newProxyInstance — the very one from 10-03.
  3. The InvocationHandler receives each call, checks whether the method has @Query (and uses that JPQL) or analyses its name (and generates the JPQL), runs it with the EntityManager and adapts the result.
graph LR
    A["manager.repository<br/>.findByStatus(ACTIVE)"] --> B["JDK proxy<br/>(Spring Data)"]
    B --> C["Analyses the name<br/>or reads @Query"]
    C --> D["Generates JPQL"]
    D --> E["EntityManager"]
    E --> F["Hibernate -> SQL"]
    F --> G[("H2")]

It is literally what you wrote in 10-03: a proxy over an interface that interprets the method's metadata and acts. The difference is the sophistication of the analysis, not the mechanism.

And there is a practical consequence: the repository can be replaced by a mock in tests with no difficulty at all, because it is an interface. That is exactly what 11-06 will do.

  1. BiblioTech: from CSV to H2

The complete result of the migration.

// domain/Material.java — it can no longer be a sealed record: it is an entity
package com.nexussoftware.bibliotech.domain;

import jakarta.persistence.*;

@Entity
@Table(name = "materials")
@Inheritance(strategy = InheritanceType.SINGLE_TABLE)   // one table for the whole hierarchy
@DiscriminatorColumn(name = "type", discriminatorType = DiscriminatorType.STRING, length = 20)
public abstract class Material {

    @Id @GeneratedValue(strategy = GenerationType.IDENTITY)
    private Long id;

    @Column(nullable = false, unique = true, length = 20)
    private String isbn;

    @Column(nullable = false, length = 200)
    private String title;

    @Column(name = "available_copies", nullable = false)
    private int availableCopies;

    @Version
    private Long version;

    protected Material() { }

    protected Material(String isbn, String title, int copies) {
        this.isbn = isbn;
        this.title = title;
        this.availableCopies = copies;
    }

    public void lendOneCopy() {
        if (availableCopies <= 0) throw new NoCopiesAvailableException(isbn);
        availableCopies--;
    }
    public void returnOneCopy() { availableCopies++; }
    // getters...
}

@Entity @DiscriminatorValue("BOOK")
public class Book extends Material {
    @Column(length = 120) private String author;
    protected Book() { }
    public Book(String isbn, String title, int copies, String author) {
        super(isbn, title, copies);
        this.author = author;
    }
}

@Entity @DiscriminatorValue("MAGAZINE")
public class Magazine extends Material {
    private Integer issueNumber;
    protected Magazine() { }
}

@Entity @DiscriminatorValue("DVD")
public class Dvd extends Material {
    @Column(name = "duration_minutes") private Integer durationMinutes;
    protected Dvd() { }
}

Here one of the ORM's real prices is paid, and it has to be said plainly: the sealed hierarchy from 10-06 disappears. A JPA entity cannot be sealed, a record or final. You lose the exhaustive switch expressions checked by the compiler and have to replace them with polymorphism or with instanceof checks. It is a conscious trade-off: you gain real persistence, queries and integrity; you lose part of the type model's expressiveness. An alternative design —and a frequent one in projects that value the domain highly— keeps the domain pure and uses separate JPA entities with a mapping between the two. That discussion belongs to 12-02.

The three inheritance strategies, so you know what you chose:

Strategy How Advantages Drawbacks
SINGLE_TABLE (the default) One table, a discriminator column Fast, no JOINs Subclass columns must be nullable
JOINED One table per class, joined by PK Normalised, no nulls A JOIN on every query
TABLE_PER_CLASS A complete table per subclass No JOIN for specific queries UNION when querying the superclass

The repositories:

public interface MaterialRepository extends JpaRepository<Material, Long> {
    Optional<Material> findByIsbn(String isbn);
    List<Material> findByTitleContainingIgnoreCase(String fragment);
    List<Material> findByAvailableCopiesGreaterThan(int minimum);
}

public interface EmployeeRepository extends JpaRepository<Employee, Long> {
    Optional<Employee> findByEmail(String email);
}

public interface LoanRepository extends JpaRepository<Loan, Long> {

    @EntityGraph(attributePaths = { "material", "employee" })
    List<Loan> findByStatus(LoanStatus status);

    long countByEmployeeAndStatus(Employee employee, LoanStatus status);

    @Query("""
           SELECT l FROM Loan l JOIN FETCH l.material JOIN FETCH l.employee
           WHERE l.dueDate < :date AND l.returnDate IS NULL
           """)
    List<Loan> overdue(@Param("date") LocalDate date);
}

Seed data for development, in src/main/resources/data.sql:

INSERT INTO employees (email, name) VALUES
  ('[email protected]',   'Marta Ruiz'),
  ('[email protected]', 'Diego Alonso'),
  ('[email protected]',  'Nuria Vidal');

INSERT INTO materials (type, isbn, title, available_copies, version, author) VALUES
  ('BOOK', '978-0000000001', 'Effective Java',   3, 0, 'J. Bloch'),
  ('BOOK', '978-0000000002', 'Design Patterns',  2, 0, 'GoF'),
  ('BOOK', '978-0000000003', 'Refactoring',      1, 0, 'M. Fowler');

Running it:

mvn spring-boot:run -Dspring-boot.run.profiles=dev
# H2 web console: http://localhost:8080/h2-console
# JDBC URL: jdbc:h2:mem:bibliotech  |  User: sa  |  No password

The before and after:

Aspect CSV (module 7) JPA over H2
Format Plain text A relational database
Queries Load everything and filter in memory SQL with indexes
Transactions None ACID
Referential integrity None Foreign keys
Concurrency Two processes corrupt it Optimistic locking with @Version
Writing Hand-written AtomicWrite Managed by the engine
Access code CsvReader + CsvWriter + mapping Three interfaces with no implementation
New queries Programme the filtering One more method in the interface
Scalability Thousands of rows Millions

  1. Common Mistakes and Tips

Mistake: javax.persistence instead of jakarta.persistence. The number one copy-and-paste error today. Spring Boot 3 and Hibernate 6 use jakarta.

Mistake: leaving @Enumerated at its default (ORDINAL). Reordering an enum silently corrupts all the historical data. Always EnumType.STRING.

Mistake: leaving @ManyToOne at EAGER. It is the default and it is bad. Set LAZY explicitly on every association and fetch what you need with JOIN FETCH.

Mistake: the N+1 problem without noticing. A loop over entities accessing a lazy relationship. Switch show-sql on in development and count the queries.

Mistake: spring.jpa.open-in-view=true. It is Spring Boot's default and it hides the N+1 until production. Set it to false.

Mistake: modifying only the inverse side of a relationship. employee.getLoans().add(l) without l.setEmployee(employee) persists nothing. Write convenience methods.

Mistake: ignoring merge's return value. merge(e) returns another instance; the original stays detached and its changes are lost.

Mistake: CascadeType.ALL without thinking. Deleting a material could delete its entire loan history.

Mistake: ddl-auto=update in production. It does not delete, does not version, does not migrate data and nobody reviews the SQL it is about to run. validate plus Flyway.

Mistake: concatenating values into a query. That is SQL injection. Parameters always, no exceptions.

Mistake: final entities, record entities or entities with final methods. Hibernate needs to generate proxy subclasses. Either the model will not compile or lazy loading will fail.

Mistake: equals and hashCode based on the generated id. Before persisting, the id is null; afterwards it changes, and an entity put into a HashSet can no longer be found. Use a business key (the isbn) or Hibernate's recommended pattern.

Tip: switch show-sql on while you are learning. Seeing the generated SQL is half the learning, and it teaches you to spot the N+1 instantly.

Tip: mark readOnly = true on query methods. It switches off dirty checking and saves real work.

Tip: use DTOs (record) for whatever leaves the service. Returning managed entities outside the transaction is the direct cause of the LazyInitializationException, and it also couples your internal API to the schema.

Tip: @Version on every entity modified concurrently. It costs one annotation and prevents lost updates.

Tip: do not do everything with the ORM. For reports and aggregations, native SQL or JdbcTemplate are better tools. It is legitimate and professional.

Tip: getReference when all you need is the reference in order to set an FK. It saves a SELECT per operation.

  1. Exercises

Exercise 1: mapping the Reservation entity

BiblioTech has reservations: when there are no copies, an employee reserves one and is notified when one is returned. This is module 10's record:

public record Reservation(Long id, String isbn, String employeeEmail,
                          LocalDate requestDate, LocalDate expiryDate,
                          Optional<Instant> notifiedAt, ReservationStatus status) { }

public enum ReservationStatus { PENDING, NOTIFIED, COMPLETED, EXPIRED }

Turn it into a JPA entity that satisfies:

  1. A reservations table with an auto-generated primary key.
  2. Lazy @ManyToOne relationships to Material and Employee, not strings.
  3. The enum stored safely.
  4. notifiedAt nullable in the database but exposed as Optional.
  5. Optimistic locking.
  6. An index over (material_id, status) for the pending-reservations query.
  7. A domain method markNotified(Clock clock) that changes the status and records the instant, with validation.
  8. The Spring Data repository with: a material's pending reservations ordered by age, reservations expired before a date, and counting an employee's pending ones.

Exercise 2: diagnosing and fixing an N+1

This service works fine with development's 5 loans and takes 9 seconds in production with 800.

@Service
public class ReportService {

    private final LoanRepository repository;

    @Transactional(readOnly = true)
    public List<String> overdueReport(LocalDate date) {
        List<Loan> overdue = repository.findByDueDateBefore(date);

        return overdue.stream()
                .map(l -> String.format("%s | %s | %s | %d days",
                        l.getMaterial().getTitle(),
                        l.getMaterial().getCategories().stream()
                                .map(Category::getName)
                                .collect(Collectors.joining(", ")),
                        l.getEmployee().getName(),
                        ChronoUnit.DAYS.between(l.getDueDate(), date)))
                .toList();
    }
}

You are asked to:

  1. Work out how many SQL queries run with 800 loans, broken down.
  2. Explain exactly why.
  3. Propose three different solutions with their code, saying how many queries each leaves.
  4. Say which one you would choose and why.

Exercise 3: the disputed copy

Marta Ruiz and Diego Alonso try to lend the last copy of "Refactoring" (978-0000000003, 1 copy) at the same time.

@Service
public class LoanManager {

    @Transactional
    public Loan lend(String isbn, String email) {
        Material material = materialRepository.findByIsbn(isbn).orElseThrow();
        Employee employee = employeeRepository.findByEmail(email).orElseThrow();

        if (material.getAvailableCopies() <= 0) {
            throw new NoCopiesAvailableException(isbn);
        }
        material.setAvailableCopies(material.getAvailableCopies() - 1);

        Loan loan = new Loan(material, employee, LocalDate.now(clock), 15);
        return loanRepository.save(loan);
    }
}

You are asked to:

  1. Describe the exact sequence of events that produces two loans from a single copy.
  2. Explain why synchronized (08-04) does not solve this in a real deployment.
  3. Solve it with optimistic locking, including handling the exception with a retry.
  4. Solve it with pessimistic locking.
  5. Compare the two in a table and recommend one.

Solutions

Solution 1

package com.nexussoftware.bibliotech.domain;

import jakarta.persistence.*;
import java.time.*;
import java.util.Optional;

@Entity
@Table(name = "reservations",
       indexes = @Index(name = "idx_reservation_material_status",   // (6)
                        columnList = "material_id, status"))
public class Reservation {

    @Id                                                            // (1)
    @GeneratedValue(strategy = GenerationType.IDENTITY)
    private Long id;

    @ManyToOne(fetch = FetchType.LAZY, optional = false)           // (2)
    @JoinColumn(name = "material_id", nullable = false)
    private Material material;

    @ManyToOne(fetch = FetchType.LAZY, optional = false)           // (2)
    @JoinColumn(name = "employee_id", nullable = false)
    private Employee employee;

    @Column(name = "request_date", nullable = false)
    private LocalDate requestDate;

    @Column(name = "expiry_date", nullable = false)
    private LocalDate expiryDate;

    @Column(name = "notified_at")                                  // (4) nullable
    private Instant notifiedAt;

    @Enumerated(EnumType.STRING)                                   // (3) NEVER ORDINAL
    @Column(nullable = false, length = 20)
    private ReservationStatus status;

    @Version                                                       // (5)
    private Long version;

    protected Reservation() { }   // required by JPA

    public Reservation(Material material, Employee employee, LocalDate request, int validDays) {
        this.material = material;
        this.employee = employee;
        this.requestDate = request;
        this.expiryDate = request.plusDays(validDays);
        this.status = ReservationStatus.PENDING;
    }

    // (7) domain behaviour with validation
    public void markNotified(Clock clock) {
        if (status != ReservationStatus.PENDING) {
            throw new InvalidReservationStatusException(
                    "Only a PENDING reservation can be notified; current status: " + status);
        }
        this.status = ReservationStatus.NOTIFIED;
        this.notifiedAt = Instant.now(clock);   // injected Clock (10-05)
    }

    public void complete() {
        if (status != ReservationStatus.NOTIFIED) {
            throw new InvalidReservationStatusException("Only a NOTIFIED reservation completes");
        }
        this.status = ReservationStatus.COMPLETED;
    }

    public boolean hasExpired(LocalDate today) {
        return status == ReservationStatus.PENDING && expiryDate.isBefore(today);
    }

    // (4) the field is nullable; the public API returns Optional
    public Optional<Instant> getNotifiedAt() {
        return Optional.ofNullable(notifiedAt);
    }

    public Long getId() { return id; }
    public Material getMaterial() { return material; }
    public Employee getEmployee() { return employee; }
    public ReservationStatus getStatus() { return status; }
    public LocalDate getRequestDate() { return requestDate; }
    public LocalDate getExpiryDate() { return expiryDate; }
}

The repository (8):

public interface ReservationRepository extends JpaRepository<Reservation, Long> {

    // A material's pending ones, oldest first (the FIFO reservation queue)
    @EntityGraph(attributePaths = { "employee" })   // avoids N+1 when reading the name
    List<Reservation> findByMaterialIsbnAndStatusOrderByRequestDateAsc(
            String isbn, ReservationStatus status);

    // Expired ones
    List<Reservation> findByStatusAndExpiryDateBefore(ReservationStatus status, LocalDate date);

    // Counting an employee's pending ones
    long countByEmployeeEmailAndStatus(String email, ReservationStatus status);
}

Notable decisions:

  • record → class: mandatory, because a record is final and has no empty constructor.
  • Relationships instead of strings: Material and Employee instead of isbn and employeeEmail. It gives referential integrity (you cannot reserve a non-existent material) and navigation. The price is having to manage laziness.
  • Optional outside, a nullable field inside: 10-04's guarantee is preserved without fighting JPA.
  • The behaviour lives in the entity: markNotified, complete and hasExpired validate the transitions. The entity is not a bag of data.
  • @EntityGraph on the first query: we know the employee's name is going to be read, so it is fetched in one go.

Solution 2

1. How many queries.

Item Queries
findByDueDateBefore 1
l.getMaterial() — a lazy proxy, 800 times 800
l.getMaterial().getCategories() — a lazy collection, 800 times 800
l.getEmployee() — a lazy proxy, 800 times 800
Total 2,401

At 3-4 ms of latency per query, between 7 and 10 seconds. That matches the symptom.

2. Why. All three associations are lazy. The initial query fetches only the loans rows, with proxies in material and employee and a lazy collection in categories. Each access inside the map fires its own SELECT. In development, with 5 loans, that is 16 queries and nobody notices.

Solution A — JOIN FETCH (1 query if limited to scalars, 2 with the collection):

public interface LoanRepository extends JpaRepository<Loan, Long> {

    @Query("""
           SELECT DISTINCT l FROM Loan l
           JOIN FETCH l.material m
           LEFT JOIN FETCH m.categories
           JOIN FETCH l.employee
           WHERE l.dueDate < :date
           """)
    List<Loan> overdueComplete(@Param("date") LocalDate date);
}

DISTINCT is necessary because the JOIN FETCH on the categories collection duplicates loan rows. 1 query. The limitation: you cannot paginate efficiently with a collection JOIN FETCH.

Solution B — @EntityGraph (declarative, same result):

@EntityGraph(attributePaths = { "material", "material.categories", "employee" })
List<Loan> findByDueDateBefore(LocalDate date);

The advantage: the service code is untouched and the same method can carry a different graph in another query. 1-2 queries.

Solution C — DTO projection (the most efficient):

public record ReportLine(String materialTitle, String employeeName,
                         LocalDate dueDate) {

    public String format(LocalDate reference) {
        return "%s | %s | %d days".formatted(materialTitle, employeeName,
                ChronoUnit.DAYS.between(dueDate, reference));
    }
}
@Query("""
       SELECT new com.nexussoftware.bibliotech.domain.ReportLine(
              m.title, e.name, l.dueDate)
       FROM Loan l JOIN l.material m JOIN l.employee e
       WHERE l.dueDate < :date
       ORDER BY l.dueDate
       """)
List<ReportLine> overdueLines(@Param("date") LocalDate date);

1 query, and only 3 columns instead of all the columns of three tables. No managed entities, no dirty checking, no risk of a later N+1. For the categories you would need an additional query or STRING_AGG in native SQL.

Comparison:

Solution Queries Data transferred Flexibility Risk of relapse
A: JOIN FETCH 1-2 Every column Medium Medium
B: @EntityGraph 1-2 Every column High Medium
C: DTO 1 Only what is needed Low None

4. Which to choose. For a read-only report, C. Reasons: it fetches only what is needed, there are no managed entities and no dirty checking, and —most importantly— it is structurally impossible for the N+1 to reappear, because there are no proxies that could fire. With A and B, somebody who adds an l.getMaterial().getPublisher().getCountry() to the format tomorrow has the problem back.

If the method had to return entities in order to modify them afterwards, B, because it is declarative and reusable.

Solution 3

1. The sequence.

Moment Marta's transaction Diego's transaction DB
t1 SELECT material → available = 1 1
t2 SELECT material → available = 1 1
t3 check 1 > 0 → OK 1
t4 check 1 > 0 → OK 1
t5 UPDATE available = 0 0
t6 INSERT loan
t7 COMMIT 0
t8 UPDATE available = 0 0
t9 INSERT loan
t10 COMMIT 0, with 2 loans

It is a lost update: Diego's UPDATE was computed from a stale value. With READ_COMMITTED, neither transaction sees anything anomalous.

2. Why synchronized is no use. synchronized synchronises threads within a single JVM. In a real deployment there are several BiblioTech instances behind a load balancer (12-06), and Marta and Diego may be served by different processes, on different machines. A monitor in JVM A locks nothing in JVM B. Besides, even with a single instance, serialising every loan of every material through one lock would be an unnecessary bottleneck. The coordination has to be where the shared data is: in the database.

3. Optimistic locking.

@Entity
public class Material {
    @Version
    private Long version;      // the only addition to the model

    public void lendOneCopy() {                 // logic INSIDE the entity
        if (availableCopies <= 0) throw new NoCopiesAvailableException(isbn);
        availableCopies--;
    }
}
@Service
public class LoanManager {

    private static final Logger log = LoggerFactory.getLogger(LoanManager.class);
    private final LoanManager self;   // so the retry goes through the proxy (11-02)

    @Transactional
    public Loan lend(String isbn, String email) {
        Material material = materialRepository.findByIsbn(isbn)
                .orElseThrow(() -> new MaterialNotFoundException(isbn));
        Employee employee = employeeRepository.findByEmail(email)
                .orElseThrow(() -> new EmployeeNotFoundException(email));

        material.lendOneCopy();           // validates and decrements

        Loan loan = new Loan(material, employee, LocalDate.now(clock), 15);
        return loanRepository.save(loan);
    }   // COMMIT: update materials ... where id=? and version=?

    // Retry: every attempt is a NEW transaction
    public Result<Loan> lendWithRetry(String isbn, String email) {
        for (int attempt = 1; attempt <= 3; attempt++) {
            try {
                return Result.success(self.lend(isbn, email));       // goes through the proxy
            } catch (OptimisticLockingFailureException e) {
                log.warn("Conflict on {} (attempt {}/3)", isbn, attempt);
            } catch (NoCopiesAvailableException e) {
                return Result.error("There are no copies left of " + isbn);
            }
        }
        return Result.error("The material is being modified; please try again.");
    }
}

What happens now at t8: Diego's UPDATE carries where id=1 and version=0, but the row already has version=1. It affects 0 rows, Hibernate throws OptimisticLockException and Spring translates it into OptimisticLockingFailureException. The retry reads again, sees available=0 and returns a correct business error.

Result<T> is 10-01's generic type, and the retry loop is your RetryProxy from 10-03 — which in a real project would be @Retryable(retryFor = OptimisticLockingFailureException.class, maxAttempts = 3).

4. Pessimistic locking.

public interface MaterialRepository extends JpaRepository<Material, Long> {

    @Lock(LockModeType.PESSIMISTIC_WRITE)
    @Query("SELECT m FROM Material m WHERE m.isbn = :isbn")
    Optional<Material> findForLending(@Param("isbn") String isbn);
}
@Transactional
public Loan lend(String isbn, String email) {
    // SELECT ... FOR UPDATE: it locks the row until the COMMIT
    Material material = materialRepository.findForLending(isbn)
            .orElseThrow(() -> new MaterialNotFoundException(isbn));
    // Diego waits here until Marta commits; then he reads available = 0
    material.lendOneCopy();   // correctly throws NoCopiesAvailableException
    return loanRepository.save(new Loan(material, employee, LocalDate.now(clock), 15));
}

It is worth adding a maximum wait time so as not to block indefinitely:

@Lock(LockModeType.PESSIMISTIC_WRITE)
@QueryHints(@QueryHint(name = "jakarta.persistence.lock.timeout", value = "3000"))

5. Comparison and recommendation.

Criterion Optimistic (@Version) Pessimistic (FOR UPDATE)
Cost without conflict None A lock on every operation
Concurrency allowed High Low: operations serialise
What happens on conflict An exception and a retry A wait
Deadlock risk None Yes, if several rows are locked
Risk of waiting indefinitely No Yes, without a timeout
Code complexity Handling the exception None apparent
Works across several instances Yes Yes
Good for Infrequent conflicts Frequent and expensive conflicts

Recommendation: optimistic locking. In BiblioTech, the chance of two people lending the same copy in the same second is tiny, and pessimistic locking would make the thousands of operations that never clash pay the cost of serialisation. @Version costs one annotation, does not penalise the normal case and works the same with one instance as with twenty.

Pessimistic locking is reserved for operations where conflict is the norm and the lost work is expensive: for example, a nightly process that recalculates every fine and does not want to retry five thousand times.

Conclusion

BiblioTech's CSVs are gone. There is a real database underneath, and you know exactly what is going on inside it.

You understand the object-relational mismatch an ORM exists to manage: double identity, inheritance that does not exist in SQL, navigation that is free in Java and costs queries in SQL, granularity, collections and types. And you know what you gain and what you lose compared with plain JDBC —which you saw in twenty lines, with its column-by-column mapping, its type conversion and its resource management— with the professional rule that follows: ORM for the domain CRUD, direct SQL for reports and bulk processing, and both coexisting without conflict.

You can tell JPA from Hibernate: specification (jakarta.persistence) versus implementation, with the recommendation to program against the specification and the version warning that today causes more compilation errors than any other — jakarta, never javax.

You know how to map entities: @Entity, @Table, @Id, the @GeneratedValue strategies with the real reason SEQUENCE beats IDENTITY (it allows insert batching), @Column with its attributes, and the requirements JPA imposes —a no-argument constructor, nothing final— that explain why a record cannot be an entity. You know the absolute rule of @Enumerated(EnumType.STRING), because ORDINAL is the default and reordering an enum silently corrupts all the historical data. You map LocalDate and Instant natively with the same distinction from 10-05, you write converters with AttributeConverter for value objects such as Isbn, and you use @Embeddable to have small objects in Java without fragmenting the schema.

You model complete relationships with their four types, you know that the owning side is the one holding the foreign key and that touching only the inverse side persists nothing —hence the convenience methods—, and you have mastered the decision with the greatest performance impact: FetchType. You know the defaults are inconsistent, that @ManyToOne comes as EAGER and that the professional rule is to set LAZY explicitly everywhere and ask for what you need in the query. You know the two consequences: the LazyInitializationException, with its four ranked solutions and the advice to leave open-in-view at false so the problem shows up on your machine and not in production; and the N+1 problem, which turns one query into 501 and is invisible with five test rows, with its solutions —JOIN FETCH, @EntityGraph, @BatchSize and, the most efficient for reports, the DTO projection, where record types finally find their place—. And you know that switching to EAGER fixes nothing: it makes it worse.

You understand the persistence context, which is JPA's conceptual heart: a first-level cache that guarantees the same row is the same object, with an entity's four states and its transitions. And with it, automatic dirty checking, which surprises everybody: you modify a managed entity, you call nothing, and on commit an UPDATE appears. You know how it works (snapshot and comparison), its four consequences —including that an accidental change is persisted— and why readOnly = true switches it off and saves real work. You know flush, which reorders the operations and explains half of Hibernate's odd behaviour, and the difference between persist and merge that causes the "I ignored the return value" bug.

You write queries: JPQL over entities and attributes, named queries validated at start-up, the Criteria API for dynamic cases, native SQL when it is the right call, and projections into record types in a single line. And you have internalised the rule that admits no exceptions: every external value goes in as a parameter, because concatenating is SQL injection —with an allow-list as the only way out when the dynamic part is a column name—.

You handle transactions with a real database: ACID, @Transactional with its two rules inherited from 11-02, readOnly, propagation with the legitimate case for REQUIRES_NEW and its cost in connections, and isolation with the three phenomena explained over BiblioTech and the advice hardly ever to touch it. And you solve the disputed-copy case with optimistic locking using @Version, understanding why module 8's synchronized is no use when there are several instances: the coordination has to be where the shared data is. With a retry over Result<T>, which is your generic from 10-01 and your RetryProxy from 10-03 turned into production code.

You know the first- and second-level caches and when the second is worth it —data read a lot and changed little, measured first—; and you know why ddl-auto=update is never used in production: it does not delete, does not version, does not migrate data and nobody reviews the ALTER TABLE it is about to run. validate plus versioned migrations with Flyway, which arrive in 12-06.

And you have watched Spring Data JPA turn eighty lines of repository into an empty interface, with methods derived from names, @Query for what does not fit in a name, and pagination. With the mechanism laid bare: Proxy.newProxyInstance over an interface, exactly the one from 10-03, analysing the method's metadata and generating the query.


BiblioTech now has real entities, relationships with referential integrity, ACID transactions, optimistic locking and indexed queries over H2. It starts with mvn spring-boot:run and its schema is created on its own.

And it still has not a single automated test.

That is now, by a long way, the most serious debt. You have just carried out an enormous migration: you have changed the domain model, turned record types into classes, replaced the entire persistence layer and introduced optimistic concurrency with retries. How do you know the fine calculation still gives the same result? That the limit of three loans per employee is still respected? That a loan falling due on a Sunday still triggers a notice on the Monday? The only honest answer today is: by starting the application and looking.

And there is something worse. That injectable Clock you introduced in 10-05, which went through Spring in 11-02 as a @Bean and which in this lesson has appeared in every LocalDate.now(clock) and every Instant.now(clock), exists exclusively so the code can be tested without waiting fifteen days. It has spent two modules waiting for somebody to make use of it.

In the next lesson it gets used. You will see what an automated test is and why you are already paying the cost of not having them; you will meet the testing pyramid and the whole of JUnit 5, from @Test and the Arrange-Act-Assert pattern to parameterised tests that verify twelve cases of the fine calculation in six lines; you will compare JUnit's assertions with AssertJ's; and you will discover what makes a design testable — and why everything you have done in the last two modules (constructor injection, interfaces at the boundaries, the injected Clock, the annotation-free domain) was pointing exactly there.

BiblioTech is about to get its first safety net.

Java Programming Course

Module 1: Introduction to Java

Module 2: Control Flow

Module 3: Object-Oriented Programming

Module 4: Advanced Object-Oriented Programming

Module 5: Data Structures and Collections

Module 6: Exception Handling

Module 7: File Input/Output

Module 8: Multithreading and Concurrency

Module 9: Networking

Module 10: Advanced Topics

Module 11: Java Frameworks and Libraries

Module 12: Building Real-World Applications

© Copyright 2026. All rights reserved