The previous lesson ended on an irony: BiblioTech has a @Transactional on lend that does absolutely nothing, because there is no database underneath that could commit or roll back. All the persistence is still what it was in module 7: hand-written CSV files.
This lesson replaces them with a real relational database.
Hibernate is the most widely used ORM in the Java world, and the reference implementation of Jakarta Persistence (formerly JPA). Its job is to translate between two worlds that do not fit together: the world of objects, with inheritance, references and navigation; and the world of tables, with rows, columns and foreign keys. That job is tedious, repetitive and easy to get wrong, and that is why a tool exists to automate it.
But an ORM is not a transparent layer you can use without understanding it. It is probably the piece most people abuse without knowing what is going on: queries that multiply by a hundred, exceptions that appear outside the transaction, entities that save themselves without anybody calling save, and production schemas destroyed by one misplaced property. All of that has an explanation, and this lesson gives it.
The advantage you arrive with is the usual one: Hibernate reads @Entity and @Column by reflection, exactly as your AnnotatedExporter read @CsvField. And Spring Data JPA will give you working repositories without writing the implementation, using the same Proxy.newProxyInstance from 10-03.
By the end you will know what an ORM automates and what it costs; you will map complete entities and relationships; you will understand the persistence context and an entity's states; you will know why the LazyInitializationException appears and what the N+1 problem is, with its solutions; you will write JPQL queries with parameters —and know why concatenating strings is a vulnerability—; you will handle transactions, propagation and optimistic locking; and BiblioTech will store its data in H2 with Material, Employee and Loan as real entities.
Contents
- The object-relational mismatch
- What you gain and what you lose compared with JDBC
- JDBC in twenty lines: the baseline
- JPA versus Hibernate: specification and implementation
- Getting JPA running with H2 in BiblioTech
- The first entity:
@Entity,@Table,@Id @GeneratedValueand its strategies@Columnand type mapping@Enumerated: whySTRINGand neverORDINALLocalDate,Instantand your own converters@Embeddableand@Embedded- Relationships: the four types
- The owning side and
mappedBy FetchType: lazy versus eager loading- The
LazyInitializationException - The N+1 problem
- Solving it:
JOIN FETCHand@EntityGraph CascadeTypeandorphanRemoval- The persistence context and the
EntityManager - The four states of an entity
- Automatic dirty checking
flushand the ordering of operations- The operations:
persist,merge,find,getReference,remove - Queries: JPQL
- Named queries, the Criteria API and native SQL
- SQL injection: why you never concatenate
- Transactions:
@Transactionalwith a real database - Propagation and isolation
- Optimistic locking with
@Version - First- and second-level caches
ddl-autoand why neverupdatein production- Spring Data JPA
- BiblioTech: from CSV to H2
- Common Mistakes and Tips
- Exercises
- The object-relational mismatch
The problem has a name of its own: the object-relational impedance mismatch. The object model and the relational model were invented for different things and they do not fit together naturally.
The concrete points of friction:
| Concept | In objects | In tables |
|---|---|---|
| Identity | == (reference) and equals (value): two concepts |
Primary key: just one |
| Inheritance | Material sealed with Book, Magazine, Dvd |
Does not exist |
| Association | A directed reference (loan.material()) |
A foreign key, bidirectional by nature |
| Navigation | Chained: loan.material().publisher().country() |
With a JOIN, all at once |
| Granularity | Small objects: Address, Money |
Columns in the same row |
| Collections | List, Set, Map |
Related rows |
| Types | LocalDate, Optional, enums, record |
DATE, VARCHAR, NUMERIC |
null |
Absence of a reference | NULL with three-valued semantics |
Look at the inheritance row, because it is the most brutal: the relational model has no inheritance. BiblioTech's sealed Material hierarchy, which in 10-06 allowed exhaustive switch expressions checked by the compiler, simply cannot be expressed in SQL. Somebody has to decide how it is represented, and there are three different strategies, each with its trade-offs.
And the navigation row is the one that causes the most performance problems: in Java, loan.material().publisher() is free; in SQL, every hop can be a query. That is where section 16's N+1 problem comes from.
An ORM (Object-Relational Mapper) is a tool that translates automatically between the two worlds. It does not remove the mismatch: it manages it. And knowing it exists is the difference between using Hibernate well and suffering it.
- What you gain and what you lose compared with JDBC
Honestly, because there are strong opinions in both directions and both are partly right.
| Aspect | Plain JDBC | ORM (Hibernate/JPA) |
|---|---|---|
| Boilerplate | Masses of it: ResultSet to object, by hand, column by column |
Almost none |
| Control over the SQL | Total | Partial; recoverable with native SQL |
| Learning curve | Low (if you know SQL) | High |
| Straightforward performance | Predictable | Good, with caches that help |
| Performance when misused | Bad, and your fault | Catastrophic (N+1, massive eager loading) |
| Portability between engines | Low: the SQL varies | High: dialects |
| Complex queries / reports | Natural | Awkward; you end up using native SQL |
| Writing an object graph | Very laborious | Trivial: cascades |
| Transparency | You see every statement | It generates SQL you do not see unless you ask |
| Caching | You build it | First and second level included |
The reasonable professional rule:
- ORM for the domain CRUD, which is 80% of the code of a business application: saving a loan, loading an employee, updating a material.
- Direct SQL (JDBC,
JdbcTemplateor jOOQ) for reports, complex aggregations and bulk processing.
The two coexist perfectly well in the same project, and that is what BiblioTech does.
- JDBC in twenty lines: the baseline
To understand what Hibernate automates, you first have to see what had to be written. This is pure JDBC reading loans:
package com.nexussoftware.bibliotech.persistence;
import java.sql.*;
import java.time.LocalDate;
import java.util.*;
public class JdbcLoanRepository {
private final DataSource dataSource;
public JdbcLoanRepository(DataSource dataSource) {
this.dataSource = dataSource;
}
public List<Loan> findByEmployee(String employeeEmail) {
String sql = """
SELECT id, isbn, employee_email, loan_date,
due_date, return_date
FROM loans
WHERE employee_email = ?
""";
List<Loan> result = new ArrayList<>();
// try-with-resources from module 6: closes the three resources in reverse order
try (Connection connection = dataSource.getConnection();
PreparedStatement statement = connection.prepareStatement(sql)) {
statement.setString(1, employeeEmail); // a parameter, NEVER concatenation
try (ResultSet rows = statement.executeQuery()) {
while (rows.next()) {
// MANUAL column-by-column mapping: this is what the ORM automates
Date returned = rows.getDate("return_date");
result.add(new Loan(
rows.getLong("id"),
rows.getString("isbn"),
rows.getString("employee_email"),
rows.getDate("loan_date").toLocalDate(),
rows.getDate("due_date").toLocalDate(),
returned == null
? Optional.empty()
: Optional.of(returned.toLocalDate())));
}
}
} catch (SQLException e) {
throw new PersistenceException("Error finding loans for " + employeeEmail, e);
}
return result;
}
}What is in there, and what Hibernate is going to eliminate:
- Hand-written SQL for every query.
- Column-by-column mapping, with
getLong,getString,getDate. - Type conversion:
java.sql.DatetoLocalDate,nulltoOptional.empty(). - Resource management: three nested
try-with-resources. - Exception translation:
SQLExceptioninto module 6's hierarchy. - Column names as strings: a misspelled
"due_date"fails at runtime.
Multiply that by the forty queries of a real application and by six entities, and you get several thousand lines that add no value and have to be maintained every time a column changes.
JDBC is not bad. It is low-level. With Hibernate, that whole query becomes:
And you do not even have to write that line, as you will see in section 32.
- JPA versus Hibernate: specification and implementation
Another classic confusion worth clearing up.
Jakarta Persistence (JPA) is a specification: a set of standardised interfaces and annotations, in the jakarta.persistence package. It defines EntityManager, @Entity, @Id, @OneToMany, JPQL... but it contains no working code.
Hibernate is an implementation of that specification. It is what actually generates the SQL, manages the persistence context and talks to the database.
graph TD
A["Your BiblioTech code"] --> B["JPA API<br/>jakarta.persistence<br/>@Entity, EntityManager, JPQL"]
B --> C["Hibernate 6.x<br/>(the implementation)"]
B -.->|alternatives| D["EclipseLink"]
B -.-> E["OpenJPA"]
C --> F["JDBC"]
F --> G[("H2 / PostgreSQL /<br/>MySQL / Oracle")]
| Aspect | JPA (jakarta.persistence) |
Hibernate (org.hibernate) |
|---|---|---|
| What it is | A specification, interfaces | An implementation with code |
| Package | jakarta.persistence.* |
org.hibernate.* |
| Examples | @Entity, EntityManager, JPQL |
Session, @BatchSize, HQL, filters |
| Portable to another implementation | Yes | No |
The practical recommendation: program against JPA whenever you can and use Hibernate extensions only when you genuinely need them. It is the same facade principle you will see with SLF4J in 11-07.
Version warning, important. Jakarta EE renamed every
javax.*package tojakarta.*. Spring Boot 3 and Hibernate 6 usejakarta.persistence. If you copy an example withimport javax.persistence.Entity, it is from Spring Boot 2 or earlier and it will not compile. It is the most frequent copy-and-paste error today.
- Getting JPA running with H2 in BiblioTech
Two dependencies added to 11-02's pom.xml:
<!-- Hibernate + JPA + Spring Data + transaction management -->
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-data-jpa</artifactId>
</dependency>
<!-- H2: an in-memory database. NOTHING to install -->
<dependency>
<groupId>com.h2database</groupId>
<artifactId>h2</artifactId>
<scope>runtime</scope>
</dependency>And the configuration in application.yml:
spring:
datasource:
url: jdbc:h2:mem:bibliotech;DB_CLOSE_DELAY=-1
username: sa
password:
driver-class-name: org.h2.Driver
h2:
console:
enabled: true # http://localhost:8080/h2-console (dev only)
path: /h2-console
jpa:
hibernate:
ddl-auto: create-drop # ONLY in development. See section 31
show-sql: true # shows the generated SQL
properties:
hibernate:
format_sql: true # formats it readably
jdbc:
batch_size: 25
open-in-view: false # IMPORTANT. See section 15That is all it takes. Run:
And at start-up you will see the conditions from section 19 of 11-02 being met: DataSourceAutoConfiguration detects H2 on the classpath, sees that you have not defined a DataSource, and creates one. HibernateJpaAutoConfiguration detects Hibernate and assembles the EntityManagerFactory and the transaction manager. Zero configuration code.
Two notes about show-sql: it is indispensable while you are learning —you are going to see exactly what SQL Hibernate generates, and that is half the learning— and it must be switched off in production, where proper logging goes through SLF4J (11-07):
logging:
level:
org.hibernate.SQL: DEBUG # the statements
org.hibernate.orm.jdbc.bind: TRACE # the parameters
- The first entity:
@Entity, @Table, @Id
@Entity, @Table, @IdWe turn BiblioTech's Book into an entity:
package com.nexussoftware.bibliotech.domain;
import jakarta.persistence.*;
import java.time.LocalDate;
@Entity // "this class maps to a table"
@Table(name = "materials", // explicit table name
indexes = @Index(name = "idx_material_isbn", columnList = "isbn", unique = true))
public class Book {
@Id // primary key
@GeneratedValue(strategy = GenerationType.IDENTITY)
private Long id;
@Column(name = "isbn", nullable = false, unique = true, length = 20)
private String isbn;
@Column(nullable = false, length = 200)
private String title;
@Column(name = "publication_year")
private Integer publicationYear;
@Column(name = "available_copies", nullable = false)
private int availableCopies;
// JPA REQUIRES a no-argument constructor (it may be protected)
protected Book() { }
public Book(String isbn, String title, int availableCopies) {
this.isbn = isbn;
this.title = title;
this.availableCopies = availableCopies;
}
// getters and (only the necessary) setters
public Long getId() { return id; }
public String getIsbn() { return isbn; }
public String getTitle() { return title; }
public int getAvailableCopies() { return availableCopies; }
// Domain behaviour: the entity is not a bag of data
public void lendOneCopy() {
if (availableCopies <= 0) {
throw new NoCopiesAvailableException(isbn);
}
availableCopies--;
}
public void returnOneCopy() { availableCopies++; }
}The requirements JPA imposes on an entity, worth knowing because the errors are cryptic:
| Requirement | Reason |
|---|---|
Annotated with @Entity |
It is what makes it mappable |
An @Id |
Without identity there is no row |
| A no-argument constructor | Hibernate instantiates it by reflection (10-03) |
A non-final class |
It needs to generate proxy subclasses for lazy loading |
Non-final fields |
It fills them in by reflection |
Methods cannot be final |
Lazy-loading interception |
And there you have why a record cannot be a JPA entity: it is final, its fields are final and it has no no-argument constructor. BiblioTech's record types (Card, SessionSummary) will still be perfect as DTOs and as query projections, but entities have to be mutable classes. It is a real ORM trade-off, not a whim.
About @Table: if you leave it out, the table is named after the class. Putting it in explicitly is good practice, because Loan would map to a table LOAN, and words such as ORDER, USER or GROUP are reserved in SQL and give bewildering errors.
@GeneratedValue and its strategies
@GeneratedValue and its strategiesWho generates the primary key. It is a decision with a real performance impact.
| Strategy | How it works | Advantages | Drawbacks |
|---|---|---|---|
IDENTITY |
The engine's auto-increment column | Simple; works everywhere | Prevents insert batching: it forces an immediate INSERT to learn the id |
SEQUENCE |
A database sequence | The best: it can reserve ids in batches and group inserts | Not every engine has sequences (older MySQL) |
TABLE |
An auxiliary table of counters | Portable everywhere | Slow; contention. Avoid it |
AUTO |
Hibernate chooses | Convenient | Unpredictable across engines |
| (none) | You assign it | Full control; natural keys | You have to guarantee uniqueness |
The recommendation with Hibernate 6 and a database that supports sequences:
@Id
@GeneratedValue(strategy = GenerationType.SEQUENCE, generator = "material_seq")
@SequenceGenerator(name = "material_seq", sequenceName = "material_seq",
allocationSize = 50) // reserves 50 ids at once: 1 query per 50 inserts
private Long id;The reason allocationSize matters: with IDENTITY, Hibernate cannot batch inserts, because it needs to execute each INSERT so the engine returns the id. With SEQUENCE and allocationSize=50, it can accumulate 50 inserts and send them as one batch. On an import of 10,000 materials, the difference is minutes.
For this lesson's H2 examples we will use IDENTITY for simplicity, but know the reason SEQUENCE is preferred in production.
@Column and type mapping
@Column and type mapping@Column describes the column. Its useful attributes:
| Attribute | What it does | Example |
|---|---|---|
name |
The column name | @Column(name = "loan_date") |
nullable |
NOT NULL in the generated schema |
nullable = false |
unique |
A uniqueness constraint | unique = true |
length |
VARCHAR length (255 by default) |
length = 20 |
precision / scale |
For BigDecimal |
precision = 10, scale = 2 |
insertable / updatable |
Exclude from INSERT or UPDATE |
updatable = false |
columnDefinition |
Raw SQL. It breaks portability | A last resort |
The types JPA maps automatically:
| Java type | Typical SQL column |
|---|---|
String |
VARCHAR |
int, Integer, long, Long |
INTEGER, BIGINT |
boolean, Boolean |
BOOLEAN |
BigDecimal |
NUMERIC(p,s) |
LocalDate |
DATE |
LocalDateTime |
TIMESTAMP |
Instant |
TIMESTAMP (UTC) |
byte[] |
BLOB / VARBINARY |
enum |
Depends on @Enumerated (section 9) |
| Anything else | Needs an AttributeConverter (section 10) |
Two concrete warnings about BiblioTech:
One: never double for money. It was said back in 01-04 and here it matters more than ever. Fines are BigDecimal with explicit precision and scale:
Two: Optional is not mapped. Optional<LocalDate> returnDate was an excellent decision in 10-04 for the public API, but JPA cannot persist Optional. The solution is a nullable field and a getter that returns Optional:
@Column(name = "return_date")
private LocalDate returnDate; // the field: it may be null
public Optional<LocalDate> getReturnDate() { // the API: Optional
return Optional.ofNullable(returnDate);
}That way you keep 10-04's guarantee —callers never receive null— and JPA can persist the field. It is the standard pattern.
@Enumerated: why STRING and never ORDINAL
@Enumerated: why STRING and never ORDINALBiblioTech has had three enums since module 4: Severity, MaterialType and LoanStatus. Mapping them has a trap with serious consequences.
// BAD: by default JPA uses ORDINAL (the index: 0, 1, 2)
@Enumerated(EnumType.ORDINAL)
private LoanStatus status;
// GOOD: always STRING
@Enumerated(EnumType.STRING)
@Column(length = 20, nullable = false)
private LoanStatus status;Why ORDINAL is dangerous. With it, the database stores 0, 1, 2. Now imagine that six months from now somebody adds a status in its logical place:
Every loan stored with 1 (OVERDUE) now reads as RENEWED, and the 2s (RETURNED) as OVERDUE. All the historical data is silently corrupted. No exception, no warning. It is discovered weeks later, when somebody is chased for a fine on a book they returned.
On top of that, ORDINAL makes the database unreadable: SELECT * FROM loans returns numbers that mean nothing.
| Aspect | ORDINAL |
STRING |
|---|---|---|
| What it stores | The index (0, 1, 2) | The name ("ACTIVE") |
| Space | Less | A little more |
| Reordering or inserting constants | Corrupts the data | No problem |
| Renaming a constant | No problem | A migration is needed (and the failure is visible) |
| Readable in SQL | No | Yes |
The rule is absolute: @Enumerated(EnumType.STRING), always. ORDINAL's space saving does not remotely compensate for the risk. And the most dangerous part is that ORDINAL is the default: if you forget the annotation, you get the bad behaviour.
LocalDate, Instant and your own converters
LocalDate, Instant and your own convertersGood news: since JPA 2.2, java.time maps natively. All the work from 10-05 carries straight over:
@Column(name = "loan_date", nullable = false)
private LocalDate loanDate; // -> DATE
@Column(name = "due_date", nullable = false)
private LocalDate dueDate; // -> DATE
@Column(name = "recorded_at", nullable = false)
private Instant recordedAt; // -> TIMESTAMP in UTCThe distinction from 10-05 is still the right one: LocalDate for business dates (a loan falls due "on the 20th", with no time and no zone) and Instant for audit timestamps (an absolute instant on the timeline).
And java.util.Date and Calendar are never used again. If you see @Temporal(TemporalType.DATE) in an example, it is pre-Java 8 code.
Converters for your own types
And what if you want to persist a type JPA does not know? That is where AttributeConverter comes in. BiblioTech has a record Isbn as a value object:
package com.nexussoftware.bibliotech.domain;
public record Isbn(String value) {
public Isbn {
if (value == null || !value.matches("\\d{3}-\\d{10}")) {
throw new InvalidIsbnException(value);
}
}
}The converter:
package com.nexussoftware.bibliotech.persistence;
import jakarta.persistence.AttributeConverter;
import jakarta.persistence.Converter;
import com.nexussoftware.bibliotech.domain.Isbn;
@Converter(autoApply = true) // applies to ALL fields of type Isbn
public class IsbnConverter implements AttributeConverter<Isbn, String> {
@Override
public String convertToDatabaseColumn(Isbn isbn) {
return isbn == null ? null : isbn.value();
}
@Override
public Isbn convertToEntityAttribute(String column) {
return column == null ? null : new Isbn(column);
}
}With autoApply = true, any Isbn field is converted on its own. Without it, you have to mark it:
@Convert(converter = IsbnConverter.class)
@Column(name = "isbn", nullable = false, unique = true)
private Isbn isbn;This lets you have value objects with validation in the domain and simple columns in the database: the best of both worlds. Other common uses: encrypting a sensitive field, storing a short list as a comma-separated string, or mapping a boolean to 'Y'/'N' in a legacy database.
@Embeddable and @Embedded
@Embeddable and @EmbeddedAn embeddable is an object that has no identity of its own and whose fields are stored in the same row as the entity containing it. It gives structure to the object model without creating tables.
package com.nexussoftware.bibliotech.domain;
import jakarta.persistence.Embeddable;
@Embeddable
public class PublicationData {
private String publisher;
private Integer year;
private String language;
protected PublicationData() { }
public PublicationData(String publisher, Integer year, String language) {
this.publisher = publisher;
this.year = year;
this.language = language;
}
public boolean isRecent(int currentYear) { // behaviour of its own
return year != null && currentYear - year <= 3;
}
}Usage:
@Entity
public class Book {
@Embedded
private PublicationData publication;
// If the same embeddable appears twice, the columns must be renamed:
@Embedded
@AttributeOverrides({
@AttributeOverride(name = "publisher", column = @Column(name = "original_publisher")),
@AttributeOverride(name = "year", column = @Column(name = "original_year"))
})
private PublicationData originalPublication;
}The resulting table has columns publisher, year, language, original_publisher, original_year... in the same row of materials. There is no publication_data table and no JOIN.
| Concept | @Entity |
@Embeddable |
|---|---|---|
| A table of its own | Yes | No: columns in the container's |
Identity (@Id) |
Yes | No |
| Lifecycle | Its own | That of the containing entity |
| Can be queried alone | Yes | No |
| Example | Book, Loan |
PublicationData, Address, Money |
It is the answer to section 1's granularity problem: you can have small, cohesive objects in Java without fragmenting the schema.
- Relationships: the four types
BiblioTech's model:
- A
Loanbelongs to oneMaterialand oneEmployee. - A
Materialhas manyLoans. - An
Employeehas manyLoans and manyReservations. - A
Materialhas manyCategoryvalues and aCategoryhas manyMaterials.
@ManyToOne: the side that holds the foreign key
@Entity
@Table(name = "loans")
public class Loan {
@Id @GeneratedValue(strategy = GenerationType.IDENTITY)
private Long id;
@ManyToOne(fetch = FetchType.LAZY, optional = false) // ALWAYS LAZY. Section 14
@JoinColumn(name = "material_id", nullable = false) // the FK column
private Material material;
@ManyToOne(fetch = FetchType.LAZY, optional = false)
@JoinColumn(name = "employee_id", nullable = false)
private Employee employee;
@Column(name = "loan_date", nullable = false)
private LocalDate loanDate;
@Column(name = "due_date", nullable = false)
private LocalDate dueDate;
@Column(name = "return_date")
private LocalDate returnDate;
@Enumerated(EnumType.STRING)
@Column(nullable = false, length = 20)
private LoanStatus status;
@Version // optimistic locking. Section 29
private Long version;
}The loans table will have material_id and employee_id columns with foreign keys.
@OneToMany: the inverse side
@Entity
@Table(name = "employees")
public class Employee {
@Id @GeneratedValue(strategy = GenerationType.IDENTITY)
private Long id;
@Column(nullable = false, unique = true)
private String email;
@Column(nullable = false)
private String name;
// mappedBy = "employee" -> the Loan field that owns the relationship
@OneToMany(mappedBy = "employee",
cascade = CascadeType.ALL,
orphanRemoval = true,
fetch = FetchType.LAZY) // LAZY is the default here
private List<Loan> loans = new ArrayList<>();
// Convenience method: it keeps BOTH sides in sync
public void addLoan(Loan loan) {
loans.add(loan);
loan.assignEmployee(this);
}
public void removeLoan(Loan loan) {
loans.remove(loan);
loan.assignEmployee(null);
}
}Those convenience methods are not optional in practice: if you add to the list without setting the field on the other side, the relationship is not persisted (section 13) and the in-memory state is left inconsistent too.
@OneToOne
@Entity
public class Employee {
@OneToOne(mappedBy = "employee", cascade = CascadeType.ALL, fetch = FetchType.LAZY)
private LibraryCard card;
}
@Entity
public class LibraryCard {
@Id @GeneratedValue private Long id;
@OneToOne(fetch = FetchType.LAZY, optional = false)
@JoinColumn(name = "employee_id", unique = true)
private Employee employee; // the owning side
}A known trap: an optional @OneToOne on the inverse side cannot be lazy with proxies, because Hibernate needs to query in order to know whether there is a row. If performance matters, make it optional = false or use a shared primary key.
@ManyToMany
@Entity
public class Material {
@ManyToMany(fetch = FetchType.LAZY)
@JoinTable(name = "material_category",
joinColumns = @JoinColumn(name = "material_id"),
inverseJoinColumns = @JoinColumn(name = "category_id"))
private Set<Category> categories = new HashSet<>();
}
@Entity
public class Category {
@ManyToMany(mappedBy = "categories")
private Set<Material> materials = new HashSet<>();
}A much-repeated piece of professional advice: a pure @ManyToMany falls short almost always. As soon as you need an attribute on the relationship —the date the category was assigned, who assigned it— you have to turn it into an intermediate entity with two @ManyToOnes. Many teams model the intermediate entity from the start.
Summary:
| Annotation | Where the FK lives | Owning side | Use in BiblioTech |
|---|---|---|---|
@ManyToOne |
In this table | This one | Loan → Material, Loan → Employee |
@OneToMany |
In the other table | The other one (mappedBy) |
Employee → Loan |
@OneToOne |
In one of the two | The one with @JoinColumn |
Employee → LibraryCard |
@ManyToMany |
An intermediate table | The one with @JoinTable |
Material ↔ Category |
- The owning side and
mappedBy
mappedByThis is a concept that confuses nearly everybody and causes the "I saved it and nothing was saved" bug.
In a relational database, the relationship is a single column: the foreign key. In Java, a bidirectional relationship is two fields. Somebody has to decide which of the two is in charge.
The owning side is the one that holds the foreign key. It is the only one Hibernate looks at in order to persist the relationship. The other side carries
mappedByand is read-only as far as persistence is concerned.
The consequence:
// BAD: only the inverse side is touched
Employee marta = employeeRepository.findById(1L).orElseThrow();
Loan loan = new Loan(book, LocalDate.now(clock));
marta.getLoans().add(loan); // only the INVERSE side
loanRepository.save(loan);
// -> The employee_id column stays NULL. The relationship has NOT been saved.// GOOD: the owning side is touched
loan.assignEmployee(marta); // the OWNING side (@ManyToOne)
loanRepository.save(loan);
// -> employee_id is saved correctly// BETTER: a convenience method that syncs both sides
marta.addLoan(loan); // sets both fields
loanRepository.save(loan);The practical rule: whenever you have a bidirectional relationship, write convenience methods that update both sides. Never manipulate the collections directly from outside.
FetchType: lazy versus eager loading
FetchType: lazy versus eager loadingWhen you load a Loan, should Hibernate also load its Material and its Employee? That decision is FetchType, and it is the decision with the greatest performance impact in the whole lesson.
FetchType |
Behaviour | When the association is loaded |
|---|---|---|
EAGER |
Always loaded, together with the entity | Immediately, with a JOIN or an extra query |
LAZY |
A proxy is placed there; it loads on access | The first time you call one of its methods |
JPA's defaults are, unfortunately, inconsistent:
| Annotation | Default | Is it a good idea? |
|---|---|---|
@ManyToOne |
EAGER |
No |
@OneToOne |
EAGER |
No |
@OneToMany |
LAZY |
Yes |
@ManyToMany |
LAZY |
Yes |
Why @ManyToOne's default EAGER is bad: imagine Loan has eager @ManyToOnes to Material and Employee, and Material has an eager @ManyToOne to Publisher, which has another to Country. Loading one loan drags in five tables. And you pay that cost always, even when all you wanted was the due date.
Professional rule: set
fetch = FetchType.LAZYexplicitly on ALL associations. And when you need the related data, ask for it explicitly in the query withJOIN FETCHor@EntityGraph.
@ManyToOne(fetch = FetchType.LAZY) // ALWAYS explicit
@JoinColumn(name = "material_id")
private Material material;How a lazy proxy works, and here module 10 returns: when you load a Loan, Hibernate does not put a Material in the field. It puts an object of a generated subclass of Material that has no data, only the id. The first time you call material.getTitle(), that subclass intercepts the call, fires the SELECT and returns the value.
It is your dynamic proxy from 10-03 again, and from it come two consequences you already know:
- An entity cannot be
final: the subclass could not be generated. - While debugging you will see objects called
Material$HibernateProxy$xYz123. That is your material, unloaded.
- The
LazyInitializationException
LazyInitializationExceptionThe flip side of lazy loading, and one of the most famous exceptions in the Java world.
@Service
public class ReportService {
@Transactional(readOnly = true)
public Loan get(Long id) {
return repository.findById(id).orElseThrow();
} // <- THE TRANSACTION ENDS HERE. The persistence context is closed.
}
// Somewhere else:
Loan loan = service.get(1L);
System.out.println(loan.getMaterial().getTitle()); // BOOMorg.hibernate.LazyInitializationException: could not initialize proxy
[com.nexussoftware.bibliotech.domain.Material#7] - no SessionWhat happened, step by step:
- Inside the transaction, Hibernate loaded the loan and put a proxy in the
materialfield. - The transaction ended and the persistence context was closed.
- On calling
getTitle(), the proxy tries to fire theSELECT... and there is no longer a session to do it with.
The solutions, from best to worst:
1. Fetch what you need in the query (the correct one):
@Query("SELECT l FROM Loan l JOIN FETCH l.material JOIN FETCH l.employee WHERE l.id = :id")
Optional<Loan> findWithDetails(@Param("id") Long id);2. Return a DTO instead of the entity (the best architecturally):
@Transactional(readOnly = true)
public LoanCard get(Long id) {
Loan l = repository.findById(id).orElseThrow();
// Everything is accessed INSIDE the transaction and a detached object is returned
return new LoanCard(l.getId(), l.getMaterial().getTitle(),
l.getEmployee().getName(), l.getDueDate());
}That LoanCard can perfectly well be a record — where they do fit beautifully.
3. Widen the transaction to cover the usage: sometimes a valid solution, but keeping transactions open longer than necessary has its own cost.
4. spring.jpa.open-in-view=true: discouraged, even though it is Spring Boot's default. It keeps the persistence context open for the whole HTTP request, so the exception never appears... and in exchange it hides the N+1 problem, keeps connections busy for longer and causes queries from the presentation layer. That is why section 5's application.yml explicitly contains:
A tip: keeping it at false makes the LazyInitializationException appear on your machine instead of the performance problem appearing in production.
- The N+1 problem
The number one performance problem of ORMs. And it is so easy to cause that almost everybody has it without knowing.
@Transactional(readOnly = true)
public void loanReport() {
List<Loan> loans = repository.findAll(); // 1 query
for (Loan l : loans) {
System.out.println(l.getMaterial().getTitle()); // 1 query EVERY TIME!
}
}With 500 loans, the show-sql log shows:
-- 1: the main query
select l.id, l.material_id, l.employee_id, ... from loans l
-- N: one per loan, on accessing its material
select m.id, m.isbn, m.title, ... from materials m where m.id=1
select m.id, m.isbn, m.title, ... from materials m where m.id=2
select m.id, m.isbn, m.title, ... from materials m where m.id=3
-- ... 500 times501 queries where there should be one. At 2 ms of latency per query, a whole second lost. And if you add the employee, it is 1001.
The insidious part: in development, with 5 test loans, it is instantaneous. The problem appears in production with real data, and by then the cause is a loop that looks harmless.
graph TD
A["findAll()<br/>1 query: 500 loans"] --> B["Loop over the 500"]
B --> C["l.getMaterial() -> proxy"]
C --> D["getTitle() fires a SELECT"]
D --> E["500 additional queries"]
E --> F["TOTAL: 501 queries<br/>instead of 1"]
An important note: switching to EAGER does not fix it. With EAGER, Hibernate always loads the associations, but it very often does so with one query per entity anyway. All you achieve is having the problem always, even when you were not going to use the materials.
How to detect it:
spring.jpa.show-sql=trueand count the queries in the log.- Hibernate statistics:
spring.jpa.properties.hibernate.generate_statistics=true. - In production: metrics and distributed tracing (12-07).
- Solving it:
JOIN FETCH and @EntityGraph
JOIN FETCH and @EntityGraphTwo solutions, both from the specification.
JOIN FETCH in JPQL
public interface LoanRepository extends JpaRepository<Loan, Long> {
@Query("""
SELECT l FROM Loan l
JOIN FETCH l.material
JOIN FETCH l.employee
WHERE l.status = :status
""")
List<Loan> findActiveWithDetails(@Param("status") LoanStatus status);
}Now a single query is generated:
select l.*, m.*, e.*
from loans l
join materials m on l.material_id = m.id
join employees e on l.employee_id = e.id
where l.status = 'ACTIVE'From 501 queries to 1.
Watch out for one detail: a JOIN FETCH over a collection can duplicate rows. If you fetch an employee with their 5 loans, the JOIN returns 5 rows and you could get the employee repeated. It is solved with SELECT DISTINCT or by asking for a Set. And a known limitation: you cannot efficiently paginate a collection JOIN FETCH; Hibernate warns that it is paginating in memory. In that case, the solution is to fetch the paginated ids in one query and the full entities in another.
@EntityGraph: declarative
public interface LoanRepository extends JpaRepository<Loan, Long> {
@EntityGraph(attributePaths = { "material", "employee" })
List<Loan> findByStatus(LoanStatus status);
@EntityGraph(attributePaths = { "material", "material.categories" })
Optional<Loan> findById(Long id);
}The advantage: the same method can have several graphs, and there is no JPQL to write.
Comparison
| Solution | When | Advantage | Drawback |
|---|---|---|---|
JOIN FETCH |
A specific, bespoke query | Full control | Hand-written JPQL |
@EntityGraph |
On derived methods | Declarative, reusable | Less flexible |
@BatchSize(size = 25) |
When laziness is unavoidable | From 501 to 21 queries | Still not 1 |
| DTO projection | Read-only queries | Fetches only the needed columns | They are not managed entities |
EAGER |
Never | — | It makes the problem worse |
The DTO projection deserves a separate mention because it is the most efficient solution for reports:
@Query("""
SELECT new com.nexussoftware.bibliotech.domain.LoanCard(
l.id, m.title, e.name, l.dueDate)
FROM Loan l JOIN l.material m JOIN l.employee e
WHERE l.status = :status
""")
List<LoanCard> activeCards(@Param("status") LoanStatus status);One query, only four columns, no managed entities and no risk of N+1. And LoanCard is a record — which finally finds its place in the persistence layer.
CascadeType and orphanRemoval
CascadeType and orphanRemovalCascades propagate operations from the parent to the children.
CascadeType |
What it propagates | Example in BiblioTech |
|---|---|---|
PERSIST |
Saving the parent saves the new children | Saving an Employee saves their Loans |
MERGE |
Merging the parent merges the children | Updating a detached graph |
REMOVE |
Deleting the parent deletes the children | Dangerous: deleting a Material would delete its history |
REFRESH |
Cascading refresh | Rare |
DETACH |
Cascading detach | Rare |
ALL |
All five | Convenient and dangerous |
@OneToMany(mappedBy = "employee",
cascade = { CascadeType.PERSIST, CascadeType.MERGE },
orphanRemoval = true)
private List<Loan> loans = new ArrayList<>();orphanRemoval = true means: if you remove a child from the collection, delete it from the database.
employee.getLoans().remove(loan);
// With orphanRemoval = true -> DELETE FROM loans WHERE id = ...
// Without it -> UPDATE loans SET employee_id = NULL (or an error if NOT NULL)The difference from CascadeType.REMOVE:
CascadeType.REMOVE |
orphanRemoval = true |
|
|---|---|---|
| Triggered when | Deleting the parent | Removing the child from the collection |
| Semantics | "When you delete the whole, delete the parts" | "A child without a parent makes no sense" |
A serious warning: CascadeType.ALL on relationships where the child has a life of its own is a source of data loss. If Material had cascade = ALL on its loans, withdrawing a book would delete its entire loan history, with its fines and its audit trail. Use cascades only when the relationship is genuine composition (the child does not exist without the parent).
- The persistence context and the
EntityManager
EntityManagerHere is JPA's conceptual heart, and what separates those who understand Hibernate from those who suffer it.
The EntityManager is the gateway into JPA. And it manages a persistence context: a cache of entities where every entity loaded or saved stays managed for as long as the transaction lasts.
Three properties of the persistence context, all with consequences:
- It is a first-level cache. If you ask for the same entity twice in the same transaction, the second time generates no query.
- It guarantees identity. Within a transaction, the same row is always the same Java object:
a == bistrue. That solves section 1's identity problem. - It detects changes automatically. That is section 21, and it is the most surprising part.
@Transactional
public void contextExample(Long id) {
Material a = entityManager.find(Material.class, id); // SELECT
Material b = entityManager.find(Material.class, id); // no query: cache
System.out.println(a == b); // true: THE SAME object
}In Spring, the EntityManager is injected like this:
@Repository
public class JpaMaterialRepository {
@PersistenceContext
private EntityManager entityManager; // Spring injects a per-transaction proxy
}That injected EntityManager is not a real instance: it is a proxy that on each call looks up the persistence context bound to the current thread's transaction. Once again, the same mechanism from module 10.
- The four states of an entity
Every entity is in one of four states, and knowing which one is the key to understanding what Hibernate is going to do.
stateDiagram-v2
[*] --> New: new Material(...)
New --> Managed: persist()
Managed --> Detached: transaction close<br/>or detach()
Detached --> Managed: merge()
Managed --> Removed: remove()
Removed --> New: persist()
Removed --> [*]: flush / commit
Managed --> Managed: changes are detected on their own
[*] --> Managed: find() / query
| State | What it means | Is Hibernate watching it? | Does it have an id? |
|---|---|---|---|
| New (transient) | Just created with new |
No | No |
| Managed | In the persistence context | Yes | Yes |
| Detached | It was managed, no longer is | No | Yes |
| Removed | Marked for deletion | Yes | Yes |
The example that clears everything up:
@Transactional
public void demonstration() {
// NEW: just a Java object. The database knows nothing about it.
Material book = new Book("978-0000000001", "Effective Java", 3);
// MANAGED: it enters the context. It will be inserted on commit.
entityManager.persist(book);
// Managed: this change will be saved ON ITS OWN, without calling anything.
book.setTitle("Effective Java, 3rd edition");
} // COMMIT: INSERT with the title already corrected
// Outside the method: DETACHED. Changes are no longer detected.And the classic error it produces:
// In a controller or service, outside a transaction
Material material = service.findByIsbn("978-0000000001"); // detached
material.setTitle("Another title"); // NOTHING happens
// The change stays in memory and is lost.For a change to a detached entity to be saved, it has to be re-attached with merge (section 23).
- Automatic dirty checking
This is the most surprising thing about JPA, and at the same time the most powerful.
@Transactional
public void renewLoan(Long loanId, int days) {
Loan loan = repository.findById(loanId).orElseThrow();
loan.renew(days); // changes dueDate
// There is NO save(). There is NO update().
} // And yet, on commit an UPDATE is executed.How? The mechanism is called dirty checking and it works like this:
- On loading the entity, Hibernate stores a snapshot of the state of all its fields.
- On
flush(normally on commit), it compares the current state with the snapshot. - For every differing field, it generates the corresponding
UPDATE.
The practical consequences, all of them important:
One: you do not need to call save on managed entities. A great deal of Spring Data code calls save() unnecessarily. It does no harm, but it is noise.
Two: any accidental modification is persisted. If inside a transaction you touch a field of a managed entity "just to calculate something", that change goes to the database. It is a real source of baffling bugs.
Three: readOnly = true switches it off and saves work.
@Transactional(readOnly = true) // no snapshots, no comparisons: faster
public List<Loan> list() { ... }On queries returning many entities the difference is measurable. It is a good habit to mark readOnly = true on every read-only method.
Four: the cost grows with the number of managed entities. Loading 100,000 entities in one transaction means every flush compares 100,000 snapshots. For bulk processing you have to clear the context periodically with entityManager.clear().
flush and the ordering of operations
flush and the ordering of operationsflush is the moment when Hibernate actually writes the pending SQL to the database. It is not the same as commit.
When it happens automatically:
- On committing the transaction (always).
- Before running a query that could be affected by the pending changes.
- When you call
entityManager.flush()explicitly.
And here comes a surprising detail: Hibernate does not execute the SQL in the order in which you call the methods. It reorders the operations like this:
- Entity
INSERTs, inpersistorder UPDATEs- Removal of collection elements
- Insertion of collection elements
- Entity
DELETEs
That is why this code can fail incomprehensibly:
@Transactional
public void replaceMaterial(String isbn, Material replacement) {
Material old = repository.findByIsbn(isbn).orElseThrow();
entityManager.remove(old); // "deletes" the one with isbn = X
entityManager.persist(replacement); // "inserts" another with isbn = X
} // On commit: FIRST the INSERT, THEN the DELETE
// -> violation of the isbn uniqueness constraintThe solution is to force the order:
entityManager.remove(old);
entityManager.flush(); // execute the DELETE now
entityManager.persist(replacement);Understanding that there is "pending SQL" executed later and reordered explains half of Hibernate's odd behaviour.
- The operations:
persist, merge, find, getReference, remove
persist, merge, find, getReference, remove| Operation | What it does | Resulting state | When |
|---|---|---|---|
persist(e) |
Marks a new entity for insertion | Managed (the same instance) | New entities |
merge(e) |
Copies a detached entity's state into the context | Managed (a different instance) | Detached entities |
find(C, id) |
Looks up by primary key | Managed or null |
You need the data |
getReference(C, id) |
Returns an unloaded proxy | Managed (proxy) | You only need the reference |
remove(e) |
Marks for deletion | Removed | Deletion |
detach(e) |
Takes it out of the context | Detached | Bulk processing |
refresh(e) |
Reloads from the database | Managed | Discarding local changes |
The difference between persist and merge (which causes real bugs)
// persist: the SAME instance becomes managed
Material book = new Book("978-0000000003", "Refactoring", 2);
entityManager.persist(book);
System.out.println(book.getId()); // it already has an id: it is the managed instance
// merge: it returns ANOTHER instance. The original stays detached.
Material detached = new Book(...); detached.setId(7L);
Material managed = entityManager.merge(detached);
managed.setTitle("A"); // IS SAVED
detached.setTitle("B"); // NOT saved: it is still detached
System.out.println(managed == detached); // falseThe rule: always use the object merge returns. Ignoring the return value is an extremely frequent mistake.
getReference: the optimisation people forget
// You need to associate a loan with a material, but you do not need its data.
// Option A: find -> executes an unnecessary SELECT
Material material = entityManager.find(Material.class, materialId);
loan.assignMaterial(material);
// Option B: getReference -> NO query; it just creates a proxy with the id
Material material = entityManager.getReference(Material.class, materialId);
loan.assignMaterial(material); // enough to set the FKAll the INSERT needs is the material_id, and getReference has it. You save one SELECT per operation. The flip side: if you access any property of the proxy outside the session, you get a LazyInitializationException, and if the id does not exist, the error arrives late (EntityNotFoundException on access, not on requesting the reference).
- Queries: JPQL
JPQL (Jakarta Persistence Query Language) looks like SQL but operates on entities and their Java attributes, not on tables and columns.
// SQL: TABLE and COLUMN names
SELECT l.* FROM loans l WHERE l.due_date < ?
// JPQL: ENTITY and ATTRIBUTE names
SELECT l FROM Loan l WHERE l.dueDate < :dateComplete examples over BiblioTech:
@Repository
public class QueryRepository {
@PersistenceContext
private EntityManager em;
// A simple query with a named parameter
public List<Loan> overdue(LocalDate date) {
return em.createQuery("""
SELECT l FROM Loan l
WHERE l.dueDate < :date
AND l.returnDate IS NULL
ORDER BY l.dueDate
""", Loan.class)
.setParameter("date", date)
.getResultList();
}
// JOIN FETCH to avoid the N+1
public List<Loan> overdueWithDetails(LocalDate date) {
return em.createQuery("""
SELECT l FROM Loan l
JOIN FETCH l.material
JOIN FETCH l.employee
WHERE l.dueDate < :date
""", Loan.class)
.setParameter("date", date)
.getResultList();
}
// Aggregation with GROUP BY, returning a record directly
public List<CountByEmployee> loansByEmployee() {
return em.createQuery("""
SELECT new com.nexussoftware.bibliotech.domain.CountByEmployee(
e.name, COUNT(l))
FROM Loan l JOIN l.employee e
GROUP BY e.id, e.name
HAVING COUNT(l) > 2
ORDER BY COUNT(l) DESC
""", CountByEmployee.class)
.getResultList();
}
// Pagination
public List<Material> page(int number, int size) {
return em.createQuery("SELECT m FROM Material m ORDER BY m.title", Material.class)
.setFirstResult(number * size)
.setMaxResults(size)
.getResultList();
}
// Bulk modification: it does NOT go through the persistence context
@Modifying
public int markOverdue(LocalDate date) {
return em.createQuery("""
UPDATE Loan l SET l.status = :overdue
WHERE l.dueDate < :date AND l.returnDate IS NULL
""")
.setParameter("overdue", LoanStatus.OVERDUE)
.setParameter("date", date)
.executeUpdate();
}
}A warning about that last one: bulk modification queries run directly on the database and do not update entities already loaded in the persistence context, which are left out of sync. After one of those, an em.clear() is advisable.
Key differences from SQL:
| JPQL | SQL |
|---|---|
FROM Loan l (entity) |
FROM loans l (table) |
l.dueDate (attribute) |
l.due_date (column) |
JOIN l.material (navigates the relationship) |
JOIN materials ON ... (explicit condition) |
SELECT l returns entities |
SELECT * returns rows |
| Portable between engines | Depends on the dialect |
- Named queries, the Criteria API and native SQL
Named queries
They are declared on the entity and validated at start-up, not on execution:
@Entity
@NamedQuery(name = "Loan.overdue",
query = """
SELECT l FROM Loan l
WHERE l.dueDate < :date AND l.returnDate IS NULL
""")
public class Loan { ... }List<Loan> overdue = em.createNamedQuery("Loan.overdue", Loan.class)
.setParameter("date", LocalDate.now(clock))
.getResultList();A real advantage: a syntax error in the JPQL stops the application from starting, instead of blowing up the day somebody runs that report.
The Criteria API, mentioned
For queries built dynamically —a search screen with five optional filters— concatenating JPQL by hand is ugly and dangerous. JPA offers a typed API:
CriteriaBuilder cb = em.getCriteriaBuilder();
CriteriaQuery<Loan> query = cb.createQuery(Loan.class);
Root<Loan> l = query.from(Loan.class);
List<Predicate> filters = new ArrayList<>();
if (isbn != null) filters.add(cb.equal(l.get("material").get("isbn"), isbn));
if (from != null) filters.add(cb.greaterThanOrEqualTo(l.get("loanDate"), from));
query.where(cb.and(filters.toArray(new Predicate[0])));
List<Loan> result = em.createQuery(query).getResultList();It is verbose, and that is why many people prefer Spring Data's Specification or QueryDSL. It is mentioned so you know it exists and what it is for; in practice it is not used much.
Native SQL
When you need an engine-specific function, a heavily optimised query or a complex report:
List<Object[]> rows = em.createNativeQuery("""
SELECT m.title, COUNT(l.id) AS total
FROM materials m LEFT JOIN loans l ON l.material_id = m.id
WHERE l.loan_date >= :from
GROUP BY m.id, m.title
ORDER BY total DESC
FETCH FIRST 10 ROWS ONLY
""")
.setParameter("from", from)
.getResultList();It is perfectly legitimate. An ORM does not force you to do everything with it, and pretending otherwise leads to convoluted JPQL that nobody understands. The only loss is portability between engines.
- SQL injection: why you never concatenate
This is not a style recommendation. It is a security vulnerability.
// NEVER. NOT EVER.
public List<Loan> findByEmployee(String email) {
return em.createQuery(
"SELECT l FROM Loan l WHERE l.employee.email = '" + email + "'",
Loan.class).getResultList();
}If email comes from a form or from an API parameter, an attacker can send:
And the query becomes:
which returns every loan of every employee. With native SQL the consequences are worse: depending on the engine and the permissions, an attacker can end up reading other tables or modifying data.
The solution is always the same, and it is trivial:
// A named parameter
em.createQuery("SELECT l FROM Loan l WHERE l.employee.email = :email", Loan.class)
.setParameter("email", email)
.getResultList();
// Or a positional one
em.createQuery("SELECT l FROM Loan l WHERE l.employee.email = ?1", Loan.class)
.setParameter(1, email)
.getResultList();With parameters, the value is never interpreted as part of the query: it travels separately and is treated as a literal piece of data, whatever it contains. It is the same reason section 3's JDBC example used a PreparedStatement with ? and setString.
The rule, without exceptions: every value coming from outside goes in as a parameter. Not even "when I know it is a number" — because tomorrow that code gets copied somewhere it is not.
The only things that cannot be parameterised are table and column names and the direction of an ORDER BY. If you need those to be dynamic, validate them against an allow-list of permitted values; never concatenate them as they come.
Application security —input validation, authorisation, secrets management, headers— is covered in 12-07. Here the rule is enough: always parameters.
- Transactions:
@Transactional with a real database
@Transactional with a real databaseNow that there is a database, 11-02's @Transactional actually does something.
A transaction satisfies the ACID properties:
| Property | What it guarantees |
|---|---|
| Atomicity | Either every operation or none |
| Consistency | Integrity constraints are respected |
| Isolation | Concurrent transactions do not tread on each other |
| Durability | What is committed survives a crash |
BiblioTech's case:
@Service
public class LoanManager {
@Transactional
public Loan lend(String isbn, String employeeEmail) {
Material material = materialRepository.findByIsbn(isbn)
.orElseThrow(() -> new MaterialNotFoundException(isbn));
Employee employee = employeeRepository.findByEmail(employeeEmail)
.orElseThrow(() -> new EmployeeNotFoundException(employeeEmail));
if (loanRepository.countActive(employee) >= properties.loan().maxPerEmployee()) {
throw new LoanLimitExceededException(employeeEmail);
}
material.lendOneCopy(); // UPDATE (through dirty checking)
Loan loan = new Loan(material, employee,
LocalDate.now(clock), properties.loan().defaultDays());
loanRepository.save(loan); // INSERT
auditLog.record(loan); // INSERT
return loan;
}
@Transactional(readOnly = true) // optimised: no dirty checking
public List<Loan> activeFor(String email) {
return loanRepository.findActiveByEmployee(email);
}
}If auditLog.record fails, everything is rolled back: the copy becomes available again and no orphan loan is left behind. Without a transaction, one fewer copy and an unaudited loan would remain, and nobody would know why.
Remember the two rules from 11-02, which still hold here: rollback only on unchecked exceptions by default, and internal calls do not go through the proxy.
- Propagation and isolation
Propagation: what to do if a transaction is already open
| Propagation | Behaviour | Use |
|---|---|---|
REQUIRED (the default) |
Joins the existing one; if there is none, creates it | 95% of cases |
REQUIRES_NEW |
Suspends the current one and opens an independent one | Auditing that must persist even if the rest fails |
SUPPORTS |
Joins if there is one; if not, no transaction | Queries |
MANDATORY |
Requires one to exist already; error otherwise | Internal methods that must never be called alone |
NOT_SUPPORTED |
Suspends the current one and runs without a transaction | Very long operations |
NEVER |
Error if there is a transaction | Rare |
NESTED |
A savepoint inside the current one | Limited support |
The case where REQUIRES_NEW is the right answer:
@Service
public class AuditLog {
// We want to record the ATTEMPT even if the loan is rolled back
@Transactional(propagation = Propagation.REQUIRES_NEW)
public void recordAttempt(String isbn, String employee, String result) {
em.persist(new AuditEntry(isbn, employee, result, Instant.now(clock)));
}
}Careful: REQUIRES_NEW consumes a second connection from the pool while the first is still open. Overusing it exhausts the pool and causes deadlocks. Use it only when independence is a genuine requirement.
Isolation: how concurrent transactions see each other
| Level | Dirty read | Non-repeatable read | Phantom read | Cost |
|---|---|---|---|---|
READ_UNCOMMITTED |
Possible | Possible | Possible | Minimal |
READ_COMMITTED |
No | Possible | Possible | Low (the default in PostgreSQL, Oracle, H2) |
REPEATABLE_READ |
No | No | Possible | Medium (the default in MySQL) |
SERIALIZABLE |
No | No | No | High |
The three phenomena, with BiblioTech:
- Dirty read: you read a copy as being on loan and that transaction rolls back. You read something that never existed.
- Non-repeatable read: you read
available_copies = 1, somebody else lends it, you read again and there is0. Two reads, two results. - Phantom read: you count 3 active loans, somebody inserts one, you repeat the query and there are 4.
@Transactional(isolation = Isolation.REPEATABLE_READ)
public MonthlyReport generate(YearMonth month) { ... }In practice, isolation is almost never touched. READ_COMMITTED with optimistic locking (the next section) solves 99% of cases at far less cost than raising the level.
- Optimistic locking with
@Version
@VersionThe concrete case, which picks up module 8 directly:
Marta Ruiz and Diego Alonso open the record for "Effective Java" at the same time; it has 1 available copy. Both click "Lend" in the same second.
sequenceDiagram
participant M as Marta
participant DB as Database
participant D as Diego
M->>DB: SELECT material 1 (available=1)
D->>DB: SELECT material 1 (available=1)
M->>DB: UPDATE available=0
D->>DB: UPDATE available=0
Note over DB: Two loans, one copy.<br/>LOST UPDATE
It is exactly the race condition from 08-04, but distributed: here synchronized is of no use whatsoever, because the two users can be on different instances of the application.
Two strategies:
| Pessimistic locking | Optimistic locking | |
|---|---|---|
| Assumption | There will be a conflict | There will almost never be a conflict |
| How | SELECT ... FOR UPDATE: it locks the row |
A version column checked on update |
| Cost | High: locks, waits, deadlocks | Very low |
| When | Frequent and expensive conflicts | Almost always |
| JPA | @Lock(LockModeType.PESSIMISTIC_WRITE) |
@Version |
Optimistic locking is a single annotation:
@Entity
public class Material {
@Id @GeneratedValue private Long id;
@Version // Hibernate manages it on its own
private Long version;
@Column(name = "available_copies", nullable = false)
private int availableCopies;
}With that, every UPDATE includes the version that was read in the condition:
If somebody else has already updated, the row has version=6 and the condition matches no row. Hibernate detects it and throws:
Handling it properly:
@Service
public class LoanManager {
@Transactional
public Loan lend(String isbn, String email) {
// ... as before
}
public Result<Loan> lendWithRetry(String isbn, String email) {
for (int attempt = 1; attempt <= 3; attempt++) {
try {
return Result.success(self.lend(isbn, email));
} catch (OptimisticLockingFailureException e) {
log.warn("Concurrency conflict on {} (attempt {}/3)", isbn, attempt);
}
}
return Result.error("The material is being modified. Please try again.");
}
}That Result<T> is the generic type you wrote in 10-01, and it fits perfectly here. And that retry loop is, conceptually, your RetryProxy from 10-03: in a real project you would declare it with Spring Retry's @Retryable.
Pessimistic locking, when the operation is expensive and conflict is likely:
@Lock(LockModeType.PESSIMISTIC_WRITE)
@Query("SELECT m FROM Material m WHERE m.isbn = :isbn")
Optional<Material> findForUpdate(@Param("isbn") String isbn);It generates SELECT ... FOR UPDATE: the second waits for the first. Correct, but with a real concurrency cost.
- First- and second-level caches
| Cache | Scope | Enabled by default | What it stores |
|---|---|---|---|
| First level | The persistence context (the transaction) | Yes, always | Managed entities |
| Second level | The EntityManagerFactory (the whole application) |
No | Entities between transactions |
| Query cache | Global | No | Query results |
You already know the first-level one: it is section 19's persistence context. It cannot be switched off and it is what guarantees object identity.
The second-level one is optional and shared by every transaction. It needs a provider (Ehcache, Caffeine, Hazelcast):
@Entity
@Cacheable
@org.hibernate.annotations.Cache(usage = CacheConcurrencyStrategy.READ_WRITE)
public class Material { ... }spring:
jpa:
properties:
hibernate:
cache:
use_second_level_cache: true
region.factory_class: org.hibernate.cache.jcache.JCacheRegionFactoryWhen to enable it: data that is read a lot and changes little. BiblioTech's material catalogue or its categories are good candidates; the loans are not.
Warnings to bear in mind:
- Coherence: if another application modifies the database, your cache goes stale and serves old data.
- In a cluster: with several instances, each one has its own cache. Either you use a distributed cache or you will get inconsistencies.
- Measure first: enabling the second-level cache "just in case" adds complexity with no demonstrated benefit. It is exactly what 10-07 said: measure before optimising.
ddl-auto and why never update in production
ddl-auto and why never update in productionHibernate can generate the schema from your entities. It is convenient and dangerous.
| Value | What it does | When to use it |
|---|---|---|
none |
Nothing | Production |
validate |
Checks that the schema matches the entities; fails if not | Production (recommended) |
update |
Tries to modify the schema to fit | Never in production |
create |
Drops and creates the schema at start-up | Tests |
create-drop |
Like create, and drops on shutdown |
Development with H2, tests |
Why update is never used in production, with concrete reasons:
- It never deletes anything. If you remove a field, the column stays there forever. If you rename
titletoname, it addsnameand leavestitle— and it does not copy the data. - It versions nothing. There is no record of which changes were applied or when, and no way to undo them.
- It is unpredictable. The SQL it generates depends on the dialect, the Hibernate version and the current state of the schema.
- It cannot be reviewed beforehand. Nobody has seen the
ALTER TABLEthat is about to run against the production database. - It cannot do data migrations. Splitting
full_nameintofirst_nameandsurnameneeds logic no automatic tool can invent.
The correct configuration per environment:
# application-dev.yml — in-memory H2, schema regenerated on every start-up
spring:
jpa:
hibernate:
ddl-auto: create-drop# application-prod.yml — the schema is managed by a migration tool
spring:
jpa:
hibernate:
ddl-auto: validatevalidate is especially valuable: if the real schema does not match the entities, the application does not start, and that is much better than finding out through an error at three in the morning.
Flyway and Liquibase, mentioned
In production, the schema is managed with versioned migrations: numbered SQL files applied in order and recorded in a control table.
-- src/main/resources/db/migration/V1__initial_schema.sql
CREATE TABLE materials (
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
isbn VARCHAR(20) NOT NULL UNIQUE,
title VARCHAR(200) NOT NULL,
available_copies INT NOT NULL DEFAULT 0,
version BIGINT NOT NULL DEFAULT 0
);-- V2__add_category.sql
ALTER TABLE materials ADD COLUMN category VARCHAR(50);
UPDATE materials SET category = 'GENERAL' WHERE category IS NULL;The advantages: the schema change is in git, it is reviewable, it is reproducible and it can be tested beforehand. It is developed in 12-06, alongside deployment.
- Spring Data JPA
The last piece, and the one that removes the most code.
With plain JPA, a repository is about eighty lines. With Spring Data JPA:
package com.nexussoftware.bibliotech.persistence;
import org.springframework.data.jpa.repository.*;
import org.springframework.data.repository.query.Param;
public interface LoanRepository extends JpaRepository<Loan, Long> {
// And that is it. No implementation.
}That already gives you save, findById, findAll, deleteById, count, existsById, pagination and sorting.
Methods derived from the name
public interface LoanRepository extends JpaRepository<Loan, Long> {
List<Loan> findByStatus(LoanStatus status);
List<Loan> findByEmployeeEmail(String email);
List<Loan> findByDueDateBeforeAndReturnDateIsNull(LocalDate date);
long countByEmployeeAndStatus(Employee employee, LoanStatus status);
Optional<Loan> findFirstByMaterialIsbnOrderByLoanDateDesc(String isbn);
boolean existsByMaterialIsbnAndStatus(String isbn, LoanStatus status);
}Spring Data analyses the method name and generates the query. The usual keywords:
| Keyword | Generated JPQL |
|---|---|
findBy, readBy, getBy |
SELECT ... |
And, Or |
AND, OR |
Between, LessThan, GreaterThan, Before, After |
Comparisons |
IsNull, IsNotNull |
IS NULL |
Like, Containing, StartingWith |
LIKE |
In, NotIn |
IN |
OrderBy...Asc/Desc |
ORDER BY |
countBy, existsBy, deleteBy |
Aggregation, existence, deletion |
First, Top3 |
LIMIT |
A practical tip: derived names are fantastic until they stop being so. findByDueDateBeforeAndReturnDateIsNullAndStatusNot is unreadable. When the name goes beyond three conditions, use @Query.
@Query
@Query("""
SELECT l FROM Loan l
JOIN FETCH l.material m
JOIN FETCH l.employee e
WHERE l.dueDate < :date AND l.returnDate IS NULL
ORDER BY l.dueDate
""")
List<Loan> overdueWithDetails(@Param("date") LocalDate date);
@Modifying
@Transactional
@Query("UPDATE Loan l SET l.status = :status WHERE l.id = :id")
int updateStatus(@Param("id") Long id, @Param("status") LoanStatus status);
@Query(value = "SELECT * FROM loans WHERE loan_date > ?1", nativeQuery = true)
List<Loan> nativeSince(LocalDate from);Pagination
Pageable page = PageRequest.of(0, 20, Sort.by("dueDate").descending());
Page<Loan> result = repository.findByStatus(LoanStatus.ACTIVE, page);
result.getContent(); // the 20 on this page
result.getTotalElements(); // the total (it runs an additional COUNT)
result.getTotalPages();
result.hasNext();If you do not need the total, Slice<T> avoids the COUNT and is faster.
How it works: dynamic proxies again
LoanRepository is an interface with no implementation. Who runs findByStatus?
At start-up, Spring Data:
- Scans the interfaces that extend
Repository. - For each one, it creates an implementation on the fly with
Proxy.newProxyInstance— the very one from 10-03. - The
InvocationHandlerreceives each call, checks whether the method has@Query(and uses that JPQL) or analyses its name (and generates the JPQL), runs it with theEntityManagerand adapts the result.
graph LR
A["manager.repository<br/>.findByStatus(ACTIVE)"] --> B["JDK proxy<br/>(Spring Data)"]
B --> C["Analyses the name<br/>or reads @Query"]
C --> D["Generates JPQL"]
D --> E["EntityManager"]
E --> F["Hibernate -> SQL"]
F --> G[("H2")]
It is literally what you wrote in 10-03: a proxy over an interface that interprets the method's metadata and acts. The difference is the sophistication of the analysis, not the mechanism.
And there is a practical consequence: the repository can be replaced by a mock in tests with no difficulty at all, because it is an interface. That is exactly what 11-06 will do.
- BiblioTech: from CSV to H2
The complete result of the migration.
// domain/Material.java — it can no longer be a sealed record: it is an entity
package com.nexussoftware.bibliotech.domain;
import jakarta.persistence.*;
@Entity
@Table(name = "materials")
@Inheritance(strategy = InheritanceType.SINGLE_TABLE) // one table for the whole hierarchy
@DiscriminatorColumn(name = "type", discriminatorType = DiscriminatorType.STRING, length = 20)
public abstract class Material {
@Id @GeneratedValue(strategy = GenerationType.IDENTITY)
private Long id;
@Column(nullable = false, unique = true, length = 20)
private String isbn;
@Column(nullable = false, length = 200)
private String title;
@Column(name = "available_copies", nullable = false)
private int availableCopies;
@Version
private Long version;
protected Material() { }
protected Material(String isbn, String title, int copies) {
this.isbn = isbn;
this.title = title;
this.availableCopies = copies;
}
public void lendOneCopy() {
if (availableCopies <= 0) throw new NoCopiesAvailableException(isbn);
availableCopies--;
}
public void returnOneCopy() { availableCopies++; }
// getters...
}
@Entity @DiscriminatorValue("BOOK")
public class Book extends Material {
@Column(length = 120) private String author;
protected Book() { }
public Book(String isbn, String title, int copies, String author) {
super(isbn, title, copies);
this.author = author;
}
}
@Entity @DiscriminatorValue("MAGAZINE")
public class Magazine extends Material {
private Integer issueNumber;
protected Magazine() { }
}
@Entity @DiscriminatorValue("DVD")
public class Dvd extends Material {
@Column(name = "duration_minutes") private Integer durationMinutes;
protected Dvd() { }
}Here one of the ORM's real prices is paid, and it has to be said plainly: the sealed hierarchy from 10-06 disappears. A JPA entity cannot be sealed, a record or final. You lose the exhaustive switch expressions checked by the compiler and have to replace them with polymorphism or with instanceof checks. It is a conscious trade-off: you gain real persistence, queries and integrity; you lose part of the type model's expressiveness. An alternative design —and a frequent one in projects that value the domain highly— keeps the domain pure and uses separate JPA entities with a mapping between the two. That discussion belongs to 12-02.
The three inheritance strategies, so you know what you chose:
| Strategy | How | Advantages | Drawbacks |
|---|---|---|---|
SINGLE_TABLE (the default) |
One table, a discriminator column | Fast, no JOINs |
Subclass columns must be nullable |
JOINED |
One table per class, joined by PK | Normalised, no nulls | A JOIN on every query |
TABLE_PER_CLASS |
A complete table per subclass | No JOIN for specific queries |
UNION when querying the superclass |
The repositories:
public interface MaterialRepository extends JpaRepository<Material, Long> {
Optional<Material> findByIsbn(String isbn);
List<Material> findByTitleContainingIgnoreCase(String fragment);
List<Material> findByAvailableCopiesGreaterThan(int minimum);
}
public interface EmployeeRepository extends JpaRepository<Employee, Long> {
Optional<Employee> findByEmail(String email);
}
public interface LoanRepository extends JpaRepository<Loan, Long> {
@EntityGraph(attributePaths = { "material", "employee" })
List<Loan> findByStatus(LoanStatus status);
long countByEmployeeAndStatus(Employee employee, LoanStatus status);
@Query("""
SELECT l FROM Loan l JOIN FETCH l.material JOIN FETCH l.employee
WHERE l.dueDate < :date AND l.returnDate IS NULL
""")
List<Loan> overdue(@Param("date") LocalDate date);
}Seed data for development, in src/main/resources/data.sql:
INSERT INTO employees (email, name) VALUES
('[email protected]', 'Marta Ruiz'),
('[email protected]', 'Diego Alonso'),
('[email protected]', 'Nuria Vidal');
INSERT INTO materials (type, isbn, title, available_copies, version, author) VALUES
('BOOK', '978-0000000001', 'Effective Java', 3, 0, 'J. Bloch'),
('BOOK', '978-0000000002', 'Design Patterns', 2, 0, 'GoF'),
('BOOK', '978-0000000003', 'Refactoring', 1, 0, 'M. Fowler');Running it:
mvn spring-boot:run -Dspring-boot.run.profiles=dev
# H2 web console: http://localhost:8080/h2-console
# JDBC URL: jdbc:h2:mem:bibliotech | User: sa | No passwordThe before and after:
| Aspect | CSV (module 7) | JPA over H2 |
|---|---|---|
| Format | Plain text | A relational database |
| Queries | Load everything and filter in memory | SQL with indexes |
| Transactions | None | ACID |
| Referential integrity | None | Foreign keys |
| Concurrency | Two processes corrupt it | Optimistic locking with @Version |
| Writing | Hand-written AtomicWrite |
Managed by the engine |
| Access code | CsvReader + CsvWriter + mapping |
Three interfaces with no implementation |
| New queries | Programme the filtering | One more method in the interface |
| Scalability | Thousands of rows | Millions |
- Common Mistakes and Tips
Mistake: javax.persistence instead of jakarta.persistence. The number one copy-and-paste error today. Spring Boot 3 and Hibernate 6 use jakarta.
Mistake: leaving @Enumerated at its default (ORDINAL). Reordering an enum silently corrupts all the historical data. Always EnumType.STRING.
Mistake: leaving @ManyToOne at EAGER. It is the default and it is bad. Set LAZY explicitly on every association and fetch what you need with JOIN FETCH.
Mistake: the N+1 problem without noticing. A loop over entities accessing a lazy relationship. Switch show-sql on in development and count the queries.
Mistake: spring.jpa.open-in-view=true. It is Spring Boot's default and it hides the N+1 until production. Set it to false.
Mistake: modifying only the inverse side of a relationship. employee.getLoans().add(l) without l.setEmployee(employee) persists nothing. Write convenience methods.
Mistake: ignoring merge's return value. merge(e) returns another instance; the original stays detached and its changes are lost.
Mistake: CascadeType.ALL without thinking. Deleting a material could delete its entire loan history.
Mistake: ddl-auto=update in production. It does not delete, does not version, does not migrate data and nobody reviews the SQL it is about to run. validate plus Flyway.
Mistake: concatenating values into a query. That is SQL injection. Parameters always, no exceptions.
Mistake: final entities, record entities or entities with final methods. Hibernate needs to generate proxy subclasses. Either the model will not compile or lazy loading will fail.
Mistake: equals and hashCode based on the generated id. Before persisting, the id is null; afterwards it changes, and an entity put into a HashSet can no longer be found. Use a business key (the isbn) or Hibernate's recommended pattern.
Tip: switch show-sql on while you are learning. Seeing the generated SQL is half the learning, and it teaches you to spot the N+1 instantly.
Tip: mark readOnly = true on query methods. It switches off dirty checking and saves real work.
Tip: use DTOs (record) for whatever leaves the service. Returning managed entities outside the transaction is the direct cause of the LazyInitializationException, and it also couples your internal API to the schema.
Tip: @Version on every entity modified concurrently. It costs one annotation and prevents lost updates.
Tip: do not do everything with the ORM. For reports and aggregations, native SQL or JdbcTemplate are better tools. It is legitimate and professional.
Tip: getReference when all you need is the reference in order to set an FK. It saves a SELECT per operation.
- Exercises
Exercise 1: mapping the Reservation entity
BiblioTech has reservations: when there are no copies, an employee reserves one and is notified when one is returned. This is module 10's record:
public record Reservation(Long id, String isbn, String employeeEmail,
LocalDate requestDate, LocalDate expiryDate,
Optional<Instant> notifiedAt, ReservationStatus status) { }
public enum ReservationStatus { PENDING, NOTIFIED, COMPLETED, EXPIRED }Turn it into a JPA entity that satisfies:
- A
reservationstable with an auto-generated primary key. - Lazy
@ManyToOnerelationships toMaterialandEmployee, not strings. - The enum stored safely.
notifiedAtnullable in the database but exposed asOptional.- Optimistic locking.
- An index over
(material_id, status)for the pending-reservations query. - A domain method
markNotified(Clock clock)that changes the status and records the instant, with validation. - The Spring Data repository with: a material's pending reservations ordered by age, reservations expired before a date, and counting an employee's pending ones.
Exercise 2: diagnosing and fixing an N+1
This service works fine with development's 5 loans and takes 9 seconds in production with 800.
@Service
public class ReportService {
private final LoanRepository repository;
@Transactional(readOnly = true)
public List<String> overdueReport(LocalDate date) {
List<Loan> overdue = repository.findByDueDateBefore(date);
return overdue.stream()
.map(l -> String.format("%s | %s | %s | %d days",
l.getMaterial().getTitle(),
l.getMaterial().getCategories().stream()
.map(Category::getName)
.collect(Collectors.joining(", ")),
l.getEmployee().getName(),
ChronoUnit.DAYS.between(l.getDueDate(), date)))
.toList();
}
}You are asked to:
- Work out how many SQL queries run with 800 loans, broken down.
- Explain exactly why.
- Propose three different solutions with their code, saying how many queries each leaves.
- Say which one you would choose and why.
Exercise 3: the disputed copy
Marta Ruiz and Diego Alonso try to lend the last copy of "Refactoring" (978-0000000003, 1 copy) at the same time.
@Service
public class LoanManager {
@Transactional
public Loan lend(String isbn, String email) {
Material material = materialRepository.findByIsbn(isbn).orElseThrow();
Employee employee = employeeRepository.findByEmail(email).orElseThrow();
if (material.getAvailableCopies() <= 0) {
throw new NoCopiesAvailableException(isbn);
}
material.setAvailableCopies(material.getAvailableCopies() - 1);
Loan loan = new Loan(material, employee, LocalDate.now(clock), 15);
return loanRepository.save(loan);
}
}You are asked to:
- Describe the exact sequence of events that produces two loans from a single copy.
- Explain why
synchronized(08-04) does not solve this in a real deployment. - Solve it with optimistic locking, including handling the exception with a retry.
- Solve it with pessimistic locking.
- Compare the two in a table and recommend one.
Solutions
Solution 1
package com.nexussoftware.bibliotech.domain;
import jakarta.persistence.*;
import java.time.*;
import java.util.Optional;
@Entity
@Table(name = "reservations",
indexes = @Index(name = "idx_reservation_material_status", // (6)
columnList = "material_id, status"))
public class Reservation {
@Id // (1)
@GeneratedValue(strategy = GenerationType.IDENTITY)
private Long id;
@ManyToOne(fetch = FetchType.LAZY, optional = false) // (2)
@JoinColumn(name = "material_id", nullable = false)
private Material material;
@ManyToOne(fetch = FetchType.LAZY, optional = false) // (2)
@JoinColumn(name = "employee_id", nullable = false)
private Employee employee;
@Column(name = "request_date", nullable = false)
private LocalDate requestDate;
@Column(name = "expiry_date", nullable = false)
private LocalDate expiryDate;
@Column(name = "notified_at") // (4) nullable
private Instant notifiedAt;
@Enumerated(EnumType.STRING) // (3) NEVER ORDINAL
@Column(nullable = false, length = 20)
private ReservationStatus status;
@Version // (5)
private Long version;
protected Reservation() { } // required by JPA
public Reservation(Material material, Employee employee, LocalDate request, int validDays) {
this.material = material;
this.employee = employee;
this.requestDate = request;
this.expiryDate = request.plusDays(validDays);
this.status = ReservationStatus.PENDING;
}
// (7) domain behaviour with validation
public void markNotified(Clock clock) {
if (status != ReservationStatus.PENDING) {
throw new InvalidReservationStatusException(
"Only a PENDING reservation can be notified; current status: " + status);
}
this.status = ReservationStatus.NOTIFIED;
this.notifiedAt = Instant.now(clock); // injected Clock (10-05)
}
public void complete() {
if (status != ReservationStatus.NOTIFIED) {
throw new InvalidReservationStatusException("Only a NOTIFIED reservation completes");
}
this.status = ReservationStatus.COMPLETED;
}
public boolean hasExpired(LocalDate today) {
return status == ReservationStatus.PENDING && expiryDate.isBefore(today);
}
// (4) the field is nullable; the public API returns Optional
public Optional<Instant> getNotifiedAt() {
return Optional.ofNullable(notifiedAt);
}
public Long getId() { return id; }
public Material getMaterial() { return material; }
public Employee getEmployee() { return employee; }
public ReservationStatus getStatus() { return status; }
public LocalDate getRequestDate() { return requestDate; }
public LocalDate getExpiryDate() { return expiryDate; }
}The repository (8):
public interface ReservationRepository extends JpaRepository<Reservation, Long> {
// A material's pending ones, oldest first (the FIFO reservation queue)
@EntityGraph(attributePaths = { "employee" }) // avoids N+1 when reading the name
List<Reservation> findByMaterialIsbnAndStatusOrderByRequestDateAsc(
String isbn, ReservationStatus status);
// Expired ones
List<Reservation> findByStatusAndExpiryDateBefore(ReservationStatus status, LocalDate date);
// Counting an employee's pending ones
long countByEmployeeEmailAndStatus(String email, ReservationStatus status);
}Notable decisions:
record→ class: mandatory, because arecordisfinaland has no empty constructor.- Relationships instead of strings:
MaterialandEmployeeinstead ofisbnandemployeeEmail. It gives referential integrity (you cannot reserve a non-existent material) and navigation. The price is having to manage laziness. Optionaloutside, a nullable field inside: 10-04's guarantee is preserved without fighting JPA.- The behaviour lives in the entity:
markNotified,completeandhasExpiredvalidate the transitions. The entity is not a bag of data. @EntityGraphon the first query: we know the employee's name is going to be read, so it is fetched in one go.
Solution 2
1. How many queries.
| Item | Queries |
|---|---|
findByDueDateBefore |
1 |
l.getMaterial() — a lazy proxy, 800 times |
800 |
l.getMaterial().getCategories() — a lazy collection, 800 times |
800 |
l.getEmployee() — a lazy proxy, 800 times |
800 |
| Total | 2,401 |
At 3-4 ms of latency per query, between 7 and 10 seconds. That matches the symptom.
2. Why. All three associations are lazy. The initial query fetches only the loans rows, with proxies in material and employee and a lazy collection in categories. Each access inside the map fires its own SELECT. In development, with 5 loans, that is 16 queries and nobody notices.
Solution A — JOIN FETCH (1 query if limited to scalars, 2 with the collection):
public interface LoanRepository extends JpaRepository<Loan, Long> {
@Query("""
SELECT DISTINCT l FROM Loan l
JOIN FETCH l.material m
LEFT JOIN FETCH m.categories
JOIN FETCH l.employee
WHERE l.dueDate < :date
""")
List<Loan> overdueComplete(@Param("date") LocalDate date);
}DISTINCT is necessary because the JOIN FETCH on the categories collection duplicates loan rows. 1 query. The limitation: you cannot paginate efficiently with a collection JOIN FETCH.
Solution B — @EntityGraph (declarative, same result):
@EntityGraph(attributePaths = { "material", "material.categories", "employee" })
List<Loan> findByDueDateBefore(LocalDate date);The advantage: the service code is untouched and the same method can carry a different graph in another query. 1-2 queries.
Solution C — DTO projection (the most efficient):
public record ReportLine(String materialTitle, String employeeName,
LocalDate dueDate) {
public String format(LocalDate reference) {
return "%s | %s | %d days".formatted(materialTitle, employeeName,
ChronoUnit.DAYS.between(dueDate, reference));
}
}@Query("""
SELECT new com.nexussoftware.bibliotech.domain.ReportLine(
m.title, e.name, l.dueDate)
FROM Loan l JOIN l.material m JOIN l.employee e
WHERE l.dueDate < :date
ORDER BY l.dueDate
""")
List<ReportLine> overdueLines(@Param("date") LocalDate date);1 query, and only 3 columns instead of all the columns of three tables. No managed entities, no dirty checking, no risk of a later N+1. For the categories you would need an additional query or STRING_AGG in native SQL.
Comparison:
| Solution | Queries | Data transferred | Flexibility | Risk of relapse |
|---|---|---|---|---|
A: JOIN FETCH |
1-2 | Every column | Medium | Medium |
B: @EntityGraph |
1-2 | Every column | High | Medium |
| C: DTO | 1 | Only what is needed | Low | None |
4. Which to choose. For a read-only report, C. Reasons: it fetches only what is needed, there are no managed entities and no dirty checking, and —most importantly— it is structurally impossible for the N+1 to reappear, because there are no proxies that could fire. With A and B, somebody who adds an l.getMaterial().getPublisher().getCountry() to the format tomorrow has the problem back.
If the method had to return entities in order to modify them afterwards, B, because it is declarative and reusable.
Solution 3
1. The sequence.
| Moment | Marta's transaction | Diego's transaction | DB |
|---|---|---|---|
| t1 | SELECT material → available = 1 |
1 | |
| t2 | SELECT material → available = 1 |
1 | |
| t3 | check 1 > 0 → OK |
1 | |
| t4 | check 1 > 0 → OK |
1 | |
| t5 | UPDATE available = 0 |
0 | |
| t6 | INSERT loan |
||
| t7 | COMMIT | 0 | |
| t8 | UPDATE available = 0 |
0 | |
| t9 | INSERT loan |
||
| t10 | COMMIT | 0, with 2 loans |
It is a lost update: Diego's UPDATE was computed from a stale value. With READ_COMMITTED, neither transaction sees anything anomalous.
2. Why synchronized is no use. synchronized synchronises threads within a single JVM. In a real deployment there are several BiblioTech instances behind a load balancer (12-06), and Marta and Diego may be served by different processes, on different machines. A monitor in JVM A locks nothing in JVM B. Besides, even with a single instance, serialising every loan of every material through one lock would be an unnecessary bottleneck. The coordination has to be where the shared data is: in the database.
3. Optimistic locking.
@Entity
public class Material {
@Version
private Long version; // the only addition to the model
public void lendOneCopy() { // logic INSIDE the entity
if (availableCopies <= 0) throw new NoCopiesAvailableException(isbn);
availableCopies--;
}
}@Service
public class LoanManager {
private static final Logger log = LoggerFactory.getLogger(LoanManager.class);
private final LoanManager self; // so the retry goes through the proxy (11-02)
@Transactional
public Loan lend(String isbn, String email) {
Material material = materialRepository.findByIsbn(isbn)
.orElseThrow(() -> new MaterialNotFoundException(isbn));
Employee employee = employeeRepository.findByEmail(email)
.orElseThrow(() -> new EmployeeNotFoundException(email));
material.lendOneCopy(); // validates and decrements
Loan loan = new Loan(material, employee, LocalDate.now(clock), 15);
return loanRepository.save(loan);
} // COMMIT: update materials ... where id=? and version=?
// Retry: every attempt is a NEW transaction
public Result<Loan> lendWithRetry(String isbn, String email) {
for (int attempt = 1; attempt <= 3; attempt++) {
try {
return Result.success(self.lend(isbn, email)); // goes through the proxy
} catch (OptimisticLockingFailureException e) {
log.warn("Conflict on {} (attempt {}/3)", isbn, attempt);
} catch (NoCopiesAvailableException e) {
return Result.error("There are no copies left of " + isbn);
}
}
return Result.error("The material is being modified; please try again.");
}
}What happens now at t8: Diego's UPDATE carries where id=1 and version=0, but the row already has version=1. It affects 0 rows, Hibernate throws OptimisticLockException and Spring translates it into OptimisticLockingFailureException. The retry reads again, sees available=0 and returns a correct business error.
Result<T> is 10-01's generic type, and the retry loop is your RetryProxy from 10-03 — which in a real project would be @Retryable(retryFor = OptimisticLockingFailureException.class, maxAttempts = 3).
4. Pessimistic locking.
public interface MaterialRepository extends JpaRepository<Material, Long> {
@Lock(LockModeType.PESSIMISTIC_WRITE)
@Query("SELECT m FROM Material m WHERE m.isbn = :isbn")
Optional<Material> findForLending(@Param("isbn") String isbn);
}@Transactional
public Loan lend(String isbn, String email) {
// SELECT ... FOR UPDATE: it locks the row until the COMMIT
Material material = materialRepository.findForLending(isbn)
.orElseThrow(() -> new MaterialNotFoundException(isbn));
// Diego waits here until Marta commits; then he reads available = 0
material.lendOneCopy(); // correctly throws NoCopiesAvailableException
return loanRepository.save(new Loan(material, employee, LocalDate.now(clock), 15));
}It is worth adding a maximum wait time so as not to block indefinitely:
@Lock(LockModeType.PESSIMISTIC_WRITE)
@QueryHints(@QueryHint(name = "jakarta.persistence.lock.timeout", value = "3000"))5. Comparison and recommendation.
| Criterion | Optimistic (@Version) |
Pessimistic (FOR UPDATE) |
|---|---|---|
| Cost without conflict | None | A lock on every operation |
| Concurrency allowed | High | Low: operations serialise |
| What happens on conflict | An exception and a retry | A wait |
| Deadlock risk | None | Yes, if several rows are locked |
| Risk of waiting indefinitely | No | Yes, without a timeout |
| Code complexity | Handling the exception | None apparent |
| Works across several instances | Yes | Yes |
| Good for | Infrequent conflicts | Frequent and expensive conflicts |
Recommendation: optimistic locking. In BiblioTech, the chance of two people lending the same copy in the same second is tiny, and pessimistic locking would make the thousands of operations that never clash pay the cost of serialisation. @Version costs one annotation, does not penalise the normal case and works the same with one instance as with twenty.
Pessimistic locking is reserved for operations where conflict is the norm and the lost work is expensive: for example, a nightly process that recalculates every fine and does not want to retry five thousand times.
Conclusion
BiblioTech's CSVs are gone. There is a real database underneath, and you know exactly what is going on inside it.
You understand the object-relational mismatch an ORM exists to manage: double identity, inheritance that does not exist in SQL, navigation that is free in Java and costs queries in SQL, granularity, collections and types. And you know what you gain and what you lose compared with plain JDBC —which you saw in twenty lines, with its column-by-column mapping, its type conversion and its resource management— with the professional rule that follows: ORM for the domain CRUD, direct SQL for reports and bulk processing, and both coexisting without conflict.
You can tell JPA from Hibernate: specification (jakarta.persistence) versus implementation, with the recommendation to program against the specification and the version warning that today causes more compilation errors than any other — jakarta, never javax.
You know how to map entities: @Entity, @Table, @Id, the @GeneratedValue strategies with the real reason SEQUENCE beats IDENTITY (it allows insert batching), @Column with its attributes, and the requirements JPA imposes —a no-argument constructor, nothing final— that explain why a record cannot be an entity. You know the absolute rule of @Enumerated(EnumType.STRING), because ORDINAL is the default and reordering an enum silently corrupts all the historical data. You map LocalDate and Instant natively with the same distinction from 10-05, you write converters with AttributeConverter for value objects such as Isbn, and you use @Embeddable to have small objects in Java without fragmenting the schema.
You model complete relationships with their four types, you know that the owning side is the one holding the foreign key and that touching only the inverse side persists nothing —hence the convenience methods—, and you have mastered the decision with the greatest performance impact: FetchType. You know the defaults are inconsistent, that @ManyToOne comes as EAGER and that the professional rule is to set LAZY explicitly everywhere and ask for what you need in the query. You know the two consequences: the LazyInitializationException, with its four ranked solutions and the advice to leave open-in-view at false so the problem shows up on your machine and not in production; and the N+1 problem, which turns one query into 501 and is invisible with five test rows, with its solutions —JOIN FETCH, @EntityGraph, @BatchSize and, the most efficient for reports, the DTO projection, where record types finally find their place—. And you know that switching to EAGER fixes nothing: it makes it worse.
You understand the persistence context, which is JPA's conceptual heart: a first-level cache that guarantees the same row is the same object, with an entity's four states and its transitions. And with it, automatic dirty checking, which surprises everybody: you modify a managed entity, you call nothing, and on commit an UPDATE appears. You know how it works (snapshot and comparison), its four consequences —including that an accidental change is persisted— and why readOnly = true switches it off and saves real work. You know flush, which reorders the operations and explains half of Hibernate's odd behaviour, and the difference between persist and merge that causes the "I ignored the return value" bug.
You write queries: JPQL over entities and attributes, named queries validated at start-up, the Criteria API for dynamic cases, native SQL when it is the right call, and projections into record types in a single line. And you have internalised the rule that admits no exceptions: every external value goes in as a parameter, because concatenating is SQL injection —with an allow-list as the only way out when the dynamic part is a column name—.
You handle transactions with a real database: ACID, @Transactional with its two rules inherited from 11-02, readOnly, propagation with the legitimate case for REQUIRES_NEW and its cost in connections, and isolation with the three phenomena explained over BiblioTech and the advice hardly ever to touch it. And you solve the disputed-copy case with optimistic locking using @Version, understanding why module 8's synchronized is no use when there are several instances: the coordination has to be where the shared data is. With a retry over Result<T>, which is your generic from 10-01 and your RetryProxy from 10-03 turned into production code.
You know the first- and second-level caches and when the second is worth it —data read a lot and changed little, measured first—; and you know why ddl-auto=update is never used in production: it does not delete, does not version, does not migrate data and nobody reviews the ALTER TABLE it is about to run. validate plus versioned migrations with Flyway, which arrive in 12-06.
And you have watched Spring Data JPA turn eighty lines of repository into an empty interface, with methods derived from names, @Query for what does not fit in a name, and pagination. With the mechanism laid bare: Proxy.newProxyInstance over an interface, exactly the one from 10-03, analysing the method's metadata and generating the query.
BiblioTech now has real entities, relationships with referential integrity, ACID transactions, optimistic locking and indexed queries over H2. It starts with mvn spring-boot:run and its schema is created on its own.
And it still has not a single automated test.
That is now, by a long way, the most serious debt. You have just carried out an enormous migration: you have changed the domain model, turned record types into classes, replaced the entire persistence layer and introduced optimistic concurrency with retries. How do you know the fine calculation still gives the same result? That the limit of three loans per employee is still respected? That a loan falling due on a Sunday still triggers a notice on the Monday? The only honest answer today is: by starting the application and looking.
And there is something worse. That injectable Clock you introduced in 10-05, which went through Spring in 11-02 as a @Bean and which in this lesson has appeared in every LocalDate.now(clock) and every Instant.now(clock), exists exclusively so the code can be tested without waiting fifteen days. It has spent two modules waiting for somebody to make use of it.
In the next lesson it gets used. You will see what an automated test is and why you are already paying the cost of not having them; you will meet the testing pyramid and the whole of JUnit 5, from @Test and the Arrange-Act-Assert pattern to parameterised tests that verify twelve cases of the fine calculation in six lines; you will compare JUnit's assertions with AssertJ's; and you will discover what makes a design testable — and why everything you have done in the last two modules (constructor injection, interfaces at the boundaries, the injected Clock, the annotation-free domain) was pointing exactly there.
BiblioTech is about to get its first safety net.
Java Programming Course
Module 1: Introduction to Java
- Introduction to Java
- Setting Up the Development Environment
- Basic Syntax and Structure
- Variables and Data Types
- Operators
- Console Input and Output
- Your First Complete Program: BiblioTech
Module 2: Control Flow
- Conditional Statements
- Loops
- Switch Statements
- Break and Continue
- Debugging and Execution Traces
- Project: The BiblioTech Interactive Menu
Module 3: Object-Oriented Programming
- Introduction to OOP
- Classes and Objects
- Methods
- Constructors
- Inheritance
- Polymorphism
- Encapsulation
- Abstraction
- The Object Class: equals, hashCode and toString
Module 4: Advanced Object-Oriented Programming
- Interfaces
- Abstract Classes
- Inner Classes
- Anonymous Classes
- Lambda Expressions
- Functional Interfaces and Method References
- Enums and Records
Module 5: Data Structures and Collections
- Arrays
- The Collections Framework
- ArrayList
- LinkedList
- HashMap
- HashSet
- Queue and Deque
- Stack
- Sorting and Searching Collections
Module 6: Exception Handling
- Introduction to Exceptions
- The Try-Catch Block
- Throw and Throws
- Custom Exceptions
- The Finally Block
- Try-with-resources and AutoCloseable
- Error Handling Strategies and Logging
Module 7: File Input/Output
- Reading Files
- Writing Files
- File Streams
- BufferedReader and BufferedWriter
- Serialization
- The NIO.2 API: Path and Files
- Interchange Formats: CSV and Properties
Module 8: Multithreading and Concurrency
- Introduction to Multithreading
- Creating Threads
- Thread Lifecycle
- Synchronization
- Concurrency Utilities
- Concurrent Collections and Atomic Variables
- Asynchronous Tasks with CompletableFuture
Module 9: Networking
- Introduction to Networking
- Sockets
- ServerSocket
- DatagramSocket and DatagramPacket
- URL and HttpURLConnection
- The Modern HTTP Client
Module 10: Advanced Topics
- Generics
- Annotations
- Reflection
- Java 8 Features: Streams and Optional
- Dates and Times with java.time
- Java 9 and Beyond
- Memory, Garbage Collection and Performance
Module 11: Java Frameworks and Libraries
- Introduction to Java Frameworks
- Spring Framework
- Hibernate
- JUnit
- Maven
- Advanced Testing with Mockito
- Essential Ecosystem Libraries
