An inventory service uses SELECT ... FOR UPDATE on the stock row before every reservation. It is correct, it has never oversold, and under normal traffic nobody notices it.
A flash sale starts. Two thousand concurrent requests all want the same SKU. Every one of them queues behind a row lock that is held for the duration of a transaction which also makes a call to a pricing service.
Throughput on that SKU collapses to roughly one reservation per transaction duration. Connections pile up waiting, the pool exhausts, and requests for entirely unrelated products start failing because there are no connections left for them.
The locking was right. The transaction it was wrapped around was too long, and the strategy stopped fitting the moment contention arrived.
Detect versus prevent
Both approaches solve the lost update problem. They disagree about when to find out.
Pessimistic locking takes the lock before reading. Nobody else can touch the row until you commit. Conflicts cannot happen because contending transactions wait.
Optimistic locking takes nothing. It reads, works, and at write time checks whether the row changed since the read. If it did, the transaction fails and the caller decides what to do.
flowchart TB
subgraph pess["Pessimistic"]
P1["lock row"] --> P2["read"] --> P3["compute"] --> P4["write"] --> P5["commit, release"]
P6["Other writers: waiting"] -.-> P1
end
subgraph opt["Optimistic"]
O1["read row + version"] --> O2["compute"] --> O3["write WHERE version = seen"]
O3 --> O4{"rows affected?"}
O4 -->|"1"| O5["commit"]
O4 -->|"0"| O6["conflict, all work wasted"]
end
style P6 stroke:#f59e0b,stroke-width:2px,color:#fff
style O6 stroke:#ef4444,stroke-width:2px,color:#fff
Optimistic locking in its simplest form needs no framework at all.
UPDATE products
SET stock = stock - 1, version = version + 1
WHERE id = 42 AND version = 7;
If that affects zero rows, somebody else got there first. The check and the write are one atomic statement, so there is no window between them.
In JPA the version column is declarative and the failure arrives as an exception:
@Entity
public class Product {
@Id private Long id;
@Version private Long version; // JPA adds the predicate and increments it
private int stock;
}
The crossover is contention, and it is measurable
The decision is usually presented as a philosophy question. It is arithmetic.
Under optimistic locking, a conflict costs everything the transaction did before the write, and then costs it again on retry. If conflicts are rare, that expected cost is close to zero and you get the benefit of never blocking.
As the conflict rate climbs, two things compound. Wasted work grows linearly with failures, and retries add load, which increases concurrency on the contended row, which increases the conflict rate. That feedback is why optimistic locking degrades sharply rather than gradually.
| Conflict rate | Reasonable choice |
|---|---|
| under 1 percent | optimistic, clearly |
| 1 to 5 percent | optimistic, with bounded retries and jitter |
| 5 to 10 percent | measure both; depends on transaction cost |
| above 10 percent | pessimistic, or redesign the contention away |
Those bands are rules of thumb rather than thresholds, and the point is that the number is knowable. Most ORMs surface optimistic failures as a distinct exception type, so counting them as a proportion of writes is a metric you can have today. On the pessimistic side, log_lock_waits in Postgres tells you how long transactions spend waiting.
Deciding once at design time and never revisiting is the actual mistake, because contention is a property of traffic and traffic changes. A strategy chosen for a steady workload meets a flash sale eventually.
Retrying is where the correctness bugs live
An optimistic failure means the row changed. Retrying blindly means re-reading and reapplying, which is only safe if the operation is a function of current state.
// Safe: the new value derives from what was just read.
@Retryable(retryFor = OptimisticLockingFailureException.class,
maxAttempts = 3,
backoff = @Backoff(delay = 20, multiplier = 2, random = true))
public void decrementStock(Long productId) {
Product p = repo.findById(productId).orElseThrow();
if (p.getStock() < 1) throw new OutOfStockException();
p.setStock(p.getStock() - 1);
repo.save(p);
}
That is fine because the decrement is recomputed from the fresh read.
What is not fine is retrying a user’s edit. If somebody opened a form, changed the description, and the save fails because a colleague changed the price, silently re-reading and reapplying the user’s stale description overwrites the colleague’s work. The version check existed precisely to catch that, and the retry throws away the answer.
For human edits, the conflict is information rather than a transient failure. Surfacing it, showing what changed, and letting the user decide is the correct behaviour, and it is the case where neither locking strategy is really the question.
The jitter on that backoff matters for the same reason it always does. Two transactions that conflicted and both retry after exactly 20 milliseconds will conflict again, which is the alignment problem that shows up everywhere a fixed interval exists.
When neither is the answer
Sometimes contention is high and the right move is to remove the contended row rather than to lock it better.
Atomic operations sidestep the read-modify-write entirely:
-- No version, no lock, no retry. The conflict cannot happen because
-- there is no gap between reading and writing.
UPDATE products SET stock = stock - 1
WHERE id = 42 AND stock >= 1;
Zero rows affected means insufficient stock. One means it worked. This is strictly better than either locking strategy when the operation can be expressed as a single statement, and a surprising number can be.
Sharding the counter works when even that is too contended. A stock count split across ten rows means writers contend on one tenth as much, at the cost of reads having to sum them, which is the same shape as splitting a hot key so it stops landing on one node.
Queueing serialises deliberately. Reservations become messages processed in order by a single consumer per SKU, which removes concurrency on that row entirely and converts a contention problem into a throughput one you can reason about.
The thing that actually caused the outage
Going back to the flash sale: the lock was held while calling a pricing service. That is the real defect, and it makes every locking strategy look bad.
A pessimistic lock held across a network call means the lock duration is the network latency, and everything queues behind it. An optimistic transaction spanning a network call has a much wider window in which to conflict, so the conflict rate rises for reasons unrelated to write volume.
Both strategies assume transactions are short. A transaction that waits on anything external is not short, and that is the same failure as holding a pooled connection across work that is not database work.
Fetch what you need, close the transaction, do the external call, then open a second short transaction to apply the result with a version check. That restructuring usually helps more than switching strategies does.
// SPONSORSHIP
If this research saved you time or improved your architecture, consider sponsoring my work on GitHub. All sponsorships go directly toward infrastructure and further technical research.
[ Become a Sponsor ]