It is 11:59:58. A flash sale is about to open, and one lakh people have the same phone on screen, thumbs hovering over "Buy". The sale starts at 12:00:00, and there are exactly 500 units in stock.
At 12:00:00 every one of those taps turns into an HTTP request, and they land on your service within roughly the same second. Your job is simple to state and surprisingly hard to get right: sell exactly 500 phones. Not 499, because then you left money on the table. Not 501, because then you have oversold: someone paid for a phone that does not exist, and now support, refunds and a very angry customer are your problem.
With Big Billion Days and every other festive sale going live this month, I wanted to write down how I would actually build this. A quick note before we start: I don't know how Flipkart or Amazon build their sale systems, and I won't pretend to. This is how I would design it with Java, Spring Boot, Postgres and Redis, and why.
How do you prevent overselling in a flash sale?
The short answer: make the stock check and the stock decrement one atomic step, in a place every server instance shares. In the database, that is a single conditional UPDATE ... WHERE stock > 0. Under heavy load, put an atomic Redis Lua script in front of it as a fast gate and keep the database update as the final word. Then treat each sale as a time-limited reservation until the payment succeeds.
The rest of this post shows why the obvious code fails, five ways to fix it, and how to handle payments that never come back.
The code most of us would write
Here is the version that shows up in most codebases, most tutorials and, honestly, most first drafts of mine:
@Transactionalpublic Order buy(long productId, long userId) { Product product = productRepository.findById(productId) .orElseThrow(ProductNotFoundException::new); if (product.getStock() <= 0) { throw new SoldOutException(productId); } product.setStock(product.getStock() - 1); return orderRepository.save(new Order(productId, userId));}It reads well. It is wrapped in @Transactional. It checks the stock before selling. It passes every unit test you will write for it, because unit tests call it one request at a time.
And under real concurrency, it oversells.
Why it oversells: two buyers, one phone
Forget one lakh buyers for a moment. Take two, Asha and Ben, and one phone left in stock.
-
Asha's request reads the product. Stock is
1. -
Ben's request, a couple of milliseconds later, reads the same row. Stock is still
1, because Asha has not written anything yet. -
Asha's request checks
1 > 0, sets stock to0, creates an order and commits. -
Ben's request checks
1 > 0(it is still holding the value it read), sets stock to0, creates an order and commits.
Two orders. One phone. And here is the nasty part: the stock column says 0, exactly as you would expect after selling out. Nothing looks wrong until someone counts the rows in the orders table. This is a classic lost update, and it is a read-check-write race.
@Transactionaldoes not prevent this. A transaction groups your writes so they commit or roll back together. It does not stop two transactions from reading the same row at the same time. Postgres, MySQL and most other databases default to the read committed isolation level, and at that level both transactions are allowed to read the committed value1.
Now scale that window from two buyers to a lakh. Every request that reads the row before the previous writer commits is a potential extra sale. The more concurrent the traffic, the more you oversell. Flash sales are the worst case by design.
The first instinct is to add synchronized to the method. That works on your laptop and fails in production the moment you run a second instance of the service, because a Java lock only exists inside one JVM. The fix has to live where all instances meet: the database, or something like Redis.
Here are five ways to do that, from the simplest to the most scalable.
Fix 1: Lock the row with SELECT ... FOR UPDATE
The most direct fix is to tell the database "I'm about to change this row, nobody else may touch it until I'm done". That is a pessimistic lock:
public interface ProductRepository extends JpaRepository<Product, Long> { @Lock(LockModeType.PESSIMISTIC_WRITE) @Query("select p from Product p where p.id = :id") Optional<Product> findByIdForUpdate(@Param("id") long id);}Swap findById for findByIdForUpdate in the service and Hibernate issues SELECT ... FOR UPDATE, which takes a row-level lock. Now when Ben's request tries to read the row, it waits until Asha's transaction commits, and then it reads the new value, 0.
The good: correct, and a two-line change.
The bad: every buyer of that phone now queues on one row lock, one at a time. Each waiting request holds a database connection while it waits, so your connection pool drains in seconds, and requests for completely unrelated products start timing out too. If you go this way, set a lock timeout in Postgres (lock_timeout) so waiting buyers fail fast instead of piling up.
Good for: admin tools, low-traffic inventory, anything where correctness matters and contention is rare. Not good for: the 12:00:00 spike.
Fix 2: Let one UPDATE do the check and the write
The bug exists because the check (in Java) and the write (in the database) are two separate steps. So remove the gap. Ask the database to decrement the stock only if there is stock left, in a single statement:
@Modifying@Query("update Product p set p.stock = p.stock - 1 where p.id = :id and p.stock > 0")int decrementIfAvailable(@Param("id") long id);@Transactionalpublic Order buy(long productId, long userId) { if (productRepository.decrementIfAvailable(productId) == 0) { throw new SoldOutException(productId); } return orderRepository.save(new Order(productId, userId));}The return value is the number of rows updated. 1 means you got a unit. 0 means it was already gone.
Why this is safe: when two UPDATEs hit the same row, the second one waits for the first to commit. In Postgres it then re-checks the condition in the WHERE clause against the new version of the row. If the first buyer took the last unit, stock > 0 is now false and the second update touches zero rows. No Java-side check, no window.
Add a database constraint as a seatbelt, so even a future bug cannot push stock below zero:
ALTER TABLE product ADD CONSTRAINT stock_not_negative CHECK (stock >= 0);The good: correct, simple, and the row is locked only for the duration of one tiny statement instead of a whole read-think-write cycle.
The bad: it is still one hot row. Every buyer still serializes on it, just for a much shorter time. For most products and most sales, this is plenty. For one hero phone with a lakh people hitting it in the same second, the database is still doing all the work.
If I had to pick one fix for a normal e-commerce app, it would be this one.
Fix 3: Optimistic locking with @Version (and why it is wrong here)
Optimistic locking is the fix people reach for next, so it is worth explaining why I would not use it for a flash sale.
@Entitypublic class Product { @Id private Long id; private int stock; @Version private long version;}With a @Version column, Hibernate adds WHERE version = ? to every update. If someone else changed the row since you read it, your update matches zero rows and you get an ObjectOptimisticLockingFailureException. No oversell. Then you retry:
@Retryable(retryFor = ObjectOptimisticLockingFailureException.class, maxAttempts = 3, backoff = @Backoff(delay = 20))public Order buyWithRetry(long productId, long userId) { return buyService.buy(productId, userId); // @Transactional lives inside}(Note the retry has to sit outside the transaction. Retrying inside a transaction that has already failed just fails again.)
Optimistic locking assumes conflicts are rare. In a flash sale, conflicts are the whole point. Of a thousand buyers who read the row at the same moment, one wins and nine hundred and ninety-nine fail and retry, then collide with each other again. You turn a traffic spike into a retry storm, and most buyers see an error that has nothing to do with stock actually running out.
Good for: two people editing the same profile or document. Not good for: a thousand people buying the same phone.
Fix 4: Put a Redis gate in front of the database
Fixes 1 to 3 all make the database absorb the full lakh of requests. But only 500 of them can ever succeed. What if the other 99,500 never reached the database at all?
Redis is single-threaded for command execution, so a Lua script runs atomically: no other command can run in the middle of it. That makes it a very fast gate:
-- KEYS[1] = stock key, KEYS[2] = set of buyers, ARGV[1] = userIdif redis.call('SISMEMBER', KEYS[2], ARGV[1]) == 1 then return -1 -- this user already bought oneendlocal stock = tonumber(redis.call('GET', KEYS[1]) or '0')if stock <= 0 then return 0 -- sold outendredis.call('DECR', KEYS[1])redis.call('SADD', KEYS[2], ARGV[1])return 1 -- you got oneprivate static final RedisScript<Long> TRY_BUY = RedisScript.of(new ClassPathResource("lua/try_buy.lua"), Long.class);public Order buy(long productId, long userId) { List<String> keys = List.of("stock:{" + productId + "}", "buyers:{" + productId + "}"); Long result = redis.execute(TRY_BUY, keys, String.valueOf(userId)); if (result == null || result == 0) throw new SoldOutException(productId); if (result == -1) throw new AlreadyPurchasedException(productId); return orderService.createOrder(productId, userId); // the database write}Two details that matter:
-
The curly braces in the keys are a Redis Cluster hash tag. They force both keys onto the same node, which a Lua script needs because it can only touch keys on one node.
-
The one-per-customer check lives in the same script, so it is atomic too. Checking it in Java first would bring the race right back.
Before the sale, you load the stock into Redis: SET stock:{42} 500. After that, Redis answers "sold out" to 99,500 people in microseconds and only 500 requests ever touch Postgres.
The catch: now stock lives in two places, and they can drift. If the database write fails after Redis said yes, you must give the unit back (INCR the stock and SREM the user). And Redis replication is asynchronous, so a failover at the wrong moment can lose a few decrements. That is why I keep Fix 2 underneath: Redis is the fast gate, the atomic database update is the final word. If Redis ever lets one too many through, the database still says no.
Fix 5: Don't race at all. Queue.
Every fix so far handles the race. The last one removes it.
Instead of processing "buy" requests the moment they arrive, accept them as intents and put them on a queue. With Kafka, use the product ID as the message key:
kafkaTemplate.send("buy-intents", String.valueOf(productId), new BuyIntent(intentId, productId, userId));Kafka sends every message with the same key to the same partition, and one consumer reads a partition in order. So all intents for this phone are processed one at a time, by one consumer, in arrival order. No two of them ever run at the same time, so there is nothing to race. The consumer can even keep the remaining stock in memory and only write winners to the database.
The trade-off is the user experience. The buyer no longer gets an instant "Order placed". They get "You're in the queue", and the result arrives a moment later over a WebSocket, a push notification or a polling endpoint. For a big sale, many users find that fairer, because it is literally first come, first served. You also need the consumer to be idempotent (use intentId to ignore duplicates), because Kafka delivers at least once.
The five fixes side by side
|
Approach |
Oversell-safe? |
Under a 1-lakh spike |
Complexity |
I'd use it for |
|---|---|---|---|---|
|
Read, check, write (naive) |
No |
Oversells |
Lowest |
Nothing that sells stock |
|
SELECT ... FOR UPDATE |
Yes |
Queues on one lock, drains the connection pool |
Low |
Low-traffic inventory, admin flows |
|
Atomic conditional UPDATE |
Yes |
Short lock per request, still one hot row |
Low |
The default for most apps |
|
Optimistic @Version |
Yes |
Retry storm, buyers see errors |
Medium |
Low-contention edits |
|
Redis Lua gate + atomic UPDATE |
Yes |
Only stock-count requests reach the DB |
Medium |
Hot items in a flash sale |
|
Queue (Kafka, keyed by product) |
Yes, by design |
Absorbs any spike, result is async |
High |
Very large drops, fairness matters |
The part most write-ups skip: the payment never comes back
Everything above assumes "buy" means "sold". In real life it means "reserved". The buyer still has to pay, and some of them will close the tab, run out of balance, or sit on a UPI screen until it times out.
If you decrement stock at "Buy" and never give it back, you will undersell: the sale shows "Sold out" while 40 units sit in abandoned carts. So stock needs a lifecycle:
A reservation row carries a status (HELD, CONFIRMED, RELEASED) and an expires_at, say ten minutes out. Buying creates a HELD reservation along with the decrement. A successful payment flips it to CONFIRMED. And a scheduled job sweeps up the ones that expired and puts their stock back:
WITH expired AS ( UPDATE reservation SET status = 'RELEASED' WHERE id IN ( SELECT id FROM reservation WHERE status = 'HELD' AND expires_at < now() ORDER BY expires_at LIMIT 500 FOR UPDATE SKIP LOCKED ) RETURNING product_id)UPDATE product pSET stock = p.stock + e.releasedFROM (SELECT product_id, count(*) AS released FROM expired GROUP BY product_id) eWHERE p.id = e.product_id;FOR UPDATE SKIP LOCKED skips rows that another transaction has already locked, which lets you run this sweeper on several instances at once without two of them releasing the same reservation.
Then comes the edge case that bites everyone eventually: the payment succeeds after the reservation was released. The buyer was slow, the sweeper ran, the unit went back on sale, and then the payment gateway calls your webhook with "success". So the confirm step must be conditional too:
UPDATE reservation SET status = 'CONFIRMED'WHERE id = :reservationId AND status = 'HELD';If that touches zero rows, the reservation is gone. Try to take a fresh unit with the same atomic decrement, and if there are none left, refund automatically. Make the webhook handler idempotent too, because payment gateways retry callbacks.
And if you use the Redis gate, every released unit must go back into Redis as well, or Redis keeps saying "sold out" while the database has stock.
What I would actually ship
-
A Redis Lua gate that answers the 99%+ of requests that cannot win, enforces one unit per customer, and never lets the DB see more than roughly the stock count.
-
An atomic conditional UPDATE in Postgres as the final word, plus a
CHECK (stock >= 0)constraint as the seatbelt. -
Reservations with a 10-minute expiry, a
SKIP LOCKEDsweeper, and a conditional confirm that handles late payments. -
Idempotency everywhere a retry can happen: the buy request, the payment webhook, the Kafka consumer.
-
A queue only if the drop is big enough that fairness and spike absorption matter more than an instant answer.
Notice what is not on that list: synchronized, a bigger database, or hoping traffic is lower than last year.
Prove it yourself in 30 lines
You don't need a load-testing tool to see the bug. A plain Spring Boot test with virtual threads will do it. (If booting a Spring context for tests feels painfully slow, here is how I cut our Spring Boot startup time from 45s to 8s.) Seed a product with stock 500 and fire 5,000 buyers at it at the same instant:
@SpringBootTestclass OversellTest { @Autowired BuyService buyService; @Autowired OrderRepository orderRepository; @Test void neverSellsMoreThanStock() throws Exception { long productId = 1L; // seeded with stock = 500 int buyers = 5_000; CountDownLatch startGun = new CountDownLatch(1); try (var executor = Executors.newVirtualThreadPerTaskExecutor()) { for (int i = 0; i < buyers; i++) { long userId = i; executor.submit(() -> { startGun.await(); // everyone waits here... try { buyService.buy(productId, userId); } catch (SoldOutException ignored) { } return null; }); } startGun.countDown(); // ...and goes at once } // close() waits for all tasks assertThat(orderRepository.countByProductId(productId)).isLessThanOrEqualTo(500); }}Run it against the naive version and it fails: you will see more than 500 orders, and the exact count changes from run to run, which is the signature of a race. Swap in Fix 2 and it passes every time. I would keep a test like this in the codebase for any code path that touches stock, seats, coupons or anything else that can run out.
The takeaway
The overselling bug is not exotic. It is the most natural code you can write, it passes every test that runs one request at a time, and it only shows up on the one day you cannot afford it. The fix is not more cleverness in Java. It is moving the check and the write into one atomic step, in the place every instance shares.
If you are working on anything that sells limited stock this festive season, I'd love to know: which of these approaches does your system use, and what broke the first time it met real traffic?

Discussion