500 Phones, 1 Lakh Buyers, 1 Second: Building a Flash Sale That Never Oversells

500 Phones, 1 Lakh Buyers, 1 Second: Building a Flash Sale That Never Oversells

14 min read

It is 11:59:58. A flash sale is about to open, and one lakh people have the same phone on screen, thumbs hovering over "Buy". The sale starts at 12:00:00, and there are exactly 500 units in stock.

At 12:00:00 every one of those taps turns into an HTTP request, and they land on your service within roughly the same second. Your job is simple to state and surprisingly hard to get right: sell exactly 500 phones. Not 499, because then you left money on the table. Not 501, because then you have oversold: someone paid for a phone that does not exist, and now support, refunds and a very angry customer are your problem.

With Big Billion Days and every other festive sale going live this month, I wanted to write down how I would actually build this. A quick note before we start: I don't know how Flipkart or Amazon build their sale systems, and I won't pretend to. This is how I would design it with Java, Spring Boot, Postgres and Redis, and why.

How do you prevent overselling in a flash sale?

The short answer: make the stock check and the stock decrement one atomic step, in a place every server instance shares. In the database, that is a single conditional UPDATE ... WHERE stock > 0. Under heavy load, put an atomic Redis Lua script in front of it as a fast gate and keep the database update as the final word. Then treat each sale as a time-limited reservation until the payment succeeds.

The rest of this post shows why the obvious code fails, five ways to fix it, and how to handle payments that never come back.

Thousands of glowing buy requests streaming toward one small crate of phones

The code most of us would write

Here is the version that shows up in most codebases, most tutorials and, honestly, most first drafts of mine:

Java
@Transactionalpublic Order buy(long productId, long userId) {    Product product = productRepository.findById(productId)            .orElseThrow(ProductNotFoundException::new);    if (product.getStock() <= 0) {        throw new SoldOutException(productId);    }    product.setStock(product.getStock() - 1);    return orderRepository.save(new Order(productId, userId));}

It reads well. It is wrapped in @Transactional. It checks the stock before selling. It passes every unit test you will write for it, because unit tests call it one request at a time.

And under real concurrency, it oversells.

Why it oversells: two buyers, one phone

Forget one lakh buyers for a moment. Take two, Asha and Ben, and one phone left in stock.

Timeline: Asha and Ben both read stock 1, both sell, and two orders are created for one phone
  1. Asha's request reads the product. Stock is 1.

  2. Ben's request, a couple of milliseconds later, reads the same row. Stock is still 1, because Asha has not written anything yet.

  3. Asha's request checks 1 > 0, sets stock to 0, creates an order and commits.

  4. Ben's request checks 1 > 0 (it is still holding the value it read), sets stock to 0, creates an order and commits.

Two orders. One phone. And here is the nasty part: the stock column says 0, exactly as you would expect after selling out. Nothing looks wrong until someone counts the rows in the orders table. This is a classic lost update, and it is a read-check-write race.

@Transactional does not prevent this. A transaction groups your writes so they commit or roll back together. It does not stop two transactions from reading the same row at the same time. Postgres, MySQL and most other databases default to the read committed isolation level, and at that level both transactions are allowed to read the committed value 1.

Now scale that window from two buyers to a lakh. Every request that reads the row before the previous writer commits is a potential extra sale. The more concurrent the traffic, the more you oversell. Flash sales are the worst case by design.

The first instinct is to add synchronized to the method. That works on your laptop and fails in production the moment you run a second instance of the service, because a Java lock only exists inside one JVM. The fix has to live where all instances meet: the database, or something like Redis.

Here are five ways to do that, from the simplest to the most scalable.

Fix 1: Lock the row with SELECT ... FOR UPDATE

The most direct fix is to tell the database "I'm about to change this row, nobody else may touch it until I'm done". That is a pessimistic lock:

Java
public interface ProductRepository extends JpaRepository<Product, Long> {    @Lock(LockModeType.PESSIMISTIC_WRITE)    @Query("select p from Product p where p.id = :id")    Optional<Product> findByIdForUpdate(@Param("id") long id);}

Swap findById for findByIdForUpdate in the service and Hibernate issues SELECT ... FOR UPDATE, which takes a row-level lock. Now when Ben's request tries to read the row, it waits until Asha's transaction commits, and then it reads the new value, 0.

The good: correct, and a two-line change.

The bad: every buyer of that phone now queues on one row lock, one at a time. Each waiting request holds a database connection while it waits, so your connection pool drains in seconds, and requests for completely unrelated products start timing out too. If you go this way, set a lock timeout in Postgres (lock_timeout) so waiting buyers fail fast instead of piling up.

Good for: admin tools, low-traffic inventory, anything where correctness matters and contention is rare. Not good for: the 12:00:00 spike.

Fix 2: Let one UPDATE do the check and the write

The bug exists because the check (in Java) and the write (in the database) are two separate steps. So remove the gap. Ask the database to decrement the stock only if there is stock left, in a single statement:

Java
@Modifying@Query("update Product p set p.stock = p.stock - 1 where p.id = :id and p.stock > 0")int decrementIfAvailable(@Param("id") long id);
Java
@Transactionalpublic Order buy(long productId, long userId) {    if (productRepository.decrementIfAvailable(productId) == 0) {        throw new SoldOutException(productId);    }    return orderRepository.save(new Order(productId, userId));}

The return value is the number of rows updated. 1 means you got a unit. 0 means it was already gone.

Why this is safe: when two UPDATEs hit the same row, the second one waits for the first to commit. In Postgres it then re-checks the condition in the WHERE clause against the new version of the row. If the first buyer took the last unit, stock > 0 is now false and the second update touches zero rows. No Java-side check, no window.

Add a database constraint as a seatbelt, so even a future bug cannot push stock below zero:

SQL
ALTER TABLE product ADD CONSTRAINT stock_not_negative CHECK (stock >= 0);

The good: correct, simple, and the row is locked only for the duration of one tiny statement instead of a whole read-think-write cycle.

The bad: it is still one hot row. Every buyer still serializes on it, just for a much shorter time. For most products and most sales, this is plenty. For one hero phone with a lakh people hitting it in the same second, the database is still doing all the work.

If I had to pick one fix for a normal e-commerce app, it would be this one.

Fix 3: Optimistic locking with @Version (and why it is wrong here)

Optimistic locking is the fix people reach for next, so it is worth explaining why I would not use it for a flash sale.

Java
@Entitypublic class Product {    @Id    private Long id;    private int stock;    @Version    private long version;}

With a @Version column, Hibernate adds WHERE version = ? to every update. If someone else changed the row since you read it, your update matches zero rows and you get an ObjectOptimisticLockingFailureException. No oversell. Then you retry:

Java
@Retryable(retryFor = ObjectOptimisticLockingFailureException.class,           maxAttempts = 3, backoff = @Backoff(delay = 20))public Order buyWithRetry(long productId, long userId) {    return buyService.buy(productId, userId); // @Transactional lives inside}

(Note the retry has to sit outside the transaction. Retrying inside a transaction that has already failed just fails again.)

Optimistic locking assumes conflicts are rare. In a flash sale, conflicts are the whole point. Of a thousand buyers who read the row at the same moment, one wins and nine hundred and ninety-nine fail and retry, then collide with each other again. You turn a traffic spike into a retry storm, and most buyers see an error that has nothing to do with stock actually running out.

Good for: two people editing the same profile or document. Not good for: a thousand people buying the same phone.

Fix 4: Put a Redis gate in front of the database

Fixes 1 to 3 all make the database absorb the full lakh of requests. But only 500 of them can ever succeed. What if the other 99,500 never reached the database at all?

Redis is single-threaded for command execution, so a Lua script runs atomically: no other command can run in the middle of it. That makes it a very fast gate:

Lua
-- KEYS[1] = stock key, KEYS[2] = set of buyers, ARGV[1] = userIdif redis.call('SISMEMBER', KEYS[2], ARGV[1]) == 1 then  return -1                      -- this user already bought oneendlocal stock = tonumber(redis.call('GET', KEYS[1]) or '0')if stock <= 0 then  return 0                       -- sold outendredis.call('DECR', KEYS[1])redis.call('SADD', KEYS[2], ARGV[1])return 1                         -- you got one
Java
private static final RedisScript<Long> TRY_BUY =        RedisScript.of(new ClassPathResource("lua/try_buy.lua"), Long.class);public Order buy(long productId, long userId) {    List<String> keys = List.of("stock:{" + productId + "}", "buyers:{" + productId + "}");    Long result = redis.execute(TRY_BUY, keys, String.valueOf(userId));    if (result == null || result == 0) throw new SoldOutException(productId);    if (result == -1) throw new AlreadyPurchasedException(productId);    return orderService.createOrder(productId, userId); // the database write}

Two details that matter:

  • The curly braces in the keys are a Redis Cluster hash tag. They force both keys onto the same node, which a Lua script needs because it can only touch keys on one node.

  • The one-per-customer check lives in the same script, so it is atomic too. Checking it in Java first would bring the race right back.

Before the sale, you load the stock into Redis: SET stock:{42} 500. After that, Redis answers "sold out" to 99,500 people in microseconds and only 500 requests ever touch Postgres.

The catch: now stock lives in two places, and they can drift. If the database write fails after Redis said yes, you must give the unit back (INCR the stock and SREM the user). And Redis replication is asynchronous, so a failover at the wrong moment can lose a few decrements. That is why I keep Fix 2 underneath: Redis is the fast gate, the atomic database update is the final word. If Redis ever lets one too many through, the database still says no.

Fix 5: Don't race at all. Queue.

Every fix so far handles the race. The last one removes it.

Instead of processing "buy" requests the moment they arrive, accept them as intents and put them on a queue. With Kafka, use the product ID as the message key:

Java
kafkaTemplate.send("buy-intents", String.valueOf(productId),        new BuyIntent(intentId, productId, userId));

Kafka sends every message with the same key to the same partition, and one consumer reads a partition in order. So all intents for this phone are processed one at a time, by one consumer, in arrival order. No two of them ever run at the same time, so there is nothing to race. The consumer can even keep the remaining stock in memory and only write winners to the database.

A crowd funnelling into a single-file line through one turnstile, first come first served

The trade-off is the user experience. The buyer no longer gets an instant "Order placed". They get "You're in the queue", and the result arrives a moment later over a WebSocket, a push notification or a polling endpoint. For a big sale, many users find that fairer, because it is literally first come, first served. You also need the consumer to be idempotent (use intentId to ignore duplicates), because Kafka delivers at least once.

The five fixes side by side

Approach

Oversell-safe?

Under a 1-lakh spike

Complexity

I'd use it for

Read, check, write (naive)

No

Oversells

Lowest

Nothing that sells stock

SELECT ... FOR UPDATE

Yes

Queues on one lock, drains the connection pool

Low

Low-traffic inventory, admin flows

Atomic conditional UPDATE

Yes

Short lock per request, still one hot row

Low

The default for most apps

Optimistic @Version

Yes

Retry storm, buyers see errors

Medium

Low-contention edits

Redis Lua gate + atomic UPDATE

Yes

Only stock-count requests reach the DB

Medium

Hot items in a flash sale

Queue (Kafka, keyed by product)

Yes, by design

Absorbs any spike, result is async

High

Very large drops, fairness matters

The part most write-ups skip: the payment never comes back

Everything above assumes "buy" means "sold". In real life it means "reserved". The buyer still has to pay, and some of them will close the tab, run out of balance, or sit on a UPI screen until it times out.

If you decrement stock at "Buy" and never give it back, you will undersell: the sale shows "Sold out" while 40 units sit in abandoned carts. So stock needs a lifecycle:

Reservation lifecycle: in stock, held for 10 minutes, then confirmed or released back to stock

A reservation row carries a status (HELD, CONFIRMED, RELEASED) and an expires_at, say ten minutes out. Buying creates a HELD reservation along with the decrement. A successful payment flips it to CONFIRMED. And a scheduled job sweeps up the ones that expired and puts their stock back:

SQL
WITH expired AS (    UPDATE reservation SET status = 'RELEASED'    WHERE id IN (        SELECT id FROM reservation        WHERE status = 'HELD' AND expires_at < now()        ORDER BY expires_at        LIMIT 500        FOR UPDATE SKIP LOCKED    )    RETURNING product_id)UPDATE product pSET stock = p.stock + e.releasedFROM (SELECT product_id, count(*) AS released FROM expired GROUP BY product_id) eWHERE p.id = e.product_id;

FOR UPDATE SKIP LOCKED skips rows that another transaction has already locked, which lets you run this sweeper on several instances at once without two of them releasing the same reservation.

Then comes the edge case that bites everyone eventually: the payment succeeds after the reservation was released. The buyer was slow, the sweeper ran, the unit went back on sale, and then the payment gateway calls your webhook with "success". So the confirm step must be conditional too:

SQL
UPDATE reservation SET status = 'CONFIRMED'WHERE id = :reservationId AND status = 'HELD';

If that touches zero rows, the reservation is gone. Try to take a fresh unit with the same atomic decrement, and if there are none left, refund automatically. Make the webhook handler idempotent too, because payment gateways retry callbacks.

And if you use the Redis gate, every released unit must go back into Redis as well, or Redis keeps saying "sold out" while the database has stock.

What I would actually ship

Architecture: a Redis Lua gate in front of Postgres, with reservations, an idempotent payment webhook and an expiry sweeper
  • A Redis Lua gate that answers the 99%+ of requests that cannot win, enforces one unit per customer, and never lets the DB see more than roughly the stock count.

  • An atomic conditional UPDATE in Postgres as the final word, plus a CHECK (stock >= 0) constraint as the seatbelt.

  • Reservations with a 10-minute expiry, a SKIP LOCKED sweeper, and a conditional confirm that handles late payments.

  • Idempotency everywhere a retry can happen: the buy request, the payment webhook, the Kafka consumer.

  • A queue only if the drop is big enough that fairness and spike absorption matter more than an instant answer.

Notice what is not on that list: synchronized, a bigger database, or hoping traffic is lower than last year.

Prove it yourself in 30 lines

You don't need a load-testing tool to see the bug. A plain Spring Boot test with virtual threads will do it. (If booting a Spring context for tests feels painfully slow, here is how I cut our Spring Boot startup time from 45s to 8s.) Seed a product with stock 500 and fire 5,000 buyers at it at the same instant:

Java
@SpringBootTestclass OversellTest {    @Autowired BuyService buyService;    @Autowired OrderRepository orderRepository;    @Test    void neverSellsMoreThanStock() throws Exception {        long productId = 1L;              // seeded with stock = 500        int buyers = 5_000;        CountDownLatch startGun = new CountDownLatch(1);        try (var executor = Executors.newVirtualThreadPerTaskExecutor()) {            for (int i = 0; i < buyers; i++) {                long userId = i;                executor.submit(() -> {                    startGun.await();     // everyone waits here...                    try {                        buyService.buy(productId, userId);                    } catch (SoldOutException ignored) { }                    return null;                });            }            startGun.countDown();         // ...and goes at once        }                                 // close() waits for all tasks        assertThat(orderRepository.countByProductId(productId)).isLessThanOrEqualTo(500);    }}

Run it against the naive version and it fails: you will see more than 500 orders, and the exact count changes from run to run, which is the signature of a race. Swap in Fix 2 and it passes every time. I would keep a test like this in the codebase for any code path that touches stock, seats, coupons or anything else that can run out.

The takeaway

The overselling bug is not exotic. It is the most natural code you can write, it passes every test that runs one request at a time, and it only shows up on the one day you cannot afford it. The fix is not more cleverness in Java. It is moving the check and the write into one atomic step, in the place every instance shares.

If you are working on anything that sells limited stock this festive season, I'd love to know: which of these approaches does your system use, and what broke the first time it met real traffic?

Discussion

?