A live auction on Symfony: six steps and six mistakes I made along the way
The auction ended. Three minutes later a bid arrived, and my code accepted it.
By the book, everything was correct: planned_end_at had passed, but the auction status was still TRADE, because moving it to CHOICE is the buyer's action and the buyer had not clicked yet. My code checked the status, saw “trading is open,” and wrote the bid. The longer the buyer waited, the longer that window stayed open for anyone who noticed.
On a platform that handles money, this is the worst class of bug: it crashes nothing, logs nothing, and surfaces in a dispute rather than in monitoring.
Below are the six steps a live reverse auction is built from, and in each one, what I got wrong the first time. The code is open under MIT and I link to the files as I go, so I will not retell the documentation.
Step 1. The domain and the state machine
My auction lives in three scenarios (AuctionTypeEnum): REDUCTION — a classic price reduction with a step from the start price (step_mode=fixed) or any price below the current one (step_mode=free); FREE_PRICE — no mandatory reduction, where current_price simply tracks the best offer; PRICE_REQUEST — collecting prices within a window, with no live steps and no extensions.
The lifecycle is described by a complete symfony/workflow definition (config/workflow/auction.yaml): 16 persisted statuses plus a virtual CREATED that never appears in the database, and 38 transitions, 23 of them named.
Two things here are worth copying. First: places and transitions reference AuctionStatusEnum::…->value through !php/enum, so renaming a status produces a compile error instead of a magic string that surfaces in production a month later. Second, the guard on the entry into trading:
!php/enum App\Auction\Entity\Enum\AuctionStatusTransition::START_TRADE->value:
from:
- !php/enum App\Auction\Entity\Enum\AuctionStatusEnum::SCHEDULED->value
to: !php/enum App\Auction\Entity\Enum\AuctionStatusEnum::TRADE->value
guard: "subject.getRulesSnapshot() !== null"
Before the start, the service calls Auction::captureRulesSnapshot() and freezes the step, the window length, the extension limit and the boundaries. From then on the whole auction reads its rules only from that snapshot, so a manager can edit platform settings mid-auction and a running auction will not notice.
What it cost me. I drew a beautiful state machine and never attached an engine to it. The SCHEDULED → TRADE transition existed in the code, but it was called from exactly one place: the console fixture for the load test. Auctions scheduled for the future never started — an auction waited for its moment in SCHEDULED and stayed there forever. There was no symptom at all: no error, no alert, just an auction that did not start.
Every transition in a state machine needs an explicit caller: a user, an event consumer, or a scheduler. A transition that will “fire on its own” does not fire.
Step 2. Anti-sniping and a timer that cannot be cheated
The rule is simple: a bid in the last window, that is, within step_duration_sec of the end, extends trading by extension_duration_sec. Without it the auction turns into a contest of network latency — whoever clicks in the last second wins. That is what sniping means.
I moved the logic into a class with no dependencies at all, AuctionTimer:
public function extendOnBid(
\DateTimeImmutable $now,
\DateTimeImmutable $plannedEndAt,
int $extensionsCount,
RulesSnapshot $snapshot,
?\DateTimeImmutable $executionStartAt,
): ?\DateTimeImmutable {
if (!$snapshot->extendOnLastStep) {
return null; // plugin rule: no extension
}
// The bid is not in the last window — no extension needed.
if ($now->getTimestamp() <= $plannedEndAt->getTimestamp() - $snapshot->stepDurationSec) {
return null;
}
// Extension limit exhausted.
if ($extensionsCount >= $snapshot->maxExtensions) {
return null;
}
$candidate = $plannedEndAt->add(new \DateInterval('PT'.$snapshot->extensionDurationSec.'S'));
// Hard boundary: trading cannot run past execution_start − lead_hours.
$cutoff = $this->tradeEndCutoff($executionStartAt, $snapshot->tradeEndLeadHours);
if (null !== $cutoff) {
if ($candidate > $cutoff) {
$candidate = $cutoff; // truncate to the limit
}
if ($candidate <= $now) {
return null; // boundary passed — forbidden
}
}
return $candidate;
}
There are three safeguards and each works independently. The window — only a bid in the last step extends, so the length of trading stays predictable. The limit — max_extensions stops anyone from dragging the auction out forever. The hard boundary — trade_end_lead_hours prevents trading from ending later than N hours before lot execution begins, and if the truncated boundary has already passed, extension is forbidden outright.
Keeping this in a dependency-free class is my advice to anyone building something similar. Anti-sniping is the one place where logic moves participants' money directly, and it should be covered by a table of cases rather than an integration test with a database.
What it cost me. I got the extension right and slept through the closing — the exact bug this article opens with.
The logic was: while the auction is in TRADE, bids are accepted. The TRADE → CHOICE transition is performed by the buyer, when they finish trading or pick a winner. Between the expiry of planned_end_at and that decision the auction kept hanging in TRADE, and the status check honestly answered “trading is open.”
It got worse. An auction with a long-expired timer also looked alive by a second signal: auctions:heartbeat kept refreshing its live state, because it looked at the status rather than the clock. The system was actively confirming that a dead auction was alive.
The fix came in two parts. The bid check stopped trusting the status:
// The trading window is closed by time (FR-1.3.3). The status is not enough here:
// TRADE → CHOICE is done by the buyer or by the auctions:finish-expired command,
// and between the expiry of planned_end_at and that transition the auction stays in TRADE.
$plannedEnd = $locked->getPlannedEndAt();
if (null !== $plannedEnd && $now > $plannedEnd) {
throw new BidRejectedException('Trading window is closed (planned end passed)', 'auction_window_closed');
}
One detail matters: planned_end_at in this comparison already includes every anti-sniping extension, so the time check and the extension rule do not conflict. The second moves the boundary, the first guards it.
The second part is the auctions:finish-expired command, which takes TRADE auctions whose window has arrived and finishes trading on behalf of the system. It does not pick a winner: that is the buyer's decision, automatic for a reduction auction and manual for the other types. The command is idempotent, and a race with a manual finish raises an exception on one auction without touching the rest. In the scheduler loop it runs every 30 seconds, and that interval directly sets the delay between the end of the window and the move to CHOICE.
If you take one thing from this article, let it be this: in an auction, the expiry of time is an event, not a state. Until someone turns it into a transition, it does not happen.
Step 3. The bid transaction
A bid starts its path at POST /auctions/{id}/bids. The service validates the price against the auction type and calls BidTransaction::commitBid() inside an active transaction.
Before the read-modify-write I take the auction row with SELECT … FOR UPDATE (LockMode::PESSIMISTIC_WRITE). Two simultaneous bids serialize on that lock: the second waits for the first to commit and reads the already updated price. That gives the main invariant: of two bids with the same price, exactly one wins. The auction_bids table is append-only; history is never rewritten.
Under the same lock a bid passes four checks, and their order matters: replay by Idempotency-Key, auction status, window time, participant admission. Replay comes first on purpose — a redelivery of an already accepted bid must return that same bid rather than run into the fact that the auction has since closed.
Money is counted in integers: prices in minor units, the VAT rate in basis points, no floats at any step. The audit records before/after for every mutation of price, timer and extension count, so an incorrect price jump shows up in the audit rather than in a conversation with a participant.
What it cost me. The target throughput is 100–200 bids per second, and against that background I found I was making two database round-trips per bid where one is enough.
My AuditService calls flush() by itself by default, which is convenient and correct for ordinary mutations. But the hot bid path writes four things at once: the bid, the updated price, the audit record and the outbox event. With the default behaviour the audit took its flush separately, and one batch became two.
A parameter fixes it, but the habit matters more than the line:
$this->audit->record(
// ...
// Task 4.10 (NFR-1: 100–200 bids/sec): the audit record is persisted without
// a separate flush — bid, audit and outbox are written in ONE batch.
flush: false,
);
A default that is good for 95% of calls is almost certainly bad for the hot path. I now check what every shared service does behind my back before calling it inside a transaction that has to fit into 100 ms.
The result is one EntityManager::flush() per bid: bid, price, audit and outbox event, atomically and in a single round-trip.
Step 4. Push without php-fpm
Delivery runs from the core, strictly asynchronously, through four links.
PHP-FPM holds no connections here. SSE streams live on a separate Go hub that scales independently of the PHP pool, while the worker wrote the bid, committed, published the event and moved on. The classic “SSE on php-fpm” trap, where workers get stuck on open connections, is avoided architecturally: PHP takes no part in delivery at all.
The topic is private, auction:{id}. Subscribing requires a JWT carrying the sub claim and the auction ID, and that token goes to admitted participants, the buyer and observers. The core publishes with a JWT holding the publish right. An outside client cannot subscribe — the hub rejects the connection before PHP sees a single request.
The outbox guarantees the link between bid and event: the event is written in the same transaction as the bid, so either both are stored or neither is.
What it cost me. Three evenings, all three on Mercure.
First: I took the latest image. The latest tag is Mercure 2.x, where authorization moved to OAuth2 authorization_details (RFC 9396) with typ: at+jwt and mandatory iss/aud. The legacy claims mercure.publish/mercure.subscribe that symfony/mercure 0.7.x generates simply do not work on a v2 hub — publish returns 401. The application does not crash, the logs say nothing interesting, and every auction SSE stream is dead. The image is now pinned to v0.16.3, the last release of the 0.x line.
Second: the publish format changed back in 0.16. The topic and data fields go into the form body as application/x-www-form-urlencoded, not into the query string. A query gets you a 400 complaining “Missing topic parameter,” which gives no hint that the problem is the transport rather than the parameter itself.
Third, a small one that cost an hour: the publish claim in the JWT accepts only ['*'], and a glob subscription such as load:* returns 401.
Pinning to 0.x has its own price: metrics. The hub already counts mercure_subscribers_connected, mercure_subscribers_total and mercure_updates_total, but the Caddy module does not register /metrics, so in the official image the endpoint returns 404. I had to build the image locally from the v0.16.3 sources with a ten-line patch. Delivery latency is unavailable in 0.x either way — the hub does not measure it.
Step 5. Failures
Live state lives in Redis; the source of truth is PostgreSQL. A Redis failure therefore does not take the auction down, it moves the auction into a managed mode.
While trading runs, auctions:heartbeat refreshes the state every 30 seconds. If there are no updates for longer than AUCTION_HEARTBEAT_TIMEOUT (300 seconds by default), the auction automatically pauses through the TRADE → PAUSED transition, and nobody trades blind against a dead timer. The heartbeat interval has to be smaller than the threshold, otherwise live trading starts pausing itself because of one line in a config file.
The remaining seconds are stored in PostgreSQL on pause, and on RESUME the countdown continues from that remainder rather than from zero. A pause steals no time from participants — for an auction that handles money this is a question of whether the result is legitimate. State can be rebuilt by auctions:recover and auctions:state:rebuild, which restore the live snapshot from PostgreSQL together with the last bid, the price, the timer and the status.
What it cost me. I wrote the recovery commands and then could not run them at exactly the moment they were needed.
The console metrics subscriber injected CollectorRegistry, and the registry opened Redis. The event dispatcher instantiated the subscriber on every bin/console invocation, so with Redis unavailable every console command failed — including auctions:recover, whose entire reason to exist is to run when Redis is unavailable.
The fix: the Redis service is lazy, the connection opens only when a metric is actually written, and a storage failure is logged and swallowed. A recovery tool must not depend on the thing that broke.
Step 6. Metrics, SLOs and numbers that lie
The domain exports metrics on equal terms with the HTTP layer. The application publishes /metrics through promphp with Redis storage; Prometheus and Grafana come up as a separate profile (make observability-up) with a ready dashboard.
The domain has seven metrics of its own: auction_bids_total — committed bids; auction_bid_latency_seconds — a histogram of the whole write path including the Redis snapshot; auction_extensions_total — anti-sniping extensions; auction_pauses_total — pauses and resumes; auction_active_trades — auctions in trading; auction_stalled_now and auction_stall_events_total — auctions stuck without bids. Next to them sits the RED pair auction_bid_attempts_total{outcome} and auction_bid_rejections_total{reason}, with an enumerable set of reasons (bid_rejected, auction_not_trade, auction_window_closed, duplicate_bid).
I framed the SLOs as an error budget with a multi-window burn rate, following the SRE Workbook: 99% of bids are written within 100 ms, 99.9% of HTTP requests come back without a 5xx. An alert fires only when the error ratio crosses the threshold on the long and the short window at once — 14.4x on the 1 hour / 5 minute pair and 6x on the 6 hour / 30 minute pair. The window pair cuts off flapping: to raise an alert, degradation has to hold on the long window instead of blinking once.
What it cost me. I made two mistakes here, and both are worth the walkthrough, because both looked like working code.
The first was the latency histogram. I wrapped the bid path in try/finally and measured the time in finally, which felt tidy. Everything lands in finally: accepted bids, rejected attempts and idempotent replays. A price rejection completes in single-digit milliseconds, a replay is faster still, and both fell into the le="0.1" bucket. That histogram feeds the SLI “99% of bids within 100 ms,” which means the more rejections participants hit, the better my SLO looked. A metric that improves when users get errors is worse than no metric.
The second was cardinality. I exported stalled auctions as auction_no_bids_alert{auction_id}, one series per auction. In promphp's Redis storage such a series lives forever, so the label set grew with the number of auctions that had ever entered trading and never shrank. On a dev stack this goes unnoticed; on a platform with thousands of auctions it is a grenade with the pin pulled.
It is now separated explicitly: a rejection increments rejected with a reason, a replay counts as neither accepted nor rejected (it creates no bid and is not a refusal), and latency and auction_bids_total are written only after the commit. Stalled auctions are a gauge without labels (auction_stalled_now) plus a transition counter (auction_stall_events_total), and the alert sits on the counter.
A category of its own is numbers that lie, and I suggest checking it before you optimize any code.
APP_DEBUG=1 adds 300–400 ms per request through container compilation and the profiler. With it a p95 under 100 ms is unreachable even on a catalog of 100 rows, where p95 grows to roughly 600 ms. I spent an evening staring at queries before I realized I was measuring the profiler.
The stock php:8.5-fpm pool ships five workers. On them, you get around 15–20 bids per second at a p95 of about 700 ms with 20 virtual users, and it looks exactly like a slow application. Setting max_children=30 takes the result to 22 bids per second at a p95 of about 550 ms. The application logic did not change by a single line.
Webhooks tell a similar story: a shared messenger:consume across all transports loops over the queues, and empty RabbitMQ/Redis queues slow the fetch of webhook jobs down to roughly 10 deliveries per second. A dedicated webhooks worker delivers 1200+ events per minute at a p95 of about 2.7 seconds.
The final numbers on the dev stack:
| Scenario | SLO | Actual |
|---|---|---|
| Bids, domain write path | p95 < 100 ms, 100–200 bids/sec | p95 ~9–13 ms, ≥30 bids/sec (smoke) |
| Bids, HTTP e2e | bid_write_ms p(95) < 1000 |
~22 bids/sec, p95 ~550 ms with max_children=30 |
| Catalog | p95 < 200 ms | p95 ~166 ms on ~100 published / 5000 total |
| SSE delivery to the client | p95 < 1 sec | node script, N subscribers |
| Webhooks | ≥10,000 events/min, latency < 5 sec | ~1200+/min, p95 ~2.7 sec |
The gap between 9–13 ms on the domain path and 22 bids per second over HTTP is the cost of transport: serialization, the worker pool, the network. Both numbers are honest and both are needed. The first speaks to the quality of the domain logic; the second shows that what you will hit when scaling is the worker pool, not the transaction.
What is left in the repository
Everything I did not write about is out in the open: the load scenarios in load/ together with a walkthrough of the findings, the full state machine, the dashboards and alerts in docker/. The project is MIT, so take anything.
Coming next in the series: the modular monolith and PHPArkitect, sealed bids encrypted until the opening moment, and an outbox with 63 JSON Schema event contracts.
While building this live auction I thought the hard part would be keeping the timer in agreement between Redis, PostgreSQL and the client. The timer turned out to be arithmetic. The hard part begins in those three minutes when trading is already over and the system does not know it yet.