Skip to content
← Work

Enterprise retail e-commerce

Taking an enterprise commerce platform from 1,000 daily errors to double digits

A composable-commerce platform for a large optical retail chain was logging around a thousand errors a day, timing out on discounts and charging the wrong amounts. We worked the defect classes down in order.

Period
2021 – 2023 · 28 months
What we delivered
Stability programme, discount engine, layered caching, order integrity, search migration
Setting
Backend delivery within a ~20-person distributed organisation

Client and project names are withheld under confidentiality. Sector, scale and outcomes are reported as delivered.

~1,000 → double digits

Error log entries per day

10 → ~3

External service calls per request path

10s → sub-second

Catalogue search latency

3 defect classes

Resolved by one architectural change

Context

A large optical retail chain selling across several European markets — prescription lenses, frames, contact lenses and lens care, a catalogue of roughly 20,000 products. The platform ran on a composable commerce stack: a headless commerce backend as the system of record, and a SaaS commerce frontend built on Symfony and React providing base capabilities like cart, discounts and payment.

The vendor platform allowed any default page to be overridden and re-implemented on both backend and frontend. The client had customised essentially every surface except login.

Problem

The platform averaged roughly 1,000 error log entries per day. Not all distinct — a single defect could log 150 times — but the profile was wide: performance collapse on data-heavy endpoints, incorrect cart totals, orders that failed to materialise, payment amounts that diverged from what was owed, and the costliest class of all, custom prescription orders manufactured wrong.

Diagnosis had a hard constraint: the SaaS subscription exposed only the last six months of logs, so every root cause had to be found inside a rolling window rather than full history.

Approach

A staged error-reduction programme

Rather than fixing errors as they surfaced, we worked through them by class, in deliberate order of tractability and blast radius, using centralised logs and APM traces to isolate causes:

  • Baseline defects — audited and eliminated the straightforward failures wherever the originating code was still live, clearing the noise that was hiding everything else.
  • Race conditions — fixed concurrency bugs producing incorrectly calculated cart totals and orders that never reached the commerce backend.
  • Timeouts — traced to multiple overlapping discount definitions on the same product categories, where the compounded calculation ran long enough to time the request out.
  • Third-party resilience — hardened the payment, search and add-to-cart paths against non-responding upstream services, which had been surfacing directly as user-facing errors.
  • Platform-generated load — reduced the errors triggered by the vendor platform’s own server-side rendering and cache warm-up spikes.

The discount engine was the common root cause

The commerce backend provided basic discount structures but no true rule engine, so promotional logic had been written into application code. A single campaign could read: take the defined discount for product A; if the customer previously bought product X, stack a second one; but if lifetime spend crosses a threshold, remove both and apply a third instead.

This forced a constant round-trip — fetch the definitions, compute in application code, push the result back as a custom discount — and the whole calculation repeated on the product page, in the cart, and again before payment. With three to five conflicting campaigns live, requests timed out outright.

We extended the platform’s admin panel into a managed discount layer: an explicit priority list resolving conflicts deterministically, active/inactive state so campaigns toggled without code changes, and caching to eliminate the repeated recomputation.

That one piece of work resolved the majority of the cart race conditions, the payment amount discrepancies and the timeouts. Three of the platform’s highest-impact defect classes all traced back to the same uncontrolled calculation path.

Payment correctness

The payment provider was already integrated but producing amount mismatches: overlapping discount rules miscalculated under race conditions and incorrect precedence ordering, so the amount charged could diverge from the amount owed. We diagnosed and fixed the underlying calculation defects — a defect class with direct financial and customer-trust consequences.

Layered caching

The vendor platform provided full-page caching but cached nothing at the API layer: every backend request to the commerce backend and other third-party services went out fresh. As default pages were progressively replaced with custom implementations, we introduced Redis caching in stages — user-independent catalogue responses, then per-customer data under customer-scoped keys, then fully rendered user-specific pages.

Prescription order integrity

The costliest defect class was structural, not a bug. A customer selected lenses and the order went straight through, where an automated manufacturing validation system picked it up asynchronously and, on failure, changed its status. Three problems compounded: that status change never propagated back to the storefront, so the customer saw a successful order regardless; customers discovered mis-entered prescriptions 5–10 minutes later with no way to intercept; and physically impossible combinations were accepted — a prescription strong enough to need lenses too thick for the chosen frame.

Orders either silently never entered production while the customer believed otherwise, or were manufactured incorrectly and returned.

We proposed the replacement architecture and flow, which the team implemented: a guided step-by-step journey replacing direct add-to-cart — lens type, then prescription data with on-screen guidance mapping each field to its position on a real prescription image, then only frames compatible with that prescription, then cart, with explicit re-confirmation before submission. Behind it, a deferred confirmation state machine: orders are not shown as confirmed on placement; after a 10–15 minute settlement window, if manufacturing validation has raised nothing, the order is promoted to confirmed — and the customer is emailed on both outcomes, closing a feedback loop that had been missing entirely.

Catalogue search was a top-line problem: the fastest queries took 10 seconds and some timed out. We proposed migrating to Algolia on the observation that the existing pipeline feeding the commerce backend could, with minor modification, also feed the search index — removing search load from the backend, cutting latency and unlocking similar-product recommendations. After feasibility review the vendor added it to their own roadmap, but the client declined to wait: we delivered the whole implementation — index design, data structure and export pipeline, React integration, and a Symfony API layer so search could also be consumed server-side.

Outcome

  • Daily error log entries fell from ~1,000 to double digits.
  • Payment amount mismatches eliminated on the path where money was actually at risk.
  • Request paths that issued ten external service calls reduced to roughly three.
  • Catalogue search went from 10-second-plus and timing out to sub-second.
  • Prescription orders could no longer be silently lost or manufactured against an impossible specification.
  • Linting introduced to a codebase that had neither tests nor static analysis — the client had declined automated testing on the grounds that a manual QA engineer exercised the system as a real user, so a linter was the strongest gate achievable within that constraint.