Skip to content

Retail and e-commerce

AI for keeping the promise your product page makes

We work across the whole span of a retail business: from the one SKU the system insists there are four of while the shelf is empty, and the rolling 30-day minimum that decides whether a discount claim is legal, to the peak week that has to run without a war room. Most of that work is data repair — censored demand, reservation drift, GTINs a supplier recycled — with a model at the end of it.

30 minutes with the engineer who would do the work — not a salesperson. No obligation, and you keep whatever we work out on the call.

Warehouse worker counting stock on a tablet in front of shelved cartons

The numbers this sector is run on

29–50%
Exact-match inventory record accuracy at store level, untreatedECR Loss research across six European retailers, measured on shop-floor counts. A single well-run DC with cycle counting sits at 95–99% and is not your problem; the gap between the two is why chain-level stock looks adequate while a store is empty and the OMS still sources from it. Accuracy within ±1 unit runs 51–71%, and phantom stock affects 8–24% of items per audit, rising towards 27% when counting slips to twice a year. If your replenishment is quietly starving fast movers, this is usually why.
10–20% / 30–50%
WMAPE at SKU × location × week: stable FMCG at DC, versus fresh and fashion at storeThe two working bands. Anyone quoting a single site-wide MAPE is quoting the wrong number — accuracy is only meaningful at the level a decision is taken, and the long tail is where the money leaks.
20–40% / 8–15%
Online return rate: apparel versus electronicsCategory ranges compiled from NRF/Appriss returns reporting and platform datasets — Littledata’s Shopify sample and Dynamic Yield’s session data measure different populations and disagree by design; footwear 17–30%, home 15–23%, beauty 4–12%, overall around 19–20%. Treat them as an order of magnitude until your own numbers replace them. Roughly half of apparel returns are fit and sizing, which means the product detail page is manufacturing them.
under 3%
Healthy WISMO contact rate, tickets ÷ orders shippedThe operational benchmark: 3–8% is normal but expensive, and above 8% is a fulfilment or tracking-data defect rather than a staffing problem. The denominator must be orders shipped — measured against tickets it flatters you as you grow.

What we build

Retail planner reviewing stock levels on a laptop

Forecasts built on demand, not on sales

Before any model, the history gets repaired: stockout intervals flagged from availability logs and on-hand snapshots, latent demand imputed for the censored periods, promotional uplift separated from baseline, and expected returns netted out by SKU and size. Then gradient-boosted models at SKU × location × day with a quantile objective, so replenishment is driven at a service level rather than at a point estimate.

  • Rolling-origin backtest against seasonal naive before any accuracy number is quoted
  • Forecast Value Added published per step, including the step we built
  • New items forecast from attribute-based analogues and scaled launch curves, not a category average
Shop manager updating shelf-edge price labels with a handheld device

Pricing where the 30-day floor is a constraint, not a report

The first component is an append-only, timestamped price history per SKU, channel and currency — the compliance spine. The rolling 30-day minimum computes from it and becomes a hard guardrail every proposed price passes through, alongside margin floor, price-ladder consistency across pack sizes, maximum daily move and KVI follow rules. Supplier recommended prices enter as an advisory flag with a logged override, never as an enforced floor: a binding minimum resale price is a hardcore restriction under Article 101 TFEU and Regulation (EU) 2022/720, and an engine that enforces one is building the evidence against you. Every automated change is logged with its rule, its inputs and its approver.

  • Elasticity estimated on stable SKUs, pooled hierarchically for the long tail
  • Competitor matching by GTIN first, fuzzy matching only as a scored fallback
  • The same guardrail engine drives the webshop and the marketplace offers — Omnibus applies to both
Team member photographing a product beside a laptop showing the listing

Product data the channels will actually accept

Supplier spreadsheets and PDFs staged into a PIM with per-channel completeness rules as a publish gate. Vision-and-text extraction fills a fixed attribute schema, normalised to a controlled vocabulary and units, with confidence thresholds routing weak fields to a human queue instead of publishing them. Internal taxonomy mapped to Shopify’s Standard Product Taxonomy, GS1 GPC and Google’s categories, GTINs validated against check digit and prefix.

  • Disapproval reasons ingested back as tickets against the attribute responsible, not treated as an ads problem
  • GPSR fields — manufacturer, EU Responsible Person, identifier, warnings — modelled as attributes, not as a legal page
  • Extraction lands around 85–91% F1 on clean benchmarks; the review queue is budgeted, not wished away
Two workers packing orders at a bench during a busy period

Integration that survives a peak week

Event-driven writes on a durable queue: idempotency keys on every write into the ERP, exponential backoff with jitter on 429s, and Bulk Operations or async bulk REST for catalogue-wide work so the interactive budget stays free. One stock authority per SKU with an explicit per-channel buffer, and a scheduled full reconciliation — hourly for stock, daily for orders — that reports a delta count every morning.

  • An exception queue with named owners: address failures, COD refusals, failed labels, payment and fulfilment mismatches
  • On Magento: indexer in schedule mode, daily reservation cleanup, inventory:reservation:list-inconsistencies as a health check
  • A change-freeze window agreed up front, typically from mid-October, and a documented rollback for anything live

Worked examples

These are designs, and the ranges we would contract against, drawn from published sector data — not Palamed results. The four systems we have actually delivered are on case studies.

Demand forecasting corrected for stockout censoring

Problem

Buyers replenish from a system trained on historical sales, so every stockout teaches it that demand was lower than it was. It orders less, stocks out again, and the fastest-selling lines are the ones it most reliably starves. Promotional weeks sit in the same table as normal demand and inflate next year’s baseline.

Approach

The history is rebuilt before any model: stockout intervals flagged from availability logs or on-hand snapshots, latent demand imputed for censored periods with a censored-regression or two-stage recovery model, promotional uplift separated from baseline using calendar, price and mechanic features, and expected returns netted out by SKU and size. Then LightGBM at SKU × location × day — lags, rolling windows, price and elasticity features, local calendars and weather — with a pinball-loss quantile objective, so replenishment runs at a service-level quantile instead of a point forecast. Hierarchy reconciled bottom-up, rolling-origin backtest against seasonal naive, FVA published per process step.

What we would target

The FreshRetailNet-50K study — 50,000 hourly store-product series across 898 stores and 863 perishable SKUs — removed a 7.37% systematic under-forecast once latent demand was recovered, and improved accuracy by 2.73%. That is fresh grocery, where stockout censoring is at its worst, and the recoverable bias is a function of your own out-of-stock hours: at 97% availability expect low single digits, at 90% expect most of it. On your catalogue we expect a bias correction sized to your censored share plus a WMAPE improvement against your incumbent. The two numbers that fund the work are downstream of it: 8–15% inventory reduction at equal or better service level, and a 1–3 percentage point availability gain — both design targets to be proven against your own agreed baseline, both stated in working capital and lost sales rather than in error metrics. Expect the first four to six weeks to be data repair rather than modelling — and if that phase concludes the problem is a cycle count, not a model, that is the finding and we say so.

Repricing built around the Omnibus 30-day floor

Problem

Prices are set in spreadsheets against competitor screenshots, discount claims are calculated against RRP rather than your own 30-day low, and nobody can reconstruct what price was live on a given day across web, marketplace and store. Competitor moves are answered days late on the SKUs that matter, and never on the tail.

Approach

An append-only, timestamped price store per SKU, channel and currency is the first component — it is the compliance artefact and the elasticity dataset at once. The rolling 30-day minimum computes from it and becomes a hard constraint on every proposal. It is also simulated forward, because the reference ratchets: every markdown you take today lowers the price you may lawfully advertise against for the next 30 days. A promo calendar is checked as a sequence before it is committed, so you find out in October that the planned December claim is not available — not on 1 December. Elasticity from log-log models with regularisation on stable SKUs and hierarchical pooling for the tail; competitor matching by GTIN with a scored fuzzy fallback. Every proposal then passes margin floor, price-ladder consistency across pack sizes, maximum daily move, KVI follow rules and a rounding policy, and every change is written with its rule, its inputs and its approver. Supplier recommended prices are modelled as an advisory flag with a logged override, never as an enforced floor — a binding minimum resale price is a hardcore restriction under Article 101 TFEU and Regulation (EU) 2022/720, and a repricing engine that enforces one is building your evidence against you.

What we would target

A 1–4% relative improvement in gross margin value on the repriced assortment — roughly 30–120 basis points of margin rate on a typical mix — is a defensible expectation; the 5–10% in vendor case studies is a ceiling, not a forecast. Competitor response latency drops from days to hours. The by-product is often worth more to the CFO than the margin: an audit trail that answers a regulator asking for 30-day evidence — the question UOKiK put to Zalando and Temu before fining them roughly €8.5m combined in January 2026.

Catalogue enrichment, taxonomy mapping and feed remediation

Problem

Supplier data arrives as inconsistent spreadsheets and PDFs. Titles carry the supplier’s internal codes, colour is free text in three languages, variants are not grouped, and GTINs are missing or recycled. Merchandisers spend the week retyping attributes, Shopping disapprovals sit unfixed, and site search returns nothing for queries the catalogue could answer.

Approach

Supplier data staged into Akeneo or Pimcore with per-channel completeness rules and a data-quality score as a publish gate. Vision-plus-text LLM extraction against a fixed attribute schema, normalised to a controlled vocabulary and units, with confidence thresholds routing weak fields into a human queue rather than publishing them. Internal taxonomy mapped to Shopify’s Standard Product Taxonomy, GS1 GPC and Google’s product categories through a maintained crosswalk, and GTINs validated against check digit and prefix. Merchant Center and eMAG disapproval reasons are ingested back into the PIM as tickets against the attribute responsible — on variant disapprovals that is usually item_group_id.

What we would target

Required-attribute completeness moves, in published implementations of this pattern, from 60–75% to above 95% inside a quarter, with disapprovals falling 70–90% — a target on your data, not a Palamed result, held to the same standard as the metrics above. Because attributes drive facets and search, category-page conversion moves 3–8% in relative terms — on a 2.5% base that is 2.58–2.70%, not 5% — and zero-result searches drop. The same attributes then make cold-start forecasting possible for a catalogue that turns over every season, which is the second payback nobody budgets for.

Order orchestration across store, ERP, courier and marketplace

Problem

Orders arrive from the webshop, from eMAG and over the phone; stock lives in the ERP; labels come from two couriers; and the glue is a nightly CSV plus a person. The symptoms are oversells at peak, duplicate orders in the ERP after a retry, stock on the site that is hours stale, and marketplace orders shipping late enough to hit seller-rating penalties.

Approach

Event-driven integration on a durable queue: webhooks for order-created and fulfilment events, idempotency keys on every write into the ERP, exponential backoff with jitter on 429s, and Shopify Bulk Operations or Magento async bulk REST for anything catalogue-wide so the interactive budget stays free. One stock authority per SKU with an explicit per-channel buffer, plus a scheduled full reconciliation — hourly for stock, daily for orders — that reports a delta count, because event streams lose messages quietly. Econt and Speedy label and COD flows sit behind the same exception queue, and COD orders are risk-scored at checkout on address quality, basket composition, customer history and office-versus-door delivery, then routed to prepayment or a confirmation call rather than blocked.

What we would target

Oversells into single digits per 10,000 orders, stock sync lag p95 from hours to under two minutes, and most manual order touches gone — targets on your own agreed baseline, not results we are reporting. Confirmed refusals are kept as demand — the customer wanted the item — but a refusal probability is modelled by region, courier, office-versus-door and basket, and netted at the replenishment stage. That keeps your pick capacity and reservations sized on gross orders while stopping replenishment from ordering against sales that never settled. Treating a refusal as a deleted sale is the common error and it starves the same fast movers all over again. Expect the reconciliation job to find drift in week one that nobody believed existed.

Order status and returns handled by tools, not by prose

Problem

A quarter to half of inbound tickets are “where is my order”, and agents answer them by copying tracking numbers between the shop admin, the courier portal and the ERP. Returns run on email: everything is approved, the warehouse grades by eyeball, refunds go out before goods arrive, and serial bracketers are invisible.

Approach

The agent is wired to tools rather than to a knowledge base: read-only functions against the store order API, the OMS and Econt, Speedy or DPD tracking, so every answer is retrieved and never composed. A defined intent set — status, ETA, address change before dispatch, cancel before pick, start a return — with a write-action allowlist and confirmation on anything with money attached. AI disclosure at first contact under AI Act Article 50, full tool-call logging, escalation on low confidence or a second failed resolution. Returns move onto a portal with a structured reason-code taxonomy. The portal first classifies the return: a statutory withdrawal under CRD Article 9 is not a negotiation — cash refund including standard outbound delivery, inside 14 days of return or proof of return, no exchange-first prompt in the path. Everything outside the statutory window or the statutory exceptions is goodwill, and that is where policy may vary: exchange or credit offered first, returnless refund below a computed recovery threshold, manual review for high-value or high-risk cases, and reason codes mined by SKU and size to find the styles whose size chart is manufacturing the returns.

What we would target

30–45% of total ticket volume handled end to end is a realistic first-year target on a well-integrated store; the 70–80% vendors quote applies only to the narrow order-status slice. Proactive dispatch, out-for-delivery and exception notifications usually do more, taking WISMO from 6–8% of shipped orders to under 3% by removing the ticket rather than answering it. On returns, refund rate falls 15–30% through exchange conversion while return rate barely moves.

Allocation, rebalancing and size curves

Problem

Chain-level stock looks adequate while stores are broken on core sizes. The DC pushes evenly against one national size curve, slow stores end the month holding the units the fast stores needed, and by week six the only lever left is markdown — 15–35% of sales in apparel, the largest single margin line in the P&L and the one nobody models. Online the same failure wears a different name: “in stock” at a node the order manager will not source from.

Approach

Stores are clustered on demand shape rather than on revenue, because two stores with identical turnover can want opposite size curves. Store-and-style size profiles are then fitted with hierarchical shrinkage, so a low-volume store borrows strength from its cluster instead of being fitted on noise. Push allocation is constrained by presentation minimums, pack sizes and transport capacity — an allocation that cannot be picked is not a recommendation. A weekly rebalance ranks transfers by expected margin recovered minus transfer cost, with a hard floor so nothing moves for less than it costs to move. The same availability picture feeds the distributed order management sourcing rules, so split shipments and cost-to-serve are priced in rather than discovered.

What we would target

These are design targets agreed with you before we start, not results we are reporting: broken-size incidence down 30–50% on core styles, a 2–5% improvement in full-price sell-through, and lower terminal stock at end of season, each measured against your own matched prior season. Because the rebalance is margin-gated, transfer volumes usually fall rather than rise, and the first thing the model tends to prove is how many transfers you were already making that never paid for themselves.

Systems we work with

We integrate with what you already run. If a platform below is missing, tell us — the pattern usually transfers.

Storefronts, marketplaces and feeds

  • Shopify and Shopify PlusGraphQL Admin API, Bulk Operations, Flow, Functions
  • Adobe Commerce / Magento 2.4.xMSI, reservations, async bulk REST, message queues
  • WooCommerceHPOS, Action Scheduler, wp wc hpos sync
  • CloudCart
  • Shopiko
  • NextBasket
  • PrestaShop
  • OpenCart
  • eMAG Marketplace APIproduct_offer, order, AWB, RMA; RO/BG/HU and FBE
  • Google Merchant Center
  • Amazon SP-API
  • Kaufland Global Marketplace
  • Allegro
  • Channable
  • Feedonomics
  • DataFeedWatch
  • Productsup
  • Shopify Standard Product Taxonomy26 verticals, 10,000+ leaf categories

ERP, POS, PIM and master data

  • SAP S/4HANA and SAP Business One
  • Microsoft Dynamics 365 Business Central, Odoo, NetSuite
  • Microinvest, Selmatic, Business Navigator, Plus MinusBulgarian ERP/POS with НАП fiscal integration
  • Akeneo PIM with Data Quality Insights, Pimcore, Salsify, inRiver
  • GS1GTIN, SSCC, GS1-128, GPC, GDSN
  • EANCOM ORDERS
  • ORDRSP
  • DESADV
  • INVOIC
  • PRICAT
  • Celigo
  • Patchworks
  • Alumio
  • n8n
  • Make
  • Workato

Fulfilment, support, returns and payments

  • Econt and Speedy/DPD APIslabels, COD, office and locker selection
  • Fluent Commerce, Manhattan Active Omni, IBM Sterlingenterprise distributed order management, integrated where you already run it rather than implemented by us
  • Loop Returns
  • Returnless
  • Narvar
  • nShift
  • ReturnGO
  • Gorgias
  • Zendesk AI agents
  • Intercom Fin
  • Freshdesk
  • Adyen, Stripe, Mollie, myPOS, Borica3DS2, SCA and TRA exemption handling
  • Forter, Signifyd, Riskified, SEONfraud, chargeback and returns-abuse scoring, wired into where you already license it

Forecasting, pricing and the pipeline underneath

  • RELEX, Blue Yonder, Slimstock, ToolsGroup, Lokadplanning suites we integrate with or benchmark against, not ones we deploy
  • Omnia Retail
  • Competera
  • Pricefx
  • Prisync
  • Dealavo
  • LightGBM, XGBoost, Nixtla StatsForecast and MLForecast, Darts, GluonTS
  • dbt with Airflow
  • Dagster or Prefect
  • Snowflake
  • BigQuery
  • ClickHouse
  • DuckDB
  • Klaviyo, Bloomreach, Emarsysdriven from scored segments, not blanket delays
  • Segment
  • RudderStack
  • Snowplow; Algolia
  • Klevu
  • Constructor

What we design against

10–25%
WMAPE reduction at SKU-store-week against the incumbent baselineA target we design against, not a Palamed result. It is only claimable after a rolling-origin backtest on your own history against seasonal naive and your current process, with Forecast Value Added published per step — including the step we built.
15–30%
Lower refund rate through exchange and credit conversionBaseline is your current refunds ÷ orders over matched eight-week windows, not return rate. Return rate barely moves under this work; what changes is whether cash leaves. Returns-portal vendors publish exchange conversion of 20–40% of returns; that is vendor-measured on their own installed base, so treat it as the ceiling.
6–8% → under 3%
WISMO contacts per order shipped, after proactive notificationsMeasured on your own ticket export against shipped orders for the same period. Most of the movement comes from dispatch, out-for-delivery and exception notifications removing the ticket, not from an agent answering it faster.

Regulation and standards in scope

  • Omnibus Directive (EU) 2019/2161 amending the Price Indication Directive 98/6/EC, Article 6a — any announced reduction must state the prior price, defined as your own lowest price for that product and channel in at least the preceding 30 days. Progressive reductions and perishables are Member State options under Article 6a(2)–(3), not automatic EU-wide allowances: in Bulgaria they arrive through the Закон за защита на потребителите and are enforced by КЗП, and that transposition is what a pricing engine implements rather than the abstract Directive. RRP may be shown as a labelled comparison but never as the basis for a discount claim, and 98/6/EC also carries the unit-price obligation — price per kilogram or litre — which is the routine КЗП finding on a webshop. Enforcement is active across the single market: Poland’s UOKiK fined Zalando and Temu almost 37m PLN combined, roughly €8.5m, in January 2026, both decisions non-final and under appeal; the Dutch ACM fined five webshops €621,000 in June 2024 over misleading “from” prices; and the Landgericht München I ruled against Amazon on 14 July 2025 (4 HK O 13950/24, also under appeal) for benchmarking Prime Deal Days against manufacturer RSP.
  • General Product Safety Regulation (EU) 2023/988, applicable since 13 December 2024 — every online offer must display the manufacturer’s name and postal and electronic address, the EU Responsible Person where the manufacturer sits outside the EU, a product identifier such as type, batch or serial, a product image and any safety warnings. In practice this is a PIM schema change, a feed-mapping change and a supplier data-collection campaign, not a legal-page edit. The same schema carries two neighbours that block offers rather than warn about them: EPR registration numbers per country and per stream — packaging, WEEE, batteries, with Germany’s LUCID and France’s UIN the ones marketplaces check — and the DSA trader-traceability data eMAG and Amazon verify before an offer goes live. Both are required fields in the offer payload, which makes them PIM attributes with an owner and an expiry date, not a back-office task.
  • Regulation (EU) 2024/1689 (AI Act) — Article 5 prohibited practices in force since 2 February 2025, which bounds how far behavioural pricing and persuasion may go, and Article 50 transparency, which has applied since 2 August 2026 and requires that a customer is told they are dealing with an AI at first point of contact, perceivably inside the interaction rather than in the terms. The Digital Omnibus on AI, applicable since 27 July 2026, deferred most Annex III high-risk obligations to 2 December 2027 and moved synthetic-content marking to 2 December 2026 — it did not touch the disclosure duty. If your support bot is live and silent about what it is, that is already a live exposure, not a 2027 project. Penalties reach €15m or 3% of global turnover for transparency breaches and €35m or 7% for prohibited practices. Moffatt v. Air Canada — a small-claims tribunal decision in British Columbia, binding nowhere and least of all here — is still the cleanest illustration of the exposure: the airline had to honour a policy its bot invented. In the EU the operative liability sits in the Unfair Commercial Practices Directive and in national contract law, and the design answer is the same either way: the agent never composes policy.
  • Наредба № Н-18/2006 in Bulgaria — an e-shop must be declared to НАП electronically with a qualified electronic signature before trading begins, covering domain, platform name and version, payment methods, virtual POS contract and receiving accounts; changes are filed within 7 days; where the shop falls outside СУПТО and is exempt from a fiscal device, a monthly sales audit file is due by the 15th of the following month; a read-only auditor profile is provided for inspectors, and each sale carries a unique order code. Shops taking only remote card payments through a virtual POS can be exempt from a fiscal device and from СУПТО. Alongside it, euro adoption on 1 January 2026 at the fixed 1.95583 BGN/EUR touched price ladders, rounding rules, fiscal receipt layouts and every historical price series used for elasticity or Omnibus evidence. Practically, the price store keeps both the original BGN value and the EUR conversion with the rate and the rounding rule applied, so a discount claim whose 30-day window straddles the changeover can still be reconstructed for КЗП — and so a rounding decision taken at conversion is not later read as a price increase.
  • Consumer Rights Directive 2011/83/EU and the European Accessibility Act, Directive (EU) 2019/882 — the 14-day right of withdrawal, refund of standard outbound delivery cost and the order button that must state a payment obligation set the floor any returns automation sits on; the Accessibility Act has applied to e-commerce services since 28 June 2025, which in practice means EN 301 549 and WCAG 2.1 AA on the storefront and checkout plus a published accessibility statement.
  • GDPR (EU) 2016/679 and the ePrivacy rules — lawful basis and a retention schedule for the order-derived features behind COD risk scoring, return-behaviour scoring and CLV; a DPIA where profiling is systematic and extensive; and Article 22, which is what stops a COD risk score becoming a solely automated decision — high-risk orders are routed to a confirmation call a person makes, never silently blocked. Consent Mode v2 determines whether your marketing data exists at all in the EEA, which is why holdout design and platform-reported conversions are not the same conversation.

Before a model can forecast demand, the stock number has to be true.

What you are probably thinking

Our data is a mess — you will spend six months cleaning before anything works.

Correct, and that is the engagement. The first deliverable is an audit that quantifies the mess in business terms: what share of your sales history is censored by stockouts, what your exact-match inventory accuracy actually is, how many SKUs have no reliable GTIN, how many orders reconcile between store and ERP. That audit is useful on its own and it tells you whether the modelling phase is worth funding. Anyone who offers to skip it is selling you a demo.

We already bought a forecasting module with the ERP and it did not beat our planners.

Usually true, and usually correct behaviour on the planners’ part, because the module forecast on sales rather than demand, ignored promotions and had no way to express a service-level quantile. The test is cheap: run a rolling-origin backtest of the module, the planners’ final numbers and a seasonal-naive baseline on the same windows. If the module does not beat naive, the planners were right to override it. If the planners do not beat the module, you have an FVA problem, not a modelling problem. Two weeks and a data extract either way.

Automated pricing will get us fined under the Omnibus rules.

Only if you build it without the constraint. The rule is mechanical: a discount claim must reference your own lowest price for that product and channel over the preceding 30 days. So the append-only price history is the first component you build, and the 30-day minimum is a guardrail every proposal passes through rather than a report someone checks afterwards. Done that way, automation makes you more compliant than a spreadsheet, because the evidence is generated as a by-product of every price change.

An AI agent will tell a customer something wrong and it will cost us.

It will if you let it compose policy. The design answer is that the agent never writes policy text — it retrieves from a versioned source or calls a tool, and anything with money attached is a confirmed write action with limits rather than a sentence. Disclosure at first contact is not optional: Article 50 of the AI Act has applied since 2 August 2026 and the Digital Omnibus did not defer it. Add it, with full tool-call logging and escalation on low confidence or a second failed resolution. Then scope it: order status, tracking, returns initiation, address change before dispatch. Not warranty interpretation.

Our margins are 3% — this has to pay back inside a year.

Then start where the payback is arithmetic rather than statistical. Inventory record correction is documented as a revenue action, not hygiene: ECR reports 4–8% uplift concentrated on fast-moving, high-discrepancy items, and an academic audit programme across ~24,000 SKUs in 11 grocery stores measured 11% store-wide. Both are physical-store grocery results — the mechanism transfers to your stores directly and to a single DC only where phantom stock is suppressing replenishment, which is what the audit measures first. Support automation on order status has a cost-per-ticket delta you can calculate from your own volume before a line of code is written, roughly €1.50–6 human-handled against €0.20–1.00 automated. Feed disapproval remediation converts traffic you already paid for. Forecasting and pricing have larger ceilings and longer proof cycles — sequence them second, funded by the first wave, with the baseline and the holdout agreed before you start.

When we are the wrong choice

  • Catalogues and order books too small to learn from. Under roughly two years of order history, or with volumes where a single promotion dominates the series, a forecasting project has nothing to fit — a disciplined cycle count, a corrected size chart and a working stock sync will move your numbers further and cost a fraction of the money.
  • Replatforming and storefront builds. We do not run a Shopify-to-Magento migration, redesign a theme or own your checkout front end. Those are separate contractors on a different critical path; we integrate with what you run and work alongside them, never instead of one.
  • Per-shopper price personalisation and persuasion design. Prices that differ by individual on the basis of automated profiling carry a disclosure obligation, and Article 5 of the AI Act prohibits exploiting vulnerability, including financial distress, outright. We do not build fake scarcity counters, drip pricing or urgency mechanics either — they are the same category of risk with a smaller upside.

Questions we get asked

We are on Shopify. Does the platform limit what you can build?

It shapes it. The Admin API is a refilling cost budget in points per second — 100 on Standard, 200 on Advanced, 1,000 on Plus — with a 1,000-point ceiling on a single query, arrays capped at 250 items and pagination at 25,000 objects. Anything catalogue-wide goes through Bulk Operations, which sit outside the bucket. On Magento the equivalent constraint is the reservation ledger behind salable quantity; on WooCommerce it is Action Scheduler throughput during an HPOS backfill.

We are heading into peak. Can we start now, or should we wait?

We can start, but we do not ship into peak. The system runs in shadow mode: it produces recommendations, your people keep deciding, and at the end of the period we compare the two sets of decisions against what actually happened. That gives you a real Forecast Value Added number on your own data before anything is automated. A change-freeze window, usually from mid-October, and a rollback plan for anything already live go into the statement of work. On cover, the honest position: we are founder-led and do not run a 24/7 rota, so anything of ours that is live through peak is engineered to degrade to your existing manual process rather than to stop — queues drain, reconciliation reports the delta, the exception queue has your named owner. What we commit to in writing is a named contact, an agreed response window across the freeze period, and a rollback any of your people can execute without us. If you need a follow-the-sun rota, say so now and we will tell you we are the wrong call.

Half our SKUs are new every season. Can you forecast without history?

Yes, as a cold-start problem with a known shape. New items are mapped into the same feature space as the catalogue — category, brand, price tier, colour, material, size range — their nearest historical neighbours are found, and their scaled launch curves become the prior, tightened as real sales arrive. This is why catalogue enrichment pays twice: the attributes that fix search and feeds are the attributes that make cold-start forecasting possible.

Cash on delivery is most of our volume. Does that break any of this?

It changes two things. Refused parcels are a real cost line — outbound shipping, return shipping, handling and stock reserved for three days — so orders get risk-scored at checkout on address quality, basket composition, history and office-versus-door delivery, and the high-risk tail is routed to prepayment or a confirmation call rather than blocked. Second, and more often missed: a refusal is not a deleted sale. The customer wanted the item, so it stays in demand; what gets modelled separately is a refusal probability by region, courier, office-versus-door and basket, netted at the replenishment stage. That way pick capacity and reservations stay sized on gross orders while replenishment stops ordering against sales that never settled.

How is this different from what our agency already does with Klaviyo and Google Ads?

Campaign tooling optimises what to send. This work changes what is true underneath it: whether the product data is complete enough to be found, whether the item is actually available at the node the order will source from, whether the segment a flow fires on is scored from your order and catalogue data rather than from a 30-day time window, and whether reported uplift survives a holdout. If your ad platform and your P&L disagree by a factor of two, no amount of creative testing closes that gap — a permanent 5–10% control per flow and geo-holdout tests do.

What do we own at handover, and can our own people run it?

The models are the least valuable artefact. What you keep is the feature store, the dbt transformations, the reconciliation jobs, the attribute schema in your PIM, the price-history table and the runbooks — in your infrastructure, in open formats, with the retraining pipeline documented and drift monitors and alert thresholds handed over as operating procedure. Handover includes training a named person on your side to retrain and re-score. If nobody there can do it at the end, the engagement failed whatever the metrics said.

Warm light ribbons on a dark field

Book a 30-minute review of your demand and stock data

Bring one month of orders, a stock snapshot from the same period and your top 100 SKUs by revenue. On the call we will tell you what share of your sales history is censored by stockouts, whether your reservations have drifted, and which of the four loops — forecast, price, catalogue or the integration layer underneath them — is costing you the most right now.

You talk to the engineer who would do the work, and if the answer is a cycle count and a corrected size chart rather than a model, we will say that too.