Why Shelf-Execution Pilots Stall on Hardware Procurement

Shelf-execution pilots usually stall at a procurement gate, not a technical one. How to sequence detection scope before any camera capital request.

Why Shelf-Execution Pilots Stall on Hardware Procurement
Written by TechnoLynx Published on 01 Sep 2026

The shelf-execution pilots we see stuck are rarely stuck on model accuracy. They are stuck at a capital gate: someone specified a store-wide camera estate before anyone scoped which detection classes the pilot actually needed, and the project now waits on a budget cycle it has no way to influence.

That sequencing decision — hardware as precondition rather than hypothesis — is the failure mode. It is worth naming precisely, because it is almost always misread as something else: a vendor problem, an IT problem, or a model that “wasn’t ready”.

Where exactly the pilot stops

The stall has a recognisable shape. A merchandising or retail-operations sponsor gets internal agreement that on-shelf availability is worth attacking. The first artefact produced is not a detection spec — it is a sensor spec, usually written with a hardware vendor’s help, priced across the estate because finance asks “and what does this cost at 400 stores?”. That number goes into a capital request. The CV pipeline is described in one paragraph as “AI analytics”.

Then nothing moves for two or three quarters.

By the time approval lands — if it lands — three things have usually changed. The sponsor has moved or been reorganised. The planogram has been reset at least once, so the reference layout the pilot was scoped against is stale. And a share of the SKUs in scope have had packaging refreshed, which matters because artwork churn is one of the main sources of silent detection drift, as we discuss in where shelf-execution pipelines quietly degrade.

A pilot gated on new hardware does not produce evidence; it produces a business case for a business case. That is the whole failure in one line.

Why hardware-first makes funding harder, not easier

The intuition behind hardware-first is reasonable: you cannot detect what you cannot see, so secure the seeing. In practice it inverts the risk profile that a capital committee is willing to accept.

An estate-wide sensor request asks for a large, irreversible spend against an unproven detection capability. There is no baseline on-shelf-availability number, no measured planogram-compliance rate, no evidence that a flagged gap actually shortens time-to-restock. The committee is being asked to fund the expensive half of the programme first and take the valuable half on faith. Reasonable committees defer that, and deferral reads internally as “the AI project is stalled”.

There is a second, quieter cost. A capture spec written before the detection classes are fixed is almost never tied to the image conditions those classes need — effective pixels per facing, viewing angle, capture cadence. So even when the hardware arrives, it may not support the task it was bought for. The trade-space for that decision is the subject of our sibling piece on what to reuse versus what to procure in shelf monitoring hardware; this article is about the sequencing error that happens upstream of it.

What existing capture can usually cover

Reuse is not a compromise position. In the retail-CV work we do, two capture sources already in the aisle carry a surprising amount of the detection load:

  • Existing loss-prevention and operations cameras. Coverage is uneven and viewing angles are rarely ideal, but a meaningful subset of fixtures in most stores is already observed at a resolution and cadence sufficient for gross on-shelf stock-out detection — an empty or near-empty facing is a large, low-frequency visual signal.
  • Store-associate mobile devices. Handheld capture during existing rounds gives near-ideal framing and resolution for the harder classes — planogram compliance at facing level, price-tag and promotional checks — at the cost of being sampled rather than continuous.

The two are complementary rather than redundant: fixed cameras give cadence, mobile capture gives fidelity. A pilot built on both can produce real numbers for both detection classes in scope without a single purchase order. Our approach to establishing what the existing footprint can support is engineering work, not procurement work — an audit of the current capture and compute estate, done before any spec is written. This is the same reasoning we apply in computer vision engineering generally: measure the constraint before buying around it.

Sequencing that produces evidence first

Step Hardware-first pilot Evidence-first pilot
1 Specify store-wide sensor estate Fix the detection classes in scope (planogram compliance, on-shelf stock-out)
2 Price rollout across estate Audit existing camera and mobile capture against those classes
3 Submit capital request Establish baseline OSA and compliance rates by manual shelf audit
4 Wait for budget cycle Deploy pipeline on reused capture in a small store cohort
5 Re-scope after planogram/packaging churn Measure lift; close the loop into a restock task
6 Begin pipeline scoping Request hardware only for aisles reuse cannot cover
Time to first evidence Multi-quarter, dependent on approval Weeks, controlled by the project team

The metric that distinguishes these is time-to-first-evidence: weeks from kickoff to one store with a baseline and a post-deployment on-shelf-availability and planogram-compliance rate. Reuse puts that number inside the project team’s control. A capital gate hands it to a calendar. Measuring the lift itself has its own methodological traps, which is why we treat it as a separate discipline in how to measure on-shelf-availability lift.

When new hardware is genuinely needed

Sometimes reuse cannot cover an aisle. Deep shelves with stacked facings, chilled cabinets behind glass, and end-cap displays outside camera arcs are common gaps. The honest response is not to pretend reuse works — it is to narrow the ask.

A pilot that has already produced measured lift on reused capture can go to the same committee with a different request: a handful of gap-fill positions, each justified by a named fixture and a named detection class that reuse demonstrably failed to serve. That is a small, reversible spend attached to evidence. It is a materially easier approval than an estate-wide spec attached to a hypothesis, and it is the difference between a hardware request that gets funded and one that gets deferred indefinitely.

The misdiagnosis: “the model isn’t accurate enough”

This failure mode frequently arrives wearing a different mask. The pilot has been scoped hardware-first, the capital request is in limbo, so someone runs a proof of concept on whatever footage is available — badly framed, wrong cadence, no baseline — and the detection numbers are poor. The conclusion drawn is that the model needs work.

Sometimes it does. More often the capture conditions were never derived from the detection task, so the model is being asked a question the imagery cannot answer. The tell is diagnostic: if accuracy varies sharply by fixture rather than by SKU, the constraint is capture geometry, not the detector. Retraining a PyTorch detector on more of the same unsuitable frames will not move it, and neither will a TensorRT optimisation pass — throughput was never the problem.

Early warning signs

  • The first document produced is a sensor specification, not a detection-class list.
  • No one on the project can state the current on-shelf-availability baseline.
  • The rollout is priced at estate scale before a single store has been instrumented.
  • Hardware vendor conversations are further along than pipeline scoping conversations.
  • The success criterion is “cameras installed” rather than “availability measured”.

Any two of these together mean the pilot is on the procurement track, not the evidence track.

A pilot that has already stalled is usually salvageable, and salvage is cheaper than restart. Detach the capital request from the pilot, re-scope to the two or three detection classes with the clearest operational consequence, and find the stores where existing capture is best — not average — to get a first measurement out. Our work with retailers on store operations and retail AI tends to start at exactly this point, and the same pattern shows up on factory floors, where industrial CV programmes stall on plant-camera procurement for identical structural reasons.

Frequently Asked Questions

What does “shelf-execution pilots stall on hardware procurement” mean in practice — where exactly does the pilot stop?

With Shelf Execution Pilots Stall, the detail that matters is this. It stops after the sensor specification and before the CV pipeline is scoped. A capital request for an estate-wide camera estate enters a budget cycle, and the pilot has no mechanism to accelerate it. Two or three quarters later the sponsor, the planogram, and often the SKU packaging have all changed, so the original scope no longer describes the store.

Why does specifying camera hardware before scoping the detection classes make the pilot harder to fund, not easier?

It asks a committee to approve a large, irreversible spend against an unproven capability, with no baseline availability number and no measured lift to anchor it. It also risks a capture spec that was never tied to the image conditions the detection classes need — pixels per facing, viewing angle, cadence — so the hardware may not support the task once it arrives.

What can existing store cameras and associate mobile devices actually cover for planogram compliance and stock-out detection?

Existing loss-prevention and operations cameras generally cover gross on-shelf stock-out detection on a useful subset of fixtures, because an empty facing is a large visual signal. Associate mobile capture during existing rounds gives the framing and resolution needed for facing-level planogram compliance and price-tag checks. Fixed cameras contribute cadence; mobile contributes fidelity.

How do we structure a shelf-execution pilot so it produces on-shelf-availability evidence before any capital request?

Fix the detection classes first, audit what existing capture can support against them, establish a baseline on-shelf-availability and planogram-compliance rate by manual shelf audit, then deploy on reused capture in a small store cohort. The target metric is time-to-first-evidence in weeks, and reuse keeps that number under the project team’s control.

When is new hardware genuinely required, and how do we narrow the ask to gap-fill positions?

When a specific fixture class cannot be served by reuse — deep stacked shelves, glass-fronted chillers, end-caps outside camera arcs. Narrow the request to those named positions, each tied to a named detection class that reuse demonstrably failed, and submit it after measured lift exists. A small reversible spend attached to evidence approves far more readily than an estate-wide spec.

What organisational signals indicate a pilot is about to stall — and what can be salvaged from one that already has?

The clearest signals are a sensor spec produced before a detection-class list, no stated availability baseline, estate-scale pricing before a single instrumented store, and “cameras installed” as the success criterion. Salvage by detaching the capital request, narrowing to the two or three highest-consequence detection classes, and taking a first measurement in the stores with the best existing capture.

What does this failure mode look like when it is misdiagnosed as a model-accuracy problem?

A proof of concept runs on whatever footage exists — wrong framing, wrong cadence, no baseline — the numbers look poor, and the team concludes the detector needs retraining. The diagnostic tell is that accuracy varies by fixture rather than by SKU, which points at capture geometry rather than the model.

Hardware lead times are now the longest pole

Camera and edge device orders stretch to 16–20 weeks in 2026, while software integration typically closes in six. Revisit it when your workload shifts.

Back See Blogs
arrow icon