Workflow-First AI™ · Frontline Retail

The work is 80% of the problem.
The model is the easy part.

Workflow-First AI™ is a working framework for making AI land on the retail floor. Ten components, built from twenty-eight years of operations and pressure-tested against the evidence. Here is what each one has established, and the question I am still working.

28 yrs operations700-person storeFull P&L

DEFINED
ACTUAL
Defined vs. actual — the gap is the delta.

The Framework

Ten components. What each one has proven — and what it hasn’t yet.

Most frameworks in this space are sold finished and never tested against the evidence. This one is different. Each component below carries the idea, where the evidence stands, and the open question I am actively working. The strongest parts are load-bearing today; the open questions are the frontier I am closing, in the open. That is the difference between a doctrine and a working framework.

Well-supported strong evidence base Emerging promising, still forming Contested honestly unsettled
01Well-supported

The delta

The gap between the documented workflow and how the work is actually done is the most valuable signal in the building. Map it before you automate, or AI just scales the version of the work no one has looked at.

The open questionThe field hasn’t settled how to observe real frontline work — across shifts, roles, and stores — affordably and at scale. Direct observation is costly; digital traces miss most of the physical floor.

What I’m testingA delta-mapping pass an operator can run in a week, without a research budget.

02Well-supported

Worker impact

AI that touches schedules, workload, or evaluation changes the job, not just the output. Designing for the people closest to the work is not a soft concern: a peer-reviewed retail scheduling study found that more stable, worker-friendly scheduling raised store productivity by 5.1%.

The open questionDoes formally assessing worker impact actually change what gets deployed — and who holds the authority to halt a rollout when the review surfaces real harm?

What I’m testingAn impact review light enough to run per use case, with a genuine veto attached.

03Well-supported

Hidden burden

A tool that saves one group time often quietly loads another. In healthcare, the electronic record was sold as productivity and left clinicians nearly two hours of desk work for every hour with patients. Retail’s task engines and alert systems carry the same risk.

The open questionNo one has measured the minutes-per-shift that AI tools actually add to a frontline associate’s day. Until someone does, burden gets scored on intuition.

What I’m testingA simple time study of AI-created work on the floor, by role.

04Well-supported

Lateral knowledge

The best routine in the building usually already runs in one aisle, with one team. The research on internal transfer is clear: the barrier isn’t motivation, it’s that the people who do it best often can’t say why it works.

The open questionCan AI-assisted capture preserve enough of that context to move a practice — or does it just industrialize cargo-cult copying of what happens to be visible?

What I’m testingCapturing a top team’s routine with enough context that a second team can actually run it.

05Well-supported

Pilot to scale

A pilot that works in three good stores tells you about three good stores. The economics of scale-up are unforgiving: effects shrink or vanish at rollout for structural reasons, not implementation mistakes. Validation is the stage most programs skip.

The open questionWhat does “validated” concretely require — which sites, how long, and how much validation for which level of risk? The word carries weight that is rarely specified.

What I’m testingA validation standard that deliberately picks unaccommodating stores, not flattering ones.

06Emerging

The translation gap & the Workflow-First AI Leader™

Strategy doesn’t reach the floor intact; it’s reinterpreted at every level until intent is lost. Someone has to own the translation from AI capability to frontline behavior. I named that role the Workflow-First AI Leader™.

The open questionDoes a single accountable translator actually close the gap, or become a bottleneck? The problem is well documented; the best structure for solving it isn’t.

What I’m testingWhether the role works better as a person, or as a system that outlives any one person.

07Emerging

Leadership codification

The habits that make a great operator — how they coach, escalate, run a huddle — are worth capturing. Recent evidence shows AI can diffuse a top performer’s patterns and lift the people working below them.

The open questionWhich leadership moves are genuinely codifiable, and which are irreducibly tacit? The claim that you can convert judgment into an explicit asset is contested in the research, in both directions.

What I’m testingCodifying one coaching routine and checking whether leaders who didn’t invent it actually improve.

08Emerging

The Operator’s Scoreboard™

AI value lags while you redesign the work, so measuring only lagging financials hides both healthy progress and real failure. The Operator’s Scoreboard™ separates early learning signals from business outcomes.

The open questionDo the leading indicators — learning velocity, delta resolution, trust — actually predict the business outcomes? And once they become targets, do they stop measuring what they were built to measure?

What I’m testingRunning the learning metrics alongside outcomes long enough to see which ones truly lead.

09Emerging

Bias & fairness

AI that shapes schedules, staffing, or evaluation inherits the bias in its own history. A serious review names that before it scales, not after.

The open questionThe math is settled and uncomfortable: you can’t satisfy every definition of fairness at once, so “fair” is a choice, not a checkbox. And no one has defined what fairness even means for allocating shifts and tasks.

What I’m testingMaking the fairness definition an explicit, owned decision per use case.

10Contested

The Five Stages

Ad-hoc, Mapped, Validated, Codified, Compounding — a shared language for locating where an operation actually is, so a team can agree on the next move instead of arguing about the destination.

The open questionStaged maturity models are popular and largely unproven — the genre has never shown that climbing the stages causes better outcomes. So I hold my own to that bar: which transitions actually move results? Until that’s answered, the stages are a diagnostic vocabulary, not a promise.

Explore the tools → Read the white paper →


Tools

Working tools, not slideware

The framework is only worth what it changes on Monday. These are free, browser-based tools that put its components to work — no login, no signup.

New · freeLive

Orchestration Canvas Studio

Design an autonomous AI agent before you build it. A team fills the canvas in under an hour — naming what the agent may never do, when it has to stop and ask a human, what one bad day costs, and who pulls the cord. Sign-off produces a real design document, so approval finally means something.

Open the Canvas → Read the guide →

FreeLive

Delta Auditor

Map the delta — the gap between the documented workflow and how the work is actually done. It is the first component of the framework and the most valuable signal in the building. Find it before you automate, or AI just scales the version of the work no one has looked at.

Try the Delta Auditor →


In Practice

The delta is the signal

Composite case

The order the model did not see

An AI ordering system recommends the daily produce buy. The veteran orderer overrides it for a cold front and a stadium match nearby, and on those days she is right. The override is not an exception to suppress. It is the variable the forecast never had.

Composite case

The phantom out-of-stock

A shelf scanner flags gaps for replenishment. A third are a planogram change the system has not caught; the shelf is not empty. The scoreboard counts flags closed. The real job is the availability the partner's triage protects. Those are not the same number.


The Long Form

The Operator’s Field Guide

The framework, codified in full — thirteen chapters on getting AI to survive contact with the floor. Chapter 1 is free to read; the rest arrive as early access through Field Notes. Pre-publication draft

  1. 01

    The Pilot That Died on a Tuesday

    A six-month AI pilot dies in a single shift — and nobody touched the model.

    Read the chapter → · 18 min · free

Twelve more chapters follow the ten components above from case to scoreboard: why most retail AI fails before it starts, the four operating pillars, responsible deployment, and the ninety-day diagnostic you can start Monday.

Field Notes

One operator’s take on making AI land in a real store, plus early-access chapters as they’re finished. Short, every other week, from the floor.

No spam. Unsubscribe anytime.

© 2026 Jordan K. Jones. All rights reserved. Pre-publication draft of Workflow-First: The Operator’s Field Guide (forthcoming).


About

Operator-academic

Jordan K. Jones

I have spent twenty-eight years in large-format grocery retail operations, including leading a 700-person store operation with full profit-and-loss responsibility. I can cite Brynjolfsson, Nonaka, and Edmondson in one paragraph and name the produce-forecasting workflow that broke on a Tuesday in the next.

That is why the framework carries its open questions in public. I pressure-test it against the peer-reviewed and practitioner literature and against the floor, and I report both what holds and what doesn’t. The frontline is where my work starts and where it is tested. Adoption is not coerced. Visibility is not weaponized. The people closest to the work keep their dignity.

Education
  • M.B.A. · St. Edward’s University2015–2017
  • B.S., Microbiology · The University of Texas at Austin1998–2001
Boards & advisory
  • Board Director · Imagine A Way2024 – present
  • Board Director · Alzheimer’s Association, Capital of Texas Chapter2023 – 2025
  • Vice President · St. David’s Advisory Board2022 – 2024

Contact

Let’s talk

The open questions above are the ones I pressure-test with operators and boards. If your team is scaling AI on the floor, hiring someone to lead operations, or bringing operating judgment onto a board, that is the conversation.