Most frameworks in this space are sold finished and never tested against the evidence. This one is different. Each component below carries the idea, where the evidence stands, and the open question I am actively working. The strongest parts are load-bearing today; the open questions are the frontier I am closing, in the open. That is the difference between a doctrine and a working framework.
01Well-supported
The delta
The gap between the documented workflow and how the work is actually done is the most valuable signal in the building. Map it before you automate, or AI just scales the version of the work no one has looked at.
The open questionThe field hasn’t settled how to observe real frontline work — across shifts, roles, and stores — affordably and at scale. Direct observation is costly; digital traces miss most of the physical floor.
What I’m testingA delta-mapping pass an operator can run in a week, without a research budget.
02Well-supported
Worker impact
AI that touches schedules, workload, or evaluation changes the job, not just the output. Designing for the people closest to the work is not a soft concern: a peer-reviewed retail scheduling study found that more stable, worker-friendly scheduling raised store productivity by 5.1%.
The open questionDoes formally assessing worker impact actually change what gets deployed — and who holds the authority to halt a rollout when the review surfaces real harm?
What I’m testingAn impact review light enough to run per use case, with a genuine veto attached.
03Well-supported
Hidden burden
A tool that saves one group time often quietly loads another. In healthcare, the electronic record was sold as productivity and left clinicians nearly two hours of desk work for every hour with patients. Retail’s task engines and alert systems carry the same risk.
The open questionNo one has measured the minutes-per-shift that AI tools actually add to a frontline associate’s day. Until someone does, burden gets scored on intuition.
What I’m testingA simple time study of AI-created work on the floor, by role.
04Well-supported
Lateral knowledge
The best routine in the building usually already runs in one aisle, with one team. The research on internal transfer is clear: the barrier isn’t motivation, it’s that the people who do it best often can’t say why it works.
The open questionCan AI-assisted capture preserve enough of that context to move a practice — or does it just industrialize cargo-cult copying of what happens to be visible?
What I’m testingCapturing a top team’s routine with enough context that a second team can actually run it.
05Well-supported
Pilot to scale
A pilot that works in three good stores tells you about three good stores. The economics of scale-up are unforgiving: effects shrink or vanish at rollout for structural reasons, not implementation mistakes. Validation is the stage most programs skip.
The open questionWhat does “validated” concretely require — which sites, how long, and how much validation for which level of risk? The word carries weight that is rarely specified.
What I’m testingA validation standard that deliberately picks unaccommodating stores, not flattering ones.
06Emerging
The translation gap & the Workflow-First AI Leader™
Strategy doesn’t reach the floor intact; it’s reinterpreted at every level until intent is lost. Someone has to own the translation from AI capability to frontline behavior. I named that role the Workflow-First AI Leader™.
The open questionDoes a single accountable translator actually close the gap, or become a bottleneck? The problem is well documented; the best structure for solving it isn’t.
What I’m testingWhether the role works better as a person, or as a system that outlives any one person.
07Emerging
Leadership codification
The habits that make a great operator — how they coach, escalate, run a huddle — are worth capturing. Recent evidence shows AI can diffuse a top performer’s patterns and lift the people working below them.
The open questionWhich leadership moves are genuinely codifiable, and which are irreducibly tacit? The claim that you can convert judgment into an explicit asset is contested in the research, in both directions.
What I’m testingCodifying one coaching routine and checking whether leaders who didn’t invent it actually improve.
08Emerging
The Operator’s Scoreboard™
AI value lags while you redesign the work, so measuring only lagging financials hides both healthy progress and real failure. The Operator’s Scoreboard™ separates early learning signals from business outcomes.
The open questionDo the leading indicators — learning velocity, delta resolution, trust — actually predict the business outcomes? And once they become targets, do they stop measuring what they were built to measure?
What I’m testingRunning the learning metrics alongside outcomes long enough to see which ones truly lead.
09Emerging
Bias & fairness
AI that shapes schedules, staffing, or evaluation inherits the bias in its own history. A serious review names that before it scales, not after.
The open questionThe math is settled and uncomfortable: you can’t satisfy every definition of fairness at once, so “fair” is a choice, not a checkbox. And no one has defined what fairness even means for allocating shifts and tasks.
What I’m testingMaking the fairness definition an explicit, owned decision per use case.
10Contested
The Five Stages
Ad-hoc, Mapped, Validated, Codified, Compounding — a shared language for locating where an operation actually is, so a team can agree on the next move instead of arguing about the destination.
The open questionStaged maturity models are popular and largely unproven — the genre has never shown that climbing the stages causes better outcomes. So I hold my own to that bar: which transitions actually move results? Until that’s answered, the stages are a diagnostic vocabulary, not a promise.