---
title: AI workflow redesign
description: The four-phase method for turning a recurring job into an AI workflow (Map, Test, Integrate, Compound), with what happens in each phase and what it produces, plus the one-shot fallacy, communication as the core skill, and the JIOPR delegation framework. Consult when guiding someone through redesigning a workflow, when someone is partway through and needs the next phase, or when a user asks how the phases fit together.
category: methodology
updated: 2026-09-25
---

# AI workflow redesign

**AI workflow redesign** is a four-phase method for turning an existing, recurring job into AI-augmented work without changing roles, titles, or violating organizational security policies. The phases run in order. Map breaks the job into steps and finds the one worth handing to AI, Test tries an AI version of that step on real work, Integrate builds the tested version into everyday work, and Compound makes it better every time it runs. Rather than replacing human work wholesale, the method moves specific steps to AI while keeping human judgment where it matters most.

The method addresses a common failure pattern: organizations either attempt complete automation of complex processes or implement AI superficially without understanding their actual work structure. The later phases exist because a redesign can also stall after the design is done: the instructions get written but never become part of the week, and nothing improves from one run to the next.

## The one-shot fallacy

A persistent misconception treats AI interaction as a vending machine: insert a prompt, receive a finished product. This one-shot mentality leads to repeated failures when practitioners attempt to compress what is fundamentally a multi-phase workflow into a single prompt, reasoning that since all the information exists, the AI should be able to produce the final output directly.

Consider transforming a three-day intensive workshop into a semester-long course. All the source content exists—topics, materials, structure. The temptation is strong to write one comprehensive prompt and expect the AI to deliver a complete course design. This approach fails not because the AI lacks capability, but because even a human expert wouldn't perform this transformation in a single sitting. Complex work involves iteration, judgment calls at multiple points, and progressive refinement that cannot collapse into one operation.

The recognition that apparent single tasks are actually hidden workflows fundamentally changes how to approach AI collaboration. What seems like "write this document" decomposes into creating an outline, drafting sections individually, reviewing for consistency, adjusting for audience, and refining voice. Each phase may involve different levels of AI assistance and different types of human judgment. Some phases benefit from heavy AI involvement while others require minimal automation.

Once a hidden workflow surfaces, the practitioner can design appropriate AI involvement for each phase, document feedback to improve future iterations, and build reusable processes rather than one-time prompts. The investment in workflow design pays dividends through repeated use, while one-shot attempts remain perpetually frustrating. See: what-is-a-workflow for a worked before-and-after example of one job with and without AI.

## Communication as the core skill

The term "prompting" misleads practitioners into thinking about magic formulas—lists of tricks that transform AI output through clever phrasing. This framing encourages copy-pasting prompts from the internet and expecting vending-machine results. A more useful frame treats AI interaction as communication, where the same skills that make humans effective communicators make them effective AI collaborators.

Structured, thoughtful communicators achieve the best AI results. Professionals trained to communicate clearly—learning designers, technical writers, compliance specialists—often discover they already possess the core skills. They know how to specify requirements precisely, anticipate misunderstandings, and articulate implicit knowledge that others might take for granted.

The critical shift involves making explicit what has previously remained tacit. Every domain contains assumptions, conventions, and contextual knowledge that practitioners never needed to articulate because human colleagues shared the same background. AI lacks this shared context. Improving AI output often requires surfacing these unstated assumptions—the unwritten rules, the industry conventions, the team preferences that shape what "good" looks like in a specific context.

When AI produces unsatisfactory output, effective practitioners ask what assumption they hold that they haven't communicated. The failure usually traces not to AI limitation but to implicit knowledge that remained unexpressed. This reframe—from "the AI failed" to "what did I not communicate?"—creates a productive feedback loop where each interaction surfaces more of the tacit knowledge that enables better future collaboration.

## The four phases

Each phase ends with something concrete, and the next phase starts from it.

| Phase | What happens | What you have afterwards |
|-------|--------------|--------------------------|
| Map | Describe the work, choose one recurring job, and break it into steps | A map of the job and one step suited to AI |
| Test | Design the AI-supported version of that step and try it on real work | A real result to judge |
| Integrate | Build it into everyday work and check it on a few more real tasks | A workflow that actually gets used |
| Compound | Work out what gets better every time it runs, and set that up | An improvement mechanism, a run log, and a plan for the coming weeks |

The phases do not have to happen on the same day. Integrate and Compound can wait until the test result has held up. Each workflow keeps one running document, where every phase records its decisions, what is still open, and, if the work paused partway, exactly where it stopped. Each phase starts by reading that document instead of redoing the one before it.

For where the AI Fluency framework's four competencies (delegation, description, discernment, and diligence) sit across these phases, see: ai-fluency-framework.

### Map

Map starts with context: the person's role and responsibilities, their function and industry, which AI tools they are allowed to use and on what budget, and constraints such as data sensitivity, compliance, or confidentiality. Naming these first prevents designing around unapproved tools or ignoring data-handling rules that would block the build later. It is also the moment to place the person on the four modes of AI work (chat, context, clockwork, colleague): which modes they already use, and which would be the natural next addition. Someone who only chats usually adds a custom assistant next, someone with custom assistants usually moves one workflow onto a schedule or trigger, and someone already running automations deepens and multiplies what they have. The workflow being redesigned does not have to take whichever mode comes next; its own mode is decided in Test. See: ai-work-modes.

The person then lists three to five recurring workflows they own or drive, such as weekly status updates, customer meeting preparation, campaign planning and reporting, monthly forecasting, support ticket triage, contract review, or hiring pipeline management. Two questions cross candidates off before any is chosen: should this work exist at all, and does an existing tool already handle 80 to 90 percent of it? See: ai-productivity-traps. From what remains, one workflow is picked: frequent, genuinely effortful rather than merely exciting to automate, and low-to-medium risk. Workflows that are too vague, too strategic, or too irregular make poor first candidates. See: choosing-a-first-workflow for the "week on repeat" exercise and the three criteria (structured, repetitive, easy to verify) used to build and trim the list.

The chosen workflow is broken down through structured questions: what triggers it, what inputs come from which sources, which steps happen in what order, where judgment calls are made and what they decide, what outputs are produced and who uses them, how quality gets checked, which tools are involved, and how often it runs and at what volume. The answers become a table with one row per step, recording each step's type (data gathering, transformation, decision, communication, coordination), how easy it is to verify, how much damage a wrong result would do, and how often it happens.

Each step is then classified by who should do it once AI is involved. A step suits **AI** when it is repetitive, structured, and easy to verify, typically pattern matching, summarizing, rewriting, classifying, or filling templates. A step is done by **AI and you** together when it needs judgment but AI can prepare drafts, options, or a first analysis. A step stays with **you** when it involves high stakes, political sensitivity, ambiguity, or tacit knowledge that resists being written down. The five-point filter (verifiable, step-wise and bounded, recurring and painful, inputs and outputs known, human in the loop) tests each candidate step. See: agent-use-case-evaluation. Map is also where the first diligence question belongs: which steps touch personal or confidential data, and is the tool approved for it? A step that fails that question is not a candidate, however well it scores otherwise. See: ai-diligence.

Map ends with two or three good candidate steps highlighted, favoring high frequency and low risk, and one of them chosen to take into Test.

### Test

Test designs the AI-supported version of the chosen step and tries it on real work. The design settles six things. It compares the current way of doing the step with the proposed AI version: what AI does, what the person does, and which approved tool it runs in. It chooses the mode this workflow will run in, which does not have to match the person's overall experience with AI. See: ai-work-modes. It places the human in the loop: AI drafts and the person reviews, AI suggests options and the person chooses, or AI triages and the person handles the exceptions. It sets how much checking the work needs, calibrated to the stakes of one bad run: high-stakes single runs push toward chat or context with every run reviewed, while low-stakes runs make clockwork or colleague safe to design. See: ai-output-verification. It carries forward the governance constraints named at the start of Map. And it names the outcome metric, the concrete thing that should change. A workflow whose owner cannot name that metric is at risk of becoming a tool-shaped object, which produces the feeling of progress without the change. See: ai-productivity-traps.

The test itself runs in the same working session, and how it runs depends on the mode. For chat or context, the instruction is written with the JIOPR framework (below) and run on real data. For clockwork, the routine or scheduled task is configured with the instruction and run once on demand. For colleague, both the instruction and the success criteria the AI checks itself against are written, and a few iterations show whether the AI's own judgment of "ready" matches the person's.

The best test uses real material, because a made-up example does not show where the workflow goes wrong. Two checks come first: the workplace's rules on which AI tools may be used and what may be shared in them, and a backup of anything the assistant is connected to and could change or delete. When the output disappoints, the useful question is which assumption went uncommunicated. See the communication section above.

Test ends when the person has seen real output and judged it.

### Integrate

Integrate builds the tested version into everyday work, and what gets built depends on the mode. A custom assistant (context mode) is mostly a decision about what to load: style guides, past examples, working context, and constraints, so the assistant arrives knowing what it needs. A scheduled or triggered run (clockwork mode) needs bounded sources, an output destination the person already checks, visible failures, and human review of the output before anything takes effect. See: scheduled-automation. An autonomous loop (colleague mode) depends on success criteria precise enough for the AI to check its own work against. See: example-image-generation. Tool-specific setup follows the official documentation for the tools involved.

Before anyone relies on it, the workflow is validated on three or four real tasks against the outcome metric from Test, along with the governance constraints. Four situations come up at this point, and each needs a different response. If the output breaks or disappoints, the fix is usually in the instruction, where something about what to make, how to work, or how to behave was assumed but never said. If the workflow works but the person dreads running it, the collaboration is the problem: too much back-and-forth per run, which gets cut, or the workflow gets dropped, since meeting the metric while costing more attention than it saves is a trap. See: ai-output-verification. If it works but never becomes routine, it needs an anchor in an existing habit, such as a recurring meeting, a weekly slot, or a trigger event; vague intentions fail, and named anchors hold. If it works and the person wants more, the next moves are hardening and rollout.

Two decisions finish the phase. Someone is named as the owner, who vouches for each output that leaves the team, and if output reaches readers, customers, or partners, how AI involvement is disclosed is settled against the organization's norm. See: ai-diligence. When colleagues could benefit, each gets their own copy of the assistant to adapt, rather than one shared assistant that nobody owns. See: personal-agents.

Integrate ends with a workflow that actually gets used.

### Compound

Compound makes the workflow improve with use, and it starts with one question: what gets better the next time this runs, and why? A specific answer names a mechanism, for example corrected outputs saved to an examples file that the instruction tells the assistant to match. A vague answer, such as "the AI learns", means there is no mechanism, only hope.

A working mechanism has three parts, and a workflow missing any one of them runs without compounding. The signal is the feedback captured after a run: corrections made, rounds needed, time taken, quality issues, and edge cases. The storage is where that feedback lives so the next run sees it: a feedback section in the workflow's document, an examples file loaded into the assistant, or updated rules in the instruction. The refinement is what specifically improves, and when: the instruction is updated when a correction repeats, an example is added when an output was unusually good or bad, and a rule is added when an edge case appears. See: compound-engineering, and example-linkedin-outreach and example-content-interview for workflows with two mechanisms each.

The mechanism has to cost under a minute per run, or it will not survive a busy week: a two-line run log entry, a habit of saving good outputs as examples, and a 15-minute monthly review of the log. The run log also answers the last diligence question, whether anything went out wrong before it was caught.

Compound closes by looking forward: which workflow is the next candidate for the same treatment, and which colleague to help get started. It ends with the improvement mechanism in place and a short checklist for the coming weeks: validate against the outcome metric, make the workflow part of the routine, and name what improves next time.

## The JIOPR delegation framework

For any individual AI task within a redesigned workflow, effective delegation follows the JIOPR pattern. Job defines the specific outcome desired, not a vague request for help but a concrete deliverable in a specified format. Inputs identifies what the AI can use—which sources, documents, and context—and equally what remains off-limits. Output specifies what done looks like through columns, sections, file types, and formatting requirements. Rules establishes what the AI must do and must not do, recognizing that constraints matter as much as goals. Proof determines how results will be verified and what evidence the AI should provide for human checking.

The difference between vague and useful delegation is stark. "Research competitors" provides insufficient guidance. "Research the top 5 competitors in this industry, providing company name, founding year, estimated revenue range, key differentiator in one sentence, and source URL as CSV output, excluding companies outside this geography, noting 'Not disclosed' with search explanation when revenue data isn't available" enables reliable execution.

## Related pages

- See: what-is-a-workflow - What a workflow is, with a before-and-after example of which steps move to AI
- See: choosing-a-first-workflow - How to find and pick the workflow Map starts from
- See: agent-use-case-evaluation - The five-point filter applied to each step during Map
- See: ai-work-modes - The four modes, and how the mode for one workflow gets chosen in Test
- See: compound-engineering - Patterns behind the Compound phase
- See: ai-diligence - The data, disclosure, and ownership questions attached to each phase
- See: productive-friction - Which steps should stay human even when AI could do them
