Gelişimİş
0

Skill vs. Workflow: Why the Future Belongs to People Who Can Design Their Own AI Workflows

A person at a sunlit workbench arranging colorful wooden blocks into a chain beside a laptop, designing a workflow step by step

TL;DR. Most AI upskilling still means learning tools: which model, which prompt, which feature. The research points somewhere else. In a field experiment with 758 Boston Consulting Group consultants, GPT-4 raised the number of tasks completed by 12.2% and quality by more than 40% on tasks the model handled well, but on a task just outside its abilities, consultants using it were 19 percentage points less likely to reach the correct answer. In a randomized study of experienced open-source developers, AI access made tasks take 19% longer, while the developers believed it had made them 20% faster. Across McKinsey’s survey of AI adoption, the redesign of workflows was the attribute most associated with bottom-line impact out of 25 tested. The common thread: the result depends less on the tool than on the decisions around it. Which tasks go to the model, where a human checks, what gets reused. That is a skill in itself, and it is learnable. This guide lays out the evidence and gives you a six-part canvas for designing your own AI workflows.

What is the difference between a skill and a workflow?

A tool skill is knowing how to operate something: writing a good prompt, using a feature, choosing a model. A workflow is the sequence of steps that turns an input into a result, including who (or what) does each step, in what order, and where the output gets checked.

Tool skills are necessary, but they depreciate quickly. The model you mastered last year has been replaced, and the prompt tricks that worked on it may not transfer. Workflow design is closer to a management capability: it transfers across tools because it is about how the work is divided, not which button is pressed. In our decision framework for which AI tools are worth learning deeply, we argued that you should go deep on very few tools. This piece is the other half of that argument: what to go deep on instead.

The evidence: same tool, very different results

The strongest evidence comes from studies where everyone had access to the same AI, and the outcome still varied widely depending on how the work was structured.

Table 1. Same AI, different outcomes: what the controlled studies found

Study Who and what Result with AI What made the difference
Dell’Acqua et al. (2023), Harvard Business School working paper 758 BCG consultants, GPT-4, realistic consulting tasks Inside the model’s capabilities: 12.2% more tasks, 25.1% faster, more than 40% higher quality. Outside them: 19 percentage points less likely to be correct Whether the task sat inside or outside the “jagged frontier” of what the model does well
Becker et al. (2025), METR, arXiv 16 experienced open-source developers, 246 tasks in their own repositories Tasks took 19% longer with AI allowed; developers forecast a 24% speedup and afterwards believed they got 20% Mature, familiar codebases and high quality standards, where checking and correcting AI output cost more than it saved
Brynjolfsson, Li and Raymond (2023), NBER working paper 5,179 customer support agents, AI assistant 14% more issues resolved per hour on average; 34% for novice and low-skilled workers; minimal impact for the most experienced Who used it: the AI spread the practices of top performers to newer agents
Noy and Zhang (2023), Science 453 professionals, ChatGPT, writing tasks Time fell 40%, quality rose 18% The structure of the task: rough drafting shrank, editing grew
Dell’Acqua et al. (2025), NBER working paper 776 Procter & Gamble professionals, product development Individuals with AI matched the performance of teams without AI AI took on part of what a teammate normally contributes

Two details from these studies matter more than the headline numbers.

First, the consultants who were given a short prompt-engineering overview in addition to the tool completed about 93% of tasks, compared with about 91% for those who received only the tool and about 82% for the control group. Tool training helped, but only marginally. The large swings came from which task the tool was applied to.

Second, in the writing study, 68% of the participants given ChatGPT submitted its first output without editing it. The tool did not decide that. The workflow did, or more precisely, the absence of one.

Why the gains stall at the level of the whole economy

If AI makes individual tasks so much faster, why are the measured gains across the workforce small? Two large studies give a consistent answer.

Table 2. AI time savings at the scale of the labor market (verified data)

Source Scope Finding
Bick, Blandin and Deming (2025), NBER working paper Nationally representative US survey Users save 5.4% of their work hours; across all workers, that equals 1.4% of total work hours. Between 1% and 5% of all work hours are assisted by generative AI
Humlum and Vestergaard (2026 revision), NBER working paper 25,000 workers in 7,000 Danish workplaces, 11 exposed occupations Users save about 3% of their work hours; effects on earnings and hours are small enough to rule out changes larger than 2%; employers “absorb AI through task reorganization”
McKinsey, The State of AI (March 2025) 1,491 respondents, global Of 25 attributes tested, workflow redesign had the biggest effect on the ability to see an EBIT impact from AI; 21% of organizations had fundamentally redesigned at least some workflows
McKinsey, The State of AI in 2025 (November 2025) 1,993 respondents, global 55% of AI high performers had fundamentally redesigned individual workflows, against 20% of other organizations
Microsoft Work Trend Index (April 2025) 31,000 workers in 31 markets 38% of leaders expect their teams to be redesigning business processes with AI within five years

The pattern is clear. When AI is dropped into an existing workflow, it saves a few percent of time, and that time often goes into other tasks: in the Danish data, most users say they reallocate the time they save. When the workflow is redesigned, the gains show up in results. A caution on the McKinsey figures: they are self-reported survey correlations, and the researchers’ model explains only a fifth of the variation in reported impact. Redesign is associated with impact, which is not the same as proven to cause it. The controlled studies in Table 1 are what make the direction plausible.

What workflow design actually involves

The consulting study gives the clearest picture of what skilled AI users do differently. The researchers observed two distinct patterns among successful users:

  • Centaurs divided the work, “dividing and delegating their solution-creation activities to the AI or to themselves.” They kept a clear line between human and machine tasks.
  • Cyborgs integrated completely, “completely integrating their task flow with the AI and continually interacting with the technology.”

The authors describe this as early analysis, and they did not test which pattern performs better. The lesson is not “be a centaur” or “be a cyborg”. It is that successful users made a deliberate choice about the division of work, rather than asking the tool to do everything or nothing.

McKinsey’s 2025 data points to the same skill at the organizational level: one of the practices that most distinguished high performers was having defined processes for when and how model outputs need human validation. That is a workflow decision, not a tool feature. We cover it in detail in human in the loop by design.

The Workflow Design Canvas

The canvas below translates the research into six design decisions you can make for any recurring piece of work. It is a CEOtudent editorial framework, not a result from a single study; each element is linked to the evidence that motivates it.

Table 3. The Workflow Design Canvas: six decisions (CEOtudent editorial framework)

# Decision Question to answer Why it matters (evidence)
1 Decompose Which distinct steps does this work contain? Gains are task-specific; the writing study saw drafting time fall while editing grew (Noy and Zhang)
2 Map the frontier For each step, is the model reliably good, unreliable, or untested? Inside the frontier, big gains; outside it, 19 points fewer correct answers (Dell’Acqua 2023)
3 Choose the handoff Split the work (centaur) or interleave continuously (cyborg)? Both patterns appeared among successful users (Dell’Acqua 2023)
4 Place the checkpoints Where does a human verify, and against what standard? Defined validation processes distinguish AI high performers (McKinsey 2025)
5 Measure honestly How will you know it is actually faster or better? Developers felt 20% faster while being 19% slower (METR 2025)
6 Capture and reuse What prompt, template or checklist should be saved for next time? Small individual savings only compound when they are built into the process (Humlum and Vestergaard)

A worked example: a weekly market briefing

Suppose you write a weekly one-page briefing for your team on developments in your market.

  1. Decompose. Collect sources, extract key facts, check the facts, interpret what they mean for the team, write, edit.
  2. Map the frontier. Summarizing long documents is well inside current models’ capabilities. Judging what matters for your specific team is outside it. Factual accuracy of summaries is unreliable, so it needs checking.
  3. Choose the handoff. Centaur-style: the model summarizes each source; you choose the three developments that matter and write the interpretation.
  4. Place the checkpoints. Every number in the final draft is checked against its source before sending. This is non-negotiable, because a wrong figure costs more trust than the time saved.
  5. Measure honestly. Log the actual time for four weeks before and four weeks after. Ask two readers whether the briefing got more useful.
  6. Capture and reuse. Save the summarization prompt and the fact-check checklist in one document you open every week.

Notice that only step 2 required knowing anything about the tool. The other five are about the work.

Tool skill vs. workflow skill: what to build

Table 4. Where to invest your learning time (CEOtudent editorial framework)

Tool skill Workflow skill
What it is Operating a specific model or app well Deciding how work is divided among you, AI and others
Shelf life Short: tools and interfaces change every few months Long: transfers to new tools
How you learn it Tutorials, documentation, practice with the tool Mapping your own work, experiments, measuring results
Evidence of payoff Small in the consulting study (about 93% vs 91% task completion with training) Large: the difference between inside- and outside-frontier results
Signal to employers “I know tool X” “I redesigned process Y and it now takes Z”

This is not an argument against learning tools. You cannot map the frontier (decision 2) without hands-on experience. It is an argument about proportion. An hour spent redesigning one recurring workflow usually teaches you more about the tool than an hour of tutorials, because it forces you to find where the tool fails.

The CEO and the student in workflow design

The CEO side of CEOtudent is ownership of the result. A CEO does not ask whether a new system is impressive; they ask whether it changes the output, the cost or the risk. Treating your own work as a set of processes you are responsible for redesigning is the same move at personal scale. It is also what employers increasingly reward: in our analysis of the new job titles of 2026, the fastest-growing roles are defined by outcomes that AI helps produce, not by tools.

The student side is the experimental habit. Workflow design is not a one-time plan but a loop: design, run, measure, adjust. The developers in the METR study are a warning here. Their sense of being faster was confidently wrong. The only protection is measuring. If you want a structured starting point, the Workflow Audit walks through mapping and scoring your current work, and the 5-layer AI workflow shows how the pieces fit into a working day. To learn a new domain the same way, see how to build a personal curriculum with AI.

Frequently asked questions

Is prompt engineering still worth learning?
Some of it, yes. In the consulting study, a short prompt-engineering overview slightly improved results. But the gains were small compared with the difference between tasks inside and outside the model’s capabilities. Learn enough to test the model on your own tasks, then invest the rest in workflow design.

What is the “jagged frontier”?
It is the researchers’ term for the uneven boundary of what AI does well. Tasks that seem equally difficult to a human can fall on different sides of it. The only reliable way to find the boundary for your work is to test the model on your actual tasks and check the results.

Why did developers get slower with AI?
The METR researchers studied experienced developers working in large, mature projects they knew well, with high quality standards. In that setting, reviewing and correcting AI output took more time than it saved. The result is specific to that setting, but the lesson that feeling faster is not the same as being faster applies everywhere.

How do I start designing my own AI workflows?
Pick one recurring task that takes at least an hour a week. Work through the six decisions in the canvas, measure the time before and after for four weeks, and keep the version that is actually better.

Do I need to know how to code?
No. Most of the canvas is about dividing and checking work. Automation tools can help later, but the design decisions come first.

Sources

  1. Dell’Acqua, F., McFowland III, E., Mollick, E., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., Lakhani, K. R. (2023). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. Harvard Business School Working Paper 24-013.
  2. Becker, J., Rush, N., Barnes, E., Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR, arXiv:2507.09089.
  3. Brynjolfsson, E., Li, D., Raymond, L. R. (2023). Generative AI at Work. NBER Working Paper 31161.
  4. Noy, S., Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science, 381(6654), 187-192.
  5. Dell’Acqua, F. et al. (2025). The Cybernetic Teammate: A Field Experiment on Generative AI Reshaping Teamwork and Expertise. NBER Working Paper 33641.
  6. Bick, A., Blandin, A., Deming, D. J. (2025). The Rapid Adoption of Generative AI. NBER Working Paper 32966.
  7. Humlum, A., Vestergaard, E. (2026). Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI. NBER Working Paper 33777 (revised March 2026).
  8. McKinsey & Company (2025). The state of AI: How organizations are rewiring to capture value (March 2025); The state of AI in 2025: Agents, innovation, and transformation (November 2025).
  9. Microsoft (2025). 2025 Work Trend Index: The Year the Frontier Firm Is Born. April 23, 2025.

This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.

This post is also available in: Türkçe Français Español Deutsch

Benzer içerikler