The work that looks easiest to automate can be among the hardest for an agent to finish. Upwork found web and software development performing best in an agent benchmark, followed by data science, while admin support had the lowest average across three models. Two models completed only 6% of admin-support jobs. For executives deciding where to automate, the useful question is increasingly whether the organization can reliably tell when a task is actually finished.

The easiest-looking work can be difficult to automate

That question is central to Upwork’s shift from calling itself a talent marketplace to a “work delivery platform.” A marketplace connects a client with a person; Upwork’s newer model starts with an outcome and determines which combination of people and machines should deliver it. Upwork has a commercial stake in this framing because a delivery model built around human-agent orchestration expands the role its platform can play. Andrew Rabinovich, who leads AI at Upwork, presented the underlying task-decomposition framing at AI4 and, drawing on two decades of platform data, estimates that work on Upwork can be decomposed into roughly 500 to 1,000 recurring “atomic units of work.”

Those units make simple automation hierarchies unreliable. Rabinovich cites building logos and writing summaries as low-level work that has already been automated, yet Upwork’s benchmark found writing completion ranging from 4% to 51% depending on the model, with Gemini at 4%. Meanwhile, web and software development led the tested categories. Work that sounds routine can still contain criteria an agent cannot consistently satisfy or an organization cannot cheaply evaluate.

The category differences expose a distinction between producing an artifact and establishing that the task is complete. An agent may create a document, spreadsheet, design or codebase while still failing requirements that determine whether the deliverable is acceptable. Once work is split into smaller units and assigned to machines, each unit needs a dependable definition of done. Automation economics depend on that definition because an output that still requires substantial work or judgment retains part of the production cost.

Upwork tested full completion

To measure that distinction, Upwork built the Human+Agent Productivity Index, or HAPI, around real client work rather than synthetic tasks. Its methodology, documented in a paper released through arXiv in December, used 322 jobs that paying clients had posted and human freelancers had already completed to client acceptance. Expert freelancers then created a job-specific rubric with between 5 and 20 acceptance criteria. They classified each criterion as critical, important, optional or pitfall, and completion required every critical and important criterion to pass.

That rule makes completion stricter than producing a plausible deliverable. Claude, Gemini and GPT-5 each performed work under those rubrics, so an output could satisfy individual criteria while still failing the job as a whole. For a buyer, that resembles the decision at delivery: whether the result can be accepted. Because Upwork operates the marketplace and benefits commercially from demonstrating a role for its platform in human-agent delivery, HAPI should be read as Upwork’s measurement framework and evaluated on its design.

The design first narrowed the workload to jobs that agents had a reasonable chance to complete. Upwork restricted the benchmark to fixed-price, single-milestone contracts with clearly defined scope, although their duration still ranged from nine hours to more than 100 days. Open-ended and complex assignments for which agents were judged to have no reasonable chance of success were excluded. Upwork describes those excluded assignments as the “vast majority” of work on its platform, so HAPI’s tested workload represents a deliberately tractable subset of client work.

Within that subset, the benchmark also constrained the agents’ operating environment. Each model could read the job post and attachments, formulate a plan, make the deliverable and stop, but it received no web search, outside data, additional tools or task-specific training. The authors describe the results as conservative relative to richer agent systems. A production agent equipped with browsers, APIs and specialized tooling could behave differently.

Those design choices affect interpretation in opposite directions. Excluding harder and more open-ended jobs gives agents a more favorable workload, while restricting tools gives them a less capable execution environment. HAPI consequently cannot provide a simple upper or lower bound for an organization’s deployment. Its useful methodological contribution is narrower: define completion in advance, then test whether the system achieves the entire required outcome.

Under that standard, agent-only full completion ranged from 4% to 68%, depending on model and category. Web and software development came first, data science followed, and admin support came last on average. No tested result reached universal completion, even when the experiment later introduced expert guidance. Because the range is broad, a generic claim about “agent capability” says little about a specific deployment unless task structure and acceptance conditions are also known.

Those acceptance conditions also change how leaders should interpret demonstrations. A system can rapidly create an artifact that looks advanced while missing a critical requirement buried in an attachment or violating an acceptance condition. HAPI counts that result as unfinished because the client’s required outcome remains incomplete. An automation business case needs the same discipline when its projected savings depend on work no longer requiring somebody else to finish it.

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.

Verification is a useful dividing line

The completion results make verification a useful workflow variable. Some outputs expose an objective check: code can compile and pass tests, while another model can check a translation. Rabinovich characterizes evaluation of qualitative work as largely unsolved because somebody with judgment must decide whether the output is good enough. He estimates that above 90% of Upwork’s platform work is qualitative, a claim that also matters commercially to Upwork because its delivery model can create demand for human judgment alongside agents.

For an organization, the first operational question is whether an automated output can be tested against an objective criterion. A test suite, reconciliation, compliance rule or numerical match can turn acceptance into a repeatable procedure. Work with such a criterion, contained within one function and format, is a strong candidate for autonomous deployment once measured reliability is high enough. Verification gives the organization a dependable way to detect failures; it does not by itself make the agent capable of avoiding them.

The next question is whether the work remains inside that contained context. A workflow that crosses backend development, frontend work, copy and brand approval introduces handoffs between functions and formats, even when individual components have objective checks. Automating those components can still make sense, but a person remains responsible for the handoffs. Because completion belongs to the whole workflow, component-level automation rates cannot substitute for end-to-end acceptance.

HAPI’s lead-generation example shows why objective criteria and execution capability have to be separated. The HAPI page provides five sample jobs with downloadable deliverables, including an assignment that required reducing a spreadsheet of mobile apps to companies headquartered in the US, finding two to four marketing contacts per company and filling nine specified columns. The rubric was objectively checkable, the work stayed in one function and format, and none of its criteria depended on taste. The agent-only deliverable still failed evaluation.

Expert guidance did not make that particular deliverable pass either. The case separates the ability to determine success from the ability to produce it: objective verification makes failure observable, but a model can still lack the data, reasoning, tool access or execution quality needed to satisfy the criteria. Verifiability is therefore a diagnostic for an automation plan. Autonomous operation still depends on measured execution reliability.

Expert intervention changes measured completion

Where expert feedback helped, the measured changes were substantial. On low-complexity work, Claude, Gemini and GPT-5 all had higher completion after one expert intervention. The exact results show both the size of the change and the different starting points.

Model Before intervention After intervention
Claude 39.8% 51.2%
Gemini 19.9% 32.3%
GPT-5 19.6% 33.5%

Those movements were absolute gains of roughly 11 to 14 percentage points, while Upwork also markets an “up to 70%” relative lift. The latter figure comes from GPT-5’s low starting point: its increase was 13.9 percentage points but was much larger when expressed relative to its initial completion rate. Both calculations are mathematically valid, yet they answer different questions. Executives estimating how many additional jobs become acceptable need the absolute completion change alongside the relative improvement.

The intervention effect also varied sharply by category. Claude’s data-science completion rose from 64% to 93% after feedback, while Gemini writing moved from 4% to 16%. Gemini data science showed no improvement. Those differences make a single headline lift unsuitable for workflow planning because both model choice and task category changed the observed result.

Across the experiment, second attempts made after expert feedback outperformed first attempts, and roughly one in five initially failed jobs was rescued under these favorable conditions. Yet the strongest model’s stated overall completion ceiling in the human-guided discussion was 51%. Human-guided runs also took 14.5 minutes compared with 3.6 minutes for agent-only runs, around four times as long. Better outcomes therefore came through a process that also consumed more measured time.

The experimental comparison supports a specific causal claim: a second pass following expert feedback performed better than the first pass. HAPI did not run a matched agent-only second attempt for failed jobs, so its design cannot separate the contribution of expert feedback from the benefit of another attempt. The full improvement therefore cannot be assigned to expert intervention alone. For deployment planning, a retry and an expert review are different production inputs and should be tested separately.

The reviewers also represent a specific level of human capability. Every reviewer had a 100% Job Success Score and Top Rated or Top Rated Plus status. Collectively, they had earned more than $1 million on Upwork and logged more than 96,000 platform hours. Organizations seeking comparable intervention effects should therefore plan around the kind of experienced human capacity used in the benchmark rather than treating review as an interchangeable role.

That reviewer population sets the scope of what the experiment demonstrates about oversight. The measured results describe feedback from those accomplished freelancers; they do not establish an equivalent effect for a junior reviewer or another reviewer population. Reviewer expertise is consequently a deployment variable that an organization must test in its own workflow. The label “human in the loop” is too broad for staffing or financial planning because the loop can contain people with very different judgment and cost.

The failed lead-generation sample reinforces the distinction. A well-defined, objectively checkable task can fail even after expert guidance, while the wider experiment can show improvement without isolating the full cause of that improvement. Organizations should therefore test verification design, retries and human expertise as separate operating variables. Combining them into a general assumption that review will repair agent failures gives the business case more certainty than the experiment establishes.

Upwork’s architecture includes orchestration and judgment

The same operating questions appear in Upwork’s own product architecture, where the company has a commercial interest in positioning orchestration as part of work delivery. UMA, its flagship AI system, interprets a client’s intent, decomposes the requested work, routes components to people or machines and checks the resulting work. Its role spans orchestration across the workflow rather than autonomous execution of every part of an assignment. For buyers, the relevant point is that Upwork’s commercial system itself assigns explicit functions to routing and checking.

Upwork’s economic modeling similarly changes the delivery method with the consequences of failure. Agent-only execution is favored for low-value work where a failure is cheap; collaboration becomes preferable in the middle; and fully human execution wins for high-value work when a bad result costs more than automation saves. Because Upwork can benefit when work continues to use people supplied through its platform, that model carries a commercial incentive toward human-agent combinations. Its useful decision variable is expected failure cost alongside the nominal cost of producing an output.

That failure-cost view makes judgment-based, self-contained work particularly important. Such work may fit cleanly inside one domain, yet acceptance still requires an informed person to assess quality, and HAPI found its biggest lifts in judgment-based work. Human input in that workflow is part of producing a shippable result. A forecast that classifies required judgment as incidental supervision will understate the resources needed to operate the system.

The planning case becomes more demanding when judgment-based work also crosses functions. Agents broke in every tested category fitting that combination, which makes human work with AI assistance the appropriate assumption for planning a deployment. Cross-functional responsibility adds decisions at the boundaries, while qualitative acceptance adds judgment at delivery. Both forms of work need explicit ownership and funding in the operating model.

Review belongs in production economics when judgment is required

Once human judgment becomes part of delivery, its accounting treatment affects the automation business case. Organizations already address these limits by placing humans in the loop, retaining people to own cross-functional handoffs and using experienced reviewers where outputs have no objective acceptance test. If an expert’s intervention is required before the work can ship, that review consumes production capacity. Calling the same capacity oversight can make the workflow appear cheaper without changing the work required to operate it.

For judgment-heavy automation, review time therefore belongs alongside model costs and other production inputs, especially when intervention requires scarce senior employees. HAPI’s longer human-guided runs make the added process time visible, while the unmatched retry comparison means the whole difference cannot be assigned specifically to feedback. The operational question remains concrete: how much expert time does the deployed workflow consume per accepted result? That measure connects review directly to unit economics.

Measuring that time requires companies to identify the people who actually hold acceptance judgment. Those people can be dispersed across seniority levels, job families and cost centers, which can hide the dependency and its replacement cost. The phrase “human in the loop” consequently leaves a critical staffing question unresolved: which human has enough experience and authority to make the acceptance decision? Production planning needs the role, capacity and authority to be explicit.

Once that role is explicit, the staffing problem extends beyond today’s reviewers. If companies remove entry- and mid-level production work while preserving senior experts as reviewers, they can reduce the opportunities through which future experts develop the judgment needed to challenge model output. The relevant horizon can arrive quickly: the workforce example here asks who becomes senior enough to overrule a model in three years. An automation plan can preserve current expertise while weakening the path that produces its replacements.

That workforce-development risk matters most where acceptance depends on judgment because a deterministic test cannot carry the full verification burden. Leaders making the automation business case consequently have to account for present review capacity and its future supply. Upwork’s estimate that more than 90% of its work is qualitative explains why this issue matters to its platform as well as to buyers. The company benefits if demand for human judgment remains part of an increasingly automated market.

Rabinovich’s future-market hypothesis extends that commercial logic to agents themselves. He expects agents could create demand for humans by hiring them in real time when they reach outputs they cannot independently verify. His examples include a medical question, a restaurant recommendation and matters of taste. In that model, the human’s deliverable changes from completing an entire project to validating the agent’s result.

Because real-time human validation could generate work for Upwork, Rabinovich’s hypothesis aligns with the company’s commercial interest and should be evaluated as a platform executive’s forecast. The operational question does not depend on whether that market develops. If human value shifts toward validation, companies need to know where judgment resides, who has authority to exercise it and how much of that capacity is available. Production volume alone then gives an incomplete measure of the human role.

Rerun the automation case around acceptance, authority and review cost

Those production economics can be tested with work already passing through an organization’s systems. Pull ten outputs from the most automated workflow and give them to the person who previously owned that work. Ask how many would have gone to a client unchanged. The resulting proportion gives an automation rate tied to shippable work, including whether hidden completion work remains.

That acceptance test should lead directly to authority. Identify the person with authority to reject automated output in every workflow, then determine whether that person can exercise the authority without escalation. A nominal reviewer who cannot stop delivery cannot provide a reliable acceptance mechanism when quality depends on judgment. The workflow needs rejection authority to be as explicit as generation authority.

Where the acceptance decision requires judgment, reclassify the time spent reviewing agent output as production and rerun the financial case. Include the level of expertise actually required rather than assuming review capacity is interchangeable. A workflow can still produce worthwhile savings under that accounting. The resulting decision will then reflect the operating system the organization actually has to staff.

The same acceptance discipline should extend to vendor contracts. Inspect the exact meaning of completion because criterion-by-criterion acceptance and an attempted task describe materially different outcomes. HAPI used the stricter definition by requiring every critical and important criterion to pass. For a deployed system, the operational requirement is a trustworthy definition of completion together with the authority and funded verification needed to apply it.

Final thoughts

The practical boundary for AI automation is not whether an agent can produce useful work. It is whether the organization can determine, reliably and economically, that the work is complete. Upwork’s benchmark shows that those are different capabilities. Even objectively verifiable tasks can fail, while expert intervention can improve completion without eliminating the need for human capacity.

For executives, that shifts the automation case from model performance to operating design. Acceptance criteria, failure costs, reviewer expertise, retry policies and cross-functional handoffs all affect the cost of an accepted result. Where human judgment is required to ship the work, that judgment is a production input and should appear in staffing plans and ROI calculations accordingly.

The strongest automation candidates are therefore not necessarily the tasks that look routine. They are workflows where success can be defined clearly, failures can be detected reliably and the cost of verification does not erase the gains from automation. Where those conditions do not hold, leaders need explicit human ownership rather than an assumption that an agent will eventually become reliable enough.

The management question is no longer simply how much work AI can perform. It is how much work AI can finish to an acceptable standard, who determines that standard and what the organization must spend to make that completion dependable.

Alexander Procter

September 28, 2026

15 Min

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.