← All papers

Research paper · v0.3

The Agent-Ready Company

Organizational complementarities and the limits of delegating work to autonomous agents

Published
August 2026
Reading time
18 min read
Topic
AI × operating models

Abstract

Firms are deploying AI agents into operating models built on the assumption that only humans can interpret context, handle exceptions and exercise judgment. We argue that what limits agent deployment is organizational rather than technical, and we ground the argument in three literatures: the measured complementarity between information technology and organizational capital, the knowledge-based theory of hierarchy, and field experiments on generative AI. They converge on a narrow claim. Measured returns concentrate where operational knowledge was already codified, and delegating past the boundary of what a system does reliably makes performance worse. We also look at why observational shortcuts fail to substitute for asking people directly, and we say which parts of the argument nobody has measured yet.

01

Individual gains do not aggregate on their own

Enterprise generative AI arrived aimed at the individual worker, and at that level it delivered. Firm-level results have been harder to find.

Take the two largest studies. A field experiment across 66 firms and 7,137 knowledge workers found that among the roughly 80% of treated workers who adopted the tool, email time fell by about two hours per week. The same study detected no shift in the quantity or composition of the tasks those workers performed. A second analysis, linking adoption surveys to Danish administrative records, found precise nulls on earnings and recorded hours at both the worker and the workplace level, ruling out effects above 2% two years after ChatGPT's release. Its authors read that as employers absorbing AI through task reorganization rather than through compensation.

Neither study says the technology does not work. What both say is that the organization around the worker decides whether individual speed turns into institutional capability. Economists have been making that point about information technology since the 1990s.

“The technology changed how long a task took. It did not change which tasks existed, or who was accountable for them.”

Sources [1] [2]

02

The complementarity between technology and organization is measured, not asserted

That point usually circulates as management folklore. Its provenance is better than that.

Bresnahan, Brynjolfsson and Hitt built a panel of roughly 300 large US firms running from 1987 to 1994 and matched it against a survey of organizational practice. A firm two standard deviations above the mean simultaneously in information technology, human capital and decentralized work organization had predicted productivity about 7% higher than average, and that figure already excludes the direct contribution of its larger IT capital. The productivity belonged to the combination, not to the technology line item.

A companion study of 1,216 firms over 1987–97 found each dollar of computer capital associated with roughly twelve dollars of market value, against about $1.18 for property, plant and equipment. The authors put the gap down to unmeasured organizational intangibles rather than to the computers. The detail that matters most here is a smaller one. Their organizational practice index correlated with computer assets (coefficient 0.179, standard error 0.055) and showed no significant relationship with plant and equipment (0.033) or other assets (0.009). Whatever this complementarity is, it attaches to information technology and not to capital in general.

None of this is causal and none of it is about agents. These are correlations on 1987–97 data about information technology. The authors themselves note that in the fully interacted specification the coefficients are not precisely estimated, and discussants at the time raised reverse causality: highly valued firms may simply buy more computers. We cite the literature as the best-measured account we have of a complementarity that exists, not as proof of a mechanism that carries over to agents on its own.

Market value associated with each dollar of capital, 1,216 firms, 1987–97

Computer capital
≈ $12
Property, plant and equipment
≈ $1.18
Other tangible assets
≈ $1.04

The same asymmetry in correlation with organizational practices

  • Computer assets0.179standard error 0.055 · significant
  • Property, plant and equipment0.033standard error 0.050 · not significant
  • Other assets0.009standard error 0.048 · not significant

The gap is attributed to unmeasured organizational intangibles, not to the computers. Correlational evidence on 1987–97 data; the authors themselves note the fully interacted specification is not precisely estimated.

Figure 1 — The complementarity is specific to information technology

Sources [3] [4]

03

Hierarchy exists because operational know-how is tacit

Why do organizations have layers at all? Garicano's model derives the answer instead of assuming it. Production workers learn to solve the most frequent problems, specialists learn the exceptions, and the further up you go the more unusual the problems become. Management by exception is the model's name for that arrangement, and the pyramid shape falls out of it.

One assumption generates the whole structure, and Garicano states it plainly: production know-how is often tacit, so it lives inside individuals, so matching a problem to whoever can solve it is expensive. Hierarchy is the efficient answer to that matching cost. It is the efficient answer only while the cost stays high.

The model's own extension to information technology follows that thread. Codification lowers the cost of acquiring the knowledge needed to solve a given share of problems, which predicts wider decision scope on the front line, larger spans of supervision and fewer layers. The direction of that inference is worth holding onto for the rest of this paper. In this literature, codified knowledge is the thing that changes structure, not a byproduct of having adopted a tool.

A qualification about how to read what follows. Garicano models knowledge that agents learn, not knowledge that organizations write down. Extending it to written operational context — explicit rules, owners, exceptions, escalation paths — is our reading, not his result.

Tacit knowledge and judgment01
Informal workflows02
System-observable activity03
Formal documentation04

Most enterprise systems capture the bottom two layers. Delegating to agents typically depends on all four.

Figure 2 — The four layers of operational reality

Sources [5]

04

Which way decision rights move depends on which cost falls

"Adopting technology" has no signed effect on how authority gets distributed. The sign depends on what got cheaper.

Bloom, Garicano, Sadun and Van Reenen tested this across manufacturing plants in the United States and seven European countries, using regulatory variation in telecommunications costs for identification. Information technologies, the ones that help a person know more, went with greater autonomy and wider spans of control. Communication technologies, the ones that make it cheaper to ask someone else, reduced autonomy for workers and plant managers alike. Cheaper knowing decentralizes; cheaper asking centralizes.

Whether any of that transfers to agents is unsettled, and the record cuts both ways. Ide and Talamàs model AI in a knowledge economy and get opposite organizational outcomes depending on capability: AI whose knowledge resembles that of existing workers produces smaller, less productive, less decentralized firms, while AI approaching the knowledge of specialists produces the reverse. Labro, Lang and Omartian, working with more than 25,000 manufacturing plants in US Census data, found predictive analytics associated with less delegation to local managers and more centralized control.

So which bucket a given system falls into is an empirical question about that system inside that organization. It is not a property of the category and cannot be read off the category. An organization that delegates to agents without distributing the knowledge needed to decide should expect centralization.

Sources [6] [7] [8]

05

Where measured returns concentrate

The headline effect sizes in this literature are less interesting than how they distribute.

Brynjolfsson, Li and Raymond studied 5,179 customer support agents given staggered access to a generative AI assistant. Access raised issues resolved per hour by 14% on average. Among novice and low-skilled workers it raised them 34%. Among experienced and highly skilled workers the effect was minimal. The authors' stated mechanism is the part to keep: the model disseminates the best practices of more able workers and helps newer workers move down the experience curve.

That is a measurement of tacit knowledge transfer, whatever else it is. The system's value came from surfacing what the best operators already knew and handing it to everyone else. The gain was largest exactly where the gap between documented practice and expert practice was widest.

A related pattern shows up at the other end of the specification spectrum. Developers asked to implement an HTTP server in JavaScript finished 55.8% faster with an AI assistant, the largest effect anywhere in this literature and on its most completely specified task. We report the association and stop there. Comparing across studies with different tasks, populations and designs cannot establish that the specification caused the effect size.

“The measured return was the diffusion of what the best operators already knew. The constraint was that nobody had written it down.”
StudyPopulationMeasured effect
Brynjolfsson, Li & Raymond (2023)5,179 support agents+14% issues per hour; +34% for novices; minimal for experts
Dell’Acqua and colleagues (2025)758 consultantsInside the frontier: +12.2% tasks, +25.1% speed, +40% quality
Dell’Acqua and colleagues (2025)Task outside the frontier−19 percentage points likelihood of a correct solution
Peng and colleagues (2023)HTTP server in JavaScript55.8% faster on the most specified task in this literature
Dillon and colleagues (2025)7,137 workers, 66 firms≈2 fewer hours of email per week; no detected change in tasks

The pattern is not the effect size but its distribution: gains are largest where the gap between documented and expert practice was widest, and turn negative beyond the boundary of reliable performance. These are studies with different tasks, populations and designs; the comparison is descriptive and does not establish that specification causes effect size.

Figure 3 — Where measured returns concentrate, and where they reverse

Sources [9] [10]

06

Delegating past the boundary subtracts

If returns concentrate inside a boundary, locating the boundary becomes an operational question. Crossing it is worse than neutral.

Dell'Acqua and colleagues ran a pre-registered experiment with 758 consultants, roughly 7% of one firm's individual-contributor workforce. On tasks inside the AI capability frontier, participants using the model completed 12.2% more tasks, 25.1% faster, at quality more than 40% above the control group. On one task chosen to fall outside that frontier, consultants using AI were 19 percentage points less likely to reach a correct solution.

Their term for this is the jagged frontier. Capability varies sharply between tasks of apparently similar difficulty, and the boundary is invisible from outside. No model card will give it to you, because where the boundary falls depends on the task as your organization actually performs it: your exceptions, your data, your decision criteria.

Which changes what readiness means. The useful question is not about the model at all; it is whether an organization can sort its own tasks onto the right side of a line nobody can see directly. That depends entirely on how well it understands its own operations.

Sources [11] [12] [13] [14]

07

Event data does not reconstruct the organization

If explicit operational knowledge is the constraint, the obvious move is to infer it from the traces systems already produce. Process mining is the mature version of that attempt. Its own literature is unusually direct about where it stops.

Bose, Mans and van der Aalst, writing from the field that invented the technique, concluded that process mining would be limited by the quality of data rather than the availability of it. Their worked example is worth the detail. A hospital intensive-care log covering 1,308 patients and 21,819 measurements, all entered by hand: 67% of actions had no recorded completion event, timestamps were corrupted because staff entered records in batches, granularity ran anywhere from days to milliseconds, and the same item turned up under several free-text spellings. A model could be recovered from the cleanest subset. It came out spaghetti-like, 67 nodes and more than 100 arcs.

A later agenda paper puts the cause outside the IT department. Log defects trace back to individual habits, to organizational rules and conventions, and to system configuration: an incentive that makes people record actions that never happened, a free-text field that generates six names for one activity. The same authors report that extraction and cleaning remain poorly supported by commercial tools, so results depend heavily on who the analyst is, and that fully automated detection and repair is not a realistic expectation, because deciding whether an anomaly is actually a problem takes contextual knowledge that lives with domain experts.

A qualitative study of 15 practitioners across five organizations found that the pre-analysis phase alone draws on seven distinct categories of expert knowledge, spread across roles. Participants described finding a single business user with the entire process view as very difficult and almost impossible. The same study reports that formal documentation, standard operating procedures included, may not reflect the process as it is actually executed, and that even codified knowledge is imperfect.

Both sources need bounding. The agenda paper is a position paper reflecting expert consensus, not a study with its own sample. The interview study covers 15 participants in large organizations already mature in process mining, and its external validity to companies of 40 to 150 employees — where a single operations lead may genuinely hold more of the end-to-end view — is assumed here rather than measured.

SourceSystem eventsInformal workTacit judgmentDecision rules
Event logs and process miningYesPartialNoNo
SOP repositoriesNoNoNoPartial
Elicitation from domain expertsNoYesYesYes

No single source covers all four columns. Logs record what happened, not under which rule or whether the outcome was correct; formal documentation may not reflect the process as executed; and elicitation depends on several people supplying different parts of the view.

Figure 4 — What each source of operational knowledge can observe

Sources [15] [16] [17]

08

Governance frameworks already require the documentation

Operational documentation is usually treated as preparatory work that can follow deployment. The prevailing risk-management framework treats it as a control.

The NIST AI Risk Management Framework organizes its GOVERN function around exactly the artifacts this paper has been describing. GOVERN 2.1 requires that roles, responsibilities and lines of communication for mapping, measuring and managing AI risk are documented and clear to individuals and teams throughout the organization. GOVERN 3.2 requires policies and procedures defining and differentiating roles and responsibilities for human-AI configurations and for oversight of AI systems. GOVERN 1.2 requires trustworthy-AI characteristics to be integrated into organizational policies, processes and procedures.

The framework is voluntary, and adopting it proves nothing about outcomes. We found no study measuring whether organizations that follow it do better. What it does establish is that documented decision rights, explicit ownership and defined human-AI oversight boundaries are already the stated expectation for anyone deploying these systems, whatever one believes about productivity.

04

Adaptive

Humans and agents continuously redesign the workflow

03

Orchestrated

Agents execute multi-step workflows with escalation

02

Automated

AI executes bounded tasks under supervision

01

Assisted

AI helps humans complete individual tasks

00

Manual

Humans execute and coordinate the workflow

Each step up widens what is delegated and, with it, what must be documented: NIST GOVERN 3.2 asks for policies that define and differentiate roles and responsibilities for each human-AI configuration. The ladder describes degrees of delegation; it is not a measured maturity sequence nor a recommendation to advance one step at a time.

Figure 5 — Degrees of delegation to agents

Sources [18]

09

The return arrives late, and the lag has an organizational cause

General purpose technologies, in Brynjolfsson, Rock and Syverson's formulation, enable and require significant complementary investments: co-invention of new processes, products, business models and human capital. Those investments are largely intangible and badly measured. The accounting consequence is a J-curve. Productivity is understated early, while intangible investment is expensed rather than capitalized, and overstated later, as those intangibles get harvested.

Recent establishment-level work gives us something firmer than an accounting artifact. Using US Census manufacturing microdata from 2017 and 2021, McElheran and colleagues report causal evidence of J-curve-shaped returns: industrial AI use raised work-in-process inventory and robot investment, and damaged productivity and profitability in the short run. Among older establishments, abandonment of structured production-management practices accounted for roughly a third of those losses.

The short-run cost was therefore not purely a measurement artifact. A substantial share of it came from letting structured operating practice lapse during the transition, which is the single finding in this paper we lean on hardest.

On magnitude the literature disagrees, and sharply. Acemoglu's task-based estimate puts total factor productivity gains from AI at under 0.53% over ten years, and warns that even this may be overstated for complex, context-dependent tasks with no objective outcome measure. That disputes the size of the aggregate prize rather than the existence of the lag. Any claim about returns should be stated next to it.

Sources [19] [20] [21]

10

On numbers that circulate without methods

Two numbers dominate this conversation. Neither can hold the weight put on it.

The 95% generative AI pilot failure rate traces to a July 2025 report from MIT NANDA. Its evidence base: 52 executive interviews, 153 survey responses, and a review of more than 300 publicly disclosed implementations. The report calls itself preliminary findings, publishes no operational definition of what counts as no measurable business return, and states that its scores reflect reported frequency rather than objective measurement. As an exploratory study that is entirely legitimate. As a measured failure rate it is nothing of the kind, and citing it that way misrepresents what was done.

The 80% of business processes are undocumented figure has no primary source we could find. The denominator would be undefined even if it did. What counts as a process? Does documented mean a diagram exists, that the diagram is complete, or that it describes what actually happens now?

The defensible version is a research question with a stated denominator. What share of the context required to execute a given workflow correctly — objectives, inputs and outputs, dependencies, decision criteria, owners, exceptions, permission boundaries, escalation logic — is structured, current and available to a system rather than resident in one person's head? That is measurable in principle. We are not aware of anyone having measured it.

Sources [22]

11

Limitations

The complementarity literature at the core of this argument runs on data from 1987 to 1997 and concerns information technology. It is the best-measured body of work on the question. Applying it to autonomous agents is still reasoning by analogy across a forty-year gap and a change in the kind of system being adopted.

Almost all of that evidence is correlational. The exceptions with stronger identification are the telecommunications-cost instrument in Bloom and colleagues and the causal design in McElheran and colleagues. Correlation-based tests of complementarity cannot separate genuine complementarity from correlated unobserved returns.

The J-curve describes co-invention as concurrent learning: productivity falls while the organization reorganizes. That supports the claim that complementary organizational redesign is required. It does not support a claim that codification has to be finished before delegation starts. Those are different propositions and only the first one is established here.

The field experiments measure task-level and worker-level outcomes over months. None of them measures whether the organizational preconditions discussed here were present, so the link between codified operational knowledge and measured return is inferred from one study's stated mechanism rather than tested.

Sources vary in strength and are labelled accordingly in the text. Peer-reviewed journal articles and working papers with stated samples and methods carry the argument; one position paper and one preliminary industry report are cited only for what they are.

12

Open questions

What the argument above establishes: organizational complementarities govern returns to information technology, measured AI returns concentrate where knowledge was already codified, and observational data does not substitute for elicitation. What it leaves open is listed below. These strike us as the productive next targets for empirical work.

  • Do firms that made processes, decision rules, exceptions and ownership explicit before deploying agents capture more value than firms that did not? No study we found measures this directly. The bridge from organizational complementarity to autonomous agents is, for now, an inference.
  • Are generative AI field-experiment gains larger for tasks whose process was already well defined? That would test the mechanism micro-foundationally instead of inferring it from one study's stated channel.
  • What share of the context required to execute a workflow correctly is structured and machine-available, measured with a consistent denominator across firms?
  • Does the finding that no single informant holds the end-to-end process view survive in companies of 40 to 150 employees, where an operations lead may concentrate more of it?
  • For a given AI system in a given firm, does it behave as an information technology or as a communication technology in the sense of Bloom and colleagues? And can that be predicted before deployment rather than observed after?
  • What are AI project failure rates under an explicit, pre-registered definition of failure, and what share is attributable to organizational rather than technical causes?

13

Conclusion

The evidence assembled here supports a narrower claim than the one usually made, and a more useful one. Returns to information technology have consistently been mediated by organizational capital rather than delivered by the technology itself. The knowledge-based theory of hierarchy explains why: structure is a response to the cost of matching problems to the people who can solve them, and that cost is a function of how much operational knowledge stays tacit. Field evidence on generative AI shows returns concentrating where tacit knowledge was made available to people who lacked it, and going negative where work was delegated past the boundary of reliable performance.

None of which means organizations should document everything before deploying anything. The productivity literature describes reorganization as concurrent with adoption, not prior to it. The narrower conclusion is this. The boundary between what can be delegated and what cannot is specific to each organization's actual operations, it is not visible in system logs, no single person reliably holds it, and governance frameworks already expect it in writing.

An organization that cannot describe how its own work happens is not positioned to decide what to delegate. That is a claim about organizational self-knowledge and it is measurable. As far as we can tell, nobody has measured it.

“Copilots make individuals faster. Agents require an organization to know what it is delegating.”

14

About this research

Plexo builds an operational intelligence layer for mid-market companies in Latin America. Voice and text agents interview employees across teams and levels about recurring tasks, handoffs, decisions, exceptions and undocumented knowledge. The results are structured into an evidence-backed operational model where every claim traces back to a quote.

That work makes several of the open questions above directly testable. We intend to publish results only under a stated methodology, with customer consent, and on a sample large enough to carry the claim.

This paper is published as part of a monthly research series. Version 0.3 revises the prose of the August 2026 edition. Its argument, evidence and citations are unchanged from version 0.2, which was the revision that strengthened them; the original draft predated most of the sources cited here.

References

  1. [1]Dillon, Jaffe, Immorlica & Stanton (2025). Shifting Work Patterns with Generative AI. NBER Working Paper 33795
  2. [2]Humlum & Vestergaard (2025). Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI. NBER Working Paper 33777
  3. [3]Bresnahan, Brynjolfsson & Hitt (2002). Information Technology, Workplace Organization, and the Demand for Skilled Labor: Firm-Level Evidence. Quarterly Journal of Economics 117(1), 339–376
  4. [4]Brynjolfsson, Hitt & Yang (2002). Intangible Assets: Computers and Organizational Capital. Brookings Papers on Economic Activity 33(1), 137–198
  5. [5]Garicano (2000). Hierarchies and the Organization of Knowledge in Production. Journal of Political Economy 108(5), 874–904
  6. [6]Bloom, Garicano, Sadun & Van Reenen (2014). The Distinct Effects of Information Technology and Communication Technology on Firm Organization. Management Science 60(12), 2859–2885
  7. [7]Ide & Talamàs (2025). Artificial Intelligence in the Knowledge Economy. Journal of Political Economy 133(12)
  8. [8]Labro, Lang & Omartian (2023). Predictive Analytics and Centralization of Authority. Journal of Accounting and Economics 75(1), 101526
  9. [9]Brynjolfsson, Li & Raymond (2023). Generative AI at Work. NBER Working Paper 31161
  10. [10]Peng, Kalliamvakou, Cihon & Demirer (2023). The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv:2302.06590
  11. [11]Dell’Acqua, McFowland III, Mollick, Lifshitz-Assaf, Kellogg, Rajendran, Krayer, Candelon & Lakhani (2025). Navigating the Jagged Technological Frontier. Organization Science
  12. [12]Agrawal, Gans & Goldfarb (2021). AI Adoption and System-Wide Change. NBER Working Paper 28811
  13. [13]Demirer, Horton, Immorlica, Lucier & Shahidi (2026). Chaining Tasks, Redefining Work: A Theory of AI Automation. NBER Working Paper 34859
  14. [14]Gans (2026). Endogenous Task Bundling, Skills and Automation. NBER Working Paper 35211
  15. [15]Bose, Mans & van der Aalst (2013). Wanna Improve Process Mining Results? It’s High Time We Consider Data Quality Issues Seriously. IEEE CIDM 2013, 127–134
  16. [16]ter Hofstede, Koschmider, Marrella et al. (2023). Process-Data Quality: The True Frontier of Process Mining. ACM Journal of Data and Information Quality 15(3), Article 29
  17. [17]Pradhan, Jans & Martin (2025). Process mining starts here: on expertise, exchange, and the gaps between. Process Science 2, Article 25
  18. [18]NIST (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1
  19. [19]Brynjolfsson, Rock & Syverson (2021). The Productivity J-Curve: How Intangibles Complement General Purpose Technologies. American Economic Journal: Macroeconomics 13(1), 333–372
  20. [20]McElheran, Yang, Kroff & Brynjolfsson (2025). The Rise of Industrial AI in America: Microfoundations of the Productivity J-curve(s). US Census Bureau CES-WP-25-27
  21. [21]Acemoglu (2024). The Simple Macroeconomics of AI. NBER Working Paper 32487
  22. [22]Challapally, Pease, Raskar & Chari (2025). The GenAI Divide: State of AI in Business 2025. MIT NANDA
The Agent-Ready Company | Plexo