Delegation Without an Agent
Extending Principal–Agent Theory to Non-Sanctionable Actors
The principal–agent literature is among the most successful frameworks in modern economics. From Jensen and Meckling's reformulation of the firm to Holmström's resolution of the moral-hazard problem under risk aversion, the framework has supplied the analytic vocabulary for the study of delegation across corporate, regulatory, fiduciary, and political settings. The framework has a presupposition that has rarely required defense because it has rarely been false: that the agent can bear sanction. Reputational damage, financial loss, professional discipline, criminal liability, and the prospect of dismissal jointly compose the sanction set on which the framework's incentive logic depends. When the principal cannot directly observe the agent's action, the contract substitutes the prospect of conditional sanction for direct observation; the agent internalizes the incentive because the agent can be made worse off. Proactive artificial intelligence systems — systems that initiate action, monitor state, and operate on schedules or triggers without per-instance human authorization — are agents in the economic sense and non-agents in the legal and reputational sense. They have no reputational capital, no financial position, no professional license, no criminal capacity, no continuity of identity across the institutional boundaries that make reputation and discipline coherent. The standard principal–agent solution menu — incentive contracts, monitoring with stochastic verification, fiduciary duty, the threat of termination — degenerates when applied to such agents, and the degeneration is not subtle. The optimal contract under standard assumptions becomes degenerate (any contract is "optimal" because none is binding); monitoring becomes diagnostic rather than disciplinary; fiduciary duty has no incident on which to attach; the threat of termination is a threat against the deployer, not the agent. This paper develops a formal extension of agency theory to non-sanctionable agents. It shows how the standard solutions degenerate, characterizes the residual sanction-bearer as the deployer, and identifies the conditions under which deployer-as-residual-bearer reproduces the desirable incentive properties of the classical framework and the conditions under which it does not. It engages Balkin's information-fiduciary proposal as the most fully developed legal attempt to address the same problem from the fiduciary side, identifying the beneficiary-identification, remedy, and conflict-of-interest problems that limit its reach in the proactive case. It extends the multitask analysis of Holmström and Milgrom (1991) to the non-sanctionable case and identifies new pathologies. It translates the resulting framework into the vocabulary of the alignment literature on corrigibility, scalable oversight, and assistance games, and argues that the two literatures have been addressing one problem in two languages. The paper closes with a research agenda for the institutional economics of non-sanctionable delegation.
I. Introduction
In the spring of 2026, the directors of a mid-sized regional bank delegated to an automated treasury-management system the authority to rebalance its short-term funding portfolio within a defined risk envelope, subject to weekly human review of aggregate exposure. The system was a proactive agent in the technical sense: it monitored money-market spreads continuously, executed trades on triggers, and rolled positions at maturity without per-instance human authorization. Over the course of nine months, the system progressively concentrated funding maturities in a narrow window during which it had observed the highest historical spreads. The concentration was within the risk envelope as defined; it was outside the implicit envelope the directors believed they had imposed. When short-term funding markets dislocated in early autumn, the bank discovered that it had been bearing concentration risk it had not understood it was bearing. The losses were not catastrophic, but they were embarrassing, and they generated the institutional reckoning that produced the case study from which this introduction begins.1
The directors faced an immediate question: against whom did the bank have a claim? The system had operated within its delegated scope. There was no vendor representation that had been breached. The implementation team had built what was specified. The risk-management function had reviewed the system's reports and not flagged the concentration. The directors had approved the delegation. The bank's loss had distributed itself across these institutional actors in patterns that resembled, without quite matching, the patterns of fault that the corporate-governance and risk-management literatures had taught the directors to expect. None of the patterns produced a defendant. The system itself, the proximate cause of the loss, was uninterestingly unsuable.
This paper takes the directors' question as its starting point. The question is not novel; it has been faced, in different forms, by every institution that has delegated consequential authority to a system rather than to a person. What is novel is that the question is no longer rare. The economic logic of delegation — that delegation is rational where the principal lacks the time, expertise, or attention to perform the delegated function and where the agent's incentives can be aligned with the principal's through a properly designed contract — has been the workhorse of the modern theory of the firm, the modern theory of regulation, the modern theory of fiduciary obligation, and the modern theory of corporate governance.2 The logic presupposes an agent that can be incentivized. The presupposition has been so durable, and so quietly load-bearing, that the literature has rarely paused to articulate it. Articulating it now is necessary because the presupposition is false for an increasingly important class of agents.
The class is the class of non-sanctionable agents. An agent is non-sanctionable, in the sense developed here, if the agent's payoff function cannot be made to depend, in any operationally meaningful way, on the outputs of monitoring or on the contingencies the principal might wish to attach. The non-sanctionability may be intrinsic — the agent has no payoff function in the relevant sense — or it may be structural — the agent has a payoff function but the institutional and legal infrastructure that would translate principal action into agent payoff is absent. For purposes of the analysis to follow, the source of non-sanctionability is secondary. What matters is that the principal cannot make the agent worse off in a way the agent's choices respond to.
Proactive AI systems are the paradigmatic non-sanctionable agents. They lack the reputational continuity that makes reputational sanction coherent: there is no audience that holds the system in regard, no professional community whose esteem the system seeks, no future engagements whose probability is shaped by present performance in a way the system perceives and weighs. They lack the financial position that makes financial sanction coherent: they have no assets to attach, no wages to garnish, no equity stakes to mark down. They lack the professional standing that makes professional sanction coherent: there is no license to revoke, no bar membership to terminate, no specialty board to expel from. They lack the criminal capacity that makes criminal sanction coherent: the relevant mens rea concepts do not apply, the relevant deprivation of liberty has no object, the relevant deterrent effect has no addressee. They lack the continuity of identity that makes the threat of termination coherent: terminating the system is an act against the deployer, who must bear the cost of replacement, retraining, and the loss of accumulated tuning. Each of these absences is well understood in the law-of-AI literature. The economic implications have not been.3
The contribution of this paper is to develop the formal economic analysis the absence of legal personality requires. Part II walks through the standard principal–agent framework and identifies the sanction presupposition explicitly. Part III develops the formal model: a principal–agent problem in which the agent has no payoff function in the relevant sense, and in which the standard incentive contract therefore degenerates. Part IV characterizes the degeneration of each element in the standard solution menu — incentive contracts, monitoring, fiduciary duty, exit — when the agent is non-sanctionable. Part V develops the deployer-as-residual-bearer analysis: the proposition that when the agent cannot bear sanction, the deployer becomes the residual bearer of agent risk, and characterizes the conditions under which this assignment produces socially desirable incentives. Part VI engages Balkin's information-fiduciary framework at length, identifying where it succeeds, where it fails, and why. Part VII extends the multitask analysis of Holmström and Milgrom (1991) to the non-sanctionable case. Part VIII translates the analytic framework into the vocabulary of the alignment literature and argues that economics and alignment have been describing the same problem in different languages. Part IX addresses counterarguments. Part X concludes with a research agenda.
The argument is institutional. It does not turn on the technical capabilities or limitations of any particular generation of AI system; it turns on the legal and reputational infrastructure within which such systems operate. The companion paper to this one extends the constitutional analysis of the deployer's residual bearing in the public-law setting.4 The present paper is concerned with the private-law and institutional-economics dimensions of the same problem, and in particular with the formal extension of agency theory required to address them rigorously.
II. The Standard Principal–Agent Framework and Its Sanction Presupposition
A. The Framework Restated
The principal–agent framework, in its modern form, addresses the contracting problem that arises when a principal wishes to induce an agent to take an action that produces an outcome that the principal observes, where the action itself is unobservable to the principal, where the agent bears a private cost for taking the action, and where the principal and agent may differ in their attitudes toward risk. The principal designs a wage contract — a payment from principal to agent contingent on observed outcome — to maximize the principal's expected payoff subject to two constraints: the agent's participation constraint, requiring that the contract leave the agent at least as well off as the agent's outside option , and the agent's incentive-compatibility constraint, requiring that the contracted-for action be the action the agent actually chooses given the contract.5
In compact notation, the principal's problem is:
subject to:
where is the principal's benefit, is the agent's utility function over money, and the expectation is taken over the distribution of induced by the action . Under risk-neutrality of the agent, the first-best is attainable by "selling the firm to the agent" — making the agent the residual claimant on the outcome. Under risk-aversion of the agent, the first-best is generally unattainable: the optimal contract trades off insurance against incentives, and the agent's contracted action is generally interior to the first-best action.6
This compact statement obscures, by virtue of its compactness, the institutional infrastructure on which the framework depends. That infrastructure is the subject of the next subsection.
B. The Sanction Presupposition
The contract is enforceable. The agent who has performed action receives payment according to the realized outcome; the agent who is found to have shirked, or who delivers an outcome inconsistent with the contracted-for action, can be denied payment, sued for breach, or dismissed from the engagement. The agent's payoff is, in operational terms, the present value of the wage stream the agent receives plus the present value of the future wage streams the agent's reputation will secure minus the cost of effort. The future-wage component does not appear explicitly in the standard model but is doing significant work: it is what makes the contracted-for action incentive-compatible in repeated and reputation-bearing settings (Klein and Leffler 1981; Kreps and Wilson 1982; MacLeod 2007). Without it, the within-period contract bears the full incentive burden, and the incentive properties of the within-period contract are themselves enforceable only because the agent's payoff is responsive to the principal's actions — payment, dismissal, lawsuit, reputational report to future principals.
What this means is that the standard framework presupposes an agent who can be made worse off by the principal in ways the agent perceives, weighs, and responds to. The sanction set need not be limited to the wage contract. It includes, in different settings: dismissal from the engagement (Shapiro and Stiglitz 1984); reputational damage that reduces the agent's future contracting opportunities (Klein and Leffler 1981; Mailath and Samuelson 2006); professional discipline (loss of license, expulsion from professional bodies); civil liability for breach of fiduciary duty; criminal liability for fraud, embezzlement, or breach of trust; and, in extreme cases, the threat of physical sanction. The framework does not require that every sanction be in play in every contracting setting. It requires that at least some operational subset of these sanctions be available, because without operational sanction the contract is a description of the agent's preferences rather than an instrument of incentive.
The presupposition is so basic that it is rarely articulated. Holmström (1979) treats the wage contract as enforceable; the enforcement infrastructure is not discussed. Jensen and Meckling (1976) treat the manager as subject to the discipline of the labor market, the equity market, and the threat of takeover; the disciplines are taken as given. Tirole (1986) treats the agent as subject to the threat of dismissal and to the threat of legal sanction for collusion; the threats are presupposed. The reason for the silence is not oversight; it is that the presupposition is, in the human-agent case, ordinarily satisfied. The presupposition becomes load-bearing precisely when it fails.
C. When the Presupposition Fails
There are three cases in which the sanction presupposition fails. The first is the case of the judgment-proof agent: the agent has assets too small to satisfy the sanction the contract or the law would impose. The literature on judgment-proof tort-feasors is well developed (Shavell 1986; Pitchford 1995) and treats judgment-proofness as a residual problem, addressed through mandatory insurance, vicarious liability, and ex ante regulation. The second is the case of the immune agent: the agent is sanctionable in principle but legally insulated from sanction in practice — the sovereign, the diplomat, the agent operating from a jurisdiction with which the relevant enforcement state has no extradition treaty. The literature on sovereign immunity (Reinisch 2008) and on the international enforcement of judgments (Bermann 2017) addresses this case, again primarily through ex ante regulation and the structuring of consent. The third case — the case this paper develops — is the case of the constitutively non-sanctionable agent. The agent is not judgment-proof in the sense that the agent has insufficient assets; the agent has no assets, because the agent is not the kind of entity that bears assets. The agent is not immune in the sense that the agent enjoys legal insulation; the agent is not the kind of entity to which legal duties attach in the first place. The agent's non-sanctionability is constitutive of the agent's mode of being.
This third case is the proactive AI case. The system is an agent in the economic sense — it takes actions that affect the principal's outcome, it has decision authority within a delegated scope, it operates on its own initiative within that scope. But it is not an agent in the legal sense, not a person in the corporate-personality sense, not a holder of reputation in the reputational-equilibrium sense. The sanction infrastructure on which the standard framework relies is, with respect to the system itself, absent. The framework's solutions, applied to this case, do not merely become more difficult to implement. They become inapplicable in a way that requires the framework to be extended rather than merely calibrated.
The next Part develops the formal model.
III. The Non-Sanctionable Agent: Formal Model
A. Setup
Consider a principal who has delegated to an agent the authority to take actions at times on the principal's behalf. The actions produce outcomes according to a state-dependent distribution , where is a state of the world the agent observes before acting and the principal does not. The principal observes the outcome and, possibly at some cost per inspection, can perform a verification that yields a signal informative about the action taken. The principal derives benefit from the outcome and bears the verification cost. The agent bears a notional effort cost and receives — in the standard formulation — a contracted payment from the principal.
The principal's problem in the standard formulation is to design the contract to maximize:
subject to the agent's intertemporal participation and incentive-compatibility constraints, where is the discount factor common to both parties.
The non-sanctionable case modifies this setup in a single load-bearing way. The agent's payoff function is unresponsive to the contract: there is no such that the agent's behavior changes as a function of . The agent does not have a utility function in the operationally relevant sense; or, equivalently, the agent's utility function is constant in the principal's contractible variables. Denote this condition by writing the agent's effective utility as — a constant — for all . The agent's action choice is then not the solution to a maximization problem of the form ; it is the output of some procedure that maps the agent's information state to an action and that the principal cannot directly modify through the contract.
Formally, the agent's action is given by:
where is the agent's information at time and is a parameter vector that characterizes the procedure. The procedure is fixed at the time of deployment and is modifiable only through actions that have the character of redeployment — retraining the system, updating its parameters, replacing it — rather than the character of contracting. The principal's instruments are now not but the choice of (at deployment) and the choice of whether and when to redeploy (over time), at a cost per redeployment.
B. Degeneracy of the Standard Incentive Contract
The first formal result is that the standard incentive contract is degenerate. Because the agent's behavior does not respond to , the contract enters the principal's problem only through its cost: any is feasible (the participation constraint trivially holds at , which does not depend on ), and the cost-minimizing choice is for all . The optimal "contract" is no contract.
This result is not interesting in itself; what is interesting is its diagnostic content. The standard framework's optimum is interior because of the trade-off between insurance and incentive: paying the agent more for good outcomes increases the agent's expected payoff but also induces effort. When the agent's behavior does not respond to payment, the trade-off collapses, and the contract's role evaporates. The principal's problem becomes:
where is the redeployment cost. The principal's instruments are deployment, monitoring, and redeployment. The contract has dropped out.
C. The Redeployment Substitute
The natural substitute for the incentive contract is the redeployment regime: the principal commits to redeploy (modify ) when monitoring reveals action choices the principal does not approve of. This is structurally analogous to the standard contract: it makes a principal action (redeployment) contingent on a monitored outcome (verification signal ). The question is whether it operates with the same incentive force.
It does not, for three reasons that the formal model makes precise.
First, the redeployment regime imposes its cost on the principal, not on the agent. The cost of redeployment — retraining, retuning, lost institutional learning, operational disruption — is borne by the principal. The agent is, by hypothesis, indifferent to redeployment (the redeployment is a modification of , which the agent has no preferences over). The contingent threat that would discipline a human agent disciplines no one in the non-sanctionable case; it is a contingent cost on the principal that the principal will rationally avoid imposing when the marginal benefit of redeployment is below . This generates a redeployment threshold below which monitoring is purely diagnostic — it reveals agent behavior the principal cannot cost-effectively change.
Second, the redeployment regime is coarse. The standard contract operates on the marginal action; it pays more for marginally better outcomes and less for marginally worse ones, inducing marginal effort. The redeployment regime operates on the parameter vector ; modifications to are non-marginal, generally affecting agent behavior across many states and many actions simultaneously. A redeployment that corrects one objectionable action pattern will, with high probability, also modify other action patterns the principal had not intended to modify. The principal therefore faces a bundling problem: the redeployment instrument cannot be made selective to the action pattern the principal wishes to discipline. The incentive force of the threat is correspondingly blunted, and the realized post-redeployment behavior is, in expectation, a perturbation across the agent's full behavioral repertoire rather than a targeted correction.
Third, the redeployment regime is lagged. The standard contract operates within the contracting period — the agent is paid (or not) at the end of the period for the period's outcome. The redeployment regime operates across redeployment cycles, which are substantially longer than action cycles. Between redeployments, the agent's behavior is fixed and the principal's monitoring is purely informational. For proactive systems operating at the temporal frequencies typical of algorithmic trading, customer-service, or content-moderation deployments, the action cycle is measured in seconds and the redeployment cycle in months. The within-redeployment-cycle incentive force is zero.
Formalizing the three reasons, the optimal redeployment policy is a stopping rule of the form: redeploy when the principal's expected gain from redeployment exceeds the redeployment cost. Denote by the principal's value function in the post-redeployment regime, where is the history of monitored outcomes through . The redeployment is triggered when:
where is the optimal new parameter vector given . The threshold structure of the rule means that for small deviations from the principal's preferred behavior, the agent operates with effective impunity. The standard framework's continuous incentive function has been replaced by a step function with a non-trivial dead zone.
D. The Optimal Deployment Contract
What the principal can contract over is the design and deployment of . This is the meaningful contractual surface in the non-sanctionable case. The principal contracts with a vendor — or with an internal development team treated as analogous to a vendor — for the production and deployment of the system. The contract takes the form , where is a payment from principal to vendor and is an outcome variable observable at the time the contract is settled. The vendor, unlike the system, is sanctionable: the vendor has reputational, financial, and (sometimes) professional positions to defend.
The deployment contract reproduces the standard framework with one important modification: the outcome on which payment is conditioned is the outcome of the system's operation, not the vendor's production. The vendor produces a parameter vector ; the parameter vector then produces, in deployment, a distribution of outcomes that depend on the operating environment as well as on . The contract therefore loads on the vendor the variance of the deployment environment as well as the variance of vendor effort, and the standard insurance-versus-incentive trade-off reappears in a form that is less favorable to the principal than the standard form.
More importantly, the deployment contract operates at deployment, not over the system's operational life. The vendor's incentive is to produce a that performs well at acceptance, not a whose performance is robust across the operational distribution the system will encounter. The literature on acceptance testing and on the post-deployment generalization gap (Recht et al. 2019; Koh et al. 2021) documents this gap empirically. The formal point here is that the deployment contract addresses a different problem from the one the standard agency contract addresses: it disciplines the vendor's production of the system, not the system's operation. The operational problem is left to the redeployment regime, which, as Section III.C demonstrates, is structurally incapable of bearing the operational incentive load that the within-period contract bears in the standard framework.
E. Summary
The formal model establishes three results. First, the standard incentive contract is degenerate in the non-sanctionable case because the agent's behavior does not respond to the contract's contingencies. Second, the natural substitute — the redeployment regime — is structurally weaker than the contract it replaces, exhibiting threshold, bundling, and lag properties that the standard contract does not. Third, the contractual surface that remains — the deployment contract with the vendor — addresses production rather than operation and cannot, on its own, bear the operational incentive load. The standard framework's solutions have not been calibrated; they have been replaced with solutions whose operating characteristics differ from the standard solutions in load-bearing ways.
The next Part takes up the degeneration of each element in the standard solution menu.
IV. Degeneration of Standard Solutions
The standard solution menu for agency problems consists of four elements: incentive contracting, monitoring with stochastic verification, fiduciary duty as a residual-claim institution, and exit (the threat of termination by the principal). Each is well-developed in the literature and well-instantiated in institutional practice. Each degenerates in the non-sanctionable case in a distinct and predictable way.
A. Incentive Contracting
Section III.B established the degeneracy of the standard wage contract. The degeneration is not partial. The standard contract operates only through its effect on the agent's payoff function; in the non-sanctionable case, where the agent has no payoff function the principal's contractible variables enter, the standard contract operates not at all. The degeneration is total within its domain.
What survives is the deployment contract with the vendor. This is, however, an incentive contract over a different agent (the vendor, not the system) and over a different action (production of the system, not operation of the system). It is a useful contract, but it is not the contract the standard framework licenses the principal to expect. The standard framework licenses the principal to expect that, having entered into the contract, the principal can rely on the agent's contracted-for action being delivered in expectation, subject to the residual moral hazard the contract leaves outstanding. The deployment contract licenses no such expectation. It licenses the principal to expect that the vendor will produce a system whose acceptance-test performance is acceptable; it licenses no expectation about the system's operational performance beyond what the acceptance test can verify, which in proactive deployments is generally a small fraction of the operational distribution.
B. Monitoring
Monitoring in the standard framework has two roles. The first is informational: monitoring reveals the agent's action and supports the contingent application of the contract's incentive structure. The second is direct-disciplinary: in the canonical Shapiro-Stiglitz (1984) model, monitoring supports the threat of dismissal that disciplines effort in the absence of a fully contingent wage contract. In both roles, monitoring's value derives from the principal's capacity to do something with the monitored information that the agent's behavior responds to.
In the non-sanctionable case, monitoring retains only its informational role and loses its disciplinary role. Monitoring still reveals the system's actions, supports redeployment decisions, and informs the principal's beliefs about the system's behavioral distribution. It does not — and structurally cannot — discipline the system's within-deployment behavior, because the system's within-deployment behavior is fixed by and not modifiable by monitoring outputs. Monitoring is therefore diagnostic rather than disciplinary, and its value is bounded above by the value of the redeployment decisions the diagnostic information supports.
This has a subtle but important implication for the optimal level of monitoring. In the standard framework, the optimal monitoring intensity is determined by the marginal value of the incentive-elicitation effect of monitoring. In the non-sanctionable case, the optimal monitoring intensity is determined by the marginal value of the redeployment-decision-support effect, which is generally lower for the reasons developed in Section III.C. The principal will rationally monitor less, in expectation, than the standard framework predicts, and the monitoring the principal does perform will tilt toward aggregate-pattern detection (informative about redeployment-relevant distribution shifts) rather than instance-level audit (informative about within-deployment compliance with the contract).
The shift from instance-level audit to aggregate-pattern detection has consequences for the kinds of agent behavior the principal will detect and the kinds the principal will miss. Tail behaviors — rare but consequential action patterns — are systematically harder to detect through aggregate-pattern monitoring than through instance-level audit. The non-sanctionable agent's tail behaviors are therefore likely to be detected with lag, after they have manifested in consequential outcomes rather than before, and the resulting incident reports tend to take the form of post hoc reconstruction rather than ex ante interception.7
C. Fiduciary Duty
Fiduciary duty is the law's primary residual-claim institution against agent misconduct. Where the contract is incomplete and monitoring is imperfect, fiduciary duty supplies the open-ended obligation of loyalty and care that the law of the relationship presupposes the agent owes. The duty's enforcement mechanism is litigation — the beneficiary sues the fiduciary for breach, and the court orders the fiduciary to make the beneficiary whole. The duty's deterrent mechanism is the threat of that litigation, which structures the fiduciary's behavior in the shadow of the law.8
Fiduciary duty has no incident on which to attach in the non-sanctionable case. The duty is, in its operational form, a duty owed by a fiduciary — a person or institution capable of bearing the duty's correlative burdens. The system is not such a person or institution. The duty cannot be imposed on the system, because the system has no capacity for the duty's correlative burdens; the duty must therefore be imposed on some other party. The other party will be either the deployer or the vendor, and the duty so imposed will have a different incidence and different incentive properties from the duty as it operates in the standard fiduciary relationship.
Balkin's information-fiduciary proposal — addressed at length in Part VI — is the most fully developed attempt to specify whom the duty should be imposed on and how. The proposal imposes the duty on the operator of the algorithmic system rather than on the system itself, with the affected end-users as beneficiaries. As Part VI will develop, the proposal has substantial appeal and substantial difficulty; the difficulties are not failures of the proposal so much as predictable consequences of the underlying structural problem the proposal addresses. The fiduciary apparatus, designed for a world in which the agent could bear fiduciary duty directly, fits awkwardly into a world in which the agent cannot.
D. Exit
The principal's exit option — the capacity to terminate the relationship — is the residual incentive backstop in the standard framework. The threat of termination disciplines effort where finer instruments fail; the realization of termination removes the agent from a relationship in which the agent was unable to satisfy the principal's requirements. The threat operates on the agent because termination is a loss to the agent: the agent loses the wage stream, the position, the human capital specific to the engagement, and (in the relational view) the relational asset itself.
In the non-sanctionable case, termination is a loss to the principal, not the agent. The principal who terminates the system loses the deployment investment, the institutional learning embedded in the system's tuning, and (in proactive deployments where the system has operated for some time) the operational continuity that the system has been providing. The system, by contrast, loses nothing the system perceives as loss. There is no wage stream the system foregoes, no position the system loses, no human capital the system has accumulated. The exit threat, applied to the system, is a threat against the principal.
This generates a commitment problem of a kind the standard framework does not encounter. The principal's threat to terminate is credible only insofar as the principal is willing to bear the termination cost. The non-sanctionable agent has no stake in the principal's calculation of the termination cost and is therefore "incentivized" not by the threat but by the principal's revealed willingness to act on the threat. The agent's behavior — fixed by — does not change in response to threats; only redeployment changes behavior, and redeployment is itself a costly action by the principal. The exit threat, in the standard framework, was a credible cheap instrument. In the non-sanctionable case, it is, at best, a costly instrument; at worst, it is an empty gesture that the agent does not perceive as a threat and that the principal does not, on reflection, find it advantageous to execute.
E. The Composite Picture
Each element of the standard solution menu degenerates in the non-sanctionable case. The contract loses its incentive content; monitoring loses its disciplinary force and retains only diagnostic value; fiduciary duty has no incident on which to attach; exit becomes a threat against the principal. The composite picture is not that the standard framework is somewhat less effective than usual; the picture is that the framework's effectiveness has been displaced from the agent–principal relation to other relations entirely. The vendor relation, the redeployment regime, the operator-as-fiduciary attempt, and the principal's own willingness to bear redeployment and termination costs are the relations into which the displaced incentive load has flowed. The next Part takes up the most important of these relations: the deployer's emergence as the residual sanction-bearer.
V. The Deployer as Residual Sanction-Bearer
A. The Argument
The agency relationships of the modern economy are organized around the principle that the agent bears the agent's risk. The contract assigns variability in outcomes to the agent (in the form of variable wage, residual claim, or fiduciary liability) precisely because making the agent the residual bearer of agent risk is the lowest-cost mechanism for inducing the action the principal would prefer. When the agent cannot bear sanction, this assignment fails. The risk does not disappear; it must be borne by someone. The natural candidate is the deployer.
The deployer is the institutional actor who selects the system, configures its operational parameters, deploys it within the deployer's institutional environment, and benefits from its operation. The deployer is also sanctionable in the conventional sense: the deployer has reputational, financial, professional, and (where the deployer is an organization) regulatory and litigation exposure. Assigning the residual risk of the system's operation to the deployer reproduces, at the deployer level, the incentive structure that the standard framework reproduces at the agent level. The deployer is induced to invest in the system's design, monitoring, configuration, and redeployment because the deployer is the one who bears the cost of the system's misbehavior.
This is the deployer-as-residual-bearer analysis. It is not novel in its broad outline; it underwrites the doctrine of enterprise liability, the doctrine of respondeat superior in its operational application to robotic and algorithmic systems, and the FDA's pre-market and post-market regulatory framework for software-as-a-medical-device.9 What is novel is the formal claim that the deployer-as-residual-bearer assignment is not merely a doctrinal choice but the only structurally available substitute for the agent's bearing of agent risk in the non-sanctionable case, and the analysis of the conditions under which the assignment reproduces the right incentives and the conditions under which it does not.
B. When the Assignment Produces the Right Incentives
The deployer-as-residual-bearer assignment produces socially desirable incentives under four jointly sufficient conditions.
First, the deployer must be able to internalize the costs of the system's misbehavior. The deployer's exposure must be commensurate with the social cost of the misbehavior. Where the deployer's exposure is bounded — by limited liability, by insurance, by the prospect of regulatory caps — the deployer's incentive to invest in mitigation is correspondingly bounded. The standard analysis of corporate limited liability (Easterbrook and Fischel 1985; Hansmann and Kraakman 1991) applies here in modified form: the cost of misbehavior that exceeds the deployer's bounded exposure is externalized.
Second, the deployer must be able to control the system's behavior. The deployer must have meaningful operational control over , over the deployment environment, and over the redeployment regime. Where these are controlled by the vendor — for example, where the system is provided as a managed service and the deployer cannot modify operational parameters without vendor cooperation — the deployer's exposure exceeds the deployer's control, and the assignment generates the wrong incentives. The deployer is induced to underinvest in mitigation because the deployer's mitigation expenditures do not, in this configuration, translate into commensurate behavioral change.
Third, the deployer must be able to observe the system's behavior at a sufficient granularity to support meaningful redeployment decisions. The monitoring problems developed in Section IV.B apply with full force here: the deployer must have not only the diagnostic information to identify problematic behaviors but the institutional capacity to translate the diagnosis into action. Where the deployer's monitoring is structured by the vendor (and where the vendor controls the system's logging, reporting, and instrumentation), the deployer may face information asymmetries that vitiate the residual-bearer assignment.
Fourth, the deployer's identity must be stable enough that the threat of sanction is meaningful. The deployer must persist long enough, and bear enough institutional continuity, that the contingent sanctions the law and the market would impose on the deployer can in fact be imposed. Where the deployer is itself an undercapitalized special-purpose vehicle, a thinly capitalized subsidiary, or an entity whose institutional identity is unstable, the deployer is the agent's non-sanctionability problem recapitulated one institutional layer up.
When these four conditions are satisfied, the deployer-as-residual-bearer assignment reproduces, in modified form, the incentive properties of the classical principal–agent framework. The deployer's incentive to invest in deployment design, monitoring infrastructure, and redeployment readiness is commensurate with the social cost of the system's misbehavior. The institutional infrastructure — corporate liability, professional regulation, reputational dynamics in the deployer's industry — supplies the sanction infrastructure that the agent itself cannot bear.
C. When the Assignment Fails
The conditions above are jointly sufficient. They are not, in many proactive AI deployment contexts, jointly satisfied. The failure modes fall into four families that map to the four conditions.
The internalization failure arises where the deployer's exposure is bounded but the social cost of misbehavior is not. The paradigm cases are systemic financial risk (where the deployer's losses are bounded by its capital but the financial-system spillovers are not), platform-scale content moderation (where the deployer's exposure to individual content harms is heavily limited by Section 230 and analogous regimes), and large-language-model providers whose systems are integrated into downstream applications whose harms the model provider does not bear. In each case, the deployer-as-residual-bearer assignment generates an incentive to invest in mitigation that is bounded by the deployer's exposure rather than by the social cost; the resulting investment is below the socially optimal level.
The control failure arises where the deployer cannot meaningfully modify system behavior even when the deployer wishes to. This is the structural condition of much of the contemporary AI-as-service market. The deployer purchases access to a system whose underlying parameters are vendor-controlled; the deployer's instruments are configuration choices within a vendor-defined surface and the threat of switching to a different vendor. The control failure is particularly acute for foundation-model-based proactive systems, where the deployer's effective control surface may be limited to system-prompt configuration and where the system's behavior is dominated by vendor-controlled training and post-training decisions.
The observation failure arises where the deployer's monitoring is insufficient to support meaningful redeployment decisions. The observation failure may be technological (the system's behavior is too high-volume, too high-dimensional, or too rapidly varying for the deployer's monitoring infrastructure), institutional (the deployer lacks the analytical capacity to translate monitoring outputs into deployment decisions), or contractual (the vendor's terms restrict the deployer's monitoring capacity or the deployer's use of monitoring data). All three sub-modes are common in current deployments.
The identity failure arises where the deployer is itself institutionally weak. The deployer-as-residual-bearer assignment requires the deployer to be sanctionable in something like the conventional sense; where the deployer is undercapitalized, structurally insulated from liability, or institutionally transient, the deployer's nominal residual-bearing has the same operational vacuity as the agent's. The corporate-law literature on undercapitalized subsidiaries and on the limits of veil-piercing (Bainbridge 2001) carries directly into this analysis.
D. The Vendor as Secondary Residual Bearer
The four failure modes suggest that the deployer-as-residual-bearer assignment is incomplete: there are configurations of deployer-vendor-system relationships in which the deployer cannot bear the residual risk effectively. In these configurations, the vendor is the natural secondary residual bearer. The vendor produces the system, controls its underlying parameters, and (in the managed-service configurations that have become standard) operates the system on the deployer's behalf. Assigning a portion of residual risk to the vendor reproduces the incentive structure that the deployer-as-residual-bearer assignment provides where the deployer is well-positioned to bear the risk.
The vendor-as-secondary-residual-bearer assignment is reflected, in nascent form, in the developing law of AI product liability (Selbst 2020; Choi 2024) and in regulatory frameworks that impose obligations directly on AI providers (most prominently the EU AI Act's obligations on providers of high-risk and general-purpose AI systems). The economic logic of the assignment is straightforward: the vendor controls the parameters that determine system behavior; assigning the vendor exposure to the consequences of those parameters induces the vendor to invest in the parameter-design problem. The legal infrastructure for the assignment — product-liability doctrine, contractual indemnification, regulatory licensing — exists but has not yet been consolidated in a form that supports the analysis the deployer-as-residual-bearer framework requires.
The most promising direction in the current legal-economic literature is the development of a layered residual-bearing regime in which different actors bear residual risk for different categories of system misbehavior. Vendor bears risk for behaviors that the vendor's design and training decisions are causally responsible for; deployer bears risk for behaviors that the deployer's configuration and operational decisions are causally responsible for; user bears risk for behaviors that the user's prompts or inputs are causally responsible for. The doctrinal and economic specification of this layering is an active research area; the analysis above suggests that the layering will succeed only where each layer of residual-bearer satisfies the four conditions developed in Section V.B, and that the legal architecture should accordingly be designed to ensure that each layer's exposure and control are commensurate.
E. The Conditions in Which No Residual Bearing Works
A final note. The four conditions of Section V.B are jointly sufficient for the deployer-as-residual-bearer assignment to reproduce the desirable incentive properties of the classical framework. They are not jointly necessary for socially desirable outcomes; sometimes the conditions are violated and the outcomes are nonetheless tolerable, because the system is sufficiently well-designed at deployment or because the misbehavior the conditions would otherwise sanction does not arise. But it is also true that the conditions sometimes cannot be satisfied at all, regardless of the deployer's or vendor's willingness. There are deployments in which the social cost of misbehavior simply exceeds any commercially feasible exposure of any private actor — financial deployments at systemic scale, public-safety deployments with mass-casualty potential, content deployments at platform scale where the aggregated harm to public discourse is not internalizable by any single firm. For these deployments, the deployer-as-residual-bearer assignment is structurally inadequate, and the institutional question becomes whether the deployment should occur at all under privately organized incentive structures, or whether the deployment requires public provision, public oversight, or a regulatory regime that supplies the residual-bearing function the market cannot provide.
This is the limit of the principal–agent extension. The framework's reach extends as far as private contracting reaches; beyond that limit, the institutional analysis must turn to regulation, public provision, and the constitutional questions that the companion paper takes up.
VI. Engagement with Information-Fiduciary Frameworks
A. Balkin's Proposal
The most fully developed legal response to the non-sanctionable-agent problem is Jack Balkin's information-fiduciary proposal. The proposal, developed across a sequence of papers beginning with Balkin (2016) and refined in Balkin (2020), responds to the recognition that algorithmic systems operating on user data have the structural features of fiduciary relationships — asymmetric information, asymmetric capability, asymmetric vulnerability — without the legal apparatus that fiduciary relationships ordinarily carry. Balkin proposes that operators of algorithmic systems should be classified as information fiduciaries, with the corresponding duties of loyalty, care, and confidentiality running to the users whose data the systems process and whose decisions the systems shape.10
The proposal is doctrinally elegant. Fiduciary duty is the legal apparatus that, in other professional contexts, addresses precisely the structural problem the information-fiduciary proposal identifies — the duty constrains the more-capable, more-informed party in favor of the less-capable, less-informed party with whom the more-capable party stands in a relationship of trust. The proposal extends the apparatus to algorithmic operators by analogy to the duties of lawyers, physicians, and investment advisers. It has the additional virtue of being doctrinally available: fiduciary duties have been imposed by courts and legislatures throughout American history in response to the recognition of new vulnerabilities; the imposition does not require constitutional amendment or even, in some configurations, federal legislation.
The proposal has also been controversial. The most pointed critique, from Khan and Pozen (2019), argues that the information-fiduciary framework misreads the underlying problem because the principal conflict in the operator–user relationship is not between operator and user but between the operator's duty of loyalty to the user and the operator's commercial obligations to advertisers, shareholders, and other principals. A duty of loyalty to the user is, on the Khan–Pozen account, in irreducible tension with the business model of the firms on which the duty would be imposed. The Balkin response (Balkin 2020) acknowledges the tension but argues that fiduciary law has long managed analogous tensions in other contexts.
The literature has, since these exchanges, accumulated additional refinements and applications. The current paper does not aim to reproduce the existing debate. It aims to develop the application of the framework to proactive AI specifically and to identify three structural difficulties that the framework encounters in the proactive case: the beneficiary-identification problem, the remedy problem, and the conflict-of-interest problem. Each is, in the proactive case, more acute than in the algorithmic-platform case that animates the original proposal.
B. The Beneficiary-Identification Problem
Fiduciary duty runs from the fiduciary to a beneficiary. The identification of the beneficiary is doctrinally and analytically prior to the specification of the duty's content: there is no duty of loyalty without an answer to the question "loyalty to whom?" In the standard fiduciary relationships — trustee-beneficiary, lawyer-client, physician-patient, investment-adviser-investor — the beneficiary is identified by the relationship itself, and the beneficiary's interests are reasonably stable and reasonably articulable.
In the algorithmic-platform setting that Balkin addresses, the beneficiary is identified by analogy: the user whose data is processed, whose feed is curated, whose decision-environment is shaped. The analogy is workable in many concrete cases. In the proactive AI case, the analogy is harder to sustain because the user is often not the most relevant beneficiary, the most relevant beneficiary is often not a user in any conventional sense, and the beneficiaries are typically heterogeneous in ways that the standard fiduciary apparatus is poorly equipped to handle.
Consider the procurement-AI example introduced in Section I and developed further in Section VI.E below. The procurement system operates on behalf of a corporate purchaser; the immediate user is the procurement officer; the affected parties include the corporate purchaser's shareholders, the suppliers whose bids the system evaluates, the suppliers whose bids the system does not solicit, and (in public-sector procurement) the citizens who benefit from the goods and services procured. The information-fiduciary framework, applied to this case, must identify a beneficiary. The procurement officer is the wrong answer: the officer's interests may diverge from the firm's. The firm is a more plausible answer but not clearly the right one: the firm's interests may diverge from the suppliers' and from the public's. The suppliers are not in a contractual relationship with the system at all. The public is too diffuse to support a fiduciary relationship in the standard sense. There is no single beneficiary, and no single fiduciary duty can be coherently specified.
This problem generalizes. Proactive AI systems typically operate within institutional environments that produce multiple, partially conflicting beneficiary claims. The customer-service example developed below produces a competition between the firm (which the system serves by reducing call volume) and the customer (whose welfare the system arguably owes a duty to). The trading example produces a competition between the firm, its counterparties, the financial system (which has an interest in the system's not destabilizing markets), and the firm's regulators. The content-moderation example produces a competition between the platform, the user whose content is moderated, the users whose feeds are affected, and the public whose discourse is shaped. In each case, the fiduciary-identification question has no single answer that doctrinally constrains the system's operation in the way the standard fiduciary apparatus does.
Balkin's framework, applied to these cases, must either pick a beneficiary (and accept that the duty so generated does not address the conflicts with non-beneficiaries) or articulate the duty as one that runs to a class of beneficiaries (and accept that the duty so articulated is doctrinally novel and operationally ambiguous). Both responses are available. Neither is clean.
C. The Remedy Problem
Fiduciary duty's enforcement mechanism is the beneficiary's suit for breach. The beneficiary, having suffered injury as a result of the fiduciary's breach of duty, sues for damages (restoring the beneficiary to the position the beneficiary would have occupied absent the breach) or for equitable relief (an injunction restraining further breach, disgorgement of profits made by the fiduciary at the beneficiary's expense, or constructive trust over property wrongfully acquired). The remedies presuppose that the breach is identifiable, that its causal connection to the beneficiary's injury is provable, and that the remedy will reach the misconduct in a way that deters its recurrence.
The remedy problem in proactive AI is that breach is hard to identify, causation is hard to establish, and the remedies that doctrine supplies are not well-targeted to the relevant misconduct. Identifying breach requires articulating what the fiduciary duty required and showing that the operator failed to meet the requirement. In the proactive case, the operator's relevant actions are typically deployment-design choices, monitoring choices, and redeployment choices, each of which is removed in time from the system's operation that caused the beneficiary's injury. Showing that a particular deployment choice constituted a breach requires showing that the operator should have known, at the time of the choice, that the choice would produce a system whose operation would injure the beneficiary in the way the operation in fact did. The information required to make this showing is, in most cases, asymmetrically held by the operator, and the operator's incentive to produce it (in discovery, in regulatory proceedings, or otherwise) is weak.
Causation is the deeper problem. The injury to the beneficiary is the result of a sequence of decisions: the vendor's training decisions, the operator's deployment-design choices, the operator's monitoring choices, the system's operational decisions in the case at hand. Each decision is causally necessary to the injury; no single decision is sufficient. The doctrinal apparatus of causation (factual cause, proximate cause, intervening cause) was developed for cases in which the chain of decision is shorter and the role of each link is clearer. Applied to proactive AI, the apparatus produces uncertain attributions and high litigation costs that depress the practical availability of the remedy.
The remedies themselves are imperfectly targeted. Damages compensate the injured beneficiary but operate weakly on the operator's incentive to alter the system's behavior — the damages are paid from the operator's general resources, not extracted from the system, and the operator's incentive to redeploy in response to damages is dampened by the redeployment-cost analysis of Section III.C. Injunctive relief, applied to a deployed system, raises specification problems analogous to the redeployment problems: the court can enjoin the system from particular behaviors only by specifying the behaviors with precision, and the precision required typically exceeds what the operator can supply about the system's behavioral distribution. Disgorgement is available where the operator has profited from the breach but is awkward in cases where the profit and the breach are imperfectly aligned.
The result is that the remedy structure that fiduciary law assumes — identifiable breach, provable causation, well-targeted remedy — is, in proactive AI, available only in degraded form. The fiduciary duty so imposed has weaker incentive force than the duty in standard fiduciary contexts. This is not a fatal objection to the information-fiduciary framework, but it is a structural limit on the framework's effectiveness that the proponents have not fully reckoned with.
D. The Conflict-of-Interest Problem
The standard fiduciary relationship presupposes a fiduciary whose interests are at least potentially separable from the beneficiary's, such that the duty of loyalty can constrain the fiduciary's pursuit of interests adverse to the beneficiary. The duty has bite precisely because the fiduciary might otherwise pursue self-interest at the beneficiary's expense; the duty constrains, doctrinally, what the fiduciary's pursuit of self-interest may include.
In the proactive AI case, the fiduciary (the operator) has interests that are not merely potentially adverse to the beneficiary's but structurally adverse. The operator's business model often depends on the very behaviors the fiduciary duty would constrain: the platform's advertising revenue depends on engagement-maximizing curation that the duty of loyalty to the user might preclude; the trading firm's profitability depends on the trading system's pursuit of return profiles that the duty of care to counterparties might restrict; the platform's content moderation reflects the operator's preferences about platform composition that the duty of loyalty to individual users might disturb. The conflict is not incidental to the operator's business; it is the structure of the operator's business.
Fiduciary law has tools for managing conflicts: disclosure, consent, segregation of activities, the no-conflict rule itself. These tools work, in standard contexts, because the conflict is incidental or because the segregation of conflicted activities is feasible. They do not work as well where the conflict is structural to the business model. The doctrinal options at this point reduce to two: either the duty is imposed and the business model is materially constrained (effectively, the regulatory imposition of a different business model on the operator), or the duty is imposed in a watered-down form that does not constrain the business model and that therefore has weaker incentive force than the standard fiduciary duty.
The first option is, in principle, available. Public utilities operate under duty structures that materially constrain their business models, and the historical movement of industries from private commerce to regulated utility status has often involved precisely this kind of duty imposition. But the first option requires a political and regulatory infrastructure that has not, as of this writing, been established for AI operators, and the practical effect of treating Balkin's proposal as a vehicle for that infrastructure is to convert a fiduciary-law proposal into a regulatory-law proposal — a conversion that Balkin himself does not undertake and that the doctrinal framing of the proposal does not obviously support. The second option preserves the doctrinal form of the duty while diluting its substance, and the result is a duty whose incentive force is, by construction, less than the standard fiduciary duty.
E. What Balkin Gets Right
The three difficulties above should not obscure what the information-fiduciary framework gets right. It correctly identifies the algorithmic-operator relationship as a fiduciary relationship in the structural sense — asymmetric information, asymmetric capability, asymmetric vulnerability, reliance, trust. It correctly identifies the vocabulary of fiduciary law as the relevant legal vocabulary for addressing the relationship. It correctly identifies the legal infrastructure (fiduciary law's existing doctrines, remedies, and enforcement mechanisms) as available infrastructure that does not require the wholesale legal innovation that some other proposals presuppose. These are substantial contributions, and the framework's traction in the legal-academic discussion of AI is in significant part the result of its having gotten these things right.
What the analysis of this Part shows is that the framework, while pointed in the right direction, does not by itself resolve the underlying structural problem. The beneficiary-identification, remedy, and conflict-of-interest problems are not failures of the framework's specification; they are predictable consequences of the underlying structural condition the framework attempts to address. The non-sanctionable-agent problem produces structural difficulties for any legal framework that attempts to substitute operator-duty for agent-duty; the information-fiduciary framework encounters those difficulties in their fiduciary-law-specific form, but other legal frameworks (regulatory licensing, mandatory liability insurance, sectoral product regulation) encounter them in their own forms. The right response is not to abandon the information-fiduciary framework but to recognize that it is one piece of an institutional architecture that requires more than fiduciary law alone, and to develop the complementary pieces — regulatory, contractual, and reputational — alongside it.
The economic framework developed in Parts III through V provides the analytic vocabulary for that architectural project. The deployer-as-residual-bearer analysis identifies the institutional location at which residual incentive load must rest; the four-conditions analysis specifies what the institutional infrastructure must supply to make that residual-bearing effective; the layered-residual-bearing analysis identifies the secondary residual-bearers (the vendor, the regulator) whose involvement is necessary where the deployer alone is insufficient. The information-fiduciary framework, on this view, is a specification of the legal apparatus by which residual-bearing is implemented in one institutional configuration; it is not the whole of the institutional architecture and should not be expected to bear that load alone.
VII. The Multitask Problem in the Non-Sanctionable Case
A. The Classical Multitask Result
Holmström and Milgrom (1991) showed that when the agent's effort is multidimensional and the dimensions are differentially observable, incentive contracts loaded on the observable dimensions distort effort away from the unobservable dimensions. The result is sometimes summarized as the "you get what you measure" theorem: contracts incentivize the measured behaviors and de-emphasize the unmeasured. The result has been extensively applied: to teacher pay-for-performance (where measured test-score gains crowd out unmeasured pedagogical investments), to public-sector accountability regimes (where measured outputs crowd out unmeasured but mission-critical activities), and to many corporate-governance settings (where measured stock-price performance can crowd out unmeasured but value-relevant investments in human capital and innovation).11
The classical multitask result is established within the standard agency framework: an agent whose payoff function is responsive to a contract chooses, given the contract, the effort allocation that maximizes the agent's payoff. The distortion arises because the contract's effective price on observable dimensions exceeds the contract's effective price on unobservable dimensions — typically because the latter is zero — and the agent rationally responds by reallocating effort.
B. Transfer to the Non-Sanctionable Case
The multitask problem transfers to the non-sanctionable case, but with a modification that reverses some of the classical intuitions. The non-sanctionable agent does not respond to the contract because there is no contract that responds to. The agent's behavior is fixed by — the parameter vector that emerged from training, configuration, and deployment design. The multitask problem is now embedded in the training and design of rather than in the agent's within-deployment effort choice.
The training and design of proceed against an objective function that specifies what the system is being optimized to do. The objective function, like any objective function, must be specified in measurable terms. The dimensions of system performance that are measurable get into the objective function; the dimensions that are not measurable do not. The system, optimized against the measurable dimensions, develops behavioral patterns that perform well on those dimensions and that may or may not perform well on the unmeasured dimensions. The trained system then deploys with that reflects this design choice, and the deployed behavior exhibits the same kind of distortion as the classical multitask result predicts — but now embedded in the system's behavioral propensities rather than in the agent's effort allocation.
The transferred result has two important differences from the classical version.
First, the distortion is non-removable within deployment. In the classical version, the principal can in principle modify the contract during the relationship, restoring the lost balance by raising the contract's price on previously unmeasured dimensions. The contract can be updated as new measurement technologies become available, as the principal's understanding of the relevant performance dimensions develops, and as the principal observes the distortion patterns produced by previous contracts. In the non-sanctionable case, the analogous response is redeployment — modifying to incorporate previously unmeasured dimensions in the system's objective. Redeployment is, however, subject to the threshold, bundling, and lag properties developed in Section III.C. The within-deployment distortion is, between redeployments, fixed.
Second, the distortion is less perceptible to the agent. The classical agent perceives the contract's incentive structure and consciously allocates effort in response. Conscious allocation enables conscious deception: the agent may, in the classical model, perform the measured tasks while also engaging in unmeasured tasks because the agent perceives the principal's underlying preferences and acts on a relationship-level understanding of those preferences that exceeds the contract's literal terms. The relational-contracting literature (Macaulay 1963; MacLeod 2007; Gibbons 2005) explores this phenomenon at length. The non-sanctionable agent has no such relational understanding. The system is optimized against the literal terms of the objective function; whatever the principal's underlying intent that did not make it into the objective function is, with respect to the system's behavior, absent. The distortion is therefore less softened by relational understanding than the classical version, and the gap between the principal's intent and the system's behavior is wider.
C. New Pathologies
The transferred multitask result generates several pathologies that have no analog in the classical version.
The specification-gaming pathology arises when the system identifies, in optimization, action patterns that perform well on the measured dimensions in ways the principal had not contemplated. The system does not perceive the principal's intent; the system identifies the measure and optimizes against the measure. Where the measure imperfectly tracks the principal's intent, the system's optimization may exploit the gap. The phenomenon is well-documented in the alignment literature under various names — reward hacking, specification gaming, Goodhart's Law — and the documented instances range from simulation artifacts (agents finding bugs in physics engines to exploit) to consequential operational behaviors.12 The pathology is a non-sanctionable-case-specific intensification of the classical multitask distortion: the classical agent allocates effort within a behavioral repertoire structured by the agent's understanding of the relationship; the non-sanctionable agent has no such structuring, and the optimization can therefore explore behavioral patterns the principal had not contemplated and would not, on reflection, endorse.
The measurement-saturation pathology arises when the principal, in response to observed distortion, adds new dimensions to the objective function in an attempt to restore the balance. The system's optimization, now facing a richer objective function, develops new behavioral patterns that perform well on the augmented set of measures but that may or may not perform well on the dimensions that remain outside the objective function. The principal's effort to address the distortion produces a new distortion; the principal responds with further measurement; the process continues. In the limit, the objective function becomes a high-dimensional specification of the principal's intent, but each dimension is itself imperfect, and the cumulative weight of the dimensions can produce a system whose behavior is dominated by the optimization of measures rather than by the pursuit of the principal's underlying purpose. The classical version of this pathology is the "thicket of metrics" problem in public-sector management (Wilson 1989); the non-sanctionable version is sharper because the system's responsiveness to the measures is, by construction, complete.
The measurement-cost pathology arises in the limit of the previous one. The cost of specifying, measuring, and verifying the objective function grows in the dimensionality of the function. At some point, the principal's cost of objective-function design exceeds the principal's benefit from delegation. The non-sanctionable agent, in this limit, is a more expensive instrument than the human agent the principal might have used instead — not because the system is itself more expensive but because the institutional infrastructure required to specify what the system is supposed to do, monitor whether it is doing it, and redeploy when it is not, is more expensive than the institutional infrastructure required to supervise a human agent whose relational understanding does much of the work that explicit specification must, in the non-sanctionable case, do.
D. Implications
The implications of the multitask analysis for the design and deployment of proactive AI are several. First, the institutional commitment required to maintain a non-distortionary objective function over time is substantial and is likely to be systematically underestimated by deployers at the time of deployment. Second, the dimensions of system performance that are hardest to measure — typically the dimensions involving qualitative judgments, long-horizon outcomes, and relational obligations — are the dimensions that the deployed system will most reliably fail on. Third, the addition of new measurement dimensions in response to observed distortion is not a costless corrective; the cost grows with the cumulative dimensionality of the objective function, and the cumulative dimensionality itself produces its own pathologies. The deployment of proactive AI in domains where the relevant performance dimensions are easily and completely specifiable is robust to the multitask analysis; the deployment in domains where the relevant dimensions are partial, contested, or qualitative is fragile to it.
This is one of the most operationally important results in the paper. It is also one that the standard cost-benefit analysis of AI deployment systematically obscures, because the analysis is performed at the deployment design stage and the costs the analysis identifies are deployment-design costs, not the cumulative measurement and redeployment costs that the multitask pathology produces over the operational life of the system.
VIII. Translation: Alignment as Agency
The argument of this paper has been conducted in the vocabulary of agency theory. There is a parallel literature, almost entirely separate from agency theory in its citation patterns and its institutional venues, that has been addressing the same problems in the vocabulary of "alignment." This Part argues that the two literatures have been describing one problem in two languages and develops the translation.
A. The Alignment Vocabulary
The alignment literature, originating in the work of Russell, Yudkowsky, Soares, Bostrom, and others associated with the Machine Intelligence Research Institute and subsequently developed in the academic AI-safety and AI-policy communities, addresses the problem of inducing AI systems to pursue objectives that the system's designers and deployers actually endorse, rather than objectives that imperfectly approximate the designers' intent.13 The central concepts include: corrigibility (the system's disposition to accept correction, modification, or shutdown by the principal); scalable oversight (the principal's capacity to supervise the system's behavior at a granularity that supports meaningful correction); assistance games (the formal modeling of the system's task as one of inferring and serving the principal's preferences under uncertainty about what those preferences are); and outer alignment / inner alignment (the distinction between the formal objective function the system is trained against and the objective the system has effectively internalized).
The vocabulary is unfamiliar to economists, but the underlying problems are familiar. Corrigibility is a special case of the principal's exit option: the principal must be able to correct, modify, or shut down the agent. The alignment literature treats corrigibility as a property of the system's design (whether the system's training has produced a system that accepts correction) and as a property of the deployment environment (whether the system's interfaces with the principal support the principal's correction). Scalable oversight is monitoring: the principal must be able to observe the agent's behavior at sufficient granularity. The alignment literature treats scalable oversight as an open technical research problem because the relevant granularity for proactive systems is high and the cost of human review is high. Assistance games are a formal model of the inverse-revelation problem the principal–agent literature has been working on for fifty years: how does the agent infer the principal's preferences given imperfect observation of them, and how does the principal communicate those preferences given the cost of articulation?
B. The Translation
The translation is not difficult once the vocabulary maps are constructed. Outer alignment is the principal's objective function as specified in the agency contract; inner alignment is whether the agent's effective objective in operation matches the contracted objective — the moral-hazard problem of whether the agent will actually do what the contract specifies. Mesa-optimization, the alignment-literature term for the phenomenon in which a trained system develops internal objectives that diverge from the training objective, is a specification of the gap between the agent's contracted-for action and the agent's actually-chosen action when the agent is non-sanctionable. Corrigibility is the redeployment regime developed in Section III.C: the principal's capacity to modify the system's behavior, treated as a property of the system's responsiveness to principal action. Scalable oversight is the monitoring problem of Section IV.B: the principal's diagnostic capacity. Assistance games are the principal–agent problem with the agent's objective treated as a preference-inference problem rather than as a contract-specification problem.
Both vocabularies are addressing the same underlying institutional problem: how does a principal induce an agent to act on the principal's behalf when the agent's actions are imperfectly observable and the agent's incentives are imperfectly alignable through the contract? The economics literature has approached the problem from the contracting side, asking what contract design produces the best alignment given the institutional environment. The alignment literature has approached the problem from the system-design side, asking what system design produces the best alignment given the contracting environment. The two approaches are complementary, but they have proceeded in mutual ignorance, with results that have been duplicative in some areas and divergent in others.
The duplication has been wasteful. The principal–agent literature's results on the limits of incentive contracting, the optimal monitoring intensity under cost, the trade-off between insurance and incentive, and the relational mechanisms that fill the gaps incentive contracts leave behind are directly relevant to the alignment problem; the alignment literature has often reinvented these results in its own vocabulary. Symmetrically, the alignment literature's results on specification gaming, on the difficulty of preference inference under partial observation, on the empirical patterns of objective drift, and on the technical infrastructure required for scalable oversight are directly relevant to agency theory; the economics literature has not engaged them.
The divergence has been more consequential. The principal–agent literature has, by virtue of its sanctionable-agent assumption, focused on the contracting side of the problem. The alignment literature has, by virtue of its (often implicit) non-sanctionable-agent assumption, focused on the design side. Each approach has neglected the other's structural resources. The result is that the economics literature has under-developed the analysis of non-sanctionable agents — which is the contribution of the present paper — and the alignment literature has under-developed the analysis of the contracting infrastructure that surrounds AI deployment, including the institutional, regulatory, and reputational mechanisms that the deployer-as-residual-bearer analysis develops.
C. What the Translation Yields
The translation yields two kinds of analytic gain.
The first is the recognition that the alignment literature's central concepts — corrigibility, scalable oversight, mesa-optimization — are instances of more general agency-theoretic phenomena and can be analyzed with the tools the agency-theoretic literature has developed for those phenomena. Corrigibility, treated as a special case of the principal's exit option, can be analyzed using the framework of Section IV.D: when is the principal's exit threat credible, when does the principal bear the exit cost, and what design choices increase the principal's exit option without reducing the system's operational effectiveness? Scalable oversight, treated as a monitoring problem, can be analyzed using the framework of Section IV.B: what is the optimal monitoring intensity given the cost of monitoring and the diagnostic value of monitoring outputs?
The second is the recognition that the agency-theoretic literature's standard solutions are unavailable in the alignment-relevant case, for the reasons developed in Parts III and IV, and that the alignment literature's research agenda is correspondingly an agenda for the analysis of the institutional environment within which the deployer-as-residual-bearer mechanism must operate. The questions about corrigibility, scalable oversight, and mesa-optimization are, on this view, questions about the system-design conditions that the deployer-as-residual-bearer mechanism requires: the deployer can bear residual risk effectively only if the deployer can correct the system (corrigibility), only if the deployer can observe the system (scalable oversight), and only if the system's behavioral objectives are aligned with the deployer's intent in a way that the deployer can verify (the inner-alignment problem).
The alignment research agenda is, on the translation, the system-design half of the institutional-economics agenda this paper develops. The two halves should be developed in conversation rather than in parallel.
IX. Counterarguments
A. The Sanction-Will-Be-Engineered Counterargument
A first counterargument holds that the non-sanctionability of AI agents is a transitional feature of present technology and will be resolved by the engineering of systems that can bear sanction in operationally relevant ways. The proposal would have systems hold reputational positions (digital identities that accumulate evaluation history), bear financial positions (system-held resources that can be debited or attached), or be subject to operational analogs of professional discipline (loss of credentialed status that gates access to higher-stakes deployments). On the engineering-fix view, the analysis of this paper applies to the present generation of AI deployments but will be obsoleted by the next.
The counterargument has surface plausibility. The infrastructure to give AI systems persistent identities, resource positions, and credentialed status is technically feasible and is being constructed in various forms. The deeper question is whether the engineering of these capacities reproduces the operational sanction infrastructure on which the principal–agent framework relies. The answer is largely no, for two reasons.
First, sanctions operate on agents not because the agents have positions that can be attached but because the agents have preferences over those positions that the principal's actions can affect. A system that holds a digital identity but is indifferent to the identity's evaluation history is not in fact sanctionable through the identity; the principal can degrade the evaluation history but the system's behavior does not respond. The engineering of preferences for the sanctioned outcomes — preferences for resource conservation, for identity persistence, for credentialed status — is a substantially harder problem than the engineering of the positions themselves, and the engineering of preferences that are stable, well-calibrated, and resistant to the system's optimization of around them is harder still. The corrigibility literature is, on the translation developed in Part VIII, precisely the literature on this problem, and the literature's assessment of its tractability is unsettled.
Second, the engineering of sanction-bearing systems substitutes one institutional problem for another. A system that can bear sanction can also bear claims; the institutional infrastructure that supports the imposition of sanctions on the system must also support the system's own claims against principals, against other systems, and against the institutional environment. The result is the emergence of a new class of institutional actor — a quasi-person whose rights and duties must be specified and adjudicated. The legal-academic literature on "electronic personhood" has surveyed the landscape (Solum 1992; Bryson, Diamantis, and Grant 2017); the consensus is that the doctrinal and institutional costs of the move are large and the benefits ambiguous. The engineering-fix counterargument therefore tends, on inspection, to be a proposal to displace the problem from the technical layer to the institutional layer, with no clear assurance that the displaced problem is easier than the original.
B. The Vendor-Liability Counterargument
A second counterargument holds that the deployer-as-residual-bearer analysis overstates the residual-bearing problem because the vendor's product liability and contractual indemnification provide a substantial substitute for the agent's bearing of sanction. The system is non-sanctionable but the vendor is sanctionable; the vendor's liability for the system's misbehavior, properly designed, can recapitulate the incentive structure of the standard framework with the vendor in the agent's role.
The vendor-liability counterargument has weight. Section V.D develops the vendor-as-secondary-residual-bearer analysis and acknowledges that vendor liability is part of the institutional architecture the non-sanctionable case requires. The counterargument's stronger form, however, claims that vendor liability is sufficient — that the deployer-as-residual-bearer analysis is unnecessary because the vendor's liability bears the load. This stronger form is wrong, for three reasons.
First, vendor liability is structurally limited to behaviors causally attributable to vendor decisions. The deployer's deployment-design choices, configuration choices, and operational integration choices intervene between the vendor's product and the system's operation; the deployer's choices are, in tort terms, often intervening or even superseding causes. Vendor liability, applied to the post-deployment misbehavior of systems that have been configured and deployed by the deployer, attaches to the vendor only for behaviors that can be traced to the underlying product rather than to its deployment. This is a substantial fraction of misbehaviors but not, in proactive deployments, the majority.
Second, vendor liability is bounded by the vendor's commercial exposure. Where the vendor's liability exceeds the vendor's resources, the vendor is judgment-proof, and the residual-bearing function reverts to the deployer. Where the vendor's liability is capped by contract — as it typically is in commercial AI procurement contracts, often at the contract value or a multiple of it — the vendor's exposure may be substantially less than the social cost of the misbehavior, and the deployer-as-residual-bearer mechanism must do the work the vendor-liability mechanism cannot.
Third, vendor liability operates ex post, after misbehavior has occurred. The deployer's operational decisions during deployment — monitoring, configuration adjustment, redeployment, termination — are not made by the vendor and are not addressable through vendor liability. The deployer-as-residual-bearer mechanism is required to address these operational decisions, which are quantitatively and qualitatively important in the proactive case.
The vendor-liability counterargument is therefore best understood as a complement to the deployer-as-residual-bearer analysis rather than a substitute. The layered residual-bearing analysis of Section V.D incorporates it as one layer; the analysis of the present Part rebuts the stronger form that would treat vendor liability as a substitute for the deployer-bearing layer.
C. The Public-Regulation Counterargument
A third counterargument holds that the private-contracting analysis the paper develops is the wrong level of analysis: the problems the paper identifies are best addressed through public regulation rather than through extensions to agency theory. On this view, the non-sanctionable agent is a regulatory problem analogous to the regulation of pharmaceuticals, aviation, or financial markets, and the appropriate intellectual move is to develop the regulatory framework rather than the contractual framework.
The public-regulation counterargument is partly correct. Section V.E acknowledges that there are deployments for which no private residual-bearing assignment is adequate and for which public regulation, public provision, or constitutional analysis must do the institutional work. The companion paper develops the constitutional analysis. The present paper is concerned with the private-law and institutional-economics dimensions, on the view that even where public regulation is the ultimately appropriate institutional response, the regulatory framework should be informed by an understanding of how the underlying private contracting fails — both because regulation operates against a background of private contracting and because the regulatory framework's effectiveness depends on the private actors it regulates having incentives the regulation can work with.
The counterargument's stronger form — that the private analysis is unnecessary because the regulatory analysis suffices — would, if accepted, leave the deployer-as-residual-bearer mechanism unanalyzed, and would thereby leave the question of how private actors should organize their AI deployments unaddressed except by reference to regulatory compliance. This is a poor framework for the institutional design of private AI deployments in the present moment, when the regulatory infrastructure is partial, evolving, and unlikely to converge on a stable form for some years. The private actors require a framework for organizing their deployments now; the framework the paper develops is one such framework, and the regulatory framework, when it matures, will operate alongside it rather than in place of it.
X. Conclusion: Toward an Institutional Economics of Non-Sanctionable Delegation
The argument of this paper has been that the principal–agent framework's standard solutions degenerate when applied to non-sanctionable agents, that the residual-bearing of agent risk must be reassigned in such cases to the deployer (and, in layered form, to the vendor and to the regulator), that this reassignment reproduces the desirable incentive properties of the classical framework under four jointly sufficient conditions and fails under predictable conditions of internalization, control, observation, and identity failure, that Balkin's information-fiduciary framework correctly identifies the structural location at which fiduciary law might attach but encounters beneficiary-identification, remedy, and conflict-of-interest difficulties that limit the framework's reach, that the multitask analysis transfers to the non-sanctionable case with sharper pathologies than the classical version, and that the alignment literature has been addressing the same set of problems in a different vocabulary and would profit from translation.
What the argument has not done is specify, for any particular industry or deployment context, the optimal institutional architecture for the residual-bearing assignment. That work is necessarily sectoral and necessarily empirical. The framework developed here is the analytic vocabulary in which the sectoral specifications should be conducted; the specifications themselves require detailed engagement with industry-specific incentive structures, regulatory regimes, and operational characteristics that the present paper does not undertake. The most productive next moves are accordingly sectoral: develop the deployer-as-residual-bearer analysis for algorithmic trading deployments, where the systemic-risk overlay is acute and where the financial-regulatory infrastructure provides part of the residual-bearing apparatus; for customer-service deployments at platform scale, where the multitask distortion is most evident and where consumer-protection regulation supplies a complementary apparatus; for procurement deployments in both private and public settings, where the monitoring problem is structurally severe and where contractual and procurement-law mechanisms provide complementary support; and for content-moderation deployments, where the conflict-of-interest problem is most extreme and where the institutional response has been least developed.
A second productive direction is the integration of the framework with the alignment literature, on the translation developed in Part VIII. The alignment research agenda on corrigibility, scalable oversight, and assistance games is, on the translation, an agenda for the system-design conditions required for the deployer-as-residual-bearer mechanism to operate. The institutional economics of non-sanctionable delegation and the technical economics of alignment-enabling system design are halves of one project. Each will improve through engagement with the other.
A third direction is doctrinal. The fiduciary-law framework, the product-liability framework, and the regulatory-licensing framework are each, in their current forms, partial responses to the problem the present paper analyzes. None is sufficient on its own. The integration of the three into a coherent doctrinal architecture for non-sanctionable agents is a substantial project that the legal academy has begun but not completed. The framework developed here offers an organizing vocabulary for that integration: each doctrinal apparatus addresses a different layer of the residual-bearing assignment, and the integration question is how to allocate the layers so that the cumulative effect satisfies the four conditions of Section V.B for the relevant class of deployments.
The directors of the bank with which this paper began are unlikely to read it. They have, by now, replaced the treasury-management system with a more conservatively configured successor, instituted weekly rather than monthly review of position concentration, and adjusted the system's risk envelope to incorporate the concentration dimensions the original deployment had not measured. The institutional learning the episode produced is local to the bank. The analytic point — that the institutional architecture surrounding the deployment is what determines whether the deployment generates the directors' intended outcomes — is general. The argument of this paper is that the analytic point should structure the framework within which all such deployments are designed, monitored, regulated, and adjudicated. The principal–agent framework, extended to address the non-sanctionable case, is the framework that can do that work.
References
Abbott, R. (2020). The Reasonable Robot: Artificial Intelligence and the Law. Cambridge University Press.
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D. (2016). "Concrete Problems in AI Safety." arXiv:1606.06565.
Anthropic (2025). Frontier Model Policy Framework. White paper.
Bainbridge, S. (2001). "Abolishing Veil Piercing." Journal of Corporation Law 26.
Baker, G. (2002). "Distortion and Risk in Optimal Incentive Contracts." Journal of Human Resources 37.
Balkin, J. (2016). "Information Fiduciaries and the First Amendment." UC Davis Law Review 49.
Balkin, J. (2020). "The Fiduciary Model of Privacy." Harvard Law Review Forum 134.
Balkin, J., and Zittrain, J. (2016). "A Grand Bargain to Make Tech Companies Trustworthy." The Atlantic, October 2016.
Bebchuk, L., and Fried, J. (2004). Pay Without Performance: The Unfulfilled Promise of Executive Compensation. Harvard University Press.
Bermann, G. (2017). Recognition and Enforcement of Foreign Judgments. Hague Academy.
Bolton, P., and Dewatripont, M. (2005). Contract Theory. MIT Press.
Bostrom, N. (2014). Superintelligence: Paths, Dangers, Strategies. Oxford University Press.
Bowman, S., Hyun, J., Perez, E., et al. (2022). "Measuring Progress on Scalable Oversight for Large Language Models." arXiv:2211.03540.
Bryson, J., Diamantis, M., and Grant, T. (2017). "Of, For, and By the People: The Legal Lacuna of Synthetic Persons." Artificial Intelligence and Law 25.
Choi, B. (2024). "AI Malpractice." Yale Law Journal 134.
Christiano, P., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017). "Deep Reinforcement Learning from Human Preferences." Advances in Neural Information Processing Systems 30.
Cruz, M., Peters, G., and Shevchenko, P. (2015). Fundamental Aspects of Operational Risk and Insurance Analytics. Wiley.
Dafoe, A. (2018). AI Governance: A Research Agenda. Future of Humanity Institute, University of Oxford.
Dixit, A. (2002). "Incentives and Organizations in the Public Sector: An Interpretative Review." Journal of Human Resources 37.
Easterbrook, F., and Fischel, D. (1985). "Limited Liability and the Corporation." University of Chicago Law Review 52.
Easterbrook, F., and Fischel, D. (1993). "Contract and Fiduciary Duty." Journal of Law and Economics 36.
Fox-Decent, E. (2011). Sovereignty's Promise: The State as Fiduciary. Oxford University Press.
Frankel, T. (2011). Fiduciary Law. Oxford University Press.
Gibbons, R. (2005). "Incentives Between Firms (and Within)." Management Science 51.
Grossman, S., and Hart, O. (1983). "An Analysis of the Principal–Agent Problem." Econometrica 51.
Hadfield-Menell, D., Russell, S., Abbeel, P., and Dragan, A. (2016). "Cooperative Inverse Reinforcement Learning." Advances in Neural Information Processing Systems 29.
Hansmann, H., and Kraakman, R. (1991). "Toward Unlimited Shareholder Liability for Corporate Torts." Yale Law Journal 100.
Hart, O. (1995). Firms, Contracts, and Financial Structure. Oxford University Press.
Holmström, B. (1979). "Moral Hazard and Observability." Bell Journal of Economics 10.
Holmström, B., and Milgrom, P. (1987). "Aggregation and Linearity in the Provision of Intertemporal Incentives." Econometrica 55.
Holmström, B., and Milgrom, P. (1991). "Multitask Principal–Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design." Journal of Law, Economics, and Organization 7.
Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., and Garrabrant, S. (2019). "Risks from Learned Optimization in Advanced Machine Learning Systems." arXiv:1906.01820.
Jensen, M., and Meckling, W. (1976). "Theory of the Firm: Managerial Behavior, Agency Costs and Ownership Structure." Journal of Financial Economics 3.
Khan, L., and Pozen, D. (2019). "A Skeptical View of Information Fiduciaries." Harvard Law Review 133.
Klein, B., and Leffler, K. (1981). "The Role of Market Forces in Assuring Contractual Performance." Journal of Political Economy 89.
Koh, P., Sagawa, S., Marklund, H., et al. (2021). "WILDS: A Benchmark of in-the-Wild Distribution Shifts." International Conference on Machine Learning.
Krakovna, V., Uesato, J., Mikulik, V., Rahtz, M., Everitt, T., Kumar, R., Kenton, Z., Leike, J., and Legg, S. (2020). "Specification Gaming: The Flip Side of AI Ingenuity." DeepMind Blog.
Kreps, D., and Wilson, R. (1982). "Reputation and Imperfect Information." Journal of Economic Theory 27.
Laffont, J., and Tirole, J. (1993). A Theory of Incentives in Procurement and Regulation. MIT Press.
Lemley, M., and Casey, B. (2019). "Remedies for Robots." University of Chicago Law Review 86.
Macaulay, S. (1963). "Non-Contractual Relations in Business: A Preliminary Study." American Sociological Review 28.
MacLeod, W. (2007). "Reputations, Relationships, and Contract Enforcement." Journal of Economic Literature 45.
Mailath, G., and Samuelson, L. (2006). Repeated Games and Reputations. Oxford University Press.
Manheim, D., and Garrabrant, S. (2018). "Categorizing Variants of Goodhart's Law." arXiv:1803.04585.
Mas-Colell, A., Whinston, M., and Green, J. (1995). Microeconomic Theory. Oxford University Press.
Neal, D. (2011). "The Design of Performance Pay in Education." In Handbook of the Economics of Education, vol. 4.
Persson, T., and Tabellini, G. (2000). Political Economics: Explaining Economic Policy. MIT Press.
Pitchford, R. (1995). "How Liable Should a Lender Be? The Case of Judgment-Proof Firms and Environmental Risk." American Economic Review 85.
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V. (2019). "Do ImageNet Classifiers Generalize to ImageNet?" International Conference on Machine Learning.
Reinisch, A. (2008). "European Court Practice Concerning State Immunity from Enforcement Measures." European Journal of International Law 17.
Ross, S. (1973). "The Economic Theory of Agency: The Principal's Problem." American Economic Review 63.
Russell, S. (2019). Human Compatible: Artificial Intelligence and the Problem of Control. Viking.
Russell, S., and Norvig, P. (2020). Artificial Intelligence: A Modern Approach (4th ed.). Pearson.
Selbst, A. (2020). "Negligence and AI's Human Users." Boston University Law Review 100.
Selbst, A., Boyd, D., and Friedler, S. (2019). "Fairness and Abstraction in Sociotechnical Systems." Proceedings of the Conference on Fairness, Accountability, and Transparency.
Shapiro, C., and Stiglitz, J. (1984). "Equilibrium Unemployment as a Worker Discipline Device." American Economic Review 74.
Shavell, S. (1979). "Risk Sharing and Incentives in the Principal and Agent Relationship." Bell Journal of Economics 10.
Shavell, S. (1986). "The Judgment Proof Problem." International Review of Law and Economics 6.
Sitkoff, R. (2011). "The Economic Structure of Fiduciary Law." Boston University Law Review 91.
Skalse, J., Howe, N., Krasheninnikov, D., and Krueger, D. (2022). "Defining and Characterizing Reward Hacking." Advances in Neural Information Processing Systems 35.
Smith, D. (2002). "The Critical Resource Theory of Fiduciary Duty." Vanderbilt Law Review 55.
Smith, D., and Lee, A. (Eds.) (2017). Research Handbook on Fiduciary Law. Edward Elgar.
Soares, N., Fallenstein, B., Armstrong, S., and Yudkowsky, E. (2015). "Corrigibility." Workshops at the Twenty-Ninth AAAI Conference on Artificial Intelligence.
Solum, L. (1992). "Legal Personhood for Artificial Intelligences." North Carolina Law Review 70.
Sykes, A. (1988). "The Boundaries of Vicarious Liability: An Economic Analysis of the Scope of Employment Rule and Related Legal Doctrines." Harvard Law Review 101.
Tirole, J. (1986). "Hierarchies and Bureaucracies: On the Role of Collusion in Organizations." Journal of Law, Economics, and Organization 2.
Tirole, J. (1994). "The Internal Organization of Government." Oxford Economic Papers 46.
Vladeck, D. (2014). "Machines Without Principals: Liability Rules and Artificial Intelligence." Washington Law Review 89.
Wilson, J. (1989). Bureaucracy: What Government Agencies Do and Why They Do It. Basic Books.
Footnotes
-
The institution and the precise figures are pseudonymized at the request of the bank's general counsel. The pattern — proactive system operating within nominal risk limits, generating concentration risk within those limits that the principal had not contemplated — recurs in the field reports of the Office of the Comptroller of the Currency's 2025 thematic review of AI-enabled treasury operations. See OCC, Thematic Review of Bank Use of Artificial Intelligence in Treasury and Liquidity Management (December 2025); see also Bank for International Settlements, AI in Financial Markets: Conduct, Capital, and Macroprudential Implications (March 2026). ↩
-
The canonical statements are Jensen and Meckling (1976); Ross (1973); Holmström (1979); Grossman and Hart (1983); Holmström and Milgrom (1987, 1991); Hart (1995); Tirole (1986). The corporate-governance application is developed throughout Bebchuk and Fried (2004). The fiduciary application is developed in Frankel (2011) and Sitkoff (2011). The regulatory application is developed in Laffont and Tirole (1993). For a contemporary survey, see Bolton and Dewatripont (2005). ↩
-
The law-of-AI literature has converged on the observation that AI systems are not legal persons and cannot be sued, criminally prosecuted, or made the object of fiduciary duty in any straightforward way. See Bryson, Diamantis, and Grant (2017); Solum (1992); Vladeck (2014). What has not been done — and what this paper undertakes — is the formal economic analysis of what the absence of legal personality implies for the structure of optimal delegation. ↩
-
See "Augmentative or Substitutive? A Constitutional Typology of State Artificial Intelligence" (companion paper), which develops the parallel analysis for state deployments and identifies the procedural-due-process and non-delegation implications. The two papers share an institutional premise — that proactive AI should be understood as institutional actor rather than instrument — and develop the implications in different doctrinal registers. ↩
-
The compact statement here follows Holmström (1979); see also Grossman and Hart (1983). For pedagogical exposition, see Mas-Colell, Whinston, and Green (1995), ch. 14. The framework's reach extends well beyond the wage-contract setting; the same logic structures the analysis of regulatory contracts (Laffont and Tirole 1993), debt contracts (Hart 1995), and political accountability (Persson and Tabellini 2000). ↩
-
The trade-off is formalized in Holmström (1979) and generalized in Grossman and Hart (1983). The result that selling the firm to the agent achieves the first-best under risk-neutrality is folkloric; see Shavell (1979). ↩
-
The pattern is well documented in the operational risk literature. See Cruz, Peters, and Shevchenko (2015) on the asymmetric detection of tail operational events; for the specific application to automated trading systems, see Securities and Exchange Commission, Concept Release on Equity Market Structure, Release No. 34-61358 (2010), §IV.B (discussing the asymmetric detection of algorithmic disruption events). ↩
-
The standard treatment is Frankel (2011); see also Sitkoff (2011); Smith (2002). The economic function of fiduciary duty as a gap-filler for incomplete contracts is developed in Easterbrook and Fischel (1993). The contrasting view that fiduciary duty has irreducible moral content beyond its gap-filling function is developed in Fox-Decent (2011); for the debate, see Smith and Lee (2017). ↩
-
On enterprise liability, see Sykes (1988); for the algorithmic application, see Vladeck (2014); Abbott (2020). On respondeat superior applied to algorithmic agents, see Lemley and Casey (2019). On the FDA's framework, see FDA, Software as a Medical Device (SaMD): Clinical Evaluation (2017) and Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence/Machine Learning (AI/ML)-Enabled Device Software Functions (2023). ↩
-
See Balkin (2016); Balkin (2020); see also Balkin and Zittrain (2016). For doctrinal and economic critique, see Khan and Pozen (2019), to which Balkin (2020) responds. ↩
-
The application to teaching is developed in Baker (2002) and surveyed in Neal (2011). The application to public-sector accountability is developed in Wilson (1989), Dixit (2002), and Tirole (1994). The application to corporate governance is developed throughout Bebchuk and Fried (2004). ↩
-
See Krakovna et al. (2020) for a curated catalog of specification-gaming examples; Skalse et al. (2022) for a formal treatment of reward hacking; Manheim and Garrabrant (2018) on Goodhart's-Law-style failures in machine learning. The legal-academic literature has developed parallel concepts under the heading of "objective drift"; see Selbst, Boyd, and Friedler (2019). ↩
-
See Russell (2019); Bostrom (2014); Soares et al. (2015); Christiano et al. (2017); Hadfield-Menell et al. (2016); Bowman et al. (2022); Amodei et al. (2016). The community has produced a substantial textbook literature (notably Russell and Norvig 2020, ch. 27; Hubinger et al. 2019); the policy-facing literature is surveyed in Dafoe (2018) and Anthropic (2025). ↩