Token Is Not a Wage: From Piecework to Agents, What Are Firms Actually Buying?
From Taylorist piece rates and Ford's high wages to platform labor and Agent procurement, this essay separates billing units, delivery promises, and risk allocation to examine firm boundaries, talent development, and the distribution of gains from commoditized intelligence.
企业天天谈降本,为什么有时反而愿意把工资抬高?1914 年福特的五美元工作日让这个问题变得具体。The Henry Ford 的馆藏说明记载,公司此前为九小时工作日支付 2.34 美元,新政策提出八小时、五美元;背景是流水线工作带来的严重人员流动。获得新待遇还要符合资格,并接受对私人生活的调查。福特档案,1914。
Song 等人的美国雇主—员工匹配数据研究显示,企业间收入差异和高收入劳动者的聚集,是理解收入不平等的重要部分。不过,企业之间的差距扩大不能全都归因于企业支付了更高溢价,劳动者构成同样重要;这项研究也不是科技公司薪酬专论。〈Firming Up Inequality〉,2019,作者机构摘要。理解高薪,需要同时看个人能力、企业资源,以及双方怎样分配收益。
对 HR 和正式员工来说,任务怎么变,比岗位名称怎么变更值得观察。Brynjolfsson、Li 与 Raymond 的研究考察了 5,172 名客服人员引入 AI 辅助工具后的表现,报告每小时解决问题数平均提高约 15%,收益在经验较少、技能较低的人员中更明显。这说明特定场景下确有辅助效应,却不能直接推出自主 Agent 能替代整个岗位,更不能据此计算裁员比例。〈Generative AI at Work〉,2025。
Adam Smith,1776,An Inquiry into the Nature and Causes of the Wealth of Nations,第一篇第一章。公开文本入口。分工的经典讨论。
Frederick Winslow Taylor,Shop Management,1903 年首次发表;所链电子版标注 1911 年版。全文。时间研究、计件和劳资激励;案例为作者自述。
Ronald H. Coase,1937,〈The Nature of the Firm〉,Economica 4(16): 386–405。DOI。市场交易与企业内部协调的比较。
Herbert A. Simon,1951,〈A Formal Theory of the Employment Relationship〉,Econometrica 19(3): 293–305。论文入口。雇佣中的权威与可接受行动范围。
Walter W. Powell,1990,〈Neither Market Nor Hierarchy: Network Forms of Organization〉,Research in Organizational Behavior 12: 295–336。作者公开全文。关系、互惠与网络协调。
The Henry Ford,1914 年五美元工作日相关馆藏说明。档案。工资、人员流动与资格审查背景。
Jae Song、David J. Price、Fatih Guvenen、Nicholas Bloom、Till von Wachter,2019,〈Firming Up Inequality〉,Quarterly Journal of Economics 134(1): 1–50。DOI、作者机构摘要。美国企业与劳动者排序研究,非科技行业薪酬普查。
ILO,2018,Digital Labour Platforms and the Future of Work: Towards Decent Work in the Online World。报告。五个微任务平台的跨国劳动调查。
Uma Rani 与 Rishabh Kumar Dhir,2024,〈The Artificial Intelligence Illusion: How Invisible Workers Fuel the “Automated” Economy〉。ILO 文章。AI 供应链中的人类劳动。
Intercom,〈Fin AI Agent Outcomes〉。官方规则。结果定义、推定解决与后续扣除;核对于 2026-09-19。
Bengt Holmström 与 Paul Milgrom,1991,〈Multitask Principal–Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design〉,Journal of Law, Economics, & Organization 7,special issue: 24–52。期刊入口。多维任务与激励失衡。
Erik Brynjolfsson、Danielle Li、Lindsey R. Raymond,2025,〈Generative AI at Work〉,Quarterly Journal of Economics 140(2): 889–942。DOI、作者机构摘要。客服辅助工具的生产率和异质性;采用期刊版本口径。
David H. Autor,2015,〈Why Are There Still So Many Jobs? The History and Future of Workplace Automation〉,Journal of Economic Perspectives 29(3): 3–30。论文。任务替代、互补与就业调整。
A company pays a monthly salary to an employee, a delivery fee to a courier, and a token charge for a model call. All three bills appear to buy completed work, but the underlying promises differ. One party sells a specified output, another accepts future changes in assigned work, and a third supplies a computing service. Before discussing an agent economy, we need to separate those promises and identify the responsibilities missing from the quoted price.
Five quotes, five allocations of risk
Imagine a company preparing to handle 100,000 customer-service inquiries. Its finance lead receives five proposals: hire employees, contract an outsourced support desk, pay freelancers by the case, call a model directly, or buy an agent service priced per resolution. This is a thought experiment, not a description of a real company, and it does not assume which quote is cheapest.
The owner wants lower cost without worse service. Employees want stable income and training. Outsourcers and software vendors need profitable business. Customers want their problems solved. Yet each party knows something different: the supplier knows its capabilities, the buyer knows the exceptions in its business, and the customer alone knows whether an answer actually helped.
If inquiries are standardized, answers are clear, and mistakes are easy to reverse, model calls may be economical. If a refund dispute depends on years of history, customer trust, and internal permissions, a cheap automated answer may create expensive follow-up work. Both outcomes fit the same technology. Distinguishing them requires real contracts, repeat-ticket records, human takeover time, and compensation data.
Before comparing quotes, separate three layers: What is the billing unit? What does the contract promise? Who bears the remaining risk? Monthly software is not thereby an employee. A worker paid by the piece may still have a formal employment relationship. Piece rates describe compensation, employment describes an organizational relationship, high salaries describe the level of pay, and tokens measure a service. Turning them into four successive eras confuses different layers of the problem.
One purchasing question applies to all of them: when the original plan fails, who must keep working, who may charge more, and who has the authority to change how the work is done?
Piecework: who defines a unit of work?
Frederick Taylor did not invent piece-rate pay. Shop Management discussed existing daily wages, piece rates, and bonuses. Taylor tried to combine time studies, standard tasks, and differential pay so managers could specify a reasonable output and method in advance. Its production gains are the reformer’s own reports, not independent experiments. See Taylor, Shop Management.
From Adam Smith’s division of labor to scientific management, organizations recorded and scheduled more production knowledge. The Wealth of Nations used a pin factory to explain how specialization can raise output. Taylor asked a further set of questions: how should a task be decomposed, how long should it take, and what should meeting the standard pay? See Smith, 1776, Book I, Chapter I.
That shift also redistributed power. Once the craft knowledge held by skilled workers became a standard procedure, a firm could use it to train others, compare output, and renegotiate pay. The firm acquired not only a finished product but knowledge about organizing production. Workers might earn more and still lose some control over the pace of work; both could happen at once.
Taylor also described workers’ fear that piece rates would be cut. If today’s extra effort merely gives management the basis for raising tomorrow’s quota, holding back some capacity becomes rational. This is a commitment problem: management promises to share efficiency gains, but workers cannot know how long that promise will last. See the relevant discussion in Shop Management.
Decomposing knowledge work into prompts, tool calls, evaluations, and retries has a structural resemblance to standardization. The analogy has limits. A machine has neither a worker’s living needs nor labor rights, and generated text may be harder to inspect than a conforming part. The durable question is: Who sets the standard, and who carries the work that the standard leaves out?
Employment preserves room to adapt
If every task could be fully specified, inspected, and bought instantly, perhaps a company would need little more than a purchasing office. Ronald Coase’s 1937 answer was that markets themselves have costs: finding a counterparty, bargaining, contracting, and rearranging the agreement when conditions change. Coordination inside a firm has costs too; the boundary of the firm depends on their comparison. See Coase, “The Nature of the Firm”.
Herbert Simon gave employment a more formal treatment in 1951. Within an acceptable set of actions, an employee allows the employer to decide later which action will be required. This helps explain why firms pay for work that is not yet fully specified. It is not a complete legal definition of every employment contract, and it does not imply unlimited obedience. See Simon, “A Formal Theory of the Employment Relationship”.
Return to customer support. At the beginning of a year, a company cannot specify the next product failure, refund policy, and important customer request. Maintaining an internal team preserves a set of capabilities that can be reassigned within agreed boundaries. The company usually carries more training and idle-capacity cost; employees still carry risks from dismissal, performance reviews, and career development. A salary does not eliminate uncertainty. It reallocates it.
Fixed salaries can also be a response to complex cooperation. Sharing knowledge, spotting a problem early, or helping a colleague may be difficult to price separately yet decisive for team performance. Forcing each contribution into an individual piece rate leaves important work without an owner.
People also coordinate outside both spot markets and managerial commands, using repeated dealings, reputation, and reciprocity. Walter Powell’s work on network organizations examines these relationships. Trust in professional services, supplier partnerships, and research teams can coordinate what prices and hierarchy cannot. See Powell, 1990. Agent interfaces and logs may help such arrangements operate, but a communication protocol cannot manufacture trust, capacity to compensate, or a long-term commitment.
Figure 1 | A conceptual comparison of what is purchased. These arrangements coexist and can be combined; the order is not a historical sequence, and high pay remains an incentive within employment.
From Ford to Big Tech: what does high pay buy?
Why would a firm that constantly seeks lower costs ever raise wages? Ford’s five-dollar day in 1914 makes the question concrete. The Henry Ford’s collection note says the company had paid $2.34 for a nine-hour day and proposed five dollars for eight hours amid severe turnover on the assembly line. The new treatment was conditional and included investigations into private life. See the Ford archive, 1914.
Higher pay can reduce departures and training losses, helping a firm maintain steady production. That is a plausible efficiency-wage mechanism, but the episode does not prove that a raise inevitably increases output. High pay and strict control coexisted at Ford, and an income improvement did not improve every working condition.
Today’s high compensation at large technology companies is still an arrangement within employment. Firms may pay for scarce expertise, for knowledge of complex internal systems, or to avoid project delays caused by departures. Equity and deferred bonuses try to connect employee rewards with long-term organizational performance. None of those mechanisms is captured by simply saying that talent is valuable.
A study using matched U.S. employer–employee data by Song and coauthors shows that differences between firms and the sorting of high-income workers matter for understanding inequality. The widening gap between firms is not entirely a larger firm-specific pay premium; workforce composition matters too. The paper is not a compensation survey of technology companies. See “Firming Up Inequality,” 2019, institutional summary. Understanding high pay requires looking at individual ability, firm resources, and the division of gains between them.
The same engineering judgment can affect far more revenue inside an organization with many users, proprietary data, and mature distribution. Compensation cannot therefore be explained only by counting individual actions. Yet excess corporate profit does not guarantee a proportional employee share; bargaining power, competition, and labor institutions still matter.
Agents will change this combination again. If generic execution becomes cheap, people who identify problems, hold context, and make consequential judgments may become more valuable. If those capabilities can also be replicated reliably, their premium may fall. High pay will not persist merely because the worker is human; it depends on the specific scarcity supporting it.
Figure 2 | Higher pay may reduce turnover and retain scarce capability, but it does not settle every working condition or the distribution of gains. This is a mechanism comparison, not a causal decomposition of the Ford case.
Platform labor as an early digital piece-rate experiment
Before agents, online platforms had already divided large volumes of work into tradable units. An ILO survey of about 3,500 workers across five microtask platforms and 75 countries examined pay, task availability, rejections, and unpaid time spent looking for tasks. The posted unit price did not show all of the time needed to obtain the work. See ILO, 2018.
Platforms lower the cost of finding work and transacting across borders. They can also shape the labor process through task allocation, ratings, and account rules. Freedom to decide when to log on is separate from control over pricing, evaluation, and appeals. The platforms and populations in one study cannot represent every freelancer.
Replacing a fixed role with per-task purchases can reduce a company’s cost when demand is weak. The worker may then bear the waiting time between jobs, equipment costs, and skill maintenance. This may be a genuine efficiency gain or merely a transfer of volatility. To tell the difference, calculate job search, waiting, rework, and dispute time alongside pay for completed orders.
Figure 3 | A platform unit price covers only part of the ledger. Net income comparisons should restore off-order hours and worker-funded inputs. The diagram does not estimate cost shares for any platform.
AI production still relies on data labeling, review, and other human work. The ILO’s discussion of “invisible labor” is a reason to inspect this supply chain. See Rani and Dhir, 2024. Their work concerns human inputs across the production chain; it does not claim a person is secretly producing each automated answer in real time. An automated interface alone reveals little about labor upstream.
One possible change remains to be tested. If stable, easily judged tickets are automated first, the tasks left for people may be harder, more urgent, and more volatile. That could create a skilled exception-handling market, or leave gig workers carrying more variation. Evidence would have to come from changes in task difficulty, total hours, income volatility, and rejection records before and after automation.
The distance between a token bill and a useful result
A model token is a sequence unit used to represent model inputs and outputs, and often to meter charges. It is not a fixed number of words or a universal unit of computing across models. Vendor credits are another billing layer; blockchain assets are another category again. Sharing a name does not give them the same economic character.
At the article’s source cutoff, Anthropic documented separate rates for input, output, caching, and related categories. Salesforce’s Agentforce page displayed alternatives involving Flex Credits per action, sessions, and user licenses. See Anthropic’s pricing documentation and the Agentforce pricing page. These pages establish that pricing models coexist, not that one has proved most effective.
The analogy between tokens and an electricity meter helps separate inputs from outputs, but it soon breaks down. Electricity has a standardized physical unit; a token’s computational cost and capability depend on the model and context. Two systems can consume the same number of tokens and deliver very different quality. A better system may need fewer calls for the same task.
A wage supports a labor relationship between a person and an organization, including pay, work arrangements, rights, and obligations. A model fee pays a company for a service. Software can receive a budget and execution authority, but the name of an agent does not supply recoverable assets or a promise to compensate losses. Saying that a company “pays an AI a salary” can imply that it bought a responsible organizational actor when it only bought a service.
A more useful denominator is the full cost per verified useful outcome. For a common period and quality requirement:
Full cost per useful outcome = (model and tool charges for all attempts + allocated orchestration and infrastructure + human review and exception handling + governance expense + estimated error losses) ÷ number of verified useful outcomes.
Charges for retries belong in “all attempts” and should not be added twice. Fixed investments need a stated allocation period. If there are no useful outcomes, the measure has no meaningful finite unit cost. Large losses that have not yet occurred are often difficult to estimate reliably and should appear as separate stress scenarios, not a precise decimal that hides uncertainty.
Figure 4 | An accounting model, not measured data. Use one period and task scope throughout; include retries in call costs and show estimated losses separately from paid expenses.
A cheap model that needs repeated calls and frequent human takeover may cost more than an expensive but stable system. Conversely, an elaborate model may be used for a task that only needed a lookup table. Procurement should match task difficulty to required reliability, then compare the full process cost.
Usage pricing can spare customers some idle-capacity expense, but idle compute has not vanished. The provider still provisions infrastructure, while prepayment and minimum-use contracts may return some risk to the customer. The contract, not the billing unit alone, determines who carries that risk.
Outcome pricing still needs an evaluator
If tokens do not measure value, is direct outcome pricing better? The problem moves from “how much was used?” to “what counts as complete?” At the source cutoff, Fin’s documentation distinguished customer-confirmed resolutions from inferred resolutions where the customer did not seek further help. If the customer returned to the same conversation, the original charge could be reversed. Certain configured handoffs could also count as billable outcomes. The word outcome therefore does not remove the need to read its definition. See Fin’s official outcome rules, checked September 19, 2026.
A system action, a contractually accepted result, and an actual improvement for the customer are three different things. They may coincide or come apart. That gap does not prove bad faith, but it explains why outcome pricing needs observation periods, reversals, and dispute rules.
Figure 5 | Action, contractual outcome, and customer improvement are not automatic equivalents. This is an evaluation mechanism, not evidence about a vendor’s performance; actual charges depend on the contract and current rules.
Holmström and Milgrom’s theory of multitask principal–agent problems offers a useful lens. When a job has several dimensions, strong incentives on an easily measured indicator can crowd out quality that is hard to measure. A weaker incentive on one metric can sometimes be preferable. See Holmström and Milgrom, 1991. Applying that mechanism to agent procurement is an extension in this essay; the paper did not study generative AI.
Consider another thought experiment. A company pays per “resolved ticket.” The supplier can improve the knowledge base or alter closing and escalation rules. Customers know whether their problem recurs, but the buyer mainly sees logs. One path genuinely reduces repeat contacts and benefits both parties. Another uses a short observation window or difficult escalation to improve the metric while worsening the experience. These are possible incentive branches, not allegations about a particular vendor.
To distinguish them, examine recurrence, human transfers, independent samples, and refunds after disputes. Then ask: Who defines the outcome, who can reverse it, and whose costs absorb failure? A supplier that accepts more execution risk may charge a premium, narrow the scope, or retain exclusions. Outcome pricing does not provide every safeguard for free.
Will agents make firms smaller or larger?
Coase’s framework permits two opposite predictions. If agents reduce the cost of finding services, coordinating interfaces, and inspecting deliveries, small teams may purchase more capability from the market. If agents mainly reduce internal communication, documentation, and management costs, large firms may operate more lines of business within one organization.
The result depends on which cost falls more and where data, customer relationships, and liability remain. External purchasing becomes more attractive when tasks are clear, independently verifiable, and suppliers are replaceable. Internal governance may matter more when work relies on sensitive context, frequent joint adjustment, or high-consequence errors. This is a decision framework; actual choices still depend on price, capability, and contract terms.
Figure 6 | A heuristic governance framework. Context dependence and responsibility are distinct in practice but combined here for readability. All four cells may use agents; the matrix is neither an automated decision nor a causal estimate.
Headcount, business scale, and market power should be separated when discussing firm boundaries. Fewer employees do not necessarily mean fewer controlled assets, customers, or transactions. A company can become less labor-intensive and more concentrated at the same time. To decide whether it has really become smaller, examine headcount, revenue, outsourcing, and scope of control separately.
Protocol-based cooperation may lower switching costs only if identity, data, and work history are genuinely portable. If long-term memory, permissions, and evaluation are bound to one platform, a smooth interface can conceal how much must be rebuilt on exit. The number of open interfaces alone does not tell us whether leaving is easy.
Who receives the efficiency gains?
Efficiency gains do not automatically flow to any one party. They may appear as lower prices, better service, higher wages, or more free time. They may remain as corporate profit or platform fees. Technology influences how much new value exists; competition, bargaining, contracts, and institutions help determine its distribution.
Figure 7 | One efficiency gain has several possible destinations; the diagram does not estimate actual shares. Productivity, corporate profit, and public welfare require separate observation.
Owners and managers need concrete answers. Which tasks may run automatically? Where does the budget stop? Who handles exceptions? What result justifies further spending? A limited trial using historical tickets can record full cost and quality before changing roles and suppliers. Multiplying the lowest token price by an ideal number of calls is likely to miss organizational cost.
For HR teams and employees, changing tasks matter more than changing titles. Brynjolfsson, Li, and Raymond studied 5,172 customer-support agents after the introduction of an AI assistant. They reported an average increase of roughly 15 percent in issues resolved per hour, with larger gains among less experienced and lower-skilled workers. That is evidence of assistance in one setting. It does not establish that an autonomous agent can replace an entire job, much less determine a layoff rate. See “Generative AI at Work,” 2025.
If a company automates all junior tasks while still relying on experienced people for exceptions, where will future experts gain experience? AI may also help novices learn through explanation and feedback. A break in apprenticeship is therefore a risk to test, not a settled result. Track whether new workers can handle unfamiliar problems independently, whether they receive real feedback, and whether senior capability remains renewable over several years.
For flexible workers, an agent may let one professional deliver a project that once needed a small team. It may also depress the price of standardized services. People who retain the customer relationship, domain judgment, and delivery responsibility may face a different market from those selling generic actions. Measure income after subscriptions, customer acquisition, rework, and waiting—not a single impressive delivery speed.
For model and platform vendors, products, distribution, and credible delivery stand between input cost and customer willingness to pay. This essay’s working hypothesis is that tokens will remain a low-level cost measure while some customer contracts move toward service levels and outcomes. Work with weak verifiability or exploratory requirements may remain better suited to usage or subscription pricing. Long-run coexistence of several models fits the mechanisms above better than total replacement by one.
More reproducible generic capability may reduce the price of some services, but training investment, compute, proprietary data, brand, and distribution still have costs. Low marginal call cost does not imply a market price of zero. New value may become consumer savings, employee compensation, corporate profit, or platform fees; competition and bargaining determine the proportions.
For society, task productivity, corporate profits, and public welfare must also be measured separately. David Autor’s historical discussion of automation emphasizes that technology replaces some tasks and complements others, while demand and job redesign shape employment. See Autor, 2015. If income increasingly comes from several projects and platforms, benefits and training may need to follow the person. Renaming workers “independent nodes” cannot resolve that institutional question.
What remains scarce when intelligence is commoditized?
From piecework to agents, people have repeatedly tried to describe work clearly enough to count the inputs and exchange the output. Those efforts expand coordination while exposing what lies beyond the price: context that cannot be specified in advance, quality that cannot be observed immediately, and promises that must still be honored when a dispute occurs.
As general-purpose capability becomes cheaper to call, a firm still has to decide what to do, why to trust the result, how to maintain customer relationships, and how to repair failure. Goal definition, trustworthy data, judgment, and accountable responsibility may become relatively scarce. Whether they command higher income still depends on whether they are truly difficult to replace and whether their holders retain bargaining power.
The essay ends at the purchase order: What does it promise to deliver? How will we evaluate it? Who completes the work left over? Writing those three answers down is the beginning of measuring the real efficiency gain—and of seeing which parts of human cooperation still merit long-term investment.
Sources and further reading
The sources below are ordered by their role in the essay. Historical texts help explain mechanisms; surveys and empirical papers retain their sample limits; vendor pages document public pricing and outcome rules at the review date. All seven diagrams are original conceptual models and do not estimate measured costs or distribution shares.
Adam Smith, 1776, An Inquiry into the Nature and Causes of the Wealth of Nations, Book I, Chapter I. Public text. The classic treatment of the division of labor.
Frederick Winslow Taylor, Shop Management, first published in 1903; the linked electronic text identifies the 1911 edition. Full text. Time study, piece rates, and labor–management incentives; cases are the author’s own reports.
Ronald H. Coase, 1937, “The Nature of the Firm,” Economica 4(16): 386–405. DOI. Market transactions compared with coordination inside firms.
Herbert A. Simon, 1951, “A Formal Theory of the Employment Relationship,” Econometrica 19(3): 293–305. Paper. Authority and the acceptable set of actions in employment.
Walter W. Powell, 1990, “Neither Market Nor Hierarchy: Network Forms of Organization,” Research in Organizational Behavior 12: 295–336. Author-hosted text. Relationships, reciprocity, and network coordination.
The Henry Ford, collection note on the 1914 five-dollar day. Archive. Wage, turnover, and qualification context.
Jae Song, David J. Price, Fatih Guvenen, Nicholas Bloom, and Till von Wachter, 2019, “Firming Up Inequality,” Quarterly Journal of Economics 134(1): 1–50. DOI; institutional summary. U.S. firm and worker sorting, not a technology-industry compensation census.
ILO, 2018, Digital Labour Platforms and the Future of Work: Towards Decent Work in the Online World. Report. A multinational survey across five microtask platforms.
Uma Rani and Rishabh Kumar Dhir, 2024, “The Artificial Intelligence Illusion: How Invisible Workers Fuel the ‘Automated’ Economy.” ILO article. Human labor in the AI supply chain.
Anthropic, “Pricing.” Official documentation. Input, output, caching, and related categories; checked September 19, 2026.
Salesforce, “Agentforce Pricing.” Official page. Action credits, sessions, and user licenses; checked September 19, 2026.
Intercom, “Fin AI Agent Outcomes.” Official rules. Outcome definitions, inferred resolution, and subsequent reversals; checked September 19, 2026.
Bengt Holmström and Paul Milgrom, 1991, “Multitask Principal–Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design,” Journal of Law, Economics, & Organization 7, special issue: 24–52. Journal page. Multidimensional tasks and incentive distortion.
Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond, 2025, “Generative AI at Work,” Quarterly Journal of Economics 140(2): 889–942. DOI; institutional summary. Productivity and heterogeneous effects of a customer-support assistant, using the journal version.
David H. Autor, 2015, “Why Are There Still So Many Jobs? The History and Future of Workplace Automation,” Journal of Economic Perspectives 29(3): 3–30. Paper. Task replacement, complementarity, and labor-market adjustment.