← 返回博客 ← Back to Blog

技术与组织

Token 不是工资:从计件制到 Agent,企业到底在购买什么?

从泰勒的计件制、福特的高工资到平台劳动与 Agent 采购,区分计费单位、交付承诺和风险归属,理解企业边界、人才培养与智能商品化后的收益分配。

Technology and Organizations

Token Is Not a Wage: From Piecework to Agents, What Are Firms Actually Buying?

From Taylorist piece rates and Ford's high wages to platform labor and Agent procurement, this essay separates billing units, delivery promises, and risk allocation to examine firm boundaries, talent development, and the distribution of gains from commoditized intelligence.

企业给员工发月薪,给骑手结配送费,也为一次模型调用支付 Token 费用。三张账单看起来都在买“完成工作”,实际买到的承诺却不一样:有人出售约定的产出,有人接受未来的任务调整,有人提供计算服务。讨论 Agent 经济之前,得先分清这些承诺,以及没有写进报价的责任。

五份报价,五种风险安排

假设一家公司准备处理十万次客服咨询,财务负责人收到了五份方案:招聘正式员工、采购外包坐席、让自由职业者按单处理、自己调用模型、购买按解决结果收费的 Agent 服务。这是一个思想实验,不对应任何真实公司,也不预设哪份报价更便宜。

老板希望降低成本、保持服务质量;员工希望收入稳定并获得培训;外包商和软件商希望业务可盈利;顾客希望问题得到解决。他们拥有的信息却不同:供应商更了解自己的能力,企业更了解业务中的例外,顾客才知道一句回答有没有真正帮上忙。

如果咨询内容标准、答案清楚、错误容易纠正,模型调用可能很划算。如果退款争议涉及多年往来、客户信任和内部权限,便宜的自动回答也可能制造昂贵的后续工作。两种结果都符合这个模型,区分它们需要真实的合同、工单复发记录、人工接手时间和赔付数据。

比较采购方案之前,应该先把报价分成三层:计费单位是什么,合同承诺交付什么,剩余风险由谁承担。按月收费的软件不会因此成为员工;领取计件工资的人也可能处于正式雇佣关系。计件制讲报酬方式,雇佣制讲组织关系,大厂高薪讲报酬水平,Token 讲服务计量。把它们串成四个时代,会把不同层次的问题误当成历史替代。

这几种制度可以用同一个采购问题来比较:当原先的计划失效时,谁必须继续投入,谁可以要求加钱,谁有权重新决定工作该怎样完成?

计件制:谁有权定义一件工作

计件工资并不是泰勒的发明。泰勒的《车间管理》讨论了早已存在的日薪、计件、奖金等制度;他试图把时间研究、标准任务与差别报酬结合起来,让管理者能够事先规定合理产量和工作方法。书中那些产量增长的案例,是改革倡导者自己的记录,不能直接当成独立实验。参见 Taylor,《Shop Management》

从亚当·斯密讨论分工到科学管理兴起,生产知识越来越多地被组织记录和调度。《国富论》以制针工场说明专业分工如何提高产出;泰勒则进一步追问,一项具体工作应该怎样分解、需要多少时间、达到标准应该获得什么报酬。斯密,1776,第一篇第一章

这种变化也重新分配了权力。熟练工人原先掌握的诀窍,一旦变成标准作业,就可以被用于培训别人、比较效率和重新议价。企业买到的除了成品,还有组织生产过程的知识。劳动者可能因此获得更高收入,也可能失去对劳动节奏的部分控制;两件事可以同时发生。

泰勒也讨论了工人对计件单价被削减的担忧:如果今天的超额努力只是明天提高定额的依据,保留一部分生产能力就有了现实动机。这是承诺问题。老板承诺分享效率收益,工人却不确定这个承诺会维持多久。《车间管理》相关讨论

今天,把知识工作拆成提示词、工具调用、评测和重试,与这种标准化存在结构上的相似。相似性也有边界:机器没有工人的生活需求和劳动权利;文本生成未必像合格零件一样容易验收。沿用至今的问题是:谁来规定标准,标准遗漏的工作又由谁承担。

雇佣制:为未来的调整留出空间

如果每个任务都能写清、验收并即时买到,企业似乎可以只剩下一间采购办公室。科斯在 1937 年给出的解释是,使用市场本身需要成本:寻找交易对象、议价、订约,以及在情况变化时重新安排。企业内部的协调也有成本,企业边界取决于两种成本的比较。Coase,〈The Nature of the Firm〉

西蒙在 1951 年进一步刻画雇佣关系:员工在可接受的行动范围内,允许雇主以后再决定具体要做什么。这个模型解释了为什么企业愿意为“尚未完全指定的工作”付薪。它不是所有劳动合同的完整法律定义,更不意味着员工交出了无限服从权。Simon,〈A Formal Theory of the Employment Relationship〉

回到客服场景。公司很难在年初写清下一次产品故障、下一轮退款政策和下一位重要客户的全部需求。保留内部团队,相当于保留一组可以在既定边界内重新分配的能力。企业通常承担更多培训和闲置成本,员工也仍承担失业、考核与职业发展的风险;工资并没有消除不确定性,而是重新安排了它。

固定工资有时是对复杂协作的回应。知识分享、提前发现隐患或帮助同事,很难逐项算出公平价格,却可能决定整个团队的表现。硬把它们拆成个人计件,会让没人能认领的工作被忽略。

市场交易与公司命令之外,人们也靠重复合作、声誉和互惠来协调,Powell 的网络组织研究讨论的正是这类关系。专业服务、供应商协作与研究团队中的信任,可以承担价格和层级难以完成的协调。Powell,1990。据此推演,Agent 的接口和日志能够帮助这些安排执行,却不能单凭通信协议生成信任、赔偿能力和长期承诺。

五类购买安排并存:计件购买约定产出,雇佣购买边界内的调整能力,高薪激励维系稀缺能力,Token 购买模型调用,结果合同购买约定成果。每类仍有未被覆盖的风险。 五类购买安排并存:计件购买约定产出,雇佣购买边界内的调整能力,高薪激励维系稀缺能力,Token 购买模型调用,结果合同购买约定成果。每类仍有未被覆盖的风险。
图一|购买对象的概念对照。各类安排长期并存,也可以组合使用;排列不表示历史替代,高薪属于雇佣关系内的激励安排。

从福特到大厂:高工资买到了什么

企业天天谈降本,为什么有时反而愿意把工资抬高?1914 年福特的五美元工作日让这个问题变得具体。The Henry Ford 的馆藏说明记载,公司此前为九小时工作日支付 2.34 美元,新政策提出八小时、五美元;背景是流水线工作带来的严重人员流动。获得新待遇还要符合资格,并接受对私人生活的调查。福特档案,1914

更高的报酬可能减少离职和训练损耗,帮助企业维持稳定生产;这是一种合理的效率工资解释。这段历史却不足以证明“涨薪必然增产”。福特案例中的高工资与严格控制同时存在,收入改善也没有让所有劳动条件随之改善。

今天的大厂高薪仍然是雇佣制内部的安排。企业可能为稀缺专业能力付钱,为熟悉复杂系统的人付钱,也为减少人员离开所造成的项目延误付钱。股权和递延奖金还试图把员工收益与组织的长期表现连在一起。这些机制不能都塞进一句“人才值钱”。

Song 等人的美国雇主—员工匹配数据研究显示,企业间收入差异和高收入劳动者的聚集,是理解收入不平等的重要部分。不过,企业之间的差距扩大不能全都归因于企业支付了更高溢价,劳动者构成同样重要;这项研究也不是科技公司薪酬专论。〈Firming Up Inequality〉,2019,作者机构摘要。理解高薪,需要同时看个人能力、企业资源,以及双方怎样分配收益。

同一个工程判断,放进拥有大量用户、专有数据和成熟分发渠道的组织,可能影响更大的收入规模。劳动报酬因此不能只从个人完成了多少动作来解释。但企业获得超额利润,也不保证员工按同比例分享;议价能力、竞争和劳动制度仍然重要。

Agent 会继续改变这种组合。若通用执行变便宜,识别问题、掌握背景和承担关键判断的人可能更有价值;如果这些能力也能被可靠复制,其溢价也可能下降。高薪不会因“人类身份”自动保留,需要观察它究竟建立在哪一种稀缺性上。

高工资的两面对照:它可能维系稀缺能力、减少离职训练摩擦并增强长期承诺,但不会自动减少组织控制、保证收益同比例分享,也不会让能力永久稀缺。 高工资的两面对照:它可能维系稀缺能力、减少离职训练摩擦并增强长期承诺,但不会自动减少组织控制、保证收益同比例分享,也不会让能力永久稀缺。
图二|高工资可能缓解人员流动、维系稀缺能力,却不会自动改善所有劳动条件或决定收益分配。图中是机制对照,不是对福特案例的因果分解。

平台劳动:数字计件制的先行实验

在 Agent 出现以前,互联网平台已经把大量工作拆成可以交易的小单。ILO 2018 年对五个微任务平台、75 个国家约 3,500 名劳动者的调查,讨论了报酬、任务供给、拒付以及寻找任务所耗费的无偿时间。单价没有显示劳动者为拿到任务付出的全部时间。ILO,2018

平台降低了寻找工作和跨地域交易的门槛,也可以通过订单分配、评价和账户规则影响劳动过程。劳动者在何时上线方面获得的自由,与其对定价、评分和申诉的控制能力,是不同维度。研究中的具体平台和人群也不能代表全部自由职业者。

企业把固定岗位换成按单采购,可以减少需求不足时的支出;接单者却可能要自己承担订单间的等待、设备投入和技能更新。这里既可能有真实的效率提升,也可能只是把波动换了一个承担者。要分清两者,不能只看成功订单的报酬,还得把找单、等待、返工和争议处理一起算进去。

平台劳动的两本账:订单内能看到成功订单报酬、数量、评分、拒付、抽成和到账金额;订单外容易漏掉寻找任务、等待、设备订阅、技能更新、返工、申诉和需求波动。 平台劳动的两本账:订单内能看到成功订单报酬、数量、评分、拒付、抽成和到账金额;订单外容易漏掉寻找任务、等待、设备订阅、技能更新、返工、申诉和需求波动。
图三|平台单价只覆盖账面上可见的一部分。比较净收入时,还要把订单外的总工时和个人投入算回来;图中不表示任何平台的实际成本占比。

AI 生产过程仍依赖数据标注、审核与其他人类工作,ILO 对“隐形劳动”的讨论提醒我们检查这条供应链。Rani 与 Dhir,2024。这项研究讨论的是整条生产链的人力投入,并非声称每次自动回答都由人在幕后实时完成。终端界面看起来多么自动,回答不了供应链用了多少人工。

还有一个尚待验证的变化:如果稳定、容易判断的工单先被自动化,人类接到的剩余任务可能更难、更急,也更不稳定。它既可能孕育高技能的异常处理服务,也可能让零工承担更多波动。答案要到自动化前后的任务难度、总工时、收入波动和拒付记录里找。

Token 账单与有效结果的距离

这里说的模型 Token,是模型表示输入与输出时使用的序列单位,也常被拿来计费。它不是固定数量的字,更不是跨模型通用的算力单位。厂商的商业积分属于另一层账单;区块链资产又是另一回事。名字相同,不代表经济性质相同。

截至本文核对日,Anthropic 文档分别列出输入、输出及缓存等计价类别;Salesforce Agentforce 的页面则同时展示按动作消耗 Flex Credits、按会话和按用户许可等方式。Anthropic 计价文档Agentforce 计价页面。这些页面证明商业模式正在并存,不能证明某一种已经最有效。

把 Token 比作电表读数,有助于区分投入与产出。这个比喻的局限也很明显:电量有统一的物理量纲,Token 的计算成本和能力含义却依赖模型与上下文。两个系统消耗同样多的 Token,交付质量可能完全不同;更好的系统也可能用更少调用完成同一任务。

工资维系的是人与组织之间的劳动关系,包含报酬、工作安排和相应权利义务;模型费用支付给提供服务的企业。软件可以被赋予预算和执行权限,但一个 Agent 名称本身并不提供可以追索的资产或赔偿承诺。因此,“给 AI 发工资”的说法,容易让采购者误以为买到了一个能够承担员工职责的主体。

更有用的比较单位是每个经验证有效结果的全成本。在同一统计期、同一质量要求下,可以写成:

每个有效结果的全成本 =(全部尝试的模型与工具费用 + 编排和基础设施分摊 + 人工复核与异常处理 + 治理费用 + 错误损失估计)÷ 经验证的有效结果数。

重试产生的调用费应计入“全部尝试”,避免重复相加;固定投入需要按说明清楚的期间分摊。有效结果为零时,这个指标不能形成有意义的有限单价。尚未发生的重大错误损失也往往难以可靠估计,应单列压力情景,不能用一个精确小数掩盖未知。

有效结果成本的计算框架:全部尝试的模型和工具费用,加上基础设施、人工复核、治理及错误损失估计,再除以经验证的有效结果数。图中不表示任何成本占比。 有效结果成本的计算框架:全部尝试的模型和工具费用,加上基础设施、人工复核、治理及错误损失估计,再除以经验证的有效结果数。图中不表示任何成本占比。
图四|完整成本的核算模型,非实测数据。所有成本须使用一致期间与任务范围;重试计入调用费用,错误损失估计与已支付费用分开列示。

一个便宜模型如果需要反复调用、频繁接管,可能比一个较贵但稳定的系统更昂贵。反过来,复杂模型也可能被用在本来只需查表的任务上。采购时要把任务难度与所需可靠性匹配起来,再比较完整流程的成本。

按量调用可以让客户少承担一部分闲置容量成本,但算力闲置并没有消失:提供方仍需配置基础设施,预付和最低用量合同又可能把风险分回客户。风险落在谁手里,最终要看合同,不能仅从计费单位推出。

按结果付费,还需要谁来验收

既然 Token 不代表价值,直接按结果付费是不是更好?问题并没有消失,只是从“用了多少”挪到了“什么算完成”。本文核对时,Fin 文档把“解决”区分为客户确认与客户未继续求助的推定情形;客户之后回到同一会话继续求助,原有解决计费会被扣除。按配置完成的特定转交也可以构成收费结果。所以,看到 outcome 这个词,还得继续读它的定义。Fin 官方结果规则,核对于 2026-09-19

系统完成一个动作,合同认可一项交付,顾客得到实际改善,是三件不同的事。它们可以重合,也可能分开。这个差距不代表供应商一定有恶意,却足以解释为什么“按结果”离不开观察期、撤销和争议规则。

结果付费的三层链条:系统完成动作不自动等于合同认可结果,合同认可结果也不自动等于顾客获得改善;需要用复发率、人工转接、独立抽样、撤销和退款规则重新连接。 结果付费的三层链条:系统完成动作不自动等于合同认可结果,合同认可结果也不自动等于顾客获得改善;需要用复发率、人工转接、独立抽样、撤销和退款规则重新连接。
图五|“动作—合同结果—顾客改善”不能自动画等号。图中是验收机制,不表示任何供应商的实际表现;具体收费仍以合同和当期规则为准。

Holmström 与 Milgrom 的多任务委托代理理论提供了分析工具:工作包含多个维度时,对容易测量的指标施加强激励,可能挤压难测量的质量。有时较弱的单项绩效激励反而更合适。Holmström 与 Milgrom,1991。把这一机制用于 Agent 采购,是本文的延伸分析,并非原论文研究了生成式 AI。

继续做一个思想实验。企业按“已解决工单”付款,供应商能选择改进知识库,也能选择改变关闭与升级规则;顾客知道问题是否复发,采购者却通常只能看到日志。一条路径是真正减少重复咨询,双方受益;另一条路径是观察窗口太短或升级太困难,账面指标改善而顾客体验恶化。这是可能的激励分叉,不是对某家厂商实际行为的指控。

要区分两条路径,采购者应检查复发率、人工转接、独立抽样以及争议后是否退费,并问一句:谁定义结果,谁能够推翻结果,失败最终进入谁的成本?供应商若承担更多执行风险,就可能要求风险溢价、限定任务范围或保留免责条款。结果定价不会免费提供一切保障。

Agent 会让企业变小,还是变大

科斯的框架允许两条相反的推演。如果 Agent 降低寻找服务、协调接口和验收交付的成本,小团队就可能更方便地采购外部能力,把更多执行留在市场。但如果它主要降低内部沟通、文档与管理成本,大公司也可能用同一套组织承载更多业务。

区别在于哪一边的成本下降得更多,以及数据、客户关系与责任留在哪里。任务清楚、可独立验收、供应商容易替换时,外部采购更有吸引力;依赖敏感背景、频繁共同调整、错误后果重大时,内部治理的价值可能更高。这是判断框架,实际选择仍要考虑价格、能力和具体合同。

治理选择示意矩阵:按结果可验证性与组织背景责任要求比较标准服务采购、结果合同及担保、阶段合作和内部团队。Agent 可以参与所有四类安排。 治理选择示意矩阵:按结果可验证性与组织背景责任要求比较标准服务采购、结果合同及担保、阶段合作和内部团队。Agent 可以参与所有四类安排。
图六|组织选择的启发式框架。背景依赖与责任要求在现实中是不同变量,此处合并展示以便阅读;四格均可使用 Agent,不代表自动决策或因果估计。

讨论企业边界时,需要把人数、业务规模和市场权力分开。一家公司减少了员工,并不意味着它控制的资产、客户或交易减少。少人化甚至可以与集中化同时出现。判断企业是否真的“变小”,应分别看人数、营收、外包比例和控制范围。

协议化协作有机会降低更换供应商的成本,但前提是身份、数据和工作记录真的可迁移。若企业把长期记忆、权限体系和验收方式都绑定在一个平台上,接口越顺畅,切换时要重建的东西反而可能越多。开放接口的数量,无法单独回答退出是否容易。

效率收益会落到谁手里

效率提高以后,收益不会自动流向任何一方。它可能变成更低的价格、更好的服务、更高的工资或更多自由时间,也可能留在企业利润和平台收费里。技术影响新增价值能有多少,竞争、议价、合同和制度参与决定它怎样分配。

同一份效率收益可能流向消费者的价格和服务、劳动者的报酬时间与培训、企业的利润和再投资,或平台与供应商的软件算力收费;分配受竞争、议价、合同与制度影响。 同一份效率收益可能流向消费者的价格和服务、劳动者的报酬时间与培训、企业的利润和再投资,或平台与供应商的软件算力收费;分配受竞争、议价、合同与制度影响。
图七|同一份效率收益有多条去路,图中不表示实际分配比例。生产率、企业利润和公众福利需要分别观察。

企业主需要回答一组具体问题:哪些任务允许自动执行,预算到哪里停止,谁接手异常,什么结果值得继续投入。可以先用历史工单或限定范围试运行,记录完整成本与质量,再决定岗位和供应商配置。拿最低 Token 单价乘以理想调用次数,很容易漏掉组织成本。

对 HR 和正式员工来说,任务怎么变,比岗位名称怎么变更值得观察。Brynjolfsson、Li 与 Raymond 的研究考察了 5,172 名客服人员引入 AI 辅助工具后的表现,报告每小时解决问题数平均提高约 15%,收益在经验较少、技能较低的人员中更明显。这说明特定场景下确有辅助效应,却不能直接推出自主 Agent 能替代整个岗位,更不能据此计算裁员比例。〈Generative AI at Work〉,2025

如果公司把初级任务全部自动化,却仍依赖有经验的人判断例外,新人从哪里获得经验?AI 也可能通过解释和反馈帮助新人学习。“学徒制断裂”目前只是待验证的风险,需要观察新人能否独立处理陌生问题、是否获得真实反馈,以及多年后高级人才是否仍可持续培养。

对灵活用工者,Agent 可能让一个专业人士交付过去需要小团队完成的项目,也可能压低可标准化服务的价格。能够保留客户关系、领域判断和交付责任的人,与只出售通用动作的人,未必面对相同的市场。应衡量扣除订阅、获客、返工和等待之后的收入,而不是展示一次惊艳的交付速度。

对模型和平台供应商,投入价格与客户愿意支付的价格之间,还有产品、分发和可信交付。本文的工作假说是:Token 会继续作为底层成本指标,一部分客户合同会向服务等级和结果靠拢;可验证性差、需求探索性强的工作,仍可能适合按量或订阅。多种定价长期并存,比单一模式全面取代更符合前面的机制分析。

可复制的通用能力供给增加,可能压低某些服务的价格;训练投入、算力、专有数据、品牌和渠道仍有成本。低边际调用成本不等于市场价格必然趋零。新增收益可能进入消费者节省的支出、员工报酬、企业利润或平台收费,分配比例取决于竞争和议价。

对社会而言,任务生产率、企业利润与公众福利也需要分别观察。Autor 对自动化历史的讨论强调,技术既替代任务,也可能与其他任务互补;就业结果还受到需求和工作重组影响。Autor,2015。如果收入越来越来自多个项目和平台,就需要讨论保障与培训怎样随人转移;这是制度议题,不能靠把劳动者改称“独立节点”来解决。

智能商品化以后,什么仍然稀缺

从计件到 Agent,人们反复尝试把工作描述清楚,让投入能够计算,让结果能够交换。这种努力扩大了协作范围,也不断暴露出计价之外的部分:难以提前说明的背景、无法立刻观察的质量,以及出现争议以后必须兑现的承诺。

当可调用的通用能力更便宜,企业仍需回答做什么、凭什么相信结果、怎样维护客户关系,以及失败后如何补救。目标定义、可信数据、判断力和责任承担可能变得相对稀缺;但谁会由此获得更高收入,仍取决于这些能力是否真的难以替代,以及能否保留议价权。

文章最后仍回到那张采购单:它承诺交付什么,我们怎样验收,剩下的工作由谁继续完成?把这三件事写清,才能知道效率究竟提高了多少,也才能看见人类协作的哪些部分仍值得长期投入。

参考资料与延伸阅读

主要资料按文中用途排列。历史文献帮助解释机制,调查和实证研究都有各自的样本边界,厂商页面只用来说明核对时公开的计费与结果规则。本文七组图均为原创概念图,不表示实测成本或收益分配比例。

  1. Adam Smith,1776,An Inquiry into the Nature and Causes of the Wealth of Nations,第一篇第一章。公开文本入口。分工的经典讨论。
  2. Frederick Winslow Taylor,Shop Management,1903 年首次发表;所链电子版标注 1911 年版。全文。时间研究、计件和劳资激励;案例为作者自述。
  3. Ronald H. Coase,1937,〈The Nature of the Firm〉,Economica 4(16): 386–405。DOI。市场交易与企业内部协调的比较。
  4. Herbert A. Simon,1951,〈A Formal Theory of the Employment Relationship〉,Econometrica 19(3): 293–305。论文入口。雇佣中的权威与可接受行动范围。
  5. Walter W. Powell,1990,〈Neither Market Nor Hierarchy: Network Forms of Organization〉,Research in Organizational Behavior 12: 295–336。作者公开全文。关系、互惠与网络协调。
  6. The Henry Ford,1914 年五美元工作日相关馆藏说明。档案。工资、人员流动与资格审查背景。
  7. Jae Song、David J. Price、Fatih Guvenen、Nicholas Bloom、Till von Wachter,2019,〈Firming Up Inequality〉,Quarterly Journal of Economics 134(1): 1–50。DOI作者机构摘要。美国企业与劳动者排序研究,非科技行业薪酬普查。
  8. ILO,2018,Digital Labour Platforms and the Future of Work: Towards Decent Work in the Online World报告。五个微任务平台的跨国劳动调查。
  9. Uma Rani 与 Rishabh Kumar Dhir,2024,〈The Artificial Intelligence Illusion: How Invisible Workers Fuel the “Automated” Economy〉。ILO 文章。AI 供应链中的人类劳动。
  10. Anthropic,〈Pricing〉。官方文档。输入、输出、缓存等计费类别;核对于 2026-09-19。
  11. Salesforce,〈Agentforce Pricing〉。官方页面。动作积分、会话与用户许可;核对于 2026-09-19。
  12. Intercom,〈Fin AI Agent Outcomes〉。官方规则。结果定义、推定解决与后续扣除;核对于 2026-09-19。
  13. Bengt Holmström 与 Paul Milgrom,1991,〈Multitask Principal–Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design〉,Journal of Law, Economics, & Organization 7,special issue: 24–52。期刊入口。多维任务与激励失衡。
  14. Erik Brynjolfsson、Danielle Li、Lindsey R. Raymond,2025,〈Generative AI at Work〉,Quarterly Journal of Economics 140(2): 889–942。DOI作者机构摘要。客服辅助工具的生产率和异质性;采用期刊版本口径。
  15. David H. Autor,2015,〈Why Are There Still So Many Jobs? The History and Future of Workplace Automation〉,Journal of Economic Perspectives 29(3): 3–30。论文。任务替代、互补与就业调整。

A company pays a monthly salary to an employee, a delivery fee to a courier, and a token charge for a model call. All three bills appear to buy completed work, but the underlying promises differ. One party sells a specified output, another accepts future changes in assigned work, and a third supplies a computing service. Before discussing an agent economy, we need to separate those promises and identify the responsibilities missing from the quoted price.

Five quotes, five allocations of risk

Imagine a company preparing to handle 100,000 customer-service inquiries. Its finance lead receives five proposals: hire employees, contract an outsourced support desk, pay freelancers by the case, call a model directly, or buy an agent service priced per resolution. This is a thought experiment, not a description of a real company, and it does not assume which quote is cheapest.

The owner wants lower cost without worse service. Employees want stable income and training. Outsourcers and software vendors need profitable business. Customers want their problems solved. Yet each party knows something different: the supplier knows its capabilities, the buyer knows the exceptions in its business, and the customer alone knows whether an answer actually helped.

If inquiries are standardized, answers are clear, and mistakes are easy to reverse, model calls may be economical. If a refund dispute depends on years of history, customer trust, and internal permissions, a cheap automated answer may create expensive follow-up work. Both outcomes fit the same technology. Distinguishing them requires real contracts, repeat-ticket records, human takeover time, and compensation data.

Before comparing quotes, separate three layers: What is the billing unit? What does the contract promise? Who bears the remaining risk? Monthly software is not thereby an employee. A worker paid by the piece may still have a formal employment relationship. Piece rates describe compensation, employment describes an organizational relationship, high salaries describe the level of pay, and tokens measure a service. Turning them into four successive eras confuses different layers of the problem.

One purchasing question applies to all of them: when the original plan fails, who must keep working, who may charge more, and who has the authority to change how the work is done?

Piecework: who defines a unit of work?

Frederick Taylor did not invent piece-rate pay. Shop Management discussed existing daily wages, piece rates, and bonuses. Taylor tried to combine time studies, standard tasks, and differential pay so managers could specify a reasonable output and method in advance. Its production gains are the reformer’s own reports, not independent experiments. See Taylor, Shop Management.

From Adam Smith’s division of labor to scientific management, organizations recorded and scheduled more production knowledge. The Wealth of Nations used a pin factory to explain how specialization can raise output. Taylor asked a further set of questions: how should a task be decomposed, how long should it take, and what should meeting the standard pay? See Smith, 1776, Book I, Chapter I.

That shift also redistributed power. Once the craft knowledge held by skilled workers became a standard procedure, a firm could use it to train others, compare output, and renegotiate pay. The firm acquired not only a finished product but knowledge about organizing production. Workers might earn more and still lose some control over the pace of work; both could happen at once.

Taylor also described workers’ fear that piece rates would be cut. If today’s extra effort merely gives management the basis for raising tomorrow’s quota, holding back some capacity becomes rational. This is a commitment problem: management promises to share efficiency gains, but workers cannot know how long that promise will last. See the relevant discussion in Shop Management.

Decomposing knowledge work into prompts, tool calls, evaluations, and retries has a structural resemblance to standardization. The analogy has limits. A machine has neither a worker’s living needs nor labor rights, and generated text may be harder to inspect than a conforming part. The durable question is: Who sets the standard, and who carries the work that the standard leaves out?

Employment preserves room to adapt

If every task could be fully specified, inspected, and bought instantly, perhaps a company would need little more than a purchasing office. Ronald Coase’s 1937 answer was that markets themselves have costs: finding a counterparty, bargaining, contracting, and rearranging the agreement when conditions change. Coordination inside a firm has costs too; the boundary of the firm depends on their comparison. See Coase, “The Nature of the Firm”.

Herbert Simon gave employment a more formal treatment in 1951. Within an acceptable set of actions, an employee allows the employer to decide later which action will be required. This helps explain why firms pay for work that is not yet fully specified. It is not a complete legal definition of every employment contract, and it does not imply unlimited obedience. See Simon, “A Formal Theory of the Employment Relationship”.

Return to customer support. At the beginning of a year, a company cannot specify the next product failure, refund policy, and important customer request. Maintaining an internal team preserves a set of capabilities that can be reassigned within agreed boundaries. The company usually carries more training and idle-capacity cost; employees still carry risks from dismissal, performance reviews, and career development. A salary does not eliminate uncertainty. It reallocates it.

Fixed salaries can also be a response to complex cooperation. Sharing knowledge, spotting a problem early, or helping a colleague may be difficult to price separately yet decisive for team performance. Forcing each contribution into an individual piece rate leaves important work without an owner.

People also coordinate outside both spot markets and managerial commands, using repeated dealings, reputation, and reciprocity. Walter Powell’s work on network organizations examines these relationships. Trust in professional services, supplier partnerships, and research teams can coordinate what prices and hierarchy cannot. See Powell, 1990. Agent interfaces and logs may help such arrangements operate, but a communication protocol cannot manufacture trust, capacity to compensate, or a long-term commitment.

Five purchasing arrangements coexist: piecework buys specified output, employment buys bounded adaptability, high pay helps retain scarce capability, tokens buy model calls, and outcome contracts buy an agreed result. Each leaves some risk uncovered. Five purchasing arrangements coexist: piecework buys specified output, employment buys bounded adaptability, high pay helps retain scarce capability, tokens buy model calls, and outcome contracts buy an agreed result. Each leaves some risk uncovered.
Figure 1 | A conceptual comparison of what is purchased. These arrangements coexist and can be combined; the order is not a historical sequence, and high pay remains an incentive within employment.

From Ford to Big Tech: what does high pay buy?

Why would a firm that constantly seeks lower costs ever raise wages? Ford’s five-dollar day in 1914 makes the question concrete. The Henry Ford’s collection note says the company had paid $2.34 for a nine-hour day and proposed five dollars for eight hours amid severe turnover on the assembly line. The new treatment was conditional and included investigations into private life. See the Ford archive, 1914.

Higher pay can reduce departures and training losses, helping a firm maintain steady production. That is a plausible efficiency-wage mechanism, but the episode does not prove that a raise inevitably increases output. High pay and strict control coexisted at Ford, and an income improvement did not improve every working condition.

Today’s high compensation at large technology companies is still an arrangement within employment. Firms may pay for scarce expertise, for knowledge of complex internal systems, or to avoid project delays caused by departures. Equity and deferred bonuses try to connect employee rewards with long-term organizational performance. None of those mechanisms is captured by simply saying that talent is valuable.

A study using matched U.S. employer–employee data by Song and coauthors shows that differences between firms and the sorting of high-income workers matter for understanding inequality. The widening gap between firms is not entirely a larger firm-specific pay premium; workforce composition matters too. The paper is not a compensation survey of technology companies. See “Firming Up Inequality,” 2019, institutional summary. Understanding high pay requires looking at individual ability, firm resources, and the division of gains between them.

The same engineering judgment can affect far more revenue inside an organization with many users, proprietary data, and mature distribution. Compensation cannot therefore be explained only by counting individual actions. Yet excess corporate profit does not guarantee a proportional employee share; bargaining power, competition, and labor institutions still matter.

Agents will change this combination again. If generic execution becomes cheap, people who identify problems, hold context, and make consequential judgments may become more valuable. If those capabilities can also be replicated reliably, their premium may fall. High pay will not persist merely because the worker is human; it depends on the specific scarcity supporting it.

High wages may preserve scarce capability, reduce turnover and training friction, and strengthen long-term commitment. They do not automatically reduce organizational control, guarantee a proportional share of gains, or make a capability permanently scarce. High wages may preserve scarce capability, reduce turnover and training friction, and strengthen long-term commitment. They do not automatically reduce organizational control, guarantee a proportional share of gains, or make a capability permanently scarce.
Figure 2 | Higher pay may reduce turnover and retain scarce capability, but it does not settle every working condition or the distribution of gains. This is a mechanism comparison, not a causal decomposition of the Ford case.

Platform labor as an early digital piece-rate experiment

Before agents, online platforms had already divided large volumes of work into tradable units. An ILO survey of about 3,500 workers across five microtask platforms and 75 countries examined pay, task availability, rejections, and unpaid time spent looking for tasks. The posted unit price did not show all of the time needed to obtain the work. See ILO, 2018.

Platforms lower the cost of finding work and transacting across borders. They can also shape the labor process through task allocation, ratings, and account rules. Freedom to decide when to log on is separate from control over pricing, evaluation, and appeals. The platforms and populations in one study cannot represent every freelancer.

Replacing a fixed role with per-task purchases can reduce a company’s cost when demand is weak. The worker may then bear the waiting time between jobs, equipment costs, and skill maintenance. This may be a genuine efficiency gain or merely a transfer of volatility. To tell the difference, calculate job search, waiting, rework, and dispute time alongside pay for completed orders.

The visible order ledger records completed-order pay, volume, ratings, rejections, platform fees, and cash received. A second ledger often misses search time, waiting, equipment and subscriptions, skill updates, rework, appeals, and demand volatility. The visible order ledger records completed-order pay, volume, ratings, rejections, platform fees, and cash received. A second ledger often misses search time, waiting, equipment and subscriptions, skill updates, rework, appeals, and demand volatility.
Figure 3 | A platform unit price covers only part of the ledger. Net income comparisons should restore off-order hours and worker-funded inputs. The diagram does not estimate cost shares for any platform.

AI production still relies on data labeling, review, and other human work. The ILO’s discussion of “invisible labor” is a reason to inspect this supply chain. See Rani and Dhir, 2024. Their work concerns human inputs across the production chain; it does not claim a person is secretly producing each automated answer in real time. An automated interface alone reveals little about labor upstream.

One possible change remains to be tested. If stable, easily judged tickets are automated first, the tasks left for people may be harder, more urgent, and more volatile. That could create a skilled exception-handling market, or leave gig workers carrying more variation. Evidence would have to come from changes in task difficulty, total hours, income volatility, and rejection records before and after automation.

The distance between a token bill and a useful result

A model token is a sequence unit used to represent model inputs and outputs, and often to meter charges. It is not a fixed number of words or a universal unit of computing across models. Vendor credits are another billing layer; blockchain assets are another category again. Sharing a name does not give them the same economic character.

At the article’s source cutoff, Anthropic documented separate rates for input, output, caching, and related categories. Salesforce’s Agentforce page displayed alternatives involving Flex Credits per action, sessions, and user licenses. See Anthropic’s pricing documentation and the Agentforce pricing page. These pages establish that pricing models coexist, not that one has proved most effective.

The analogy between tokens and an electricity meter helps separate inputs from outputs, but it soon breaks down. Electricity has a standardized physical unit; a token’s computational cost and capability depend on the model and context. Two systems can consume the same number of tokens and deliver very different quality. A better system may need fewer calls for the same task.

A wage supports a labor relationship between a person and an organization, including pay, work arrangements, rights, and obligations. A model fee pays a company for a service. Software can receive a budget and execution authority, but the name of an agent does not supply recoverable assets or a promise to compensate losses. Saying that a company “pays an AI a salary” can imply that it bought a responsible organizational actor when it only bought a service.

A more useful denominator is the full cost per verified useful outcome. For a common period and quality requirement:

Full cost per useful outcome = (model and tool charges for all attempts + allocated orchestration and infrastructure + human review and exception handling + governance expense + estimated error losses) ÷ number of verified useful outcomes.

Charges for retries belong in “all attempts” and should not be added twice. Fixed investments need a stated allocation period. If there are no useful outcomes, the measure has no meaningful finite unit cost. Large losses that have not yet occurred are often difficult to estimate reliably and should appear as separate stress scenarios, not a precise decimal that hides uncertainty.

A full-cost framework adds model and tool costs for every attempt, infrastructure, human review, governance, and estimated error losses, then divides by verified useful outcomes. The diagram does not imply any cost share. A full-cost framework adds model and tool costs for every attempt, infrastructure, human review, governance, and estimated error losses, then divides by verified useful outcomes. The diagram does not imply any cost share.
Figure 4 | An accounting model, not measured data. Use one period and task scope throughout; include retries in call costs and show estimated losses separately from paid expenses.

A cheap model that needs repeated calls and frequent human takeover may cost more than an expensive but stable system. Conversely, an elaborate model may be used for a task that only needed a lookup table. Procurement should match task difficulty to required reliability, then compare the full process cost.

Usage pricing can spare customers some idle-capacity expense, but idle compute has not vanished. The provider still provisions infrastructure, while prepayment and minimum-use contracts may return some risk to the customer. The contract, not the billing unit alone, determines who carries that risk.

Outcome pricing still needs an evaluator

If tokens do not measure value, is direct outcome pricing better? The problem moves from “how much was used?” to “what counts as complete?” At the source cutoff, Fin’s documentation distinguished customer-confirmed resolutions from inferred resolutions where the customer did not seek further help. If the customer returned to the same conversation, the original charge could be reversed. Certain configured handoffs could also count as billable outcomes. The word outcome therefore does not remove the need to read its definition. See Fin’s official outcome rules, checked September 19, 2026.

A system action, a contractually accepted result, and an actual improvement for the customer are three different things. They may coincide or come apart. That gap does not prove bad faith, but it explains why outcome pricing needs observation periods, reversals, and dispute rules.

A system action is not automatically a contractually accepted result, and an accepted result is not automatically a customer improvement. Recurrence, human handoff, independent sampling, reversals, and refunds reconnect the three layers. A system action is not automatically a contractually accepted result, and an accepted result is not automatically a customer improvement. Recurrence, human handoff, independent sampling, reversals, and refunds reconnect the three layers.
Figure 5 | Action, contractual outcome, and customer improvement are not automatic equivalents. This is an evaluation mechanism, not evidence about a vendor’s performance; actual charges depend on the contract and current rules.

Holmström and Milgrom’s theory of multitask principal–agent problems offers a useful lens. When a job has several dimensions, strong incentives on an easily measured indicator can crowd out quality that is hard to measure. A weaker incentive on one metric can sometimes be preferable. See Holmström and Milgrom, 1991. Applying that mechanism to agent procurement is an extension in this essay; the paper did not study generative AI.

Consider another thought experiment. A company pays per “resolved ticket.” The supplier can improve the knowledge base or alter closing and escalation rules. Customers know whether their problem recurs, but the buyer mainly sees logs. One path genuinely reduces repeat contacts and benefits both parties. Another uses a short observation window or difficult escalation to improve the metric while worsening the experience. These are possible incentive branches, not allegations about a particular vendor.

To distinguish them, examine recurrence, human transfers, independent samples, and refunds after disputes. Then ask: Who defines the outcome, who can reverse it, and whose costs absorb failure? A supplier that accepts more execution risk may charge a premium, narrow the scope, or retain exclusions. Outcome pricing does not provide every safeguard for free.

Will agents make firms smaller or larger?

Coase’s framework permits two opposite predictions. If agents reduce the cost of finding services, coordinating interfaces, and inspecting deliveries, small teams may purchase more capability from the market. If agents mainly reduce internal communication, documentation, and management costs, large firms may operate more lines of business within one organization.

The result depends on which cost falls more and where data, customer relationships, and liability remain. External purchasing becomes more attractive when tasks are clear, independently verifiable, and suppliers are replaceable. Internal governance may matter more when work relies on sensitive context, frequent joint adjustment, or high-consequence errors. This is a decision framework; actual choices still depend on price, capability, and contract terms.

A governance matrix compares standard service purchasing, outcome contracts with guarantees, staged collaboration, and internal teams by result verifiability and the need for organizational context and responsibility. Agents can participate in all four. A governance matrix compares standard service purchasing, outcome contracts with guarantees, staged collaboration, and internal teams by result verifiability and the need for organizational context and responsibility. Agents can participate in all four.
Figure 6 | A heuristic governance framework. Context dependence and responsibility are distinct in practice but combined here for readability. All four cells may use agents; the matrix is neither an automated decision nor a causal estimate.

Headcount, business scale, and market power should be separated when discussing firm boundaries. Fewer employees do not necessarily mean fewer controlled assets, customers, or transactions. A company can become less labor-intensive and more concentrated at the same time. To decide whether it has really become smaller, examine headcount, revenue, outsourcing, and scope of control separately.

Protocol-based cooperation may lower switching costs only if identity, data, and work history are genuinely portable. If long-term memory, permissions, and evaluation are bound to one platform, a smooth interface can conceal how much must be rebuilt on exit. The number of open interfaces alone does not tell us whether leaving is easy.

Who receives the efficiency gains?

Efficiency gains do not automatically flow to any one party. They may appear as lower prices, better service, higher wages, or more free time. They may remain as corporate profit or platform fees. Technology influences how much new value exists; competition, bargaining, contracts, and institutions help determine its distribution.

The same efficiency gain can flow to consumers through price and service, to workers through pay, time, and training, to firms through profit and reinvestment, or to platforms and suppliers through software and compute fees. Competition, bargaining, contracts, and institutions shape the split. The same efficiency gain can flow to consumers through price and service, to workers through pay, time, and training, to firms through profit and reinvestment, or to platforms and suppliers through software and compute fees. Competition, bargaining, contracts, and institutions shape the split.
Figure 7 | One efficiency gain has several possible destinations; the diagram does not estimate actual shares. Productivity, corporate profit, and public welfare require separate observation.

Owners and managers need concrete answers. Which tasks may run automatically? Where does the budget stop? Who handles exceptions? What result justifies further spending? A limited trial using historical tickets can record full cost and quality before changing roles and suppliers. Multiplying the lowest token price by an ideal number of calls is likely to miss organizational cost.

For HR teams and employees, changing tasks matter more than changing titles. Brynjolfsson, Li, and Raymond studied 5,172 customer-support agents after the introduction of an AI assistant. They reported an average increase of roughly 15 percent in issues resolved per hour, with larger gains among less experienced and lower-skilled workers. That is evidence of assistance in one setting. It does not establish that an autonomous agent can replace an entire job, much less determine a layoff rate. See “Generative AI at Work,” 2025.

If a company automates all junior tasks while still relying on experienced people for exceptions, where will future experts gain experience? AI may also help novices learn through explanation and feedback. A break in apprenticeship is therefore a risk to test, not a settled result. Track whether new workers can handle unfamiliar problems independently, whether they receive real feedback, and whether senior capability remains renewable over several years.

For flexible workers, an agent may let one professional deliver a project that once needed a small team. It may also depress the price of standardized services. People who retain the customer relationship, domain judgment, and delivery responsibility may face a different market from those selling generic actions. Measure income after subscriptions, customer acquisition, rework, and waiting—not a single impressive delivery speed.

For model and platform vendors, products, distribution, and credible delivery stand between input cost and customer willingness to pay. This essay’s working hypothesis is that tokens will remain a low-level cost measure while some customer contracts move toward service levels and outcomes. Work with weak verifiability or exploratory requirements may remain better suited to usage or subscription pricing. Long-run coexistence of several models fits the mechanisms above better than total replacement by one.

More reproducible generic capability may reduce the price of some services, but training investment, compute, proprietary data, brand, and distribution still have costs. Low marginal call cost does not imply a market price of zero. New value may become consumer savings, employee compensation, corporate profit, or platform fees; competition and bargaining determine the proportions.

For society, task productivity, corporate profits, and public welfare must also be measured separately. David Autor’s historical discussion of automation emphasizes that technology replaces some tasks and complements others, while demand and job redesign shape employment. See Autor, 2015. If income increasingly comes from several projects and platforms, benefits and training may need to follow the person. Renaming workers “independent nodes” cannot resolve that institutional question.

What remains scarce when intelligence is commoditized?

From piecework to agents, people have repeatedly tried to describe work clearly enough to count the inputs and exchange the output. Those efforts expand coordination while exposing what lies beyond the price: context that cannot be specified in advance, quality that cannot be observed immediately, and promises that must still be honored when a dispute occurs.

As general-purpose capability becomes cheaper to call, a firm still has to decide what to do, why to trust the result, how to maintain customer relationships, and how to repair failure. Goal definition, trustworthy data, judgment, and accountable responsibility may become relatively scarce. Whether they command higher income still depends on whether they are truly difficult to replace and whether their holders retain bargaining power.

The essay ends at the purchase order: What does it promise to deliver? How will we evaluate it? Who completes the work left over? Writing those three answers down is the beginning of measuring the real efficiency gain—and of seeing which parts of human cooperation still merit long-term investment.

Sources and further reading

The sources below are ordered by their role in the essay. Historical texts help explain mechanisms; surveys and empirical papers retain their sample limits; vendor pages document public pricing and outcome rules at the review date. All seven diagrams are original conceptual models and do not estimate measured costs or distribution shares.

  1. Adam Smith, 1776, An Inquiry into the Nature and Causes of the Wealth of Nations, Book I, Chapter I. Public text. The classic treatment of the division of labor.
  2. Frederick Winslow Taylor, Shop Management, first published in 1903; the linked electronic text identifies the 1911 edition. Full text. Time study, piece rates, and labor–management incentives; cases are the author’s own reports.
  3. Ronald H. Coase, 1937, “The Nature of the Firm,” Economica 4(16): 386–405. DOI. Market transactions compared with coordination inside firms.
  4. Herbert A. Simon, 1951, “A Formal Theory of the Employment Relationship,” Econometrica 19(3): 293–305. Paper. Authority and the acceptable set of actions in employment.
  5. Walter W. Powell, 1990, “Neither Market Nor Hierarchy: Network Forms of Organization,” Research in Organizational Behavior 12: 295–336. Author-hosted text. Relationships, reciprocity, and network coordination.
  6. The Henry Ford, collection note on the 1914 five-dollar day. Archive. Wage, turnover, and qualification context.
  7. Jae Song, David J. Price, Fatih Guvenen, Nicholas Bloom, and Till von Wachter, 2019, “Firming Up Inequality,” Quarterly Journal of Economics 134(1): 1–50. DOI; institutional summary. U.S. firm and worker sorting, not a technology-industry compensation census.
  8. ILO, 2018, Digital Labour Platforms and the Future of Work: Towards Decent Work in the Online World. Report. A multinational survey across five microtask platforms.
  9. Uma Rani and Rishabh Kumar Dhir, 2024, “The Artificial Intelligence Illusion: How Invisible Workers Fuel the ‘Automated’ Economy.” ILO article. Human labor in the AI supply chain.
  10. Anthropic, “Pricing.” Official documentation. Input, output, caching, and related categories; checked September 19, 2026.
  11. Salesforce, “Agentforce Pricing.” Official page. Action credits, sessions, and user licenses; checked September 19, 2026.
  12. Intercom, “Fin AI Agent Outcomes.” Official rules. Outcome definitions, inferred resolution, and subsequent reversals; checked September 19, 2026.
  13. Bengt Holmström and Paul Milgrom, 1991, “Multitask Principal–Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design,” Journal of Law, Economics, & Organization 7, special issue: 24–52. Journal page. Multidimensional tasks and incentive distortion.
  14. Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond, 2025, “Generative AI at Work,” Quarterly Journal of Economics 140(2): 889–942. DOI; institutional summary. Productivity and heterogeneous effects of a customer-support assistant, using the journal version.
  15. David H. Autor, 2015, “Why Are There Still So Many Jobs? The History and Future of Workplace Automation,” Journal of Economic Perspectives 29(3): 3–30. Paper. Task replacement, complementarity, and labor-market adjustment.