How to measure returns on AI
怎样衡量人工智能投资回报
使用更多人工智能,是否就创造了更多价值?本文从投入指标转向结果与组织能力,借讽刺、研究实例和管理选择,说明速度、质量、调整成本和学习能力为何必须一起看。
原文来源:The Economist
学习目标
- 区分投入、结果与组织指标
- 解释节省时间如何转化为实际收益
- 拆解插入语、长主语与指代
- 用条件和比较基准讨论回报
中英对照阅读
IN THE BEGINNING there was AI, and bosses everywhere lost their minds and said you must all use this technology no matter whether it is useful and no matter how much it costs. And they introduced AI leaderboards and came up with stupid words like “tokenmaxxing”. And lo, usage did indeed rise. And then the bosses remembered some very basic concepts like “budgets” and realised that this might not be such a great idea after all. And then a different question rang through the boardrooms, and this question was about returns, and it was harder to answer.
起初,有了人工智能,各地老板都失去了理智,说你们必须全都使用这项技术,不论它是否有用,也不论它花多少钱。 于是,他们推出人工智能使用排行榜,还想出了“tokenmaxxing”之类的蠢词。 看哪,使用量果真上升了。 然后,老板们想起了“预算”这类极其基本的概念,意识到这毕竟可能不是什么好主意。 接着,董事会会议室里响起了另一个问题,一个关于回报、也更难回答的问题。
This potted Genesis of AI adoption is not entirely fair. As long as organisations are trying to encourage usage, it can still make sense to look at things like employees’ token consumption. “I would expect software engineers to spend more money than a finance analyst,” says the boss of one software firm. “But anyone within finance, it shouldn’t be zero.”
这段人工智能应用的迷你《创世记》,并不完全公允。 只要组织还在努力鼓励使用,关注员工词元消耗之类的指标,仍可能说得通。 一家软件公司的老板说:“我会预期软件工程师花的钱比财务分析师多。” “但对财务部门的任何人来说,这个数字也不该是零。”
But usage (also known as input) measures don’t tell you whether any value is being created. And even when AI is producing benefits, points out Mika Ruokonen of LUT University in Finland, lots of them accrue quietly to individuals in the form of scattered time savings. If an extra hour or two of an employee’s time has been freed up to shop on Vinted, someone has benefited (mainly Vinted) but it isn’t your company (unless you work at Vinted).
可是,使用量指标——也称投入指标——并不能告诉你是否创造了价值。 芬兰拉彭兰塔-拉赫蒂工业大学的米卡·鲁奥科宁指出,即使人工智能正在产生好处,其中许多也会以零散节省时间的形式,悄悄落到个人手里。 如果员工因此多出一两个小时去Vinted购物,确实有人受益,主要是Vinted,但不是你的公司,除非你就在Vinted工作。
Shifting the focus to returns on AI investment means concentrating more on outcomes-based measures, such as team productivity or customer satisfaction. But these require careful thought. Scientists, for example, are using AI to produce more papers (yay!) of dubious merit (boo!). A recent study by Tulsi Suchak of the University of Surrey and her co-authors found that papers derived from a health and nutrition data set ticked along at an average of four a year between 2014 and 2021; in the first nine months of 2024, there were 190.
把关注点转向人工智能投资回报,就意味着更多关注以结果为基础的指标,比如团队生产率或客户满意度。 但这些指标需要仔细斟酌。 例如,科学家正在使用人工智能产出更多论文,真棒!可论文的价值却令人怀疑,糟糕! 萨里大学的图尔西·苏查克及其合著者最近的一项研究发现,基于某个健康与营养数据集的论文,在2014年至2021年间,每年平均发表四篇,数量一直不多;而2024年的前九个月,就有190篇。
So outcomes-based measures have to be designed to capture both positive and negative effects of AI. At Google, for example, engineers have long been measured on three dimensions. Speed is one: how long does a code review take, say. Ease is another: how much friction is there in the system for things like onboarding new developers. The last, critically, is quality: “I don’t want to just go fast and make it easy to ship terrible software,” says Richard Seroter of Google Cloud.
因此,以结果为基础的指标必须被设计成能同时捕捉人工智能的积极与消极影响。 例如,谷歌长期以来从三个维度衡量工程师的工作。 速度是一项:比如,一次代码审查要花多久。 便利程度是另一项:在新开发者入职上手等事情上,系统存在多少阻力。 最后一项尤其关键,是质量;谷歌云的理查德·塞罗特说:“我可不想只是图快,又让糟糕的软件更容易发布。”
Complicating matters further is the fact that it often takes a while for the benefits of AI to show up. People have to spend time learning how to use the technology. Increased output can create bottlenecks elsewhere. A paper published last year by Kristina McElheran of the University of Toronto and her co-authors looked at AI adoption among American manufacturing firms, and confirmed the existence of a “J-curve” in which productivity dips before it leads to an improvement. This effect is especially pronounced at established firms and seems to be partly explained by changes to management practices that had previously helped firms to keep tabs on performance. Any kind of calculation about projected returns has to take this initial period of disruption into account.
更让事情复杂的是,人工智能带来的收益常常要过一段时间才会显现。 人们必须花时间学习怎样使用这项技术。 产出增加,还可能在其他环节造成瓶颈。 多伦多大学的克里斯蒂娜·麦克埃尔赫兰及其合著者去年发表的一篇论文,研究了美国制造企业采用人工智能的情况,证实存在一条“J曲线”:生产率先下滑,随后人工智能的采用才带来改善。 这种影响在已有成熟业务的企业中特别明显,部分原因似乎是管理惯例发生了变化,而这些惯例过去曾帮助企业掌握经营表现。 任何关于预期回报的计算,都必须考虑最初这段受扰动的时期。
Bosses also have to take a view on something much more fundamental: what kind of return are they interested in? One way to show financial returns on AI is to seize on time savings as an excuse to cut headcount. For some jobs and in some circumstances, that might make sense. But chainsaws and employee morale go together about as well as salt and slugs. And people who are no longer there cannot be redeployed in more productive ways. A better option is to calculate the amount of money that AI saves on expanding headcount. (Working out how much extra revenue can be attributed to the technology is even more of an art; at the very least, you need good baseline data and a culture of A/B testing to isolate its effects.)
老板们还必须想清楚一个更根本的问题:他们究竟关心哪一种回报? 展示人工智能财务回报的一种办法,是抓住节省时间这一点,作为裁员的理由。 对某些岗位、在某些情形下,这样做或许说得通。 但电锯与员工士气的相容程度,就如同盐与蛞蝓一样。 而已经不在公司的人,也无法被重新安排去从事更有生产力的工作。 更好的办法,是计算人工智能在扩充人员方面省下多少钱。 至于算清多少额外收入可以归因于这项技术,就更加是一门艺术了;至少,你需要良好的基线数据,以及开展A/B测试的习惯,来识别它的影响。
On top of everything, there is wild uncertainty about where AI is heading. That is an argument for a third set of measures, to go alongside ones on inputs and outcomes. These measures, which Mr Seroter calls “organisation-based”, are more geared towards capturing how well change is being managed. They could include data on levels of employee satisfaction with the way AI has been implemented, or progress towards building in-house expertise in areas that AI will not automate. A one-eyed focus on usage ignores the importance of returns. A rigid focus on returns ignores the need to keep learning.
除此之外,人工智能将走向何方,还存在极大的不确定性。 这就为第三组指标提供了理由,使其与投入和结果指标并列。 塞罗特把这组指标称为“基于组织的指标”,它们更着眼于衡量变革管理得怎么样。 这可以包括员工对人工智能实施方式的满意程度数据,也可以包括在人工智能不会自动化的领域建立内部专长的进展。 只盯着使用量,会忽略回报的重要性。 死盯着回报,又会忽略持续学习的必要性。