向前沿实验室出售数据的淘金热 ✍ Vikram Aditya🕐 2026-07-17📦 24.9 KB 🟢 已读 𝕏 文章列表 文章探讨了向前沿AI实验室出售训练数据这一暴利行业的现状。虽然部分公司单季营收可达数千万美元,但该业务缺乏持续性,本质是出售专家判断力、模拟环境等。由于模型能力迭代极快,数据供应商面临订单随时消失的风险,是一场巨大的淘金游戏。 训练数据AI实验室商业模式数据标注模拟环境AnthropicOpenAILLM数据交易 # The goldmine business of selling data to frontier labs **作者**: Vikram Aditya **日期**: 2026-07-16T07:02:55.000Z **来源**: [https://x.com/viks_rum/status/2077650169265590727](https://x.com/viks_rum/status/2077650169265590727) ---  I’ve spoken to 3 founders of different companies playing this game over last 10 days. Their companies sell training data to frontier AI labs, and all of them talk the way people talk when the ground is moving under them. It goes something like this. > We started in April. First quarter we closed $30M in orders. There are open purchase orders sitting on my desk for . By December we should land somewhere north of . None of it is recurring but all of it is growing. This month might end up in $20M for us. We’re less than 12 people, and maybe some interns. Every conversation I have in this market sounds like that now. For a while I kept thinking this is a rocket ship, why are more people not talking about it? Then it dawned on me that the founders themselves are asking a better question. They know the cash is real. They know the contracts are not forever. What should you do in such a situation? ## What is actually being sold The category hides 6 different products under one label. Read the invoices and the market splits cleanly. Some companies sell hours: humans labeling images and rating chatbot answers, the assembly-line era product, already dying. Some sell judgment: doctors, lawyers, and physicists writing down how they reason, at $100 to $500 an hour, because the models exhausted what amateurs could teach them. Some sell worlds: simulated Salesforce instances, fake banks, replica hospitals where agents practice a job across millions of repetitions. The unit here is expert judgment wrapped into a task, a world to act in, a rubric that defines good, and a verifier that scores it. Some sell verdicts: benchmarks, evaluations, red teams, the referees of the race. Some sell bodies: sensor rigs, tactile gloves and camera harnesses on real workers, because robots need to watch hands. And some sell rights: licensed archives, the Reddit-style deals worth tens of millions a year, institutions converting decades of accumulated text into an annuity.  Now look at how the money actually arrives, because the revenue mechanics matter more than the revenue. Almost everything is a purchase order against a deliverable: a dataset accepted, a batch of tasks passed QA, an environment shipped. Nothing renews by default. The headline numbers you read are annualized, usually the best month multiplied by 12, in a business where a lab can double or zero its orders inside a quarter. And everyone inside knows gross is not net. Marketplaces pass 60-70% of billings through to the experts doing the work. The exception is structural and companies that run their delivery from lower-cost geographies keep 70-80%+ of every dollar, which is why some of the most profitable names in this market are ones the valuation lists barely track. Same invoice, wildly different business underneath. The invoice anyway does not care where the judgment was manufactured, at least for now. The P&L of the vendor definitely does. ## The accidental giants Almost nobody at the top of this market set out to build it. Mercor started as a marketplace matching freelance engineers to companies, with an AI interviewer doing the vetting. Micro1 started the same way, an AI recruiter named Zara. Turing spent years as a remote-developer marketplace. Handshake spent a decade as a college recruiting network and pivoted after noticing that labs were poaching PhD annotators out of its own member base. It stopped renting out its network and started selling the work itself, and went from 0 to roughly $1B in gross annualized revenue in about 16 months. Even Scale began life as an API for Mechanical Turk before finding self-driving cars. The pattern tells you what the product really is. These companies did not win because they understood data. They won because they had already built machines for verifying strangers at scale such as who is actually a doctor, which engineer can actually code, whose judgment can be trusted without meeting them. When labs suddenly needed vetted experts by the thousand, the recruiting companies were the only ones holding supply. The data was never the product. Verified judgment was and the incumbents of verified judgment were job platforms.  ## Why the labs keep paying The reason labs sign 9-figure purchase orders is a war they cannot exit. No lab holds a durable capability lead anymore. Nobody keeps the crown for a full season, open models trail the frontier by months, and every price tier keeps collapsing. It takes all the running they can do to stay in the same place, and on a treadmill the only thing everyone keeps buying is shoes. Data vendors sell the shoes. Their revenue does not require picking a winner. It is a tax on nobody winning. Alex Karp has spent this month accusing Silicon Valley of overselling AI, telling the public not to believe its lying eyes. Look at this market and a funny thing happens: the purchase orders agree with him. If models were nearly finished, labs would not be paying this much for human judgment. Every invoice in this industry is a confession about what the models still cannot do. But the same treadmill keeps executing its own suppliers. In 2023 the product was crowd workers rating responses. Once models outgrew the raters, the ratings became noise, and 2024 belonged to credentialed experts. Then reasoning models learned to grade themselves against checkable answers, and 2025 moved the money to environments and rubrics. Each generation of models graduates past the data that trained it. The rungs below the frontier keep dissolving. The frontier keeps paying. I spoke to a friend at a frontier lab this weekend and asked how many data vendors he works with directly. Seven, he said. All seven are tasked with producing the same type of datasets. It goes without saying that a year from now, some of them will watch that PO vanish. That is the whole market in one anecdote: enormous demand, deliberately duplicated supply, and a buyer who owns the clock. ## The clock inside every contract Environments deserve one economics lesson because they are the strangest product in the market. Researchers at Epoch AI interviewed vendors and published the price sheet - a simple website replica for agent training runs about $20k, and one lab reportedly bought hundreds of them, once, the way you buy cones for a driving school. A high-fidelity clone of an enterprise tool with expert-written tasks runs . Individual tasks price between $200 and $2k, and exclusivity multiplies everything 4-5x, because a task your rival also trains on teaches you nothing about beating them. But here is the twist, once models pass a task about 70% of the time, the task is discarded. The product depreciates by succeeding. You are selling homework to a student who graduates past it every quarter. That guarantees repeat orders, which is why the revenue curves look vertical, and it also guarantees that nothing annuitizes on its own. Everything must be rebuilt harder, forever. I’ve a sense that the founders in this space are bullish about the data business for next 3-4 yrs at least and maybe they should but the buyers here, the frontier labs are chosing to work both sides of the counter. Anthropic reportedly discussed spending over $1B on environments in a year while working with a dozen-plus vendors and making all of them conform to its frameworks, commoditization by procurement. OpenAI has reportedly trademarked an internal data platform aimed at reducing reliance on the very vendors it enriches, and has asked contractors to upload artifacts of real past work, the politest way of saying we would like the source, not the reseller. xAI cut a third of its in-house annotation team to grow specialist tutors instead. Karpathy, bullish on environments as a concept, is publicly bearish on the training technique the whole category monetizes. This has happened before, inside this same industry. Between 2016 and 2021 a generation of data companies fed on self-driving programs, then the surviving carmakers pulled labeling in-house and the purest suppliers were absorbed or shut. Scale lived because it jumped to the LLM wave in time. Consider Appen. An Australian company, once a $4 billion listed darling supplying human data to big tech, with, at its peak, 80% of revenue from five clients. In January 2024 Google canceled its contract without warning. The stock is down more than 95% from that peak. One customer email, one technique shift, and the incumbent of the entire category became a case study. Pharma went the other way, never took drug trials back in-house, and 40 years later the outsourced trials industry still compounds. Both endings are possible here. Which one you get is decided by one law. But what is the law? Whatever a machine can verify, machines will eventually learn without you. Whatever still needs a human to say this is good keeps paying humans. Code and math fell first because correctness is checkable, and labs now mine their own training tasks from public repositories by the tens of thousands. Taste, ambiguity, regulated judgment, and the physical world fall last, maybe never. There is no unit test for what a senior surgeon sees, and you cannot unit-test a folded shirt. Verification is the scarcity. Sell against it and the clock works for you instead of against you.  ## What the cash should buy None of this means the data wave is fake. Money is real, growth is real, and the physics of the treadmill guarantees demand for harder homework for years. It means the wave rewards a very specific shape of company, and punishes the clones, in a niche where a company fully bootstrapped can ship product and an offshore delivery team can undercut any price you quote. When a market’s number one customer is building your replacement while paying your invoices, your product is not the moat. Your position is. So here is the actual question, the one the founders printing this money ask over dinner. Nobody running a business that generates $100M-500M of PO cash at these margins is going to stop. Nor should they. Take every order. Run the machine flat out. The only mistake available at this stage is treating the windfall as the business instead of as the financing for the business. PO income is a great fuel but what follows is the menu of what it can buy, and an honest read on each option. Go deeper into data, not wider. The lazy move is horizontal which is more domains, more generalist supply, competing with 4 giants who own the trust. The compounding move is vertical such as pick one domain where verification stays hard, hire the 200 best experts in it as your own, and become the only counterparty labs call for it. One young company owns audio. One owns chip design. One owns advanced mathematics. New rungs will keep appearing as models advance, and labs generating their own data does not end this demand, it moves it up the difficulty curve toward whoever owns the top of a domain. Works when you truly own scarce experts. Fails when your experts are interchangeable with a rival’s spreadsheet. Go physical, and own the whole loop. The mistake in physical data is thinking the gloves are the business. Hardware capture is the cheap part. The companies that will matter run the collection operation end to end - they employ the workers, build the rigs, hire in-house industry experts who know what a correct weld or suture or lockout procedure looks like, encode how an industry actually operates, and sell the annotated exhaust with exclusivity terms. The emptiest squares on the map that I’ve managed to create in my head are industrial rigs, refineries, factory floors, mines, places where no dataset exists at any price while everyone crowds into retail and finance and health-care related demos. However, this works when you control capture, quality, and rights and fails when you are a middleman for other people’s cameras. Keep building environments, but sell them up the stack. The $20k website replica tier is already commoditizing into open-source hubs. The durable tier is high-fidelity, expert-graded, exclusive, and it points at 2 buyers, not 1. Labs today. Enterprises tomorrow, and that second buyer changes everything. Satya Nadella has been telling every firm that they pay for intelligence twice, once in money and once in the proprietary judgment that leaks out through every prompt, so they must build their own evals and their own learning environments inside their own walls. Read that as a product spec. The exact skill you built for lab work, turning a messy workflow into a world with rubrics and verifiers, becomes private training gyms behind a customer’s firewall such as their claims process, their trading desk, their hospital, simulated so their agents can learn without their judgment ever leaving the building. It multiplies your buyer count from 5 to 5,000. Works because it rides the same muscle. Fails only if you wait until the lab POs slow before building it. Enter enterprise workflows with open eyes. Deploying agents inside companies is forward deployment work, mapping how invoices really move, discovering the SOP is fiction, sitting with the team until exceptions stop (I recently wrote a full piece on this). It is a real destination, and some data companies will build real businesses there. But know the physics before committing the cash to this. Data revenue arrives as 25M POs signed in weeks, enterprise revenue arrives as $500k-$2M pilots signed in quarters, and roughly 95% of enterprise AI pilots today show no measurable return. The move works as a separately run unit with separate expectations and its own leadership. It fails as a side project staffed by whoever the data business can spare, because the muscle is different and needs patience, embedding, and glue code instead of throughput. Buy compute only if compute feeds your product. More than one founder in this market is asking whether the cash should become GPUs and a hosted RL platform. The honest answer is renting out raw compute is a commodity squeezed between hyperscalers and neoclouds, and a treasury full of depreciating silicon is not a moat. The version that works is narrower, hosting the training loops that run inside your own environments, where utilization is yours to guarantee and the customer is buying the world plus the gym plus the compute as one product. Prime Intellect already runs this play in the open. It gave away a hub of 2,500+ community environments and sells the compute and hosted training that run on top. The environments are the storefront. The GPUs are the checkout. That is a venture bet, not a cash-parking decision. If I was a founder doing this, I would make the decision deliberately or not at all. Acquire the next rung instead of building it late. The most instructive capital allocation in this market so far is that one giant used its PO windfall to buy 2 environment startups within 5 months, purchasing its way onto the new rung while competitors were still hiring for it. In about 18 months, the model will likely position companies with real environment engineers who will love to get acqui-hired. Speed is the only advantage here. You’re sitting on piles of cash - so a war chest plus a clear map of which rung comes next beats organic speed in a market that re-deals every 18 months. Sell to governments. There is a new customer class arriving. Governments buying sovereign AI programs will need national data pipelines, native-language corpora, local evals, and physical data from their own factories and fields, for the same reasons they buy their own grids. And convert what you can into revenue that renews. POs are like weather. Some of it can be turned into climate such as evaluation subscriptions instead of one-time benchmark sales, environment maintenance contracts instead of one-time builds, data refresh retainers, certification programs that bill annually. None of it will look as spectacular as a $50M PO and that’s the tricky part to hedge yourself with less shiny pieces. Because, all of it survives the quarter when the PO doesn’t arrive. And I’ve failed as a founder to know this - there are two purchases the cash should never make. Entry into the giants’ generalist lane, where the trust premium cannot be replicated from zero. And a moderate raise at an immoderate multiple, which buys obligations priced like software on economics that are anything but, while closing the two exits that actually exist here like staying private and rich, or becoming infrastructure someone must own.  ## Trust is the asset that compounds Every option on that menu above runs through the same gate. Enterprises will not hand you their claims process, labs will not hand you frontier training priorities, and governments will not hand you national corpora unless trust has been built deliberately, and trust in this market is not a vibe, it is a stack of verifiable commitments. The companies that get this build it like a product. First comes convincing the company you’re taking the data from that you are not taking anything sensitive out and that they are not making any mistake as far as the law goes that will put them in trouble. Security and residency certifications before the customer asks are the norm. Public benchmarks are another form of trust machine here. On the other end, the labs buying this data also want provenance rails such as camera-verified sessions, credential attestation, proof that a specific human did the thinking, because the supply chain’s dirty secret is annotators pasting model output back as human work. It helps to have neutrality covenants for example no lab on the cap table, no single buyer above a set share of revenue, learned the hard way by everyone who watched a rival’s customers flee the day a lab bought half of it - though for the Scale AI team maybe it was a brilliant outcome. Expert certification programs help if you can build a brand, so that “rated by your network” starts to mean something an industry recognizes. Every one of these is an asset that compounds while task formats die or change. When the format changes, and it will, roughly every 2 years, the trust is what transfers to the next product. ## The 50th company Scale and Mercor got there first and got there huge, so what should the 50th company do? Start with what Mercor’s rise actually teaches, because everyone copies the wrong part. The visible part is speed. Scale took about 4 years to reach its first . The next cohort took 2. Mercor took under 20 months, Micro1 and AfterQuery closer to a year, and one environments startup went from $1M to $63M in 6 months. Founders read this as the market getting kinder. It is the opposite. Each rung is steeper and shorter, and the same acceleration that pulls a newcomer to $100M in a year pulls the rung out from under them just as fast. Speed is a property of the wave, not the boat - think about this and you will have second thoughts about riding in that boat because this game is not for everyone. The part worth copying is quieter. Mercor built its verification engine before the demand existed, for a different business entirely, so when the wave arrived it onboarded trustworthy experts faster than anyone. It never needed to embed engineers inside customers or run services teams, the marketplace stayed the machine, and when the next rung appeared it bought its way there instead of building from behind. And the bootstrapped leader in this market teaches the inverse lesson with the same moral by staying profitable and never selling equity, it kept the option everyone else sold, the option to say no, to any customer, any deal structure, any quarter. In a market where your customers are your future competitors, optionality is not a luxury. It is what your margins are buying. So the 50th company enters where the ladder is still being built such as one hard domain owned completely, rubrics and verifiers and environments sold instead of hours, benchmarks published from day one, the second buyer class built before it is needed, capital story decided on day one, bootstrap and keep the option or raise big and buy rungs, never the middle. And if you are not founding one but deciding whether to join one, ask the same questions from inside that which of the 6 products does this company actually sell, whose trust does it hold, what clock is its current format on, where is the PO cash going, and who is the second customer after the labs. A company with good answers to those is worth joining because because on rocket-ships often teaches you things on a compressed timeline. ## 5 years out What’s the point of writing all this if I’m not incredibly right or horribly wrong about some of these things. So here is my 5-yr view. The gross market grows for years. The demand mechanism does not pause while the lab race stays unresolved, and there is one scheduled stress test on the calendar - the first lab IPOs (which is very close today as of July 2026), when data spend becomes a line item public analysts question every quarter. My hunch is the composition rotates violently underneath the growth. Hours die first, and are mostly dead already. Generic environments commoditize into open hubs. Value concentrates in frontier judgment, verification and provenance, referees, physical capture, and private gyms for enterprises, and if I had to rank those, verification and enterprise gyms first because both get stronger as the labs get stronger, physical second because it is the only segment where supply rather than demand is the bottleneck. Of the 100+ companies selling into labs today (I created a list and gave up midway when I realized that the attempt was futile because it gets outdated the next minute), I expect fewer than 10 still independent and at scale in 2031. Most would cease operations with some rich founders. The rest get absorbed, by the giants buying rungs, or by the labs themselves, quietly, for the people. The winners are legible if you watch what they are already doing. The bootstrapped quality leader becomes the standard setter, the name whose acceptance is itself a certification. The acquisitive giant becomes an exchange where expert work is priced, verified and sold, whoever the buyer is, and if labs are ever displaced as customers, employers stand next in line. The environment builders that survive wake up as the enterprise-simulation industry. The referees, if they stay unowned, end the decade looking like rating agencies, written into procurement rules and maybe into law. And somewhere in the physical world, a company collecting sensor-fused industrial data is compounding toward being the Scale of the embodied era, 5 years earlier on that curve than everyone crowding the digital one. One founder in this market has argued that human data becomes a trillion dollar a year affair, and he gets the deepest thing right, models learn from humans at every stage, forever. What the trillion misses is that it prices human time, not the intermediary. The intermediary’s cut is decided by whether it owns something scarcer than a spreadsheet of contractors. The good news for everyone building here is that the scarce things are now known, and every one of them is buildable with exactly the cash this market is throwing off - owned expert networks, provenance rails, referee franchises, closed loops in the physical world. Every gold rush ends one of two ways, the gold runs out or the miners industrialize. This one ends a third way. The gold learns to mine itself. When it does, the suppliers left standing will be the ones who sold the mine the one thing it can never dig up, the answer to the question every model asks and none can settle: what does good look like? Hold that answer in one narrow domain and you have a company. Hold it credibly enough, for long enough, and you stop being a vendor in someone else’s race. You become part of how the race is scored. ## 相关链接 - [Vikram Aditya](https://x.com/viks_rum) - [@viks_rum](https://x.com/viks_rum) - [384K](https://x.com/viks_rum/status/2077650169265590727/analytics) - [$100M-500M](https://x.com/search?q=%24100M-500M&src=cashtag_click) - [full piece](https://open.substack.com/pub/thoroughlyintrigued/p/forward-deployment/) - [$500k-](https://x.com/search?q=%24500k-&src=cashtag_click) - [Upgrade to Premium](https://x.com/i/premium_sign_up) - [3:02 PM · Jul 16, 2026](https://x.com/viks_rum/status/2077650169265590727) - [384K Views](https://x.com/viks_rum/status/2077650169265590727/analytics) - [View quotes](https://x.com/viks_rum/status/2077650169265590727/quotes) --- *导出时间: 2026/7/17 11:26:24* --- ## 中文翻译 # 向前沿实验室出售数据的金矿生意 **作者**: Vikram Aditya **日期**: 2026-07-16T07:02:55.000Z **来源**: [https://x.com/viks_rum/status/2077650169265590727](https://x.com/viks_rum/status/2077650169265590727) ---  在过去 10 天里,我与三家不同公司的创始人进行了交谈,他们都在玩这个游戏。他们的公司向前沿 AI 实验室出售训练数据,他们所有人说话的语气,都像是脚下的大地正在移动的人一样。情况大概是这样的。 > 我们是在 4 月开始的。第一季度我们完成了 3000 万美元的订单。我的桌子上放着总额达 [金额] 的未完成采购订单。到 12 月,我们的业绩应该会落在 [金额] 之上。这些收入都不是经常性的,但都在增长。这个月我们可能会做到 2000 万美元。我们不到 12 个人,可能还有一些实习生。 现在,我在这个市场的每一次谈话听起来都是这样。有一阵子我一直想,这是一艘火箭,为什么没有更多人谈论它?然后我突然意识到,创始人自己在问一个更好的问题。他们知道现金是真实的。他们也知道合同不会永远持续下去。在这种情况下你应该做什么? ## 实际出售的是什么 这个类别在一个标签下隐藏了 6 种不同的产品。看看发票,市场就会清晰地划分开来。 一些公司出售时间:人工标注图像和评价聊天机器人的回答,这是流水线时代的产品,已经在消亡。一些公司出售判断:医生、律师和物理学家写下他们的推理过程,每小时 100 到 500 美元,因为模型已经耗尽了业余爱好者能教给它们的东西。一些公司出售世界:模拟的 Salesforce 实例、假银行、复制医院,代理在其中数百万次地重复练习一项工作。这里的单位是包裹在任务中的专家判断、一个行动的世界、定义“好”的标准,以及一个评分的验证器。一些公司出售裁决:基准、评估、红队测试,这场比赛的裁判。一些公司出售躯体:传感器装备、触觉手套和真实工人身上的摄像头背带,因为机器人需要观察手部。还有一些公司出售权利:授权存档,类似于 Reddit 那样每年价值数千万美元的交易,机构将几十年积累的文本转化为年金。  现在看看钱实际上是怎么来的,因为收入机制比收入本身更重要。几乎所有的东西都是针对可交付成果的采购订单:被接受的数据集、通过质量检查的一批任务、交付的环境。没有任何东西会默认续约。你读到的头条数字是年化数字,通常是最好月份的 12 倍,而在这种业务中,一个实验室可以在一个季度内将其订单翻倍或归零。里面的每个人都知道毛利不是净利。市场将 60-70% 的账单转交给做这项工作的专家。例外是结构性的,那些从低成本地区运营交付的公司每一美元能留下 70-80% 以上,这就是为什么这个市场一些最赚钱的名字是那些估值榜单几乎追踪不到的公司。同样的发票,底下的业务却截然不同。反正发票也不在乎判断是在哪里制造的,至少目前是这样。供应商的损益表(P&L)肯定在乎。 ## 意外的巨头 这个市场顶端的几乎没有人一开始就是为了建立它而出发的。 Mercor 最初是一个将自由职业工程师与公司匹配的市场,由 AI 面试官进行审查。Micro1 也是这样开始的,一个名为 Zara 的 AI 招聘人员。Turing 花了数年时间作为一个远程开发者市场。Handshake 作为大学招聘网络花费了十年时间,在注意到实验室正在从其成员库中挖走博士标注员后进行了转型。它停止出租网络,开始直接出售工作,并在大约 16 个月内从 0 增长到大约 10 亿美元的年化总营收。甚至 Scale 在找到自动驾驶汽车之前,也是作为一个 Mechanical Turk 的 API 起家的。 这种模式告诉了你产品到底是什么。这些公司获胜不是因为他们理解数据。他们获胜是因为他们已经构建了大规模验证陌生人的机器,比如谁实际上是医生,哪个工程师真的会写代码,谁的判断可以在不见面的情况下被信任。当实验室突然需要数以千计的经过审查的专家时,招聘公司是唯一持有供应方的人。数据从来不是产品。经过验证的判断才是,而经过验证的判断的现有者正是招聘平台。  ## 为什么实验室继续付费 实验室签署九位数采购订单的原因是一场他们无法退出的战争。没有任何实验室再拥有持久的能力领先优势。没有人能保住王冠整整一个赛季,开源模型落后于前沿数月,而且每个价格层级都在不断崩塌。为了留在原地,他们必须竭尽全力奔跑,而在跑步机上,每个人唯一不断购买的就是鞋子。数据供应商卖的就是鞋子。他们的收入不需要押注赢家。这是对“没有赢家”的一种征税。 Alex Karp 这个月一直在指责硅谷过分吹嘘 AI,告诉公众不要相信他们撒谎的眼睛。看看这个市场,一件有趣的事情发生了:采购订单同意他的观点。如果模型几乎完成了,实验室就不会在人类判断上花费这么多钱。这个行业里的每一张发票都是对模型仍然无法做的事情的忏悔。 但是同一台跑步机也在不断处决自己的供应商。2023 年的产品是众包工人评价回答。一旦模型超越了评价者,评价就变成了噪音,2024 年属于持有证书的专家。然后推理模型学会了根据可验证的答案给自己打分,2025 年资金转移到了环境和标准。每一代模型都会超越训练它的数据。前沿之下的梯级不断溶解。前沿继续付费。 这个周末我和前沿实验室的一位朋友谈过,问他直接与多少家数据供应商合作。七个,他说。所有七个都被指派生产相同类型的数据集。不用说,一年后,他们中的一些人将眼睁睁看着那份采购订单消失。这就是整个市场的一个缩影:巨大的需求,故意重复的供应,以及一个拥有时钟的买家。 ## 每份合同里的时钟 环境值得上一堂经济学课,因为它们是市场上最奇怪的产品。Epoch AI 的研究人员采访了供应商并发布了价格表——一个用于代理训练的简单网站副本大约 2 万美元,据报道有一个实验室一次性购买了数百个,就像你给驾校买圆锥筒一样。一个带有专家编写任务的高保真企业工具克隆版运行 [价格]。单个任务的价格在 200 美元到 2000 美元之间,独家性会将所有价格乘以 4-5 倍,因为你的对手也在训练的任务无法教会你如何击败他们。 但这里有个转折,一旦模型在 70% 的时间里通过了任务,该任务就会被丢弃。产品通过成功而贬值。你在给一个每学期都会超越作业的学生卖作业。这保证了重复订单,这也是为什么收入曲线看起来是垂直的,但这也就保证了没有任何东西能自动年金化。一切必须永远被重建得更难。 我有一种感觉,这个领域的创始人们对接下来 3-4 年的数据业务持乐观态度,也许他们应该如此,但这里的买家,即前沿实验室,选择在柜台的两边都进行操作。据报道,Anthropic 讨论了一年内在环境上花费超过 10 亿美元,同时与十多家供应商合作,并让所有供应商都遵守其框架,这是通过采购进行商品化。据报道,OpenAI 已经为一个内部数据平台申请了商标,旨在减少对其所资助的供应商的依赖,并要求承包商上传真实过去工作的工件,这是最礼貌的说法:我们想要源数据,而不是转售商。xAI 裁减了三分之一的内部标注团队,转而培养专业导师。虽然 Karpathy 对环境这一概念持乐观态度,但他公开对整个类别赖以变现的训练技术持悲观态度。 这以前发生过,就在同一个行业里。2016 年到 2021 年间,一代数据公司靠自动驾驶项目为生,然后幸存的汽车制造商将标注业务收回内部,最纯粹的供应商被收购或关闭。Scale 活下来是因为它及时跳上了 LLM 的浪潮。想想 Appen。一家澳大利亚公司,曾经是 40 亿美元的上市宠儿,向大型科技巨头提供人类数据,在巅峰时期,80% 的收入来自五个客户。2024 年 1 月,谷歌在没有警告的情况下取消了合同。股价较峰值下跌了 95% 以上。一封客户电子邮件,一次技术转变,整个类别的现任者就成了一个案例研究。制药业走了另一条路,从未将药物试验收回内部,40 年后,外包试验行业仍在复利增长。这两种结局在这里都是可能的。你会得到哪一种,由一条法则决定。 但是法则是什么?凡是机器可以验证的东西,机器最终都会学会不需要你。凡是仍然需要人类说“这是好的”东西,人类就会继续获得报酬。代码和数学最先倒下,因为正确性是可验证的,实验室现在从公共存储库中挖掘数以万计的自己的训练任务。品味、歧义、受监管的判断和物理世界最后倒下,也许永远不会。对于高级外科医生看到的东西,没有单元测试,你也无法对折叠好的衬衫进行单元测试。验证是稀缺的。针对它进行销售,时钟就会为你工作,而不是对你不利。  ## 现金应该买什么 这并不意味着数据浪潮是假的。钱是真的,增长是真的,跑步机的物理学保证了数年对更难作业的需求。这意味着浪潮奖励一种非常特定形态的公司,并惩罚克隆者,在这个细分市场中,一家完全自力更生的公司可以发布产品,一个离岸交付团队可以以低于你报出的任何价格进行竞争。当一个市场的第一大客户在支付你发票的同时正在建造你的替代品时,你的产品不是护城河。你的地位才是。 所以这是真正的问题,那些印钞票的创始人们在晚餐上问的问题。没有人会停止经营一家以这些利润率产生 1 亿到 5 亿美元采购现金的业务。也不应该停止。接受每一个订单。让机器全速运转。在这个阶段唯一可用的错误是把横财当作生意,而不是当作生意的融资。采购订单收入是很好的燃料,但随之而来的是它能购买的菜单,以及对每个选项的诚实评估。 在数据上做得更深,而不是更广。懒惰的做法是横向扩展,即更多领域,更通用的供应,与拥有信任的 4 个巨头竞争。复利的做法是纵向深耕,比如选择一个验证仍然困难的领域,雇佣其中 200 名最好的专家作为你自己的员工,成为实验室在该领域唯一致电的对手。一家年轻的公司拥有音频。一家拥有芯片设计。一家拥有高等数学。随着模型的发展,新的梯级将不断出现,实验室生成自己的数据并不会结束这种需求,它将需求沿难度曲线向上移动,无论谁拥有该领域的顶端。当你真正拥有稀缺专家时有效。当你的专家与竞争对手的电子表格可以互换时失败。 走向物理,并拥有整个闭环。物理数据的错误在于认为手套就是生意。硬件捕获是廉价的部分。重要的公司端到端地运营收集操作——他们雇佣工人,建造装备,雇佣知道什么是正确的焊接、缝合或锁定程序的内部行业专家,编码行业实际运作的方式,并以独家条款出售带有注释的废料。我在脑海中设法绘制的地图上最空的方块是工业装备、炼油厂、工厂车间、矿山,这些地方根本不存在任何价格的数据集,而每个人都挤进零售、金融和医疗保健相关的演示中。然而,当你控制捕获、质量和权利时,这行得通;当你作为其他人摄像头的中间人时,这会失败。 继续构建环境,但要向上层销售。2 万美元的网站副本层级已经正在商品化进入开源中心。持久的层级是高保真、专家分级、独家的,它指向 2 个买家,而不是 1 个。今天的实验室。明天的企业,第二个买家改变了一切。Satya Nadella 告诉每家公司,他们为智能付费两次,一次是用钱,一次是通过每个提示泄露出来的专有判断,所以他们必须在自己的墙内建立自己的评估和自己的学习环境。把这读作产品规格。你为实验室工作建立的完全相同的技能,将混乱的工作流程变成一个有标准和验证器的世界,变成了客户防火墙背后的私人训练健身房,比如他们的索赔流程、他们的交易台、他们的医院,进行模拟,以便他们的代理可以在学习时让他们的判断永远不会离开大楼。这会将你的买家数量从 5 个乘以到 5000 个。有效,因为它使用相同的肌肉。失败仅当你在实验室采购订单减速之前才构建它。 带着开放的眼光进入企业工作流。在公司内部部署代理是前沿部署工作,绘制发票实际如何移动,发现标准作业程序(SOP)是虚构的,与团队坐在一起直到异常停止(我最近写了一篇完整的文章谈论这个)。这是一个真正的目的地,一些数据公司将在那里建立真正的业务。但在投入现金之前要了解物理学。数据收入以几周内签署的 2500 万美元采购订单的形式到来,企业收入以季度内签署的 50 万到 200 万美元试点的形式到来,而且大约 95% 的企业 AI 试点今天没有显示出可衡量的回报。这一举措作为一个单独运营的单元、拥有单独的期望和自己的领导层是有效的。作为一个由数据业务能抽调的人员配备的副项目是失败的,因为肌肉是不同的,需要耐心、嵌入和胶水代码,而不是吞吐量。 只有当计算能滋养你的产品时才购买计算。这个市场不止一位创始人在问现金是否应该变成 GPU 和托管 RL 平台。诚实的答案是出租原始计算是一种夹在超大规模云提供商和新型云之间被挤压的商品,而且装满贬值硅片的金库不是护城河。有效的版本更狭窄,托管在你自己环境内运行的训练循环,利用率由你来保证,客户购买的是世界加上健身房加上计算作为一个产品。Prime Intellect 已经公开开展了这种玩法。它赠送了一个拥有 2500 多个社区环境中心,并销售运行在其上的计算和托管训练。环境是店面。GPU 是收银台。这是一个风险赌注,而不是现金停放决定。如果我是这样做的创始人,我会故意做出决定,或者根本不做。 收购下一个梯级,而不是晚点去构建它。这个市场迄今为止最具指导意义的资本配置是,一家巨头利用其采购订单横财在 5 个月内收购了 2 家环境初创公司,在竞争对手还在为此招聘时,就通过购买进入了新的梯级。大约 18 个月后,模型可能会让那些拥有真正环境工程师的公司变得抢手,他们会很乐意被收购 hire。速度是这里唯一的优势。你坐拥成堆的现金——所以在每 18 个月重新洗牌一次的市场中,战金库加上一张关于下一个梯级是什么的清晰地图,胜过有机速度。 卖给政府。一个新的客户阶层正在到来。购买主权 AI 项目的政府将需要国家数据管道、本土……
A AI商业模式的危机 文章深入分析了AI产业面临的商业模式危机。尽管AI产品需求旺盛、用户增长迅速,但巨额的资本开支与实际收入之间存在巨大鸿沟。模型公司夹在中间,面临重工业般的投入成本、大宗商品般的定价压力,却享受着软件公司的高估值,这种结构性错配导致其长期盈利能力存疑。 技术 › LLM ✍ snowboat🕐 2026-07-21 AI商业模式OpenAIAnthropic资本开支估值亏损模型公司重工业大宗商品
Z Zero to AI Engineer — The Roadmap Nobody Explains Properly 本文提供了一个为期14周的实战型AI工程师学习路线图,旨在解决初学者“只学不做”的困境。文章从环境搭建开始,详细列出了从AI基础、机器学习、深度学习到现代LLM工程及Agent开发的最佳免费资源(如OpenAI/Anthropic官方课程、Karpathy的教程等)。路线强调通过GitHub项目实践来理解原理,最终掌握部署与评估技能,真正从零开始构建可用的AI系统。 技术 › LLM ✍ Shruti Codes🕐 2026-05-17 AI工程师学习路线LLMAgent深度学习RAG实战教程机器学习OpenAIAnthropic
深 深度拆解:AI Agent Harness 的构造 文章深入探讨了“AI Agent Harness”的概念,即包裹在大语言模型之外、使其转变为智能体的完整软件架构。作者详细拆解了生产级 Harness 的 12 个核心组件(如编排循环、工具、记忆、上下文管理等),对比了 Anthropic、OpenAI 和 LangChain 的不同实现路径,并指出 Harness 工程是决定 AI 应用性能的关键。 技术 › Harness Engineering ✍ 宝玉🕐 2026-05-12 AI AgentLLMHarness架构设计OpenAIAnthropicLangChain工程化
A AI Agent 从零开发指南:构建你的第一个智能体 本文是一篇从零开始构建 AI Agent 的实战教程。作者整合了 Anthropic 和 OpenAI 等资源,详细介绍了 Agent 的工作原理(核心循环)、五种核心工作流模式(如提示链、路由、并行化等),并提供了从构思、设计工具与记忆机制到落地的完整步骤。文章还探讨了如何利用 LLM 自身辅助设计 Agent,帮助开发者快速构建实用型的自动化智能体。 技术 › Agent ✍ hoeem🕐 2026-04-28 AI AgentLLM开发教程AnthropicOpenAI提示工程工作流自动化LangChainClaude
T Token计算:下一个十年的成本战争 文章指出,随着AI技术的发展,单一的“每百万Token成本”已不再是衡量支出的唯一标准。OpenAI、Anthropic等厂商引入了Session Runtime、Cache、Web Search及Outcome等多元化计费维度。这意味着企业必须从单纯的模型比价,转向针对不同任务的综合成本考量。AI经济的价值正在分层,底层资源作为“公用事业”商品化,而封装了行业Know-how和结果交付的上层服务,则将成为价值沉淀的新高地。 技术 › LLM ✍ 华尔街财经 | WSInsights 【Zenzhe 】🕐 2026-04-28 Token经济成本分析AI商业化OpenAIAnthropic定价模型AgentLLM
深 深度拆解2026年最新风口:API中转站(Token进口)行业解析 文章深入解析了2026年AI领域兴起的“API中转站”(俗称Token进口)赛道。分析了该模式基于全球AI服务价格差与访问壁垒的套利本质,揭示了其“低价资源供给、用户数据变现、模型偷换降级”的三层盈利结构。文章同时指出了用户面临的数据泄露、欺诈及法律风险,预测随着国产模型崛起,此类灰色套利窗口将逐渐消失。 技术 › LLM ✍ Vincent Logic🕐 2026-04-21 API中转Token套利LLM行业分析AI代理数据安全OpenAI商业模式灰色产业
M Mitchell Hashimoto 与 AI 工程演进:从聊天机器人到 Harness Engineering 文章详细介绍了 Terraform 之父 Mitchell Hashimoto 的 AI 工程实践六个阶段,从早期的聊天窗口尝试,到最终实现后台常驻 Agent 的愿景。核心亮点在于第五阶段提出的“Harness Engineering”(工程约束),即通过系统性设计(如 AGENTS.md 和验证工具)来防止 Agent 重复犯错。文章结合 Anthropic 和 OpenAI 的行业案例,指出工程师的核心技能正从手写代码转向设计 Agent 的运行约束环境。 技术 › Harness Engineering ✍ Jason Zhu🕐 2026-04-07 AI工程Mitchell HashimotoAgentHarness Engineering开发流程自动化LLM王树义AnthropicOpenAI
我 我今天想构建一个人工智能代理(完整课程) 本文是一篇关于从零开始构建AI智能体的完整指南。作者基于Anthropic和OpenAI的最佳实践,整理了一套适合初学者的课程。文章详细介绍了代理的核心工作循环、五种核心工作流模式(如提示链、路由、并行化等),以及如何通过定义角色、目标、工具和规则来打造第一个代理。此外,还深入探讨了工具的高效使用、短期与长期记忆的赋予方法,并对比了使用Claude与OpenAI SDK构建代理的不同路径,强调了从简单工作流入手,逐步迭代至复杂代理的重要性。 技术 › Agent ✍ hoeem🕐 2026-03-28 AI AgentLLMOpenAIAnthropic教程工作流提示工程Claude智能体设计工具调用
A Anthropic CEO Dario Amodei 访谈:我们正在接近指数的终点 Anthropic CEO Dario Amodei 在 Dwarkesh Podcast 的访谈中,深入探讨了 AI 扩展定律(Scaling Laws)的延续、强化学习的规模化以及“天才之国”(Country of Geniuses)的预测时间表。对话涵盖了 AI 编程工具带来的实际生产力提升(当前 15-20%)、AI 公司的商业盈利模型(算力供需平衡),以及与 OpenAI 在 Stargate 项目上截然不同的风险策略。 技术 › LLM ✍ 宝玉🕐 2026-02-15 AnthropicDario AmodeiScaling LawsLLM访谈AI编程Stargate行业分析DeepSeekOpenAI
C ChatGPT Agent Loop 优化技术解析 本文深入解析了 ChatGPT 如何通过 Harness、API 和 Inference 三层架构优化 Agent 循环,重点介绍了持久化 WebSocket、增量 Token 化、KV 缓存管理和推测解码等技术,以降低成本并提升效率。 技术 › Harness Engineering ✍ Bytebytego🕐 2026-07-30 Agent优化LLM架构ChatGPTOpenAI性能成本控制WebSocketTokenization
B BestBlogs 早报|实现周期骤缩后,创业者如何重选问题 本期早报探讨了 AI 智能体缩短实现周期后,创业者的机遇与挑战。文章涵盖 Sam Altman 对创业窗口的判断、GPT-5.6 的效率工程实践,以及如何通过 Skill Harness 将模型能力封装为可维护的产品功能。 技术 › Skill ✍ ginobefun🕐 2026-07-30 GPT-5.6Agent创业效率工程ProductHarnessSkillLLMOpenAI
H How To Prompt Claude 5 Models 本文介绍了如何针对 Claude 5 系列模型(Fable, Opus & Sonnet)进行高效提示。文章涵盖了通用提示原则、Fable 的自主任务处理、Opus 的日常优化以及 Sonnet 的高效应用,帮助用户最大化模型生产力。 技术 › Claude ✍ AI Edge🕐 2026-07-30 Claude 5Prompt EngineeringFableOpusSonnetAnthropicAILLM