AI 数据遗产全景图 · AI 的粮食与祖传的管道 · 2026-07 · 地基层Data & Legacy × AI · The Feedstock and the Pipes · Jul 2026 · The Foundation Layer

AI 的粮食,与祖传的管道 AI’s feedstock — and the ancestral pipes

模型层的叙事最响,但真正决定 AI 能否落地的,是它脚下的两块地基:数据——AI 的粮食,与遗留系统——AI 要接的老管道。标注的价值从体力众包暴力迁移到博士级判断力,大模型第一次让「读懂无人能懂的祖传代码」规模化可行——这门生意没有被 AI 消灭,而是被 AI 重新定价 The model layer makes the loudest story, but what decides whether AI lands is the ground beneath it: data — AI’s feedstock — and legacy systems — the old pipes AI must connect to. Labelling value migrated violently from crowdsourced muscle to doctoral judgment, and LLMs made «reading the ancestral code nobody understands» scalable for the first time — this business was not destroyed by AI; it was repriced by it

本图相信:企业 AI 的第一死因被反复证实在地基层——权威研究显示 >80% 的 AI 项目失败(约为传统 IT 两倍),但其中仅 23% 归因技术因素,头号根因是「团队对项目意图的误解与沟通不畅」;分析师称仅 12% 机构拥有 AI 就绪的数据质量(B)。在可控的工程变量里,数据质量与治理排在模型选型之前;真正的天花板往往是组织与业务。This map believes: enterprise AI’s first cause of death is repeatedly proven to lie in the foundation — authoritative research finds >80% of AI projects fail (about twice the legacy-IT rate), yet only 23% for technical reasons, the top root cause being «misunderstood intent and poor communication»; analysts count just 12% of organisations with AI-ready data quality (B). Within controllable engineering variables, data quality and governance outrank model choice; the true ceiling is usually organisational.
本图不相信:「95% 的 AI 项目全垮了」这类二传读法——名校研究的原意是 95% 的生成式 AI 试点未产生可衡量的 P&L 回报(企业已投 300–400 亿美元),并非「项目废弃/技术不可用」;「40% agentic 项目将被取消」是前瞻预测非既成事实,且数千家自称 agentic 的厂商中仅约 130 家名副其实(agent washing)。诚实的读法比耸人的数字有用。It does not believe the relayed reading that «95% of AI projects collapsed» — the study’s actual claim is that 95% of GenAI pilots produced no measurable P&L return (against $30–40B invested), not that they were abandoned or unusable; «40% of agentic projects cancelled» is a forecast, not a fact, and of thousands of self-styled agentic vendors only ~130 qualify (agent washing). The honest reading beats the sensational number.

两个半球,两个最小单元:数据半球的「一次清洗与喂养」+ 遗留半球的「一段没人敢动的代码」——企业 AI 项目里二者串联:AI 要吃数据,也要接老系统取数、回写。主脊:双半球生命周期七环节(采集授权→清洗标注→合成与枯竭→治理确权→存储与检索→AI 考古→迁移现代化),②⑥标红;三条结构带:标注工厂的诚实层、POC 死亡谷解剖、遗留战场;判断层:数据质量与治理 > 模型选型;用 AI 做「读懂」,把「敢改」留给懂业务的人。姊妹:标注劳动→gig,代码→code,财税入表→ledger,确权合规→liability,政务系统→gov Two hemispheres, two atoms: the data hemisphere’s «one wash-and-feed» plus the legacy hemisphere’s «one stretch of code nobody dares touch» — wired in series inside every enterprise AI project: AI must eat data, and must tap the old systems to fetch and write it back. The spine: a twin-hemisphere lifecycle in seven links (sourcing & licensing → labelling → synthesis & exhaustion → governance & rights → storage & retrieval → AI archaeology → migration), ② and ⑥ flagged; three bands: the labelling factory’s honesty layer, the POC valley of death dissected, the legacy battlefield; the judgment: data quality and governance > model choice; let AI do the reading, keep the daring with people who know the business. Siblings: labelling labour → gig, code → code, accounting → ledger, rights → liability, government systems → gov.

数万
标注价格 U 型曲线的跨度:谷底不足 1 美分/单(众包图片标注)→右端 500–1000 美元/时(专家级「稀缺判断力」);中间的通用标注被商品化掏空——价值从劳动时间迁移到经过验证的判断力(B)The span of labelling’s U-curve: under 1 cent per unit at the crowdsourced bottom → $500–1,000/hour for scarce expert judgment at the top; the generic middle hollowed out by commoditisation — value migrated from labour time to verified judgment (B)
95%
名校研究:生成式 AI 试点未产生可衡量 P&L 回报的比例(企业已投 300–400 亿美元)——⚠️指「无可衡量回报」,不是「项目全垮」;与「>80% 失败、仅 23% 技术原因」互为印证:死因在地基与组织(B)The share of GenAI pilots with no measurable P&L return per the university study ($30–40B already invested) — ⚠️ «no measurable return», not «all collapsed»; converging with «>80% fail, only 23% technically»: death sits in the foundation and the organisation (B)
~55
COBOL 开发者平均年龄(每年约 10% 退休、85%+ 高校停开、60% 机构称招不到人);而大型机并未将死——一主机厂商该业务 2025 财年营收 +51.7%:AI 需求反而给老平台续命(B/A)The average COBOL developer’s age (~10% retiring yearly, 85%+ of universities no longer teach it, 60% of shops cannot hire) — yet the mainframe is not dying: one vendor’s line grew +51.7% in FY2025, AI demand extending the old platform’s life (B/A)
1600 亿元
2024 年中国数据市场交易规模(官方口径,+30%,场内超 300 亿翻番)——「数据从成本项变资产项」的国家工程:第五要素、三权分置、资产入表、交易所;确权法律真空是最大变量(A)China’s 2024 data-market transaction volume (official basis, +30%; on-exchange past ¥30B, doubled) — the state project of turning data from cost into asset: fifth factor, three-way rights, balance-sheet entry, exchanges; the rights vacuum is the biggest variable (A)
⚠️ 口径裁判(先读):① COBOL 代码存量三口径相差 3 倍以上(2200 亿行=上世纪 90 年代咨询估算经通讯社传播、8000 亿=2022 年问卷主观估算、2500 亿=开源组织独立估),「95% ATM 交易跑 COBOL」等系 2017 年报道的约十年前估算,引用务必标注年代;② 名校「95%」指无可衡量 P&L 回报、非项目废弃;「40% agentic 被取消」是前瞻预测;③ 厂商披露≠独立审计(理解时间 −79%、生产力 +60% 均为厂商口径;一云厂商的主机现代化服务未披露任何具名客户量化数字);④ 中国「数据要素市场/交易规模/产业规模」三概念常被混用(2030 年 7.5 万亿是最宽口径),交易所数字为累计口径且成立时间不一、不可横比;⑤ 数据耗尽年份是置信区间(2026–2032,计算最优路径约 2028),中文报道的单一年份均出自同一区间;⑥ 一标注巨头的战略入股金额有 143 亿/148 亿美元两口径(本页取与「49%→290 亿估值」自洽的主流财经口径);⑦ 「另一合成数据公司被收购」未获独立信源证实,已剔除;⑧ 四份深度研究交叉整理(一份带源核验为主力);渗透%为编辑估值。 ⚠️ Basis rulings (read first): ① COBOL line-count bases differ 3×+ (220B lines = a 1990s consultancy estimate spread by a wire service; 800B = a 2022 survey’s subjective guess; 250B = an open-source body’s independent count), and «95% of ATM transactions run COBOL» is a decade-old estimate from a 2017 report — always date it; ② the university «95%» means no measurable P&L return, not abandonment; «40% of agentic projects cancelled» is a forecast; ③ vendor disclosure ≠ audit (−79% comprehension time and +60% productivity are vendor claims; one cloud vendor’s mainframe service names no quantified customer); ④ China’s «data-factor market / transaction volume / data industry» are three routinely conflated concepts (the ¥7.5T-by-2030 figure is the widest), and exchange numbers are cumulative from different founding dates — never compare across; ⑤ data-exhaustion dates are a confidence interval (2026–2032, ~2028 on the compute-optimal path) — single-year headlines all derive from that one interval; ⑥ the labelling giant’s strategic investment carries two bases ($14.3B/$14.8B; this page takes the mainstream financial-press figure consistent with 49% → ~$29B); ⑦ a rumoured second synthetic-data acquisition failed verification and is excluded; ⑧ cross-compiled from four deep-research reports (one URL-verified as backbone); penetration %s are editorial.
中心装置 · 两个半球The central device · two hemispheres
AI 要吃数据,也要接老管道:POC 死在哪一环,都能在这张图上定位AI must eat data and tap the old pipes: wherever the POC dies, it dies somewhere on this map
数据半球的生命周期:产生→采集→清洗标注→治理确权→存储→喂给模型→反馈回流→合规退役;遗留半球:祖传系统→文档失传→维护者老去→AI 考古→迁移/缠绕/冻结→退役。两半球在企业里串联——任何一环断,全链停。The data hemisphere’s lifecycle: generation → collection → cleaning and labelling → governance and rights → storage → feeding the model → feedback → compliant retirement; the legacy hemisphere’s: ancestral system → lost documentation → ageing maintainers → AI archaeology → migrate/strangle/freeze → retirement. Inside the enterprise the two run in series — one broken link stops the chain.
数据半球(粮食被重新定价)The data hemisphere: feedstock repriced
标注从体力活变脑力稀缺品(U 型两端相差数万倍);公开高质量文本将在 2026–2032 区间耗尽,逼出合成/私有/具身三条军备线;数据从成本项变资产项(入表、交易所、1600 亿交易规模)——粮食链的每一环都在被重估(B/A)。 Labelling turned from muscle into scarce judgment (the U-curve’s ends differ by tens of thousands of times); public high-quality text exhausts within 2026–2032, forcing the synthetic, private and embodied arms races; data flips from cost line to asset line (balance-sheet entry, exchanges, ¥160B traded) — every link of the feedstock chain is being repriced (B/A).
遗留半球(管道第一次可读)The legacy hemisphere: the pipes turn readable
核心系统跑在退休潮中的老语言上(均龄 ~55 岁、年退 10%);大模型第一次让「逆向工程祖传代码」规模化(厂商称应用理解时间 −79%,D);但 AI 考古解决「读懂」,不解决「敢不敢改」——隐性业务逻辑、等价性验证、监管风险仍是人的问题(B)。 Core systems run on a language amid its retirement wave (average age ~55, 10% retiring yearly); LLMs made reverse-engineering ancestral code scalable for the first time (a vendor claims −79% comprehension time, D); but AI archaeology solves «reading», not «daring to change» — implicit business logic, equivalence testing and regulatory risk stay human problems (B).
判词Verdict地基层是企业 AI 的第一死因,也是最被低估的生意。在工程变量里「数据质量与治理 > 模型选型」成立;但更诚实的一层是:头号死因是组织沟通,不是技术——数据墙的根源是部门权力,祖传代码的风险是隐性逻辑。AI 把两个半球的「读懂成本」都打了下来,把「敢改的决断」留给了人。The foundation layer is enterprise AI’s first cause of death — and its most underrated business. Within engineering variables, «data quality and governance > model choice» holds; the honester layer beneath: the top killer is organisational communication, not technology — data walls root in departmental power, and ancestral code’s risk is implicit logic. AI has crushed the cost of reading in both hemispheres, and left the daring to humans.
Reading the Map

从这张图带走的五条规律Five patterns to take away

立场声明:本页是批判性、祛魅的行业结构分析,用 A–D 角标区分官方文件/法规、带源研究、二传媒体与厂商口径;广传统计(COBOL 行数、95% ATM、95% 试点)的溯源脆弱性已逐一标注;厂商效率数字未经第三方审计一律标 D;被误读的研究给出正确读法。本页提供行业结构信息,不构成技术选型、投资或合规建议。核心判断一句话:这门生意被 AI 重新定价而非消灭——数据的价值从劳动时间迁到经过验证的判断力,遗留系统的「读懂」成本崩塌而「敢改」依旧属人;地基不牢,模型再强也落不了地。 Stance: a critical, demystifying structural analysis; A–D badges separate official documents/statutes, sourced research, relayed media and vendor claims; the fragile provenance of viral statistics (COBOL line counts, the 95% ATM figure, the 95% pilot figure) is flagged item by item; unaudited vendor efficiency numbers are graded D; misread studies get their correct reading. Structure only — no technology, investment or compliance advice. One line: AI repriced this business rather than destroying it — data’s value moved from labour time to verified judgment, legacy’s reading cost collapsed while the daring stays human; and with a weak foundation, no model lands.