Supporters of Marcus Endicott’s Patreon can access weekly or monthly consultations on this topic.
China has assembled one of the world's most extensive national ecosystems for building, deploying, and governing digital humans, and the striking feature of that ecosystem is that no single technology explains it. What sets China apart is convergence: state industrial policy, municipal experimentation, telecommunications scale, and platform-economy incentives have been fused into a unified deployment environment, one in which a synthetic news anchor, a livestream sales host, and a municipal service assistant all draw on the same underlying stack. The commercial stakes are large enough to justify the coordination. The core market for virtual digital humans was estimated at roughly twenty billion yuan in 2023 and projected to more than double by 2025, while the broader industry surrounding it — live commerce, virtual influencers, AI-powered customer service — was forecast to approach six hundred billion yuan in the same window. Understood as infrastructure rather than as a product, the sector is a layered structure: a policy architecture at the top, a physical network beneath it, a constrained compute base under that, and a fast-industrializing platform layer that turns all of it into a service. The layers reinforce one another, and the frontier has now shifted from whether the technology can be built to how it will be governed.
The policy architecture is most revealing in what it withholds. At the national level, China's central plans establish sweeping digital ambitions without ever naming digital humans as such. The Fourteenth Five-Year Plan devotes an entire section to accelerating digital development and singles out real-time three-dimensional graphics, motion capture, and fast rendering among its technology priorities, yet the term "digital human" appears nowhere in it; the companion plan for the digital economy sets a target for core digital industries to reach a tenth of GDP by 2025 but likewise stops short of the category. National AI planning supplies scale targets measured in hundreds of billions of yuan, but the operational specificity lives one tier down. Beijing produced the country's first policy document explicitly aimed at digital humans in August 2022, an action plan that set a fifty-billion-yuan target for the city's digital human industry, called for cultivating leading firms, and mandated shared platforms for cloud rendering and interactive driving; by 2023 the city reported having roughly met it. Shanghai, Guangdong, and Zhejiang followed with their own metaverse and AI plans. The pattern is consistent: the center sets ambition in broad strokes, and provinces and municipalities translate it into named industrial targets.
Beneath the policy layer sits the physical network that makes real-time digital humans possible at scale. A conversational avatar is not a pre-rendered video but a live pipeline — speech recognition, language understanding, synthesis, lip-sync, and gesture generated continuously for every interaction — and that pipeline depends on low-latency transport between the user and the compute that drives it. China's 5G network supplies exactly that. By the end of 2025 the country operated nearly 4.84 million 5G base stations, the largest such network in the world, reaching every town and the overwhelming majority of administrative villages while carrying well over a billion 5G connections. Layered on top are the hyperscale data centers built out by the platform giants, whose accelerator clusters furnish the training capacity and, more consequentially, the inference capacity that live avatars consume. The combination of dense wireless coverage and concentrated compute is what allows a digital human to be deployed not as a showpiece in a single studio but as an always-available service spread across a city, a retail platform, or a government portal. Coverage this broad also means the deployment environment is not confined to a handful of coastal megacities; a service built once can, in principle, reach interior provinces and rural districts on the same network fabric, which is part of what makes the convergence a national ecosystem rather than a cluster of local pilots.
The layer that most directly meets the market is the platform tier, where a handful of cloud providers have turned avatar creation from a bespoke craft into a commodity service. What once required a specialist studio and months of work can now be produced from a few minutes of reference footage and a short script, delivered within a day, at a starting price on the order of a thousand yuan. The leading platforms advertise lip-synchronization accuracy in the high nineties across dozens of languages and support round-the-clock autonomous livestreaming with real-time responses to viewer comments. The economic consequence is the point. When a synthetic host can run a sales stream continuously without a human presenter, operating costs for e-commerce livestreaming fall by something approaching ninety percent, and the calculus of who can afford a digital storefront changes for hundreds of thousands of small merchants. Production speed and price, more than any single visual breakthrough, are what have moved digital humans from spectacle into ordinary commercial infrastructure, and they are the reason the deployment curve has bent so sharply upward in the space of a few years.
Every layer above ultimately rests on compute, and it is there that the structure is most exposed. Successive rounds of United States export controls have progressively fenced off the most advanced AI chips from Chinese buyers, and the aggregate effect has been steep: China's share of global AI computing power fell from more than a third in early 2022 to roughly a seventh three years later. The constraint bites digital humans in a specific way, because their workloads are overwhelmingly inference-intensive — the expensive part is not training a model once but running synthesis, lip-sync, and dialogue for millions of live interactions — so the binding limit is inference capacity at scale rather than raw research horsepower. The response has run along three tracks at once: algorithmic efficiency that wrings more from less, exemplified by the cost-conscious training methods that recently drew global attention; accelerating domestic silicon anchored by Huawei's Ascend line and supplemented by the in-house chip efforts of the largest platforms; and system-level clustering that lashes many domestic accelerators into super-nodes to compensate for the ceiling on any individual chip. The progress is genuine but constrained, and compute access remains the quietest and most decisive variable in the entire stack.
As the capability matured, the center of gravity in the sector shifted from engineering toward governance. China's regulation of digital humans began inside general AI rules — provisions on algorithmic recommendation, deep synthesis, and generative AI that already required consent for editing a real person's biometric features and labels on synthesized media — and has narrowed steadily toward the category itself. A nationwide labeling mandate requiring both visible and embedded markers on AI-generated content took effect in September 2025, and in April 2026 the cyberspace regulator circulated draft measures aimed squarely at digital human services: continuous, prominent labeling throughout a display, explicit consent before using anyone's likeness or voice, a prohibition on offering virtual intimate relationships to minors, and a life-cycle framework spanning creation, operation, and dissemination. The analytically important effect of this cascade is structural rather than merely restrictive. Mandatory labeling, consent management, and content controls impose compliance costs that large platform providers with established governance apparatus can absorb far more easily than the hundreds of startups that currently crowd the market, so regulation is quietly reshaping the industry's competitive contours even as it fixes its rules of conduct.
Taken together, these layers describe an infrastructure that has reached deployment maturity, where the binding constraints are no longer whether a convincing digital human can be built but how it is governed, what compute it can reach, and whether its commercial model survives contact with regulation. The physical network is in place, the platforms can manufacture avatars in hours, and the domestic compute base is advancing under pressure; what remains genuinely uncertain sits at the institutional edges. The state, for its part, has signaled that it means to accelerate rather than merely restrain. The Fifteenth Five-Year Plan, released in early 2026, elevates artificial intelligence to a top national priority and explicitly calls for building open-source AI communities, a posture that reads as investment in the ecosystem's next stage rather than caution about it. The distinctive quality of China's digital human infrastructure was always the convergence of policy, network, compute, and platform into a single deployment environment. What has changed is that the frontier of that convergence is now as much institutional as it is technical.