White Paper on the Development of China’s Virtual Digital Human Industry
The 2024 White Paper on the Development of China’s Virtual Digital Human Industry (2024年中国虚拟数字人产业发展白皮书) was published by iiMedia Research (艾媒咨询), a Guangzhou-based commercial research firm, and released on 19 April 2024 at the firm’s own industry forum in Guangzhou (announcement and download). It runs to 95 pages with roughly forty exhibits, written in Chinese with bilingual exhibit titles, and is sold at ¥2,699 under the code IM43577 with a short free preview. Its market figures come from in-house monitoring and modelling; its consumer and enterprise findings come from a single survey of 1,628 respondents fielded in March 2024; its company material is case-study profiling, some of it promotional. This is vendor-adjacent commercial research, not peer-reviewed work. Four parts follow.
Section one lays the groundwork. The argument is that everything necessary for a virtual human industry has arrived at once. Government policy encourages the sector at both central and provincial levels. The digital economy is large enough — around 56 trillion yuan, heading past 70 trillion by 2025 — that digital consumption no longer needs justifying to consumers, and enterprises are already committed to digital transformation for their own reasons. Culturally, spending on entertainment keeps rising and live streaming has become mass behaviour: 816 million users by the end of 2023, three quarters of everyone online, with virtual anchoring the fastest-growing thing inside it. Underneath sits 5G and AI infrastructure buildout. Then the section delivers the numbers everyone cites: a core market of 20.5 billion yuan in 2023 against a driven or peripheral market of 333 billion, forecast to reach 48 billion and 640 billion by 2025. The reason given for continued growth is plasticity — virtual humans can be pointed at whatever is trending, so the sector never runs out of new hooks.
Section two maps the industry and then walks through who is actually doing the work. The chain runs upstream to downstream: tooling and AI capability at the base (Autodesk, NVIDIA, Microsoft, SenseTime, Huawei, Tencent, iFlytek), platforms in the middle (Baidu, JD, Xiaoice, NetEase Fuxi, XMOV and a cluster of motion-capture houses), and applications at the end sorted by use rather than technology — retail, financial endorsement, entertainment. The organising distinction for the whole report appears here: service-type virtual humans exist to replace labour and cut cost in industries that already exist, while identity-type virtual humans exist to carry a persona and create new value where there wasn’t a market before.
Upstream is where the report is most concrete about how these things are built: exterior form from modelling, physics and rendering; interior driven by either AI or a live human operator, treated as interchangeable options rather than as old and new. The exemplar is Original Force, a Nanjing 3D studio founded in 2010 with 2,200-plus staff that does CG work for DreamWorks, EA, Sony, Warner and Tencent — and, in June 2023, produced a 1:1 digital reconstruction of the Ming-era Bao’en Temple porcelain pagoda with academic collaborators. That’s the heritage-reconstruction template two years before Jingdezhen, and it’s worth noting the studio is a contractor here, not a platform owner.
The Top 100 IP ranking is the report’s proprietary artifact. It scores companies 0 to 100 on enterprise strength, traffic heat, commercial value, media reach and online reputation, computed on iiMedia’s CMDAS system with an analyst adjustment layered on top — which makes it unreproducible from public data by design.
Virtual idols get sorted twice over: by look into anime-style, hyperreal and digital-twin, and independently by drive method into fully AI-synthesised versus puppeteered by a hidden performer. The cases are iFlytek Music’s AI singer Luya, who debuted in 2022 and went straight into a joint research centre with the Shanghai Conservatory before working through NFT drops, anime insert songs and provincial gala appearances; Baidu’s Du Xiaoxiao and Xi Jiajia, claimed as China’s first interactive virtual idol and the world’s first AIGC virtual idol respectively; and Yuanyuan Technology’s Tianyu, a traditional-style figure positioned as a cultural export.
Virtual live streaming splits anchors three ways by who they face. Media-type anchors read news and host galas, and the pitch is that they never sleep and can work in multiple languages — CCTV’s Xiao Xiao Sa, Xinhua’s Xin Xiaowei, People’s Daily’s Guoguo. Entertainment-type anchors live on social platforms making short video and building relationships with audiences, the Bilibili singers and gamers and Kizuna Ai. E-commerce-type anchors sell things, and the pitch is coverage at hours when no human would stream — Kuaishou’s Guan Xiaofang, JD’s Yanxi. The case study is NetDragon Tianuo, a subsidiary carrying its parent’s gaming and IP experience into virtual characters, whose Douyin musician Xueyu Renyun had done 558 livestreams and gathered 1.35 million followers by March 2024.
Virtual employees are the least glamorous and the most operational part of the report. A sector table specifies actual functions: chasing credit card payments and reconciling accounts in banking, welding and sorting in manufacturing, ticketing and itinerary queries in tourism. The cases are Mengniu’s Naisi, Vanke’s Cui Xiaopan — credited with a 91 percent collection rate and seven-times-faster response than the system she replaced — Sequoia China’s investment analyst Hóng, reading a hundred business plans a second, and a motion-analysis figure built to help freestyle skier Xu Mengtao train.
Section three is the survey, and it’s one instrument throughout: 1,628 respondents, fielded March 2024. Awareness is high and gets scored rather than described — 9.00 for acceptance, 8.39 for favourability, both on ten-point scales. Virtual anchors are what people actually encounter, at 81 percent, ahead of idols and employees. On the enterprise side, half of firms surveyed already use the technology and a third more plan to; the obstacle they name is cost, at 44 percent, followed by technical difficulty and unclear use cases. What they almost never worry about is whether users will accept it — 6.6 percent — which is the more interesting result and one the report passes over. Consumer money is modest: the largest spending band on virtual idols is 201 to 500 yuan a month, and almost nobody exceeds 2,000. But 40 percent say they’d spend more on a virtual idol than a human one against 12 percent who’d spend less. Sentiment toward anchors is similarly warm, with people wanting to see them expand into media, culture-tourism and film. For virtual employees, advisory and assistant roles lead usage, and the sector where respondents most often meet one is culture and tourism at 58 percent, ahead of media and education — which puts heritage and tourism at the front of consumer-facing digital human exposure well before the current wave of cultural projects.
Section four turns all of that into a strategic frame and two forecasts. Strengths: an industry on the upswing, growing fast, broadly applicable, with real enterprise demand. Weaknesses, and these are franker than you’d expect from a house that sells to this sector: the technology still has bottlenecks, emotional interaction is weak, and production and operating costs are high. Opportunities: state support for new cultural formats, plus the combined lift of the metaverse concept and the rise of AIGC. The single threat is competitive intensity — tech companies of every size plus startups, all crowding in.
The two forecasts are where the report dates itself. First, policy will keep maturing, and the report is explicit that Chinese virtual digital human policy is still at a preliminary planning stage — written, of course, two years before the CAC anthropomorphic AI measures. Second, AI breakthroughs are accelerating commercialisation, with ChatGPT and Ernie Bot named as the moment when deep-learning language generation became combinable with a virtual body. That’s the tell: language models are treated as an external accelerant arriving from outside the virtual human stack and bolted on, rather than as the thing the stack is now built around. The 2026 generation of real-time avatar systems has inverted that completely.