Supporters of Marcus Endicott’s Patreon can access weekly or monthly consultations on this topic.
When China's National Radio and Television Administration issued GY/T 411-2024 in late November 2024, it produced the country's first sector-specific technical specification for the digital virtual humans that had already begun presenting the news, interpreting sign language, and hosting livestream commerce across Chinese broadcasting and online audiovisual services. Formally titled Technical Requirements for Digital Human, the standard moved through its approval process with conspicuous speed. The administration posted the approval draft for public notice in mid-November, opened a ten-day comment window, and — the committee review already behind it — issued the finished standard two days after that window closed, with the publication notice following almost immediately. Its designation marks it as a recommended industry standard for the broadcasting sector, one tier below the mandatory national standards, a status that ultimately matters less than the fact of its existence: for the first time, a Chinese regulator had committed to paper a shared technical vocabulary for what a digital human is and what it must be able to do. The lead drafter was the China Broadcasting Design and Research Institute, working alongside a coalition of broadcasting research bodies, universities, and technology firms. Tencent Cloud's presence among the co-drafters proved commercially telling, as the company soon announced that its digital human products complied with the new standard in full.
The standard's conceptual core is a classification system that sorts every digital human along three independent axes. By appearance, a digital human is either two-dimensional or three-dimensional; by interaction mode, either interactive or non-interactive; by driving mode, either algorithm-driven or real-person-driven. Any given avatar occupies one position on each axis, so that a photorealistic three-dimensional news anchor scripted entirely by software and a flat animated avatar puppeteered by a hidden operator are described in the same grammar rather than treated as unrelated species. Over that scaffold the standard lays two further structures. It sorts the uses of digital humans into four application scenarios — content broadcasting, interactive customer service, virtual performance, and content creation — and it decomposes the machinery itself into five capability layers, each governed by its own chapter: appearance, algorithm-driven capabilities, real-person driving, platform, and security. The cumulative effect is to convert a diffuse and fast-moving market of bespoke products into a legible set of specifications that a regulator, a broadcaster, or a vendor can name, measure, and point to when a dispute or a procurement decision requires it.
Where the standard becomes concrete, it becomes demanding. The appearance requirements for realistic three-dimensional humans reach down to the anatomy of the rendered head and the texture of individual strands of hair, and every type of digital human, however it is built, is forbidden from distortion, frame skipping, audio-video lag, or lip-sync failure that would betray its artifice. The algorithm-driven layer sets a hard performance floor: at high-definition resolution, video synthesis must run at least as fast as real-time playback and at no fewer than twenty-five frames per second, while speech synthesis must be controllable down to the volume and duration of individual words, switchable across registers as different as news broadcasting, narration, and customer service, and capable of shifting emotional tone automatically as the meaning of the text demands. Rendering must run on either of the two dominant game engines, and the multimodal integration layer insists on pronunciation without missed or mistaken sounds, lip movement that tracks speech naturally, body language that fits the moment, and lighting computed in real time against the physics of the scene. The real-person layer governs the motion-capture performer — the person inside, in the industry's own phrase — whose actions, expressions, and voice must map onto the avatar in real time and at a sampling precision fit for actual use. One provision in particular quietly anticipates how these systems are run in practice rather than in theory: human driving and algorithmic driving must be capable of mixing, with mutual takeover, so that a live operator can seize control of an otherwise autonomous avatar or return it to the algorithm mid-stream without any visible seam. It is a technical requirement that encodes an operational reality — that the line between a human performer and an AI is, on air, deliberately blurred.
It is in its security chapter that the standard reaches beyond technique and touches the law. Alongside conventional data-security requirements — collection bounded to lawful purposes, configured access controls, encrypted transmission, and an explicit prohibition on producing or disseminating false content — sits a single provision that carries real weight. When a real person's face, voice, or other biometric information is edited, the entity operating the digital human must notify that individual and obtain their separate consent. This imports the elevated consent threshold that China's Personal Information Protection Law reserves for sensitive personal data, and in doing so it quietly stitches a technical broadcasting standard into the country's broader framework for biometric privacy. What the standard conspicuously declines to address is equally telling. It does not venture into who may say what through a digital human, how such a figure must be disclosed to an audience, or what purposes it may not serve. That restraint is not an oversight. It is the clearest signal of what the document was built to be.
For GY/T 411-2024 was never meant to stand alone. It arrived into, and deliberately deferred to, a governance architecture that had been thickening since the Cyberspace Administration's deep synthesis provisions took effect in January 2023 — rules whose reach over deep-learning-generated media captured digital human technology almost in its entirety, and which pointedly instructed online audiovisual services to comply additionally with broadcasting regulators, laying the jurisdictional bridge the NRTA standard would later occupy. Those foundational provisions already required consent for biometric editing, invisible watermarking, and explicit visible labels wherever synthetic content might mislead the public — obligations the broadcasting standard would echo rather than invent. The interim measures for generative AI services followed in the summer of 2023, extending oversight to the underlying models and, by late 2025, drawing hundreds of registered services into a mandatory filing regime. A national standard requiring both visible and embedded labels on all AI-generated content took effect in September 2025, reaching every digital human deployment at once. Two months later, a national standard for virtual digital humans led by a major AI firm and the country's standards institute did what the broadcasting document had not, attaching hard numbers to the category — minimum polygon counts for ultra-realistic models, floor thresholds for lip-sync accuracy and emotional-response success. By early 2026 the industrial ministry had opened public comment on a mandatory identity scheme proposing a single unique code for every commercially deployed digital human in the country, a response to more than a million registered digital-human enterprises and the fraud that had grown up among them.
The two instruments that most sharply define the trajectory are draft measures from the cyberspace regulator. One, circulated at the end of 2025, targets anthropomorphic, companion-style AI with continuous-use limits, mandatory pause reminders, and a prohibition on emotional manipulation designed to retain users. The other, released in April 2026, addresses digital virtual human information services directly and comprehensively: persistent on-screen labeling throughout a service, a ban on using avatars to defeat facial or voice recognition, protection for the rights of motion-capture performers, restrictions around minors and the biometric traces of the deceased, and graduated fines for violations. That April draft calls explicitly for building out a digital virtual human technical standards system — language that retroactively frames GY/T 411-2024 not as a finished edifice but as a foundation stone. And that is finally what the broadcasting standard represents. It did one narrow thing, and did it first: it fixed a technical baseline for what a digital virtual human must be and must be able to do inside China's media system, and left everything else — content, identity, behavior — to the widening cascade of rules around it. China's approach to governing its digital humans is not converging on a single regime but layering technical specifications, labeling mandates, identity codes, and behavioral limits into an increasingly dense mesh. GY/T 411-2024 was the opening move, and the architecture it opened is still being assembled around it.