Technical requirements of large model enhanced digital human system

T/BIA 30-2025 defines a technical framework and requirements set for “large model enhanced” digital human systems, positioning large models as the enabling layer across the full pipeline from character and scene creation (2D/3D generation, textures, scene composition, HDRi/skybox imagery), through animation and control (lip-sync, gesture generation, instruction-driven actions, expression and behavior tagging mapped to outputs), to user-facing interaction (text dialogue, voice, visual and multimodal interaction). It specifies a modular architecture that typically combines modality-specific and multimodal large models with supporting components such as domain knowledge bases, tool libraries exposed via APIs, memory storage, an agent for task planning/execution, and input/output content safety screening, and then translates that architecture into measurable capability expectations and performance targets for quality, latency, accuracy, naturalness, and robustness across each stage, aiming to standardize how such digital human systems are built, evaluated, and accepted for practical deployment.

Front Matter

1 Scope (范围)

2 Normative References (规范性引用文件)

3 Terms and Definitions (术语和定义)

4 Abbreviations (缩略语)

5 Overall View: How Digital Human Systems Use Large Models (数字人系统应用大模型总体视图)

6 Technical Framework for Large-Model-Enhanced Digital Human Systems (基于大模型的数字人系统技术框架)

7 Technical Requirements (技术要求)

References (参考文献)