Evaluating Multimodal Identity Consistency Across Text, Voice and Video
A digital human can look consistent while feeling inconsistent. This framework evaluates whether text, voice, facial behavior and visual style express the same underlying identity.