Synthesia created an interactive digital version of TechCrunch journalist Dominic-Madori Davis after capturing multiple images of her and recording two minutes of her voice. The result was not merely an image reciting written text, but an Avatar that listens to questions and converts them into text, then uses a language model to formulate a response and converts the answer into speech, while a video model animates the character’s facial features as it speaks.
The journalist chose to train the interactive version on a story she wrote about the reasons for the higher incidence of fraud among venture capital-backed startups compared with those that are not venture capital-backed. As a result, the twin was “deterministic”; that is, it answers only within the scope for which it was prepared and redirects questions outside that scope to the original story.
How Was the Digital Twin Built?
The Synthesia team needed a few days to create the models. The journalist received personal versions that read the texts she provided, and interactive versions that could listen and respond, with options either with or without glasses. The system uses Synthesia’s video and audio models, while also allowing users to select alternative models from Cartesia, ElevenLabs, Google, and OpenAI, in addition to options for hosting the twin on the client’s chosen cloud or with Synthesia.
The company places these capabilities into three main categories: a platform for creating and distributing video using traditional Avatars, the interactive Sessions platform for surveys and simulation exercises, and an API that allows audio and video models to be integrated with other services to build interactive Avatars or different products.
What Does the Experiment Reveal?
The journalist found that the voice was close to hers in the personal version, but her friends thought the resemblance in the interactive version was less convincing. Nevertheless, the presence was sufficient to produce an uncanny feeling, to the extent that her parents tried asking personal questions whose answers were known only to family members, without receiving a response because of the restrictions imposed by the training.
This result highlights an important difference between a narrowly controlled Avatar and a version powered by an open conversational model. The former is relatively predictable, but of limited use outside the material on which it was trained. The latter may appear more flexible, but it creates room for fabricated answers or behaviors that are difficult to control.
Why Does This News Matter?
The writer believes that the potential use in journalism is the most important question: Will audiences accept receiving news from a digital personality? Could an Avatar replace a journalist or allow chief executives to speak to a digital version of the journalist instead of the person herself? Her position is clear on one fundamental point: The essence of journalism is trust, and she does not believe this element can easily be delegated to artificial intelligence.
Outside media, the idea may be more appealing to companies and individuals; a digital version could answer questions while its owner is absent or assist with training and serving employees and customers. But the experiment does not prove that these models are ready to replace humans. Rather, it shows that controlling the scope of the answer and obtaining consent to use the voice and image remain central parts of their design.
The broader questions of use remain open: When does the audience know that it is speaking to an artificial version? Who bears responsibility for an incorrect answer? And what boundaries prevent the digital twin from turning from a practical interface into a personality that gives the user the impression of autonomy or knowledge it does not possess? These questions arise from the nature of the experiment itself, not from a promise that the technology will inevitably develop in any particular direction.