In an article published on TechCrunch, senior venture capital and startup reporter Dominic-Madori Davis describes her personal experience creating an interactive digital avatar in her likeness at startup Synthesia. Davis recounts that her initial encounter with the technology began when Alexandru Voica, head of corporate affairs at Synthesia, sent her a link to an interactive virtual avatar of himself, which had been trained to answer common press questions about the company, its activities, and how it works. The day before, Davis had participated in a panel where she was asked whether PR outreach using AI-generated text bothered her, but Voica's avatar seemed to her like a far more advanced stage of using AI in PR.
Synthesia Company Figures and Product Offerings
In September, Davis was invited to Synthesia's new office space in New York. The company, originally based in the U.K., operates in the digital avatar space alongside other companies such as D-ID, HeyGen, and Colossyan. According to the report, Synthesia reached a $4 billion valuation earlier this year and stated last year that it had crossed the $100 million threshold in annual recurring revenue (ARR).
Synthesia enables enterprises to build interactive training videos with AI avatars, and recently launched a product named Roleplay Sessions, which allows employees to practice tasks such as sales pitches in front of an interactive AI avatar that responds to and scores their performance. Overall, Synthesia builds three main types of products:
- A video-creation and distribution platform with classic avatars, where a user types a script and the avatar repeats it.
- An agentic platform named Sessions, where users can interact with avatars in surveys or roleplay.
- An API platform that allows people to take Synthesia's video and voice models and combine them with other technology services to build interactive avatars or other products.
Digital Avatar Creation Process and Technical Specifications
When Davis was offered the opportunity to create her own avatar at the office opening event, she agreed immediately. According to Davis, prior to meeting her digital twin, she had felt indifferent toward avatars, but anticipated that they would inevitably become part of everyday online life, partly after hearing about Instagram users creating avatars in their likeness to produce content. This was the first time Synthesia had created a digital avatar for a journalist, or for anyone outside Voica himself.
The creation process involved entering a mini film studio inside the company's offices, where numerous photos of her were taken and a two-minute voice sample was recorded. Davis was required to provide explicit consent for the avatars to be made. The team created for her a personal avatar (which reads inputted scripts) in versions with and without glasses, and two interactive avatars (capable of listening and responding), also with and without glasses.
Davis's interactive avatar was built around an article she had previously published on the reasons why venture-backed startups commit more fraud than non-VC-backed startups, and it was designed to answer solely questions regarding that article and its findings.
The technical stack of the avatar is based on a combination of four components:
- A Voice-to-Text model that converts user speech into text.
- An agentic language model that makes sense of the text and can take actions based on it.
- A Text-to-Voice model that converts the generated response into audio.
- A video model built by Synthesia, which animates the avatar as it talks.
Synthesia's system architecture includes its own video and voice models, but the company allows enterprise customers to choose alternatives from other labs, such as Cartesia, ElevenLabs, Google, or OpenAI. Additionally, enterprises can choose to host their avatars on any cloud infrastructure of their choice, or pay Synthesia for hosting services.
Avatar Testing and Initial Reactions
Building the avatars took a couple of days. Davis first tested the personal avatar by inputting a script about the arrival of autumn in New York, finding that the voice was fairly accurate and did not pick up the hoarseness present during her original recording. Non-tech friends found the result interesting but somewhat creepy.
Next, the interactive avatar was tested, which is deterministic—meaning it responds only to topics on which it was pre-trained. When asked personal questions about Davis's professional background or where she lived, the model redirected the conversation back to the venture fraud article. Her parents attempted to ask questions only they would know about her, but the avatar refused to diverge from the topic and redirected back to the article.
Questions on the Future of Journalism and Public Trust
The experience led Davis to raise questions regarding the future of journalism. Among the questions raised: Would the public be willing to watch news broadcasts presented by avatars? One investor she spoke with answered with an immediate no, while others expressed uncertainty, against the backdrop of existing pushback against poor-quality AI content (AI slop) circulating across social networks and news platforms. Davis notes that the central element of journalism is trust, and in her view, it does not seem this component can be outsourced to AI.
Outside of journalism, Davis assesses that digital cloning could appeal to the corporate sector, for instance to enable ongoing work coverage during absences or vacations. However, she expresses mixed feelings and notes that non-deterministic models—where a chatbot is free to generate responses without restriction—could trigger a sense of AI psychosis. In conclusion, Davis predicted that members of Gen Z will likely struggle to get used to digital avatars due to their sci-fi feel, but noted that she finds them less jarring than humanoid robots, since with a digital avatar one can always simply log off the screen.