This article covers how to put "yourself" on camera explaining a product without re-recording every clip in person. You'll find it under AI Presenter in the top navigation once you're signed in; if the entry isn't there, your user group doesn't have it enabled yet.
Three steps to a finished clip
- Voice cloning. Upload a clean recording of your own voice and it becomes a private voice profile. One-time cost of 50 credits, and it usually goes from "cloning" to "ready" in one to two minutes. If you'd rather not clone anything, the shared system voices work directly.
- Voice synthesis. Pick a voice, paste your script (up to 10,000 characters per run), and get an mp3 back at roughly 3 credits per thousand characters. You can preview and download it.
- Avatar and video generation. Upload a piece of real on-camera footage as the avatar, and the system lip-syncs it to your audio. Billed by audio length, roughly 45 credits per minute on the standard tier.
For the fast path, use ✨ One-click video: enter the product and the angle you want, and the AI drafts a shot list you can edit before each shot is voiced, rendered and assembled automatically. For step-by-step control, use Manual pipeline: extract copy → rewrite it (you can specify the language, and synthesis and titles follow) → synthesize speech → generate the presenter → edit (captions, silence removal, background music, picture-in-picture) → title and cover → package and export.
Source material requirements
| Purpose | Requirement | Formats |
|---|---|---|
| Voice cloning | 10-20 seconds of your own voice; one speaker, no background noise | mp3 / wav / m4a |
| Presenter avatar | Front-facing talking footage with the mouth clearly visible, 10 seconds or longer, one person on camera | mp4 / mov |
| Narration audio | Either "my synthesized voice" or your own recording | mp3 / wav |
Sample quality largely determines output quality. Noisy rooms, overlapping speakers, profile angles or a partially covered mouth all raise the failure rate noticeably.
How long it takes
Lip-sync rendering is slower than image generation, and longer audio takes longer still — a one-minute track typically means several minutes of waiting. Before you submit, the page shows the estimated charge and the expected wait; nothing is deducted until you confirm. Jobs run in the background and survive a closed tab, so you can pick the result up later under Task History. Failures release the held credits automatically. Beyond the standard tier there is an HD tier, whose availability depends on what the platform currently has enabled; the price shown in the on-page estimate is the one that applies.
Three compliance gates
- Before cloning, you must confirm you have the voice rights of the person being cloned. Without it, the submission is blocked.
- When adding an avatar, you must confirm you hold both likeness and voice rights for that person. See likeness confirmation.
- The platform provides no music library, so background music has to be a licensed track you supply. See Background music and copyright.
There is no auto-publishing. The last step of the pipeline packages the finished video together with its title, hashtags and description for download, and you upload it to the destination platform yourself. Shots that include your original product photo are subject to the same product fidelity constraint, and billing follows the rules in Credits and pricing.