Help · Tool guides

Drama Ads: One Product Photo, One Spoken Clip

One product photo plus one presenter gives you a 6-10 second sales clip. How to write a scene prompt that holds up, how to pin down the product's form, and what it costs.

This article solves one problem: you want to post short sales videos, but you don't want to book a location or put a real person on camera.

What you need

A product photo is required; a presenter photo is optional. Without one, the model generates the person from your scene prompt. If you want a consistent presenter, generate an AI portrait on the platform first, save it to Assets and reuse it — no talent booking, and one less layer of real-person likeness to manage.

For the scene prompt there are three ready-made patterns you can tap straight in: tasting, unboxing, and in-use. You can also write your own, up to 500 characters.

How to write a scene prompt that doesn't fall apart

In one line: the simpler the action, the more reliable the clip.

Continuous multi-step actions are where video models are weakest. Ask a presenter to "open the cap, spray twice, put the cap back, then unscrew a second bottle" and things break in obvious ways: a duplicate bottle appears out of nowhere, the spray doesn't match the nozzle, the cap changes color. We hit this repeatedly in our own testing, and the conclusion was simply to cut the action down to one.

The dependable formula is one primary action, one selling point, one clear camera move. For example: "close-up on hands slowly opening the packaging, revealing the exterior then the detail, camera follows the hands smoothly."

  • If you can say "show," "hold up" or "hand toward the camera," don't say "disassemble," "spray" or "pour."
  • Keep one product on screen. Skip "two bottles side by side" or "a table full of them."
  • Don't count on small text rendering accurately. Anything that has to be exact belongs in a post-production caption.

Describe the product yourself, in one line

The panel has a product appearance field. Fill it in and the system writes that line into the prompt as a named anchor, explicitly forbidding the model from switching category or simplifying the form. Leave it blank and the backend falls back to recognizing the image itself — workable, but your own sentence is more accurate.

This field earns its keep most on unpackaged forms. Boxes usually survive; the bare product, once it's out of the box, is the thing most likely to be repainted into a shapeless blob. Spell out traits like "whole, legs and claws intact, textured shell" and the failure rate drops noticeably. The system separately adds its own constraint that packaging must sit naturally on a surface and may not float or stick to the person.

Fidelity works the same way here: every request that carries your original photo gets the fidelity constraint injected. In rare cases the model still drifts — re-run it, and failures release the held credits automatically. The long version is in How product fidelity works.

Pricing, length and aspect ratio

  • Two lengths: 6 seconds for 12 credits, 10 seconds for 20 credits (CN¥1 = 10 credits).
  • Resolution 480p or 720p; aspect ratio 9:16 (default), 2:3, 1:1 or 16:9.
  • Jobs run in the background and survive a closed tab. Finished clips land in My Works, and a failed generation releases the held credits.

If you want an actual narrative arc, there's a longer route: generate a script (2 credits), then a nine-panel storyboard (20 credits), then the finished clip in one pass, at 5, 10 or 15 seconds. The three stages are confirmed and charged separately, and each can be re-rolled on its own — the advantage being that you get to read the script and check the shot list before deciding whether to keep going.

Two compliance points

The moment you upload a presenter photo, "I confirm this likeness is me or that I have the rights to it" becomes mandatory; without it the backend rejects the request. This applies regardless of where the image came from: a real person's photo needs it, and so does an AI portrait generated on the platform — it's the same gate. The reasoning is in Why the likeness confirmation is mandatory.

Background music is yours to bring, licensed; the platform does not provide a music library. Marketplace review and advertising-law compliance for the finished clip are governed by the Terms of Service — give it your own pass before you publish.

Still stuck? Email [email protected]