TL;DR: Synthesia runs two separate prompt-to-video systems today: the newer, beta Assistant, and an older Legacy AI Video Assistant powered by OpenAI. Neither is available on the Free plan. A Synthesia script is read aloud by an avatar, not edited like a video clip, and creating anyone's likeness needs that person's own direct consent — not just their permission relayed by someone else.
What does "prompting Synthesia" actually mean today?
Four different things, and the schema difference between them is exactly where a copied prompt goes wrong.
| Feature | Surface | Takes a prompt? | What it produces |
|---|---|---|---|
| Assistant (homepage prompt box, beta) | New script + visuals, from a written prompt, files, or URLs | ||
| Import a script (homepage) | Your own script, used word-for-word, no rewriting | ||
| Import a file (Legacy AI Video Assistant) | New script generated from a document; original wording discarded | ||
| Import a script (Legacy, Create → Video → Import a file) | Your own script, word-for-word — same job as above, different line-break rule |
The trap is the second row. "Import a script" sounds like it should live next to "Import a file," and it does in the menu, but it takes no prompt at all: paste your own words and Synthesia uses them exactly as written. Upload that same script as a file instead, and Synthesia's documentation is direct about what happens: "Synthesia reads the file and writes a new script from its content." It does not carry over your original wording or structure, even if the file already reads like a script. If you want your own sentences in the final video, paste them; don't upload them.
The two "Import a script" paths, one from the homepage and one from the older Create menu, also disagree on formatting: the newer flow needs two soft returns (Shift+Enter twice) between scenes, and the legacy one needs only one. Getting this wrong doesn't error out — it silently merges or splits scenes in a way you won't notice until you're reviewing the storyboard.
The same script-vs-prompt split shows up on other avatar-video tools, not just Synthesia. HeyGen's own avatar-video endpoint takes a script field and has no prompt field at all, rejecting one outright if you send it — a different mechanism arriving at the same lesson covered in HeyGen's own prompt templates. If a "prompt pack" for any talking-head tool reads like generic scene direction rather than dialogue, it was probably written for the wrong field.
Which Synthesia plan do you need before you can type a prompt at all?
None of this works on the Free plan. Synthesia's own documentation draws the line plainly: "On Free plans, Start with a template, Import a PowerPoint, Start from blank, and AI Dubbing are available. Start with a prompt, Import a script, and Import a file require a paid plan, since they run through Assistant." That's Starter ($29/mo, or $18/mo billed annually) at minimum, up through Creator and Enterprise.
Assistant itself carries the same floor, stated directly: "Assistant is available on Starter, Creator, and Enterprise plans. Using a brand kit with Assistant is an Enterprise-only feature." The in-script "Edit with AI" tool, which rewrites, lengthens, shortens, changes tone, or summarizes a highlighted passage without leaving the script box, has the identical gate: "Edit with AI is available on Starter, Creator, and Enterprise plans. Workspace admins can also disable it from workspace settings." A Free account can still build a video from a template or a PowerPoint deck, just not from a written instruction.
Why does a script read differently than it reads on the page?
Because an avatar speaks it, and there is no editor sitting between the words and the delivery the way there is with recorded footage. Every craft decision that a video editor would normally fix in post has to be right in the text instead, because pacing here is set by the script, not by cutting it afterward.
Four things change specifically because the output is spoken, not read:
- Sentence length compounds. A written sentence a reader can re-scan in a second becomes, out loud, a run a listener has to hold in memory in one pass. Long, clause-stacked sentences that read fine on a slide sound like a run-on lecture from an avatar.
- Numbers and acronyms need a plan. "Q3" reads differently depending on whether the avatar says "Q-three" or "third quarter," and Synthesia's pronunciation tool exists precisely because the default guess is sometimes wrong for a brand name, a technical term, or an acronym nobody in the room says letter-by-letter.
- Punctuation is a delivery instruction, not decoration. A comma is a micro-pause a reader skips past silently; spoken, it's an actual gap in the audio. A sentence with three commas reads as measured; the same words with none of them run together.
- There is no visual context unless you write it in. A slide deck lets a viewer's eyes fill gaps a presenter leaves unsaid. A Synthesia avatar has whatever's on screen and whatever the script says, and nothing else — if a number or a name matters, the script has to carry it, because there's no cutaway to a chart doing that work for free.
None of this is unique to Synthesia; it's the same discipline behind why AI-written scripts flop on any avatar or voiceover tool. What's specific to Synthesia is which of these problems it gives you a real control for, and which it doesn't.
What can you actually control inside a Synthesia script?
More than "type text and hope," but the controls are narrower and more mechanical than a video editor's toolkit.
Pause. Insert one anywhere in the script box: "Adjust the pause duration to a value between 0.1s and 99s (the default is 1 second)." Pauses can be dragged to a new position or copied elsewhere in the script.
Pronunciation. Highlight a word or phrase and either type one, following Synthesia's own instruction to “Enter a phonetic spelling that sounds right when read naturally (for example, "sin-THEE-zhuh" for Synthesia).” Or record yourself saying it and let Synthesia use that recording as the reference. Save it to your workspace's Glossary and it applies automatically the next time anyone on the team types that word — brand names and technical terms are the obvious use, since guessing an acronym's pronunciation is exactly where a generic script sounds off.
Voice speed. Independently adjustable per paragraph or across a whole video, documented as "Minimum speed: 0.8× (slower, more deliberate delivery)" up to "Maximum speed: 1.2× (faster, more energetic delivery)". Dense, technical paragraphs can run slower than a casual intro without touching the words themselves.
Per-paragraph voice and language. Each hard-return paragraph carries its own speaker pill, so one scene can mix speakers or languages by paragraph rather than needing a separate scene per voice.
Speech regeneration. Because "Synthesia's AI voices are non-deterministic," the same script and voice can render slightly differently take to take. Regenerating gives you an alternative delivery of identical words without touching the script, which is the fix when the wording is right but the read feels flat.
One naming detail worth knowing if a voice suddenly sounds different: Synthesia's newer "remastered" voices default to an in-house variant and unlock pronunciation controls, but you can switch a voice back to the older "ElevenLabs Turbo v2" variant if you preferred its sound — a reminder that some of what powers a Synthesia voice is licensed the way ElevenLabs' own voice tools work, not built from scratch, and switching variants trades away pronunciation controls to get it.
How does Assistant turn a prompt into a video?
You type a prompt, optionally attach files or URLs, and Assistant returns an editable outline before it generates a single scene.
Set a duration (Short, Medium, or Long) and a delivery style: Dynamic for shorter scenes with frequent cuts, or Presentation for longer, denser ones. Language has no dedicated setting; instead: "If you write a prompt, the script defaults to your prompt's language, even if it differs from any attached files." If you're attaching files without writing anything, the output follows the files' language instead — so a multi-language rollout needs the target language stated in the prompt itself, not assumed from context.
Once the outline generates, further changes happen through chat rather than a fresh prompt. Synthesia's own documented examples of this iteration are worth copying directly: "Let's make this shorter", "Let's make this snappier", and, for a genuine second opinion rather than a rewrite, "Act like a VP of learning and development. What would you change about this video?" Brand kit application inside Assistant is Enterprise-only; on Starter and Creator, template choice controls the visual layout instead.
Templates for corporate training and explainer scripts
Every template below assumes a paid plan. Fill in the bracketed placeholders and paste the result into Assistant's prompt box, the Legacy AI Video Assistant, or Import a script depending on which surface you're using.
1. New-hire onboarding module (Assistant, prompt-to-video)
Create a 3-minute onboarding video for new [department] hires at [company].
Cover: what the team does, who they'll work with day to day, and where to
find the team wiki. Delivery style: Dynamic. Duration: Short. Tone: warm,
not corporate.
2. Compliance training explainer (Assistant, with source files)
Using the attached policy document, create a training video explaining
our [policy name] policy to all employees. Cover what's required, what
happens if it's not followed, and who to contact with questions.
Delivery style: Presentation. Write this in English.
3. Software feature walkthrough (Assistant)
Create a 90-second walkthrough of the new [feature name] in our product,
aimed at existing customers who haven't tried it yet. Explain what
problem it solves before showing how to use it. Delivery style: Dynamic.
4. Iterating on a generated outline (Assistant chat)
Let's make this shorter and cut the third section entirely.
Act like a training manager reviewing this for a new-hire audience.
What would you change?
5. Verbatim script for a compliance-reviewed announcement (Import a script)
Paste your already-approved wording directly; do not run it through Assistant or Import a file, both of which will rewrite it. Use two soft returns between scenes on the homepage flow, one on the legacy flow.
[Scene 1 — approved wording, exactly as written]
[Scene 2 — approved wording, exactly as written]
6. Multi-language training rollout (Assistant)
Write this in German. Create a 2-minute onboarding video covering our
data handling policy for new hires in our Berlin office, using the
attached policy PDF as source material.
7. Executive announcement, single speaker (Import a script, verbatim)
[Speaker pill: CEO avatar]
Thank you all for your work this quarter. I want to share three updates
that affect how we'll operate starting next month.
8. Sales enablement explainer (Assistant)
Create a 2-minute video for our sales team explaining the new
[pricing tier] we're launching. Cover who it's for, what changed from
the previous tier, and the one objection reps hear most often, with a
suggested response. Delivery style: Dynamic.
9. Pronunciation glossary entry for a recurring brand term
Not a video prompt, but the fix for a name every training video in a workspace will need. Highlight the term once, choose Pronunciation, enter a phonetic spelling such as sin-THEE-zhuh, and apply it to the workspace Glossary so every future script gets it right without re-teaching it per video.
10. Pause-and-emphasis pass on a dense paragraph
[Scene text, unpaused]
Before you submit, check three things: the budget code, the manager
approval, and the vendor's tax ID.
Add a 0.4–0.6 second pause after "three things:" and after each item in the list; a script this dense reads as a wall of words without a beat between items.
Who has to consent before you can put words in their mouth?
The person whose face and voice you're using, directly, and Synthesia enforces this rather than leaving it to good judgment.
Its Photo Avatar policy is explicit: "Photo Avatars must be created directly by the person whose likeness is being used. You may not create or submit a Photo Avatar on behalf of another person, even if you have their permission or consent." A manager submitting a colleague's photo and a recorded "yes, that's fine" from that colleague still fails this rule, because the submission itself has to come from the person being depicted. The same policy blocks a specific category outright: "Celebrities, historical figures, or deceased people." Every submission also goes through an identity-match check between the uploaded photo and a live consent video before it's approved.
Synthesia's Acceptable Use Policy backs this with a general instruction to "Moderate the use of Avatars, including the use of voices and likenesses, to ensure that the rights and interests of individuals are being respected and the public is being protected from harm." None of this is a workflow to route around — a prompt or a script that implies putting real words in a real, non-consenting person's mouth isn't a corner case Synthesia tolerates quietly, it's the exact thing this policy exists to catch.
This isn't a Synthesia-specific quirk, either. Descript's Custom Avatars feature, part of the surface covered in its own prompt templates, sits under the same basic logic: a presenter's face and voice are not a generic asset a prompt can conjure permission for. Whichever avatar tool a team standardizes on, the consent step is the one part of the workflow that has to happen before the first prompt, not alongside it.
Can a Synthesia video be used commercially without restriction?
No, and the limits are specific enough to check rather than assume. Synthesia's Acceptable Use Policy singles out paid social advertising specifically: putting a Stock Avatar into promoted, boosted, or paid advertising on any social platform requires Synthesia's own written consent first, and doing the same with a Custom Avatar is only allowed "with appropriate consent from the individual." The policy also prohibits removing any watermark or provenance marker the platform adds. What a specific plan, avatar type, and use case actually clears for commercial use is a question for Synthesia's terms and your account's plan, not a blanket yes this guide can hand you.
One data point worth knowing if a workspace's AI usage policy asks: Synthesia's own Legacy AI Video Assistant documentation discloses that "The AI Video Assistant is powered by OpenAI." It adds two more sentences worth reading in full: "Only prompts are transferred, and OpenAI doesn't use this data to train its models." And: "Synthesia retains prompts for abuse and misuse monitoring for a maximum of 30 days, then automatically deletes them unless otherwise required by law." That's a specific, named third-party model behind one part of the product, not an assumption — the newer Assistant doesn't publish the same disclosure, so don't carry that detail over to it.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.
Create An AccountThe prompt is the easy part once you know which of Synthesia's four systems you're actually writing for. The harder discipline is upstream of any tool: write for a listener, not a reader, and never treat a real person's likeness as raw material you can approve on their behalf.