Build your host as a set of reference images first. Then do the voice and the lip sync in one place. ElevenLabs handles both, so there's no exporting audio to a second tool and hoping the sync lands.
Publish the full episode as audio to your RSS feed and animate three to five short clips for social. That's not a cost decision, it's a behaviour one. Nobody watches twenty minutes of an animated face.
You have a podcast idea. You know your subject. You have no interest whatsoever in setting up a camera every week, doing your hair, and filming yourself.
Eighteen months ago the answer to that involved stitching three or four tools together and a lot of file management. Most of that has collapsed into one workflow, which is the single biggest reason this format went from fiddly to genuinely doable around a job.
What changed
The old process was: write in one place, generate voice in another, export the audio, upload it to a video tool, hope the lip sync landed, export again. Every handoff was a place to lose time or quality.
Now the voice and the lip sync happen in the same place. You feed in your character image and your script, and ElevenLabs generates the speech and syncs the mouth to it in one pass. No separate audio export, no third-party sync step, and tighter timing because the two were never pulled apart in the first place.
Two things about this matter more than they sound.
The character is still yours. You're not picking a stock avatar off a shelf. You bring your own reference images, the ones you generated and locked yourself, and the tool animates those. That distinction is the whole ball game: a stock avatar makes your show look like everyone else's, and reference images you built make it look like nobody else's. It also means the consistency work is on you rather than the platform, which is why locking your reference pack properly matters before episode one rather than after episode six.
Flows chains the steps. The node-based canvas connects image and video models with the audio stack in one pipeline, so a script runs through to a finished clip without you shepherding files between tabs. Change one node, swap a voice, try a different model, re-run without rebuilding. For a weekly show, that repeatability is the difference between a format you sustain and one you abandon in March.
Audio for the feed, clips for social
Even with full-length video now easy, I'd still publish audio as the episode and animate only short clips. This is about how podcasts are actually consumed.
People listen to a podcast while doing something else. Driving, walking the dog, loading the dishwasher. That's a twenty minute audio behaviour. Watching a talking head for twenty minutes is not a behaviour anyone has, animated or otherwise.
So the split I'd run:
| Output | Where it goes | Why |
|---|---|---|
| Full episode audio | RSS, Apple, Spotify | The show itself. Background listening |
| 3 to 5 clips, 30 to 60 sec | Reels, TikTok, Shorts | Discovery. This is what brings new listeners |
| Static or light video episode | YouTube, optional | Searchability, not watch time |
If you'd rather animate the full episode now that it's straightforward, nothing stops you. Just be honest that you're doing it for completeness rather than because anyone is watching it to the end.
Decide who your host is before you open anything
The most common mistake is jumping straight to a generator. Generic input, generic output.
Settle three things first:
- How she looks. Age presentation, colouring, build, wardrobe register, the world she sits in.
- How she sounds. Warm and conversational, or measured and authoritative. This has to match how she looks. A mismatch between a face's apparent age and a voice's age is one of the few things that still reads as uncanny even when everything else is right.
- What she thinks. The perspective that shows up in every episode. That part is yours and no tool supplies it.
Get these down before you build the avatar, because the avatar is easy to make and hard to un-choose once you've published six episodes with her.
Build your host in about ten minutes
The Avatar Quiz settles all three decisions in twelve questions and hands you a finished prompt pack: her reference images plus her first podcast and yap stills. Built for exactly this, and free.
The weekly workflow
1. Write it spoken, not written
Short sentences. Contractions. The way you'd say it to one person. Formal written prose sounds noticeably worse through text-to-speech and you'll hear it immediately.
Read the whole thing aloud before you generate anything. Every awkward sentence surfaces the moment you say it yourself, and fixing it at draft stage costs nothing. If you're doing a two-hander, label your lines so voices get assigned cleanly.
2. Lock your character images once
Generate your host's reference set in one session and save it somewhere you'll find it: a front-facing portrait, a three-quarter turn, and a couple of in-world shots. Flat lighting, neutral expression, plain background. This is the one-off. Every episode from here reuses those same files.
Name them properly. In three months you will not remember which of eleven images was the one that worked.
3. Generate voice and lip sync together
Feed in your character image, your script and your voice, and let it generate the speech and sync the mouth in one pass.
Design that voice rather than taking one from the library. You write a description of how your host sounds, her age, accent, tone and pace, and it generates a few versions to choose between. A library voice is one other shows are already using, and on a podcast the voice is the show, so it's worth the two minutes.
Designing is not cloning, incidentally. Cloning copies a real person and brings consent paperwork with it. Designing invents a voice that never existed, so there is nobody to ask.
Preview a short section before committing to a long render.
4. Level the audio
Export your master at WAV, 48kHz. Then check loudness: Apple Podcasts sits around -16 LUFS with peaks below -1 dBFS, Spotify around -14 LUFS. Auphonic or Descript will normalise it without you needing to learn audio engineering.
5. Cut your clips
Pull three to five moments of 30 to 60 seconds. The strongest opinion, the most surprising fact, the bit where you disagree with something popular. Those three tend to outperform anything you'd pick for being "useful".
6. Export the right files
- MP3, 128 to 192 kbps, mono, 44.1kHz for the RSS feed. This is the published file.
- WAV master archived separately as your quality reference, never the published one.
- MP4, 1920x1080 for YouTube, and 9:16 vertical for the social clips.
Add captions. Most people watch clips with the sound off, and on a podcast clip that's the whole message gone if there's nothing on screen.
What it costs
ElevenLabs Starter is around $6 a month and Creator around $22, with 100,000 characters, professional voice cloning and generations up to twenty minutes. Creator is the tier most weekly shows will want.
Video generation consumes credits on top of your character allowance, and how far a plan stretches depends on how much video you make and at what length. Check the current numbers before you commit, because this pricing has moved several times in the past year and anything I write here has a shelf life.
You'll also want an image tool for generating the character set itself, though that's a one-off rather than a monthly cost once your references are locked.
Podcast hosting is often free. Spotify for Creators, Libsyn and Podbean all take a standard MP3 and distribute to Apple, Spotify and Amazon automatically.
The UK legal bit, briefly
Short section, real risks, don't skim it.
Disclosure on the feed. There's no single blanket UK duty to label AI content, but transparency is strongly advisable anywhere a listener could reasonably think they're hearing a real person. A line in your show notes and a brief mention in the intro is enough. In my experience it builds trust rather than costing you any, and audiences in this space are already familiar with the format.
Disclosure on the clips, which is not optional. The moment your clips go to Reels and TikTok you're under platform policy rather than general advice. TikTok requires the AI-generated toggle for photorealistic AI people and AI voices, detects synthetic media automatically through C2PA credentials whether you disclose or not, and issues immediate strikes: removal for the first, seven days off posting for the second, thirty for the third. Instagram has an account-level AI creator label in profile settings plus per-post disclosure. Turn both on and stop thinking about it.
Consent for cloned voices. If you clone anyone's voice, including your own, keep written consent covering the specific use, the duration and the platforms. Performers' rights and copyright can both apply to source recordings, and UK GDPR principles can apply where voice or likeness data is involved.
Never put words in a real person's mouth. An AI voice attributing statements to someone real, or implying an endorsement they never gave, creates defamation and false endorsement exposure. Only make your host say things you wrote and authorised. The more convincing the output, the more seriously this applies.
That's roughly twenty minutes of admin once, then a line in your show notes forever.
The short version
- Voice and lip sync in one tool now. No exporting audio between platforms.
- Your own reference images are the character. Not a stock avatar, which is what stops your show looking like everyone else's.
- Flows chains script to finished clip in one pipeline, which is what makes weekly sustainable.
- Audio for the feed, clips for social. Nobody watches twenty minutes of a face.
- Decide how she looks, how she sounds and what she thinks before you build her.
- Disclose in your show notes. Consent for any cloned voice. Never attribute words to a real person.
The whole point of this format is that there's no filming day to wait for. The episode gets made on a Tuesday night after the kids are in bed, which is exactly when mine get made.