Bulk ad voiceovers with Claude and Fish Audio ⚡
<- Back to Resources
MARKETING & ADS

Bulk ad voiceovers with Claude and Fish Audio ⚡

Master this high-ROI AI workflow: This is the exact setup , upgraded into the full pipeline: Claude writes ten genuinely different ad scripts, Fish Audio v

AEO SUMMARY Quick Overview & Execution Blueprint

Master this high-ROI AI workflow: This is the exact setup , upgraded into the full pipeline: Claude writes ten genuinely different ad scripts, Fish Audio voices them, and then you give each voiceover a video to live in. Test a bunch, run the winner. Studio quality, no studio.

GROWTH & ACQUISITION NOTE Somya's Strategic Takeaway

Content and paid ads are distribution engines for your core value proposition. Don't create content just for engagement—engineer every hook to qualify high-intent buyers, demonstrate immediate proof, and lead into a frictionless capture mechanism. Track cost-per-qualified-lead (CPQL) over vanity impressions.

Most ad voiceovers either cost a fortune at a studio or sound like a robot reading a receipt. This skips both. And because a voiceover without a video is only half an ad, this guide now covers the whole thing: script, voice, visuals, and assembly.

THE STACK

The full stack at a glance 🧠

  1. ✍️ Claude writes and runs it. Ten ad scripts with ten different hooks, then it calls the Fish Audio API and generates the voiceovers itself, no extra apps.
  2. 🎙️ Fish Audio is the voice. It performs the emotion you tag inline, so the ads actually sound human.
  3. 🎥 The video layer comes last (or first). Talking-face or b-roll visuals from an AI video tool, and it works in both directions, voice-first or video-first. Step 4 covers both.

TEN REAL HOOKS

Step 1, the ad-script prompt 📝

This is where most bulk-generated ads die: you ask for ten scripts and get the same ad wearing ten hats. This prompt forces ten genuinely different hooks and a proper short-form structure (hook in the first 3 seconds, one idea per ad, CTA word for word). Paste it into Claude with your details dropped in:

📝 Ten scripts, ten hook types
You are a direct-response copywriter who writes for the ear, not the page. Here is my product and offer: [paste yours here]. My audience: [who it's for]. The one action I want: [e.g. click the link / DM us / buy today].

Write me 10 ad voiceover scripts, each 15 to 30 seconds read aloud (roughly 40 to 75 words). Rules:

1. Structure every script as: hook in the first 3 seconds, the problem in one line, my product as the fix with ONE concrete benefit, one line of proof or believability, then the call to action word for word.
2. Every script opens with a DIFFERENT hook type. Use each of these exactly once: direct callout (name the exact person or pain, like "if you run ads and your costs just doubled..."), bold claim, question, myth-bust, curiosity gap, mini story, stat shock, mistake warning ("stop doing X"), before-and-after, social proof.
3. Make the benefit land like Hormozi's value equation: name the dream outcome, make it believable, and shrink the time and effort ("in minutes, not weeks").
4. Write for the ear: short words, contractions, no jargon, nothing you wouldn't actually say out loud.
5. Mark emotion tags inline exactly where the read should change, like [excited] on the hook, [serious] on the problem, [whisper] on the close. Vary the emotional arc across the ten, they should not all peak in the same place.
6. Label each script with its hook type. If two scripts could swap hooks without anyone noticing, rewrite one.

Output: numbered 1 to 10, spoken words only, no camera directions.

Why the hook rule matters: on short-form, the first three seconds decide everything, and ad fatigue is really hook fatigue. Ten scripts with ten hook types means you're testing ten actual hypotheses, not one ad ten times.


ONE CONNECTION

Step 2, connect Fish Audio to Claude 🔌

Claude can call Fish Audio directly through its API, so it generates the audio without you touching another tool.

  1. 🔑 Make a free account at fish.audio and grab your API key. Put it in a .env file, never paste it into the chat.
  2. 🔌 Then ask Claude to connect and generate:
🔌 Generates all ten voiceovers
Connect to the Fish Audio API using the key in my .env file, and never print the key. For each of the 10 ad scripts above, generate a voiceover with the free S2.1 Pro model (model s2.1-pro-free), keep the emotion tags inline so the reads are not flat, and save them as numbered audio files I can listen through.
Heads up: the Fish Audio S2.1 API is free for developers until July 31, so you can test all of this for nothing right now.

THE EMOTION LAYER

Step 3, make it sound human 🎭

The difference between obviously-AI and wait-that's-a-real-ad is the emotion tags. You type the feeling straight into the script:

🔥 [excited] on the hook to grab attention
🤫 [whisper] on the close to pull people in

Fish Audio performs the tag instead of reading flat, so every one of your ten reads has real energy.


THE VISUAL LAYER

Step 4, give the voiceover a video to live in 🎥

A voiceover is half an ad. Here are both directions, pick the one that matches where you're starting from.

🅰 Voice-first (you have your ten MP3s):

  • 🗣 Want a talking face? Use Open Generative AI, a free open-source studio: open its Lip Sync Studio, upload a portrait (a real photo or an AI-generated presenter) plus your Fish Audio MP3, and it outputs a talking video synced to your read. The app is free; generations run on your own prepaid key at cents per clip.
  • 🎬 Want product shots or cinematic b-roll instead? Generate the visuals in Higgsfield: its Marketing Studio builds product-ad visuals from your product image, and Cinema Studio gives you cinematic shots with real camera controls. Generate one clip per script beat, then cut them under the voiceover in step 5.

🅱 Video-first (you already have a visual that works):

PROMPT
Here are the beats of my video with timestamps: [e.g. 0-3s product close-up, 3-8s hands using it, 8-14s result shot, 14-18s logo]. Write an ad voiceover script that hits each beat at the right moment, using the same hook and emotion tag rules as before. Keep it under [X] seconds total, then generate it in Fish Audio.

Claude times the script to your footage, Fish Audio reads it, and the voiceover lands on your cuts instead of fighting them.

Which order should you use? Voice-first when the message carries the ad (offers, testimonials, explainers). Video-first when the visual is the hook (a product demo, a transformation, anything that stops the scroll on its own).


PUT IT TOGETHER

Step 5, assemble and ship 🎬

  1. 📥 Drop the video and the voiceover into CapCut (or your editor), one script per cut.
  2. 💬 Auto-captions on, always. Most feeds play muted until you earn the sound-on.
  3. 🔊 One sound effect per visual reveal, nothing extra.
  4. 📱 Export 9:16 for Reels and TikTok, then test three variants at a time, changing one thing each (usually the hook).

GO FURTHER

3 bonus prompts to run next 🎁

🌍 Localize a winner.

🌍 Make a winner sound native
Rewrite my top 3 ad reads for [country or market], adjusting the slang, references, and tone so they sound native, then regenerate them in Fish Audio.

🎯 Scroll-stopping hooks.

🎯 Fifteen more openers
Give me 15 different opening lines for this ad, each under 4 seconds, designed to stop the scroll. Mark the emotion tag for each.

📊 Pick what to test first.

📊 Rank before you spend
Based on my product and audience, rank these 10 ad angles by which is most likely to convert, and tell me which 3 to test first and why.

USE IT WELL

How to get the most out of it 🎓

🌱 New to this generate just 3 reads first, not 10, and listen back so you can hear the difference the emotion tags make.
🔁 Already running ads feed Claude your current best ad and ask for 10 variations on that winner, then test them head to head.

WHAT TO WATCH

The honest part 🫶

AI voiceovers are genuinely good now, but they are not magic. The script still has to be a good ad, which is exactly why Claude writes ten different hooks instead of one, so you can test and let the winner emerge. On the video side, budget honestly: Higgsfield is a paid tool, and Open Generative AI charges cents per generation through your own key, so test cheap (one image-based clip) before rendering the expensive version. Always watch the full ad back with sound on and off before you publish, your own eye and ear are the final check.

The links 🔗

🎙 Fish Audio: fish.audio
🎬 Higgsfield (Marketing Studio + Cinema Studio for the visual layer)
🗣 Open Generative AI: hosted version · official repo (Lip Sync Studio for talking-face ads)
✂️ CapCut: capcut.com (assembly + captions)

⚡ SCALING AI & GROWTH SYSTEMS?

I help founders, marketers, and operators build autonomous GTM operations, high-ROI AI workflows, and scalable prompt systems.

Book a Growth Consultation ->
Hey, I'm Mini Somya. Click me to chat.