Fish Audio Image & Video Generation: Lip-Sync, Creative Suite 2026
Fish Audio now generates images, videos, and lip-synced talking-head clips. How the Fish Creative suite works, what it costs, and whether the lip-sync is good…
12 min read

Fish Audio was a voice tool. Text-to-speech, voice cloning, maybe sound effects if you dug around. Then they bought Thank You AI and now the same account that clones your voice also generates images, makes videos, and lip-syncs a talking head to your audio. They call it Fish Creative.
I have been using Fish Audio for TTS on YouTube for a while. When the creative suite dropped, I spent a weekend testing it. Some of it surprised me. Some of it did not. Here is what I found.
Try Fish Audio FreeWhat Fish Creative includes
Fish Creative is not a separate product. It lives inside your existing Fish Audio account. When you log in, you get access to:
- Text-to-speech — the S2.1 Pro model with emotion tags, 83 languages, 2M+ community voices
- Voice cloning — 10-15 seconds of audio, about two minutes to clone
- Image generation — text-to-image for character portraits, backgrounds, and scenes
- Video generation — turn images or text prompts into short video clips
- Lip-sync — sync any audio to an image or video of a face, producing a talking-head clip
- Sound effects and music — generate audio assets from text prompts
- Voice changer — transform existing audio into a different voice
- Story Studio — multi-speaker, long-form audio production for audiobooks and narration
- Asset Library — every upload and generation lands here, reusable across tools without re-uploading
- Avatars — save a character as a persistent avatar inside Image & Video for repeat lip-sync jobs
The voice side I already covered in my Fish Audio review. This article focuses on the visual tools: image generation, video generation, and lip-sync.
Image generation
The image generator inside Fish Creative produces character portraits and scenes from text prompts. It is not trying to compete with Midjourney or Flux on artistic quality — the images are meant to be used as lip-sync targets and video frames.
What I noticed after generating maybe 50 images:
Straight-on portraits work well. If you need a face looking at the camera, the generator produces clean, consistent results. Skin texture, lighting, and facial proportions are solid. These are the images that feed best into the lip-sync pipeline.
Full-body and action poses are hit or miss. Hands still get weird, and complex poses sometimes produce distorted limbs. Same problem every image generator has, but Fish Creative has not solved it either.
Style control is limited. You can describe what you want, but there is no style selector, no negative prompts, no aspect ratio picker visible in the UI (that I found). You get what the model gives you. For production work where you need a specific look, you might still want to generate the image in Flux or Midjourney and upload it to Fish Creative for the lip-sync step.
Speed is decent. A single image takes about 10-20 seconds. Not instant, not painful. There is also a Topaz HD upscale option for single images when you want more detail out of a generation.
My honest take: the image generator is good enough for talking-head frames. If the face is what matters and the background is secondary, it does the job. If you need a polished, stylized illustration, generate it elsewhere and bring it in.
Video generation
Fish Creative can generate short video clips from text prompts or from a reference image. The video generation is where the Thank You AI integration shows most — it was their core product before the acquisition.
What you can do:
- Text-to-video: describe a scene, get a short clip (a few seconds)
- Image-to-video: upload a still image, animate it into a video
- Lip-sync video: take an image with a face, sync it to audio, get a talking-head video
The general video generation (text-to-video, image-to-video) produces short clips — think 3 to 6 seconds. Quality is on par with other mid-tier AI video tools. Motion is smooth for simple scenes (a person turning their head, a flag waving) but gets choppy on complex movement.
For the kind of content I make — YouTube narrations, explainer clips — the lip-sync pipeline is the real draw. I rarely need a generated video of a fox running through a forest. I need a face that talks.
What I actually use
I generate the voice in Fish Audio (S2.1 Pro, emotion tags), create or upload a character image, run lip-sync, and export. That is a complete talking-head video without leaving the platform. The text-to-video and image-to-video are nice extras, but lip-sync is the feature that matters for content creators.
Lip-sync
This is the feature I was most interested in, and the one I tested the most. Lip-sync takes an image of a face and an audio clip, and produces a video where the mouth moves in sync with the audio.
How it works in practice:
- Generate or upload a face image (straight-on works best)
- Generate or upload your audio (TTS, cloned voice, or recorded audio)
- Hit lip-sync
- Wait — takes about 30 seconds to a couple of minutes depending on audio length
- Download the video
What works:
- Mouth movement tracks the audio well. Jaw, lips, and tongue motion look natural for normal speech.
- The face stays stable. No flickering, no background drift, no sudden face swaps.
- Works with any audio source — your cloned voice, a community voice, uploaded recordings, even music (though the result is more entertaining than useful for music).
- Straight-on talking-head content looks professional enough for YouTube, social media, and presentations.
What is rough:
- Extreme head angles produce artifacts. If the face in your image is turned more than about 30 degrees, the lip-sync starts looking off.
- Wide-open expressions (shouting, big laughs) sometimes distort the jaw or cheeks.
- The generated video is short — limited by the audio length, and there is a practical ceiling around 1-2 minutes per clip.
- No body animation. Only the face moves. If you need hand gestures or head nods, you need a different tool.
Compared to HeyGen:
HeyGen is the tool most people know for AI talking-head videos. I have used both. HeyGen has better body animation, more avatar options, and a more polished editor. But HeyGen starts at $24/month and charges per minute of video. Fish Creative is included in your Fish Audio account, and the lip-sync works with any face image you bring — not just their pre-built avatars.
If you need a polished, corporate-style talking head with full body movement, HeyGen is still better. If you need a face that talks to match your TTS audio and you want it cheap, Fish Creative does the job.
Asset Library and Avatars
Two quieter additions that matter more than they sound. Everything you upload or generate now lands in the Asset Library — reference photos, generated portraits, finished clips — so you stop re-uploading the same files between tools. An image you just made feeds straight into video generation or lip-sync without a download round-trip.
Avatars take it further: inside Image & Video you can save a character as a persistent Avatar and reuse that same face across lip-sync jobs. If you are building a recurring presenter for a channel, this replaces the old generate-a-face-and-pray-it-matches loop.
The workflow I actually use
Here is my real workflow for making a talking-head clip:
- Write the script
- Generate the voice in Fish Audio with S2.1 Pro and emotion tags
- Generate a character portrait using the image generator (or upload a photo)
- Run lip-sync with the audio and image
- Download the MP4
- Drop it into my video editor as one layer
Before Fish Creative, this was: ElevenLabs for voice, Midjourney or Flux for the face, HeyGen or D-ID for lip-sync, and then syncing everything in a video editor. Four tools, four subscriptions, four logins. Now it is one.
The quality is not dramatically better or worse than the old stack. The win is convenience and cost. One subscription, one login, one dashboard.
Pricing
Fish Creative tools are part of the Fish Audio account. You do not pay separately for image generation or lip-sync — they consume credits from your existing plan.
| Plan | Price | What you get |
|---|---|---|
| Free | $0 | Limited TTS, community voices, limited creative tool access |
| Starter | ~$15/month | More generations, voice cloning, creative tools |
| Pro | ~$45/month | Higher limits, commercial rights, faster processing |
| Enterprise | Custom | Dedicated support, SLA, custom integrations |
The free tier gives you enough to test the lip-sync pipeline a few times. If you are producing content regularly, the Starter plan covers TTS plus the creative tools.
For developers, programmatic access to the creative tools finally exists, in two forms. The official MCP server (https://api.fish.audio/mcp, OAuth sign-in — claude mcp add --transport http fish-audio https://api.fish.audio/mcp in Claude Code) lets an agent generate and edit images and videos directly. It spends your plan’s package credits rather than API credits, quotes a price before each job, and asks you to confirm — video generation needs a paid plan. Separately, the newer Open API surface (fishaudio.org, not the classic api.fish.audio/v1 docs) exposes async job endpoints for image generation, lip-sync (POST /api/open/v1/media/lip-sync/jobs), and video dubbing — TTS plus lip-sync in one call. It is jobs-and-polling rather than the clean single POST that TTS uses, but the automation path is there now.
Programmatic access exists now
The classic api.fish.audio docs still only list TTS, speech-to-text, and voice-model endpoints. Creative tools are reachable through the MCP server (OAuth, plan credits, price confirmed before each job) and through the Open API’s async job endpoints for images, lip-sync, and video dubbing.
What I would change
Add aspect ratio control to image generation. Right now you describe what you want and hope the aspect ratio is close to 16:9. A simple picker would save time.
Document the creative APIs on the main docs site. Lip-sync and video dubbing now exist on the Open API surface and the MCP server handles image and video generation, but docs.fish.audio does not cover them yet. The plumbing is in — the docs just need to catch up.
Improve the lip-sync on angled faces. Straight-on works. Anything beyond 30 degrees starts looking wrong. This is a hard problem, but it is the main quality ceiling right now.
Longer video support. The practical limit around 1-2 minutes per lip-sync clip means you need to chop longer narrations into segments and stitch them. Not a dealbreaker, but adds friction.
Better documentation. The TTS docs at docs.fish.audio are solid. The creative suite docs are thin. I had to figure out the workflow by trial and error.
Who should use this
YouTubers and content creators who need talking-head videos without recording themselves. Generate a voice, generate a face, lip-sync, export. No camera, no lighting, no editing.
Social media creators making short-form content. The lip-sync produces clips that work for TikTok, Reels, and Shorts.
Anyone already using Fish Audio for TTS. If you are paying for Fish Audio, the creative tools are already there. Test them before paying for a separate lip-sync tool.
Developers and hobbyists experimenting with AI video. The free tier is enough to build a proof of concept.
Who should look elsewhere
If you need full-body animation, hand gestures, or scene composition, HeyGen or Synthesia are more mature. If you need cinematic video generation at production quality, look at Veo or Kling through Kie.ai. And if you need a fully documented REST API for the creative pipeline, the MCP server and Open API job endpoints cover the basics but the main docs site has not caught up.
Getting started
- Go to Fish Audio and sign up (free)
- Generate a voice or clone your own (15 seconds of audio, 2 minutes to clone)
- Navigate to the Creative section in the dashboard
- Generate or upload a face image
- Run lip-sync with your audio
- Download the video
The whole first test takes about 10 minutes. Free tier, no credit card. If the lip-sync quality works for your use case, the Starter plan at ~$15/month covers everything.
Try Fish Creative FreeIs Fish Creative free?
The free tier includes limited access to the creative tools — enough to test image generation and lip-sync a few times. For regular use, the Starter plan at ~$15/month covers TTS, voice cloning, and the creative suite.
Can I use Fish Creative for commercial videos?
Yes, on paid plans. The free tier is for testing. The Pro plan and above include commercial rights for generated content.
Fish Creative vs HeyGen: which is better?
HeyGen has better body animation, more avatar options, and a polished editor. Fish Creative is cheaper (included in your Fish Audio plan), works with any face image, and integrates directly with your TTS voices. For talking-head content on a budget, Fish Creative is the better value. For corporate presentations with full-body avatars, HeyGen wins.
Does lip-sync work with any audio?
Yes. You can use Fish Audio TTS, your cloned voice, uploaded recordings, or audio from other sources. The lip-sync engine only needs an audio file and a face image.
Is there an API for the creative tools?
Partially. The MCP server at api.fish.audio/mcp gives agents image and video generation with a price confirmation step built in. The newer Open API (fishaudio.org) has async job endpoints for images, lip-sync, and video dubbing. The main docs site still only documents TTS, ASR, and voice models.
My view
I was skeptical when I heard Fish Audio was adding image and video generation. Tools that try to do everything usually end up mediocre at all of it. But Fish Creative is not trying to replace Midjourney or Veo. It picks one problem — turn your TTS audio into a talking-head video — and does it without leaving the platform.
And for that, it works. The lip-sync is good enough for YouTube and social media. The image generator produces usable faces. The whole pipeline takes minutes instead of the 30-45 minutes I used to spend bouncing between four different tools.
The rough edges are real: thin creative-suite docs, limited aspect ratio control, lip-sync that degrades on angled faces. The API gap I complained about after launch is mostly closed — MCP covers image and video generation, and the Open API has lip-sync and dubbing job endpoints. The foundation is there, and having it all in one dashboard with my existing TTS setup is the part that keeps me coming back.
If you are making content with AI voices and want to add a face to those voices, try Fish Creative before you pay for a separate lip-sync tool. The free tier is enough to decide.
Related
- Fish Audio Review 2026 — full review covering TTS, voice cloning, emotion tags, pricing
- Fish Audio vs ElevenLabs 2026 — side-by-side comparison
- Clone your voice with Fish Audio in 2 minutes — voice cloning walkthrough
- Kie.ai Video Generation Guide — Veo, Kling, Seedance behind one API


