Kitta Audio
Products
Products
Product overview
Explore the all-in-one creative suite
Voice library
Voices for any role or character
Create
Text to Speech
Generate human-like AI speech
Multi-speaker Dialogue
Create dialogue audio from multi-character scripts
Speech to Text
Transcribe audio and video
Voice Design
Generate custom voices
Voice Changer
Coming soon
Output audio in any voice
Voice Isolation
Extract clear speech
Voice Cloning
Clone your voice
Sound Effects
Coming soon
Generate any sound
Dubbing
Coming soon
Localize audio content
Music
Coming soon
Turn ideas into songs
Images
Generate images from text
Video
Generate video from text or images
S2
Fish Audio S2.1 Pro
Multi-speaker, multi-turn generation with natural language control over voice performance.
API
Platform
OverviewDocsAPI referenceAPI keysAPI pricingAPI Playground
API
Text to Speech
Generate speech through the API
Music
Coming soon
Create songs through the API
Speech to Text
Batch transcribe speech
Sound Effects
Coming soon
Generate sound effects through the API
Real-time Speech to Text
Transcribe speech in real time
Voice Cloning
Clone voices for TTS
Speech Engine
Coming soon
Give agents voice capabilities
Agents
Coming soon
Deploy voice agents in minutes
Dubbing
Coming soon
Translate video and audio through the API
API
Fish Audio API quick start
Debug voice generation, transcription, and account keys online
Text to Speech
Convert text to natural speech with Fish Audio, MiniMax, Qwen, and more
Speech to Text
High-accuracy transcription from uploaded audio
Voice Cloning
Clone your voice in about a minute from short samples
Voice Gallery
Browse public models and pick a reference voice
AI Image
Generate images from prompts with leading models
AI Video
Create video from text descriptions and styles
Lip-sync & digital human
Align speech to video for avatars and presenters
Voice Workspace
Voice synthesis workspace to create and manage your voice projects
Short video & dubbing
Fast voiceover for social, ads, and UGC
Audiobooks & podcasts
Long-form narration with natural pacing
Education & training
Clear narration for courses and internal comms
Company
AboutBlog
Resources
Coze
Tavo
SillyTavern
Dify
Open WebUI
AnythingLLM
Home Assistant
n8n
Affiliate program
Coming soon
API Playground
Try REST endpoints online with your API key
API keys
Create and manage API keys in your account
Pricing

Make every voice feel more alive

Sign up

Fish Audio, MiniMax, Qwen and more leading voice models in one workspace. Compare, switch, clone and export—a more flexible, cost-effective AI voice solution for creators, developers, and teams.

10/1000
Credit consumption preview
Fish Audio S2.1 Pro Flash

1 Chinese character = 1 credit, other characters = 0.5 credits

Emotion tags (e.g. (happy), (sad)) are excluded from credit calculation

Characters24
Multiplierx0.5
Total6
10/1000
Credit consumption preview
Fish Audio S2.1 Pro Flash

1 Chinese character = 1 credit, other characters = 0.5 credits

Emotion tags (e.g. (happy), (sad)) are excluded from credit calculation

Characters24
Multiplierx0.5
Total6

Generated Audio

No generated audio yet

Powered by Fish Audio S2.1 Pro
Unlock all audio features

Kitta Audio Demo

Experience Kitta Audio's ultra-realistic AI voice cloning for your own or licensed audio, powered by Fish Audio's AI voice technology

Create, edit, and localize AI voice content in one workspace

Try it now

Generate lifelike speech, turn scripts into voiceovers, clone expressive voices, and prepare audio for videos, audiobooks, podcasts, and global campaigns.

Welcome to today's episode. In the next few minutes we'll walk through how AI voice tools are changing the way creators produce audiobooks, explainer videos, and podcasts.

Stay with us for a quick demo: cloning a voice from a short sample, then generating narration in another language.

Describe the scene and generate a cinematic voiceover ...

Unified AI editor

Write scripts, design scenes, and generate polished voiceovers from the same creation flow.

On the ancient Eudoria plains, the sky burned gold while the forest wind whispered secrets. A dragon named Zephyros watched the horizon, calm, wise, and bright as an old star.

Chinese
Amy
Play

Ultra-realistic speech

Create controllable, expressive voices in 83 languages for every story.

Music

Generate background music by style, scene, mood, voice, or instrument.

Sound effects

Create custom effects, ambience, and transitions, or search your sound library.

Voice color

Clone voices, design character tones, or explore thousands of voice styles.

Image and video

Prepare voiceovers and localized audio for video, short drama, and animation.

Kitta Audio API

Build every voice workflow with one powerful API

View docs

Text to Speech API

Production-ready speech synthesis with model selection, stability controls, low latency, and multilingual output.

Fish Audio S2.1 Pro

Expressive, controllable voices with broad language coverage for premium content.

MiniMax Speech 2.8

Expressive, stable voices for character dialogue and content narration.

Qwen Audio 3.0

Cost-effective voices for large-scale generation.

Speech to Text API

Accurate ASR for audio and video, with speaker-aware transcripts and minute-level billing.

Fish Audio ASR

Reliable multilingual transcription for audio and video.

Lip Sync API

Create synchronized mouth movement from audio and video for dubbing and avatar workflows.

const response = await fetch('https://kittaai.com/api/open/tts', {
  method: 'POST',
  headers: {
    Authorization: 'Bearer YOUR_API_KEY',
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({
    reference_id: 'YOUR_VOICE_MODEL_ID',
    text: 'Turn this product walkthrough into a clear, natural voiceover.',
  }),
});

const audioBlob = await response.blob();
S2.1 Pro
MiniMax Speech 2.8
Qwen Audio 3.0
const response = await fetch('https://kittaai.com/api/open/v1/media/lip-sync/jobs', {
  method: 'POST',
  headers: { Authorization: 'Bearer YOUR_API_KEY' },
  body: JSON.stringify({
    video_url: 'https://example.com/video.mp4',
    audio_url: 'https://example.com/audio.mp3',
  }),
});

const task = await response.json();

Kitta Audio Core Features

Professional Voice Cloning Technology

Create a reusable voice from a short sample while preserving tone, delivery, and character details for branded or role-based audio.

Smart Text to Speech

Turn scripts into clear speech for narration, courses, podcasts, commercial voiceovers, and long-form reading.

Multilingual AI Voiceover

S2.1 Pro supports 83 languages, so one voice can power localized videos, audiobooks, and course content.

Professional Audio Processing

Built-in noise reduction, loudness balancing, and enhancement reduce post-production work and keep output consistent.

Fast Generation

Cloud generation and batch processing fit high-volume workflows, with short clips ready in about 20 seconds.

Wide Applications

Use it for short dramas, comics, video narration, audiobooks, courses, podcasts, and game characters.

Flexible Pricing

Choose the best plan for your text-to-speech needs

Free Plan

 
$0/chars
Free
3 daily guest trial generations
1000 credits on registration
Basic system voices
AI voiceover max length 200 chars
No credit card required

Plus Monthly

 
$4.49/month
Member benefits
Voice cloning and custom voices
All system voices
AI voiceover max length 10K characters
All advanced AI features including multi-speaker scripts
Priority support
Subscription credits
20K credits monthly (balance can keep accumulating)
At least 40K AI voiceover characters
At least 100 speech-to-text minutes
At least 400 lip-sync video seconds
At least 125 AI images
At least 30 AI videos
Popular

Pro Monthly

 
$14.99/month
Member benefits
Voice cloning and custom voices
All system voices
AI voiceover max length 10K characters
All advanced AI features including multi-speaker scripts
Priority support
Subscription credits
100K credits monthly (balance can keep accumulating)
At least 200K AI voiceover characters
At least 500 speech-to-text minutes
At least 2000 lip-sync video seconds
At least 625 AI images
At least 150 AI videos

Max Monthly

 
$34.99/month
Member benefits
Voice cloning and custom voices
All system voices
AI voiceover max length 10K characters
All advanced AI features including multi-speaker scripts
Priority support
Subscription credits
300K credits monthly (balance can keep accumulating)
At least 600K AI voiceover characters
At least 1500 speech-to-text minutes
At least 6000 lip-sync video seconds
At least 1875 AI images
At least 450 AI videos

Need higher quota or customization? Contact our business support

Compare every public service, model, billing unit, and credit rate before you generate. View all credit rules →

Kitta Audio FAQ

Learn about languages, voice cloning, API access, pricing, and production use cases

Powered by Fish Audio S2.1 Pro, Kitta Audio supports text to speech and multilingual voiceover in 83 languages, including English, Chinese, Japanese, Korean, Spanish, French, German, Russian, Arabic, and more. In most cases, you can provide text in the target language and let the model handle language detection and generation.

Use clear audio from your own voice or a voice you are licensed to use. Reduce background noise, room echo, and overlapping speakers. Short samples are useful for quick testing, while longer and more consistent samples usually help preserve tone, pacing, and delivery.

Kitta Audio is useful for short dramas, comic dubbing, video narration, YouTube or TikTok content, audiobooks, podcasts, courses, game characters, and multilingual localization. It is especially helpful when scripts change often or when teams need batch generation across many voices or languages.

AI text to speech is faster for drafts, batch production, localization, and repeated script revisions because you do not need to schedule studio time for every change. Human voice actors are still valuable for final performances that require precise acting direction. Many teams use AI first for testing and scale, then reserve human recording for selected final assets.

Commercial use depends on the current plan, usage policy, and voice rights. You can use the free allowance to evaluate quality and workflow. For ads, courses, games, films, client work, or other production use, use voices you own or have permission to use and follow the active paid plan and terms.

Yes. Developers can integrate text to speech, voice models, and audio generation through the Fish Audio API workflow. In most cases, you select the target model in the request and use it for prototypes, content tools, automated dubbing, or multilingual product experiences.

Use clean paragraph breaks, natural punctuation, and clear context for the speaking style. Avoid very long unstructured input. Start with a short preview, then adjust text, tone prompts, speed, volume, and voice selection before generating the full asset.

Product

  • Text to Speech
  • Multi-speaker Dialogue
  • Speech to Text
  • Voice Design
  • Voice ChangerComing soon
  • Voice Isolation
  • Voice Cloning
  • Sound EffectsComing soon
  • DubbingComing soon
  • MusicComing soon
  • Images
  • Video

Solutions

  • Short video & dubbing
  • Audiobooks & podcasts
  • Education & training

Research

  • Fish Audio S1
  • Fish Audio S2 Pro
  • Fish Audio S2.1 Pro

Resources

  • Docs
  • API reference
  • Model library
  • Voice clone tutorial
  • Product comparison

Company

  • About
  • Blog
Kitta AudioPowered by Fish Audio
© 2026 Kitta Audio. All rights reserved.Kitta AI|Privacy Policy|Terms of Service|Report abuse|support@kittaai.com
support@kittaai.com

Kitta Audio is independently operated by Kitta AI, not the official Fish Audio service.