Anomaly Detection in Rotating Machinery Using Self-Supervised Learning
A university research presentation turned into a narrated video for academic sharing.
Voice: ElevenLabs Eleven v3 / George
Listen to real Japanese and English clips from ElevenLabs v4 Turbo/v4 and Gemini 3.8 Flash-Lite/Flash. Compare models with the same script and voice, plus methodology, use cases, and API pricing.
Compare real MP3 samples from the ElevenLabs v4 Turbo/v4 and Gemini 3.8 Flash-Lite/Flash text-to-speech APIs. Switch between Japanese and English, and compare each provider’s two models using the same voice and script. No account or API key is needed.
Shared script: With this service, you can create a presentation video from slides in a short time.
ElevenLabs v4 Turbo · Janet
Open audio fileElevenLabs v4 · Janet
Open audio fileGemini 3.8 Flash-Lite · Leda
Open audio fileGemini 3.8 Flash · Leda
Open audio fileYou can switch among five ElevenLabs voices and 30 Gemini voices here. The providers do not share voice identities, so cross-provider listening is not a strict same-voice test. Start with Janet and Leda, then change voices within each provider to separate model differences from voice character.
These are prerecorded files generated through the APIs used by SpeechSlide AI. Within each language, all four models received the same short sentence without an extra speaking-style instruction. Each pair within a provider uses the same voice setting. The page does not regenerate audio on every visit.
What This Test Can and Cannot Show
| Model | Vendor positioning | Where to start |
|---|---|---|
| ElevenLabs v4 Turbo | Expressive real-time synthesis with low latency | Interactive apps and responsive narration |
| ElevenLabs v4 | Voice quality and emotional expression | Carefully produced explainers and stories |
| Gemini 3.8 Flash-Lite | Speed, throughput, and cost efficiency | Large slide libraries and frequently updated training |
| Gemini 3.8 Flash | Voice fidelity, acting, and dialects | Narrative and brand videos with directed delivery |
This table translates vendor positioning into likely use cases; it is not a measured ranking from the clips above. Both providers support Japanese.
The published prices below were checked on October 3, 2026. ElevenLabs charges by input characters; Gemini uses input text tokens and generated audio tokens. These numbers cannot be compared directly as a per-clip cost.
| Model | Published API rate on Oct 3, 2026 | Listed rate after offer |
|---|---|---|
| ElevenLabs v4 Turbo | $0.011 / 1,000 characters | $0.04 / 1,000 characters (Oct 12) |
| ElevenLabs v4 | $0.022 / 1,000 characters | $0.08 / 1,000 characters (Oct 12) |
| Gemini 3.8 Flash-Lite | $0.50 input + $6 audio output / 1M tokens | $1 input + $12 audio output / 1M tokens (Jan 1, 2027) |
| Gemini 3.8 Flash | $0.50 input + $9 audio output / 1M tokens | $1 input + $18 audio output / 1M tokens (Jan 1, 2027) |
The Gemini figures are from Google Cloud Agent Platform. Its offer through 2026 is delivered as a 50% credit on eligible spend. Billing may differ on Gemini Developer API or other contracts. At roughly 25 audio tokens per second, the listed output rates correspond to about $0.0015 for 10 seconds on Flash-Lite and $0.00225 on Flash, excluding input. ElevenLabs cost depends on script length. Check current official pricing, plan terms, taxes, and add-ons before purchasing.
ElevenLabs positions v4 for quality and emotional delivery, and v4 Turbo for expressive real-time use. Hold the voice and script fixed, then alternate between the two clips above.
Start with Flash-Lite for high-volume, cost-sensitive work and Flash when expressive delivery matters, then compare using your real script.
The new-model samples use Janet. SpeechSlide AI maps its legacy Rachel voice ID to Janet, so the two are not counted as separate voices here.
For a full voice gallery, visit the voice comparison page. SpeechSlide AI lets you choose any of these four models in a slide’s voice settings, then narrate your generated script. New slides default to Gemini 3.8 Flash-Lite with Leda.
Try the voices and models with a slide deck you already use.
Create a Video for FreeJust upload your slides to get a narrated presentation video like this one.
A university research presentation turned into a narrated video for academic sharing.
Voice: ElevenLabs Eleven v3 / George
Explore slide-to-video workflows across education, healthcare, training, sales, and creator use cases.
View Use CasesYes. The free plan lets you create up to 2 projects and 4 videos per month. No credit card required.
PDF and PowerPoint (PPT/PPTX) files are supported. Upload the slides you already have.
AI narration is available in Japanese, English, Chinese, Korean, German, Spanish, and more.
Yes. Videos can be used for training, sales, lectures, and marketing. Paid plans allow watermark-free MP4 downloads.
A complete guide to speech generation in the Gemini API: using it in Google AI Studio, Python code examples, embedded Japanese audio samples, the free tier and pricing model, and how it compares with ElevenLabs and others — as of August 2026.