Skip to content
Transcriwise
One platform, many engines

9 ASR families. The right engine for every file.

Transcriwise is not locked to a single provider. Select manually or let the system pick the engine best suited to the context — volume, language, domain, and privacy.

9 integrated families6 Live optionsAutomatic context-based selection

Guided selection

Filter by the capability that actually matters.

The ideal provider changes with the recording and expected result. Use filters as a starting point, not a quality guarantee.

Whisper Large V3 Turbo

Fast & Accurate

Whisper model optimized for speed

High-performance Whisper model with an integrated diarization pipeline. Ideal for long audio with technical vocabulary and results in minutes.

  • Long audio (>60 min) with high fidelity
  • Diarization integrated into the pipeline
  • Hallucination filter (Gemini Flash)
  • Technical vocabulary via initial_prompt
  • Batch support and result streaming

Best for:

Court hearings, long lectures, extended interviews

Speakers: Unlimited

AssemblyAI

Multi-speaker

High-accuracy universal model

Cloud API with a high-accuracy universal model. Stands out for speaker identification by name or role and smart result caching.

  • Speaker identification by name or role
  • File-hash caching — avoids reprocessing
  • Streaming upload with parallel processing
  • Advanced diarization with editable labels
  • Keyterms prompt for specialized vocabulary

Best for:

Corporate meetings, testimony with multiple participants

Speakers: Unlimited

ElevenLabs Scribe

Audio Events

Advanced sound event detection

Cloud API with Scribe v2 that detects non-verbal events (laughter, music, applause) and supports up to 32 speakers. Ideal for subtitle generation.

  • Event detection: laughter, music, applause
  • Up to 32 simultaneous speakers
  • Word-level timestamps for precise sync
  • Files up to 3 GB / 10 hours
  • Keyterms (Scribe v2) for specialized terms

Best for:

Subtitle generation, multimedia content, podcasts

Speakers: 32

Soniox

Built-in Translation

Simultaneous transcription and translation

Async API with automatic language identification and simultaneous translation into the target language. Domain context for specialized vocabulary.

  • Automatic translation into any target language
  • Automatic language identification
  • Domain hints (medical, legal, financial, engineering)
  • Precise segmentation with per-word timestamps
  • Ideal for multilingual content

Best for:

Multilingual content, translated subtitles, international meetings

Speakers: Unlimited

Deepgram Nova-3

Domain Models

Domain-specialized models

Cloud API with domain-specialized models (medical, legal, financial, engineering). Automatic smart formatting adds punctuation, paragraph breaks, and speaker turns.

  • Dedicated models: medical, legal, financial, engineering
  • Automatic smart formatting (punctuation and paragraphs)
  • Nova-3 — Deepgram's newest, highest-accuracy model
  • Automatic language detection
  • High-capacity batch processing

Best for:

Medical reports, financial documents, formal minutes

Speakers: Unlimited

Gladia

Multilingual

Code-switching and 100+ languages

API with native code-switching (automatic language switching), diarization, and 100+ languages. All features included at no extra cost.

  • Native code-switching across 100+ languages
  • Diarization included in every plan
  • Batch (up to 135 min) and live via WebSocket
  • PII redaction and sentiment analysis
  • Custom vocabulary and SRT subtitles

Best for:

Multilingual meetings, language-switching content, international events

Speakers: Unlimited

Voxtral (Mistral)

High Fidelity

Native diarization and context biasing

Mistral's file-transcription model with native diarization, per-word timestamps, and context biasing with up to 100 terms.

  • Mistral's proprietary model, tuned for difficult audio
  • Native diarization with per-word timestamps
  • Context biasing with up to 100 terms
  • Native support for 13 languages including Portuguese
  • Audio up to 3 hours per request

Best for:

High-fidelity transcriptions, difficult audio, technical vocabulary

Speakers: Unlimited

iFlytek

Chinese

Specialist in Mandarin and Cantonese

Engine focused on Mandarin and Cantonese. Supports batch and live transcription with Chinese→English translation in compatible flows.

  • Speech recognition specialized in Mandarin Chinese
  • Live simultaneous Chinese→English translation
  • Support for Chinese dialects and Cantonese
  • Basic and advanced (LLM) engines available
  • Batch with chunked upload for long audio

Best for:

Chinese content, conferences in China, zh→en simultaneous translation

Speakers: Unlimited

Alibaba Cloud

Asia/Global

Qwen3-ASR, Paraformer, and SenseVoice

Alibaba's model family with Qwen3-ASR for multilingual content, Paraformer optimized for Chinese, and self-hosted SenseVoice.

  • Qwen3-ASR: 30 languages + 22 Chinese dialects
  • Paraformer: a low-cost option for Chinese
  • SenseVoice: self-hosted with emotion detection
  • Audio up to 12 hours per request
  • Distinct models for Chinese, English, and multilingual content

Best for:

Low-cost Chinese, multilingual Asia, emotion detection

Speakers: Unlimited

Feature comparison

Compare what each engine supports.

A quick view of each engine and model's declared capabilities. Availability also depends on mode and plan.

FeatureWhisperAssemblyAIElevenLabsSonioxDeepgramGladiaVoxtraliFlytekAlibaba
Speaker diarization
Per-word timestamps
Audio events
Built-in translation
Domain models
Custom vocabulary
Speaker ID by name
Live transcription
Available modelsTurbo (default)Universal-3 Pro (default)Scribe v2 (default)STT Async v4 (default)Nova-3 (default)Pre-recorded v2 (default)Mini Transcribe V2 (default)Advanced LLM (default)Qwen3-ASR (default)

Formatted flows may apply dictionaries, normalization, and punctuation according to mode. Raw and Live follow their own rules.

Configurable post-processing

Formatting complements the engine when the flow requires editorial output.

Post-processing depends on job mode and options; it does not change the raw transcript or replace human review.

PT-BR Legal Dictionary

50+ automatic corrections of common errors in legal terminology

Acronym Normalization

STF, STJ, TRF, TRT, OAB, LGPD, CLT, and 30+ recognized acronyms

Punctuation Restoration

Automatic punctuation around statute articles, paragraphs, and legal citations

Recurring-phrase flags

Heuristics flag recurring ASR phrases for review (e.g., "thanks for watching")

Upload your first recording.

Automatic routing weighs language, mode, and availability — or pick a compatible provider yourself. Usage depends on plan, duration, provider, and optional features.

Transcriwise