Skip to content

Speech Recognition Providers in Hedy

What are Speech Recognition Providers?

Hedy supports multiple speech recognition options, giving you flexibility to choose between complete privacy with local processing or cloud-based alternatives. You can switch providers anytime based on your current needs - use local for offline sessions and cloud services when you prefer their specific features.

Getting Started

  1. Open the Hedy app

  2. Navigate to Settings (tap your profile icon)

  3. Scroll to “Speech Recognition Options”

  4. Select your preferred provider from the dropdown menu

  5. Configure provider-specific settings if needed

  6. Your selection takes effect in the next session

Available Providers

Hedy offers five speech recognition options, each with unique characteristics:

  • Whisper (local): Default option - 100% private, works offline, no usage costs. Your audio never leaves your device for transcription. Available on every platform Hedy runs on.

  • Nemotron (local): A newer on-device streaming engine with live transcripts and on-device speaker labels. You choose between an English-only mode (the fastest option) and a multilingual mode that covers a broad set of major languages. Available on every platform Hedy ships a native app for: Apple Silicon Macs, iPhone XS (or newer), iPads with a Neural Engine (iPad Pro 2018, iPad Air 3, iPad mini 5, iPad 8th generation, or newer), Windows, and Android. On Apple hardware it runs on the Neural Engine and labels speakers live; on Windows and Android the labels are added at the end of the session. Requires a one-time model download (about 0.6 GB for English-only, 0.7 GB for multilingual).

  • Deepgram: Cloud-based service with real-time streaming and smart formatting features. Uses Nova-3, which supports dozens of languages. Hedy exposes every language Nova-3 offers, so you can transcribe meetings in any supported language without switching providers, including Swiss German, Armenian, Gujarati, Nepali and Punjabi, plus regional variants of Dutch, French, Portuguese, Spanish and Chinese. Requires your own API key.

  • OpenAI: Cloud transcription with Voice Activity Detection and automatic language detection. Live sessions run on OpenAI’s current generation, gpt-live-transcribe, and imported files on gpt-transcribe; the older models stay selectable if you had one saved. Hedy automatically continues long sessions past OpenAI’s 60-minute per-connection cap by rotating connections behind the scenes, so hour-plus meetings keep going without interruption. Requires your own API key.

  • xAI: Cloud transcription on xAI’s grok-voice-transcribe, with live streaming for sessions and batch transcription for imported files. It covers 25 languages, and your custom vocabulary is sent along as hints, so names and product terms come back spelled the way you want them. Requires your own API key.

Configuring Whisper (local)

When using Whisper, you can optimize for your device and needs:

For macOS Users:

  • Small Model: Fastest processing, recommended for Intel Macs

  • Regular Model: Balanced speed and accuracy for most users

  • Large Model: Enhanced capabilities for non-English languages (requires 1.5GB download)

For iOS/Android Users:

  • Standard Model: Default option suitable for most devices

  • Large Model: Alternative model option (iPhone 12+ or 2024+ Android recommended)

Voice Activity Detection (VAD):

VAD automatically filters out silence and background noise to improve transcription quality. This feature is enabled by default for Whisper.

  • Enable/Disable: Toggle VAD on or off based on your recording environment

  • Sensitivity: Adjust from “High Sensitivity” (captures more speech, including quieter sounds) to “Maximum Filtering” (only captures clear speech, filters more background noise)

Transcript Speed Settings:

  • Slower: Waits for complete sentences before displaying

  • Normal: Balanced speed and display timing

  • Faster: Near real-time display with more frequent updates

Configuring Nemotron (local)

It transcribes entirely on-device and shows live transcripts as you talk. It’s available on every platform Hedy ships a native app for: iOS, iPadOS, macOS, Windows, and Android. On Apple hardware it runs on the Neural Engine.

Device requirements:

  • Apple Silicon Mac (M1 or newer), or

  • iPhone XS or newer, or an iPad with a Neural Engine (iPad Pro 2018, iPad Air 3, iPad mini 5, iPad 8th generation, or newer), or

  • Windows, or

  • A 64-bit Android phone

English-only or multilingual:

In the provider dropdown, Nemotron appears as two choices, so you can pick the one that matches your meetings:

  • Nemotron English Only (local): streaming English transcription, the fastest option.

  • Nemotron Multilingual (local): on-device streaming across a broad set of major languages, for when you need more than English.

Both run fully on-device, and both identify language from the audio rather than from your meeting language setting.

First-time setup:

  1. Select Nemotron English Only (local) or Nemotron Multilingual (local) from the provider dropdown

  2. Tap Download Nemotron model (about 0.6 GB for English-only, 0.7 GB for multilingual) - we recommend Wi-Fi

  3. Once the download finishes, Nemotron is used automatically in your next session

Speaker labels and the temporary audio cache:

Nemotron labels who’s speaking. On Apple hardware the labels appear live as you talk; on Windows and Android they are added at the end of the session. To make those speaker labels more accurate, Hedy keeps each session’s audio in a temporary on-device cache while it processes, then deletes it. This audio stays on your device. The setting, Temporary audio cache (Nemotron), is on by default; you can turn it off in Hedy’s settings, though leaving it on gives Nemotron the best speaker attribution.

Setting Up Cloud Providers

Deepgram Setup:

  1. Create an account at console.deepgram.com

  2. Generate an API key from your dashboard

  3. In Hedy Settings, select Deepgram from the dropdown

  4. Paste your API key and tap “Test” to verify

  5. Choose your model and language preferences

  6. Set maximum session duration to control costs

OpenAI Setup:

  1. Get your API key from platform.openai.com/api-keys

  2. In Hedy Settings, select OpenAI from the dropdown

  3. Enter your API key and test the connection

  4. Choose your preferred model. Live sessions default to gpt-live-transcribe, and imported files use gpt-transcribe

  5. Optionally enable Voice Activity Detection with adjustable sensitivity

  6. Set maximum session duration for cost control

xAI Setup:

  1. Get your API key from console.x.ai

  2. In Hedy Settings, select xAI from the dropdown

  3. Paste your API key

  4. Set maximum session duration for cost control

Choosing the Right Provider

Select based on your priorities and use case:

  • Privacy First: Use a local engine (Whisper or Nemotron) - audio never leaves your device for transcription

  • Offline Use: All local engines work without internet

  • Cloud Features: Deepgram, OpenAI and xAI offer cloud-based processing

  • Voice Detection: Whisper and OpenAI include Voice Activity Detection features

  • Smart Formatting: Deepgram offers automatic formatting options

  • No Usage Costs: Local engines (Whisper, Nemotron) have no per-minute charges

  • Faster On-Device Transcription: Nemotron typically delivers a lower-latency transcript than Whisper

  • Multilingual On-Device Streaming: Nemotron Multilingual gives you on-device transcription across a broad set of languages

  • Best On-Device Engine for Your Language: See the table below

  • Fully Private Analysis: On macOS (Apple Silicon) or Windows, you can pair local speech recognition with Local AI Processing to keep both transcription and AI analysis fully on-device.

Which On-Device Engine to Use for Your Language

Whisper and Nemotron are not equally accurate in every language. Find your Meeting/Class Language below to see which one to pick. Where both engines handle a language well, we recommend Nemotron, because it is lighter and faster.

A few notes that apply to the whole table:

  • Nemotron means Nemotron Multilingual (local). For English, pick Nemotron English Only (local) instead.

  • Whisper Large means Whisper with Large selected as the model on a Mac or Windows PC. It is the most accurate on-device option for several languages, but it uses more memory and runs slower than Nemotron. If it is too slow on your computer, Nemotron is the next best choice for those languages.

  • Whisper Large isn’t available on phones and tablets. Where the table says Whisper, turn on Use Large Model if your device supports it.

  • If your device can’t run Nemotron (an Intel Mac or an older iPhone, for example), use Whisper.

  • Accents, dialects, people talking over each other and distant microphones lower accuracy in every language.

Meeting languageMac or Windows PCiPhone, iPad or Android
EnglishNemotron English OnlyNemotron English Only
ArabicNemotronNot available on mobile
DutchNemotronNemotron
FrenchNemotronNemotron
GermanNemotronNemotron
HindiNemotronNot available on mobile
ItalianNemotronNemotron
JapaneseNemotronNemotron
KoreanNemotronNemotron
PortugueseNemotronNemotron
RussianNemotronNemotron
SpanishNemotronNemotron
TurkishNemotronNemotron
UkrainianNemotronNemotron
VietnameseNemotronNemotron
Chinese (Mandarin)Whisper LargeNemotron
CroatianWhisper LargeNemotron
CzechWhisper LargeNemotron
DanishWhisper LargeNemotron
FinnishWhisper LargeNemotron
HungarianWhisper LargeNemotron
NorwegianWhisper LargeNemotron
PolishWhisper LargeNemotron
RomanianWhisper LargeNemotron
SlovakWhisper LargeNemotron
SwedishWhisper LargeNemotron
AfrikaansWhisper LargeWhisper
CantoneseWhisper LargeWhisper
CatalanWhisper LargeWhisper
GreekWhisper LargeWhisper
HebrewWhisper LargeWhisper
IndonesianWhisper LargeWhisper
MalayWhisper LargeWhisper
SerbianWhisper LargeWhisper
SlovenianWhisper LargeWhisper
TagalogWhisper LargeWhisper
ThaiWhisper LargeNot available on mobile

Expect noticeably lower accuracy in Afrikaans, Cantonese, Hebrew and Serbian than in the other languages, whichever engine you use.

Cost Considerations

Understanding the cost implications of each provider:

  • Whisper (local): Free - no usage charges

  • Nemotron (local): Free - no usage charges (one-time model download, about 0.6-0.7 GB)

  • Deepgram: Pay-per-minute pricing (check current rates on their dashboard)

  • OpenAI: Usage-based pricing (check current rates on their platform)

  • xAI: Hourly pricing, charged separately for live sessions and imported files (check current rates on their platform)

The maximum session duration setting helps prevent accidental overnight recordings and manage API costs.

Best Practices

  • Start with Whisper (local) to familiarize yourself with the feature, then try Nemotron if your device is supported

  • Test cloud providers with short recordings before important sessions

  • Monitor your API usage on provider dashboards to track costs

  • Use different providers for different scenarios based on your needs

  • Switch to local when traveling or in areas with limited internet

  • Set appropriate maximum session durations (60-120 minutes for typical meetings)

Troubleshooting

API Key Not Working

  • Ensure you copied the complete key without spaces

  • Verify your account has available credits

  • Check the API key has necessary permissions

  • Try regenerating the key from provider dashboard

Connection Test Failed

  • Check your internet connection stability

  • Verify firewall isn’t blocking WebSocket connections

  • Ensure API key is active with sufficient quota

  • Wait a moment and try again (temporary service issues)

Transcription Issues

  • For Whisper: Try a different model size

  • For Whisper on Windows: If transcription lags far behind the conversation, check slow transcription GPU settings

  • For specialized terms, names, and acronyms: Add them via the custom vocabulary feature

  • For Nemotron: Use the English Only mode for English meetings; for other languages, use the Multilingual mode or switch to Whisper with the language set explicitly

  • For Cloud: Check internet connection stability

  • Ensure microphone is properly configured

  • Minimize background noise during recording

Settings Not Saving

  • Wait for the “Saved” indicator to appear

  • Don’t switch screens while saving

  • Restart the app if issues persist

  • Ensure you have a stable internet connection

Your API keys are stored securely in your device’s encrypted keychain and never transmitted to Hedy’s servers. For maximum privacy with sensitive conversations, always use a local engine (Whisper or Nemotron).