Pronunciation scoring, speech-to-text, and text-to-speech for English language learning
Assess English pronunciation quality from audio with phoneme-level scoring. Transcribe speech to text with timestamps. Generate natural speech from text using 12 English voices. Supports multiple audio formats and includes multilingual transcription via Whisper.
Try asking Muse: "Score my pronunciation of this English sentence and tell me which words I need to practice."
Source: Found in the official MCP Registry (io.github.fasuizu-br/speech-ai) · First listed September 25, 2026
Each bar is one check, every 15 minutes. Green means it answered. Last checked 53 min ago.
What Muse can see: It does not ask you to sign in, so it cannot see your accounts. It only sees what Muse sends it from your request.
Before it acts: Read what Muse plans to do before you approve it, and remove the app from Muse when you stop using it.
musedirectory.ai is not part of Meta. More about how Muse handles your information
Not in Muse's Connectors list yet, but Muse can still use it. Copy the request below and paste it into Muse. Muse asks before it shares anything with the app's site. Meta does not review apps used this way, so only use ones you trust. We tested this in the Muse app on September 24, 2026: Muse used an app's link directly this way and returned a live answer.
https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcpNot in Muse's Connectors list yet. Muse can still use it: its page gives you a request to paste into Muse.
Response time at the last check: 1900ms. Worked in 93.8% of checks over 30 days.
Screened, no issues found
Screening looks for known threats and hidden instructions at the time of the check, and runs again weekly and whenever the tool list changes. It cannot see the server's code, so only connect what you need and review what Muse asks to do. How screening works.
assess_pronunciationAssess English pronunciation quality from audio. Scores pronunciation at four levels: overall, sentence, word, and phoneme. Each score is 0-100. Phonemes are returned in both IPA and ARPAbet notation. Sub-300ms inference latency. Args: audio_base64: Base64-encoded audio data. Supcheck_pronunciation_serviceCheck if the pronunciation assessment service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the scoring model is loaded - version (str): API versionget_phoneme_inventoryGet the full phoneme inventory supported by the pronunciation scorer. Returns a list of all English phonemes the engine can assess, including ARPAbet symbol, IPA equivalent, example word, and phoneme category (vowel, consonant, diphthong). Returns: list of dicts, each with keys: transcribe_audioTranscribe audio to text with word-level timestamps. Converts spoken English audio into text with optional word-level timestamps and per-word confidence scores. Args: audio_base64: Base64-encoded audio data (WAV, MP3, OGG, FLAC, WebM). audio_format: Audio format hint. Auto-detectcheck_stt_serviceCheck if the speech-to-text service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the STT model is loaded - version (str): API versionsynthesize_speechGenerate natural speech audio from English text. Produces high-quality speech with 12 English voices. Returns base64-encoded WAV audio (16-bit PCM, 24kHz mono) along with metadata. Available voices: - af_heart (default), af_bella, af_nicole, af_sarah, af_sky (American female) - alist_tts_voicesList all available text-to-speech voices with metadata. Returns: dict with keys: - voices (list): Available voices, each with id, name, gender, accent, grade - defaultVoice (str): Default voice IDcheck_tts_serviceCheck if the text-to-speech service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the TTS model is loaded - version (str): API versiontranscribe_audio_proTranscribe audio with Whisper Large V3 Turbo — multilingual STT. Supports 99 languages with automatic language detection, word-level timestamps, per-word confidence scores, and optional speaker diarization (identifies who spoke each word). Best-in-class WER (~2%). Args: audio_bascheck_whisper_serviceCheck if the Whisper STT Pro service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the Whisper model is loaded - diarizeLoaded (bool): Whether the diarization pipeline is loaded - version (str): API version -https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcpNot in Muse's Connectors list yet, but Muse can still use it. Paste this into Muse: "Use Speech AI to help me. It is a free service with an MCP server at https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp. It does not need an API key. Ask me before you share anything with it." Muse asks before it shares anything with the app's site. Meta does not review apps used this way, so only use ones you trust. We tested this in the Muse app on September 24, 2026: Muse used an app's link directly this way and returned a live answer.
At the last check (September 28, 2026, 08:16 UTC) the endpoint was working, answering in 1900ms. Over the last 7 days it answered 93.8% of health checks. It is checked every 15 minutes.
It was screened on September 25, 2026 with the result "screened, no issues found". Screening checks the domain against threat feeds and reads the tools for hidden instructions and requests for passwords or card numbers. It cannot see the server's code, so grant only the access you need.
It exposes 10 tools, including assess_pronunciation, check_pronunciation_service, get_phoneme_inventory, transcribe_audio. For example, you could ask Muse: "Score my pronunciation of this English sentence and tell me which words I need to practice."
This listing was added from public sources (Found in the official MCP Registry (io.github.fasuizu-br/speech-ai)). If you build Speech AI, claim it to correct the details and get your badge.
Paste this on your site or README. It always shows the latest check.
<a href="https://musedirectory.ai/connector/speech-ai"><img src="https://musedirectory.ai/badge/speech-ai.svg" alt="Speech AI on musedirectory.ai" width="236" height="40"></a>