# Speech AI Muse connector

From musedirectory.ai, the independent directory of Meta Muse connectors. Not affiliated with Meta.

## Speech AI

Pronunciation scoring, speech-to-text, and text-to-speech for English language learning

- Record: https://musedirectory.ai/connector/speech-ai
- Category: Lifestyle
- Developer: fasuizu-br (https://brainiall.com)
- Muse status: Extra setup. Not in Muse's Connectors list yet. Muse can still use it: its page gives you a request to paste into Muse.
- Health: Working, 1900ms, checked 2026-09-28T08:16:13Z
- Endpoint: https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp
- Auth: No account needed; Pricing: unknown
- Screening: Screened, no issues found (2026-09-25T07:21:17Z)
- Source: Found in the official MCP Registry (io.github.fasuizu-br/speech-ai) https://registry.modelcontextprotocol.io/v0/servers?search=io.github.fasuizu-br%2Fspeech-ai

Assess English pronunciation quality from audio with phoneme-level scoring. Transcribe speech to text with timestamps. Generate natural speech from text using 12 English voices. Supports multiple audio formats and includes multilingual transcription via Whisper.

Example request: "Score my pronunciation of this English sentence and tell me which words I need to practice."

How to connect: Not in Muse's Connectors list yet, but Muse can still use it. Paste this into Muse: "Use Speech AI to help me. It is a free service with an MCP server at https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp. It does not need an API key. Ask me before you share anything with it." Muse asks before it shares anything with the app's site. Meta does not review apps used this way, so only use ones you trust. We tested this in the Muse app on September 24, 2026: Muse used an app's link directly this way and returned a live answer.

Tools:
- assess_pronunciation: Assess English pronunciation quality from audio. Scores pronunciation at four levels: overall, sentence, word, and phoneme. Each score is 0-100. Phonemes are returned in both IPA and ARPAbet notation. Sub-300ms inference latency. Args: audio_base64: Base64-encoded audio data. Sup
- check_pronunciation_service: Check if the pronunciation assessment service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the scoring model is loaded - version (str): API version
- get_phoneme_inventory: Get the full phoneme inventory supported by the pronunciation scorer. Returns a list of all English phonemes the engine can assess, including ARPAbet symbol, IPA equivalent, example word, and phoneme category (vowel, consonant, diphthong). Returns: list of dicts, each with keys: 
- transcribe_audio: Transcribe audio to text with word-level timestamps. Converts spoken English audio into text with optional word-level timestamps and per-word confidence scores. Args: audio_base64: Base64-encoded audio data (WAV, MP3, OGG, FLAC, WebM). audio_format: Audio format hint. Auto-detect
- check_stt_service: Check if the speech-to-text service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the STT model is loaded - version (str): API version
- synthesize_speech: Generate natural speech audio from English text. Produces high-quality speech with 12 English voices. Returns base64-encoded WAV audio (16-bit PCM, 24kHz mono) along with metadata. Available voices: - af_heart (default), af_bella, af_nicole, af_sarah, af_sky (American female) - a
- list_tts_voices: List all available text-to-speech voices with metadata. Returns: dict with keys: - voices (list): Available voices, each with id, name, gender, accent, grade - defaultVoice (str): Default voice ID
- check_tts_service: Check if the text-to-speech service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the TTS model is loaded - version (str): API version
- transcribe_audio_pro: Transcribe audio with Whisper Large V3 Turbo — multilingual STT. Supports 99 languages with automatic language detection, word-level timestamps, per-word confidence scores, and optional speaker diarization (identifies who spoke each word). Best-in-class WER (~2%). Args: audio_bas
- check_whisper_service: Check if the Whisper STT Pro service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the Whisper model is loaded - diarizeLoaded (bool): Whether the diarization pipeline is loaded - version (str): API version -

Screening checks:
- MCP handshake: pass (Answered in 527ms)
- Domain against threat feeds (Cloudflare security DNS): pass (apim-ai-apis.azure-api.net, brainiall.com not flagged)
- Published packages against the OSV malicious-package database: n/a (No npm or PyPI package published)
- Hidden instructions or invisible characters in tool text: pass (10 tools read, nothing found)
- Inputs asking for passwords, card numbers or seed phrases: pass (None found)
- Domain and redirects: pass (No redirects off the domain)
- AI review of purpose and tool behavior: pass (No concerns)
