Speech & pronunciation assessment MCPYour agent can hear them.Now it can grade them.
Chivox MCP turns raw speech into a dense, agent-ready payload — phoneme scores, stress, tone and fluency in one MCP call, ready for any LLM.
From speech to the next best practice
Chivox handles the acoustic judgment. Your LLM receives structured evidence it can explain, reason over and turn into the learner’s next action.
- 01/Speech inCapture the learner
- 02/Evidence outReturn acoustic detail
- 03/Next actionLet the agent respond

Score guided speech
Stream live audio or post a file. Get overall, word and phoneme-level evidence in one response.

Evaluate open dialogue
Score free-flow responses across fluency, content, grammar, accuracy and rhythm — turn by turn.

Diagnose English and Mandarin natively
Inspect tones and pinyin in Chinese; stress, rhythm and CEFR-aligned evidence in English.

Turn evidence into the next practice
Give the structured JSON to any LLM to coach, route or generate targeted drills for the next turn.
Acoustic depth you can inspect. Scale you can trust.
Twenty years of speech-assessment R&D, exposed through one stable contract. Toggle zh / en to inspect the same pron.* / details[] structure; use the benchmarks beside it to sanity-check Chivox against your own evaluation harness.
你好,今天天气……
Scores align with certified human expert rubrics at 95%+ correlation. Validated by national standardized speaking tests in 100+ cities.
- Per-dimension rubrics: pron, fluency, completeness, prosody.
- Calibration corpus refreshed quarterly across L1/L2 cohorts.
- Stable across mic quality, room noise and child voices.
Get the first structured score in 3 steps
Paste the config, connect Chivox, then call one assessment tool from your agent loop.
Add one block to your MCP configrunning
Paste the snippet into Cursor, Claude Desktop, or your custom agent — pick a tab on the right.
Call a tool from your LLM
Hand your model the audio. It gets back nested JSON: pron sub-scores, fluency + WPM, audio SNR, and details[] with ms ranges, stress, liaison and per-phoneme rows.
API referenceStart with the learner you are building for
English, Mandarin and Kids share one MCP contract, while each product path gives your agent the language and learner context it needs.
Simple points for successful speech evaluations.
One point for a word or sentence. Two for a paragraph. Failed calls use zero points, so you only pay when Chivox returns an assessment.
- Successful evaluations only
- Shared across every API key
- Points stay valid for 30 days
From free trial to production
- 1600 pts
Start free
Valid for 30 days
- 2
Evaluate successfully
Points are deducted only when an assessment returns.
word · sentence−1 ptparagraph−2 pts - 3+20%
Top up as you grow
Higher packs lower your unit cost.
Same payload. Your agent. Your production loop.
Drop Chivox MCP into Cursor, Claude Desktop, or any agent SDK. One npx and you’re reading the same JSON you just saw above.
Free trial · spend caps · low-balance alerts · zero audio retention



