Home/Case Studies/AI Speech Assessment
Education

A score tells a learner they were wrong. It does not tell them what to do.

Real-time speech assessment where the engineering went into feedback that changes behaviour, and into not penalising learners for having an accent.

AI pronunciation and speech assessment for language learners
SectorEducation & EdTechAssessmentPhoneme level, real timeFeedbackCorrective, not just a scoreFairnessAccent-tolerant by designLearnersGlobal, many first languagesLatencyFast enough to feel immediate
The situation

Two ways speech assessment fails learners.

The first is uselessness. A percentage score tells a learner they were wrong without telling them which sound, in which position, or what to do differently. Learners repeat the same error and conclude they cannot improve.

The second is unfairness, and it is worse. Systems trained predominantly on one accent penalise learners for regional variation that is entirely acceptable in real speech. For a platform serving learners across many first languages, that is both a pedagogical failure and a commercial one; learners who feel unfairly marked leave.

Latency compounds both. Feedback that arrives after the learner has moved on is feedback about something they are no longer thinking about, and it does not change anything.

What we built

Assess at phoneme level, feed back correctively, and define fairness first.

Assessment operates at phoneme level, not on the utterance as a whole, so the system can identify which specific sound in which position was produced differently from the target. That granularity is the prerequisite for feedback that means anything.

Feedback is corrective rather than evaluative: which sound, where in the word, and a concrete articulatory instruction. The interface shows the learner's production against the target, not reporting a number.

Accent fairness was specified as a requirement before modelling started. The system distinguishes between variation that is acceptable in fluent speech and error that impedes comprehension, and only marks the second. That distinction is a pedagogical decision, not a technical one, and it was made with the educators.

Latency was treated as a feature with a budget. Feedback has to arrive while the learner is still attending to the attempt, which put a hard ceiling on the processing pipeline and shaped the architecture accordingly.

Inside the system

What the platform does.

01

Phoneme-level assessment

Per-sound evaluation in context, which is the only grain at which corrective feedback is possible.

02

Corrective feedback

Which sound, where, and what to do, not a percentage.

03

Accent tolerance

Acceptable variation distinguished from comprehension-impeding error, defined with educators.

04

Real-time pipeline

Latency budgeted so feedback lands while the learner is still attending to the attempt.

05

Progress tracking

Per-phoneme progress over time, so improvement is visible where it is happening.

06

Adaptive practice

Practice targeted at the sounds a specific learner struggles with, instead of a fixed curriculum.

07

Educator view

Cohort-level patterns, so teaching can respond to what a group finds hard.

Built with

Built with.

Speech

Forced alignmentPhoneme-level scoringStreaming audio pipeline

Modelling

Acoustic modelsAccent-aware evaluationConfidence calibration

Application

Learner practice interfaceVisual articulation feedbackProgress tracking

Platform

Low-latency servingEducator dashboardsCohort analytics
What changed

What changed.

  • Feedback became actionable. Phoneme-level identification with a concrete instruction, instead of a score the learner could not act on.
  • Learners stopped being penalised for their accent. Acceptable variation separated from comprehension-impeding error, defined with educators, not inherited from the training data.
  • Feedback arrives while it still matters. A latency budget treated as a product requirement rather than as an optimisation to do later.
  • Practice targets the actual weakness. Per-phoneme tracking drives what a learner is asked to practice next.