FunFluenLearn

Pronunciation Practice with a Screen Reader

Use a screen reader and keyboard-only Word dictation workflow to check one pronunciation target without treating speech-to-text as an accent score.

The short answer

Use one vendor-tested Windows workflow—Word for the web dictation with a screen reader and keyboard—then treat the transcript as a recognition clue, not a pronunciation score.

The awkward part is often not the pronunciation target. It is being handed a “self-check” whose microphone button, waveform, and result exist only visually. So this guide does not give you a giant list of supposedly accessible pronunciation apps. It stays inside one lane that Microsoft explicitly says it has tested with Narrator, JAWS, and NVDA: keyboard-controlled dictation in Word for the web.

The method: SAY → HEAR THE TRANSCRIPT → CLASSIFY → RETRY

The whole practice loop is deliberately small:

  1. Choose one word, phrase, or short sentence and one narrow word-recognition hypothesis to watch.
  2. Dictate it once in Word for the web using the keyboard-only sequence below.
  3. Use your screen reader to read the resulting transcript.
  4. Classify the result instead of scoring yourself.
  5. Retry once under the same conditions if the result is unclear.

The central rule is simple: the transcript is a witness, not a judge. If Word writes the expected word, that means the recognizer recovered that word in that take. It does not prove that every sound or accent feature was correct.

A plain transcript cannot show lexical stress, rhythm, pitch, or intonation. If one of those is your target, use a different independently accessible feedback source—such as trusted model audio or qualified human feedback—instead of trying to infer prosody from whether Word typed the expected word.

What you need for this verified lane

Microsoft’s current accessibility instructions say Word for the web dictation requires a Microsoft 365 subscription, a headset, and a reliable internet connection. Microsoft recommends Edge for Word for the web. It also notes that Microsoft 365 features roll out gradually, so Dictate may not appear identically for every account.

  • Windows with Narrator, JAWS, or NVDA.
  • Word for the web in an editable document.
  • Microsoft 365 subscription.
  • Headset or working microphone.
  • Reliable internet connection.
  • Microsoft Edge is the browser Microsoft recommends for this workflow.

Microsoft says that with Narrator and JAWS, full-screen mode should be used for Word for the web; with NVDA, regular screen mode can be used. Its page also contains a legacy Windows 10 Narrator note about turning off scan mode while editing. Do not generalize that one old-version note into a universal screen-reader rule—follow the instructions for your current screen reader and Windows version.

The keyboard-only Word for the web pronunciation check

This section uses only the Word for the web commands from Microsoft’s current accessibility page. Desktop Word uses different commands, so do not copy this web checklist into the desktop app; use Microsoft’s desktop instructions from the same source instead.

  1. Open an editable Word document and place focus in the document.
  2. Press Alt + Windows logo key + H, then D. Microsoft says this opens the Dictate menu.
  3. Press D. Wait for the audio cue that tells you dictation is listening.
  4. Say your short target sentence once.
  5. To stop dictating, press Alt + Windows logo key + H, D, then D. Wait for the stop cue.
  6. Use your normal screen-reader document-reading/navigation commands to read the text that Word inserted.

Those Dictate commands come from Microsoft’s accessibility documentation, which explicitly says the Word for the web workflow was tested with Narrator, JAWS, and NVDA. The exact way your browser first asks for microphone permission can vary by browser, Windows version, and previous permissions, so this guide does not invent a universal permission-dialog sequence.

Read Microsoft’s current screen-reader dictation instructions for Word.

Classify the transcript—do not turn it into a score

After the screen reader reads your transcript, choose the description that best fits the two takes. This is the decision step that prevents one recognition mismatch from becoming instant learner guilt.

Which result best describes your test?

Try it with one sentence

Use this sentence:

I need three tickets for Thursday.

For this particular transcript check, use one narrow hypothesis: the initial TH in three. You are not asking Word to prove that the sound is correct. You are checking whether a repeated recognition pattern gives you a reason to verify that sound independently.

If your actual target is the stress pattern in Thursday, stop this test there. A plain text transcript cannot tell you whether you placed lexical stress correctly. Use an independently accessible audio model or qualified human feedback for that job.

  1. Dictate the sentence once.
  2. Stop dictation and have your screen reader read the transcript.
  3. Clear the line or move to a new line.
  4. Say the same sentence again under roughly the same recording conditions.
  5. Compare the two transcripts only for the chosen word-recognition pattern.
Example A: “three” appears correctly twice

Classify this as recognized in this condition. Do not upgrade the result to “my TH is definitely correct.” Speech recognition can use the whole word and sentence context, so successful transcription is not sound-level acoustic analysis.

Example B: “three” changes the same way twice while the rest stays stable

Classify this as a possible pronunciation target, not a confirmed error. Independently verify the target sound before drilling it.

Example C: the two transcripts disagree in several places

Classify this as recognition uncertainty. An unstable recognition pattern is poor evidence for assigning yourself a pronunciation problem.

Why one wrong transcript is not proof that you said it wrong

Speech-to-text is a recognition system, not a phonetics examiner. A transcription mismatch does not tell you which part of the speech-recognition chain caused the mismatch. Microphone input, sentence context, recognition behavior, language settings, and your production can all matter, so one mismatched word is not enough to establish a pronunciation error.

Imagine you say, “My meeting starts at four,” and Word writes “My meeting starts at five” once. You would not immediately invent a pronunciation drill for four. You would repeat the same task, check whether the pattern is stable, and verify the word independently if the mismatch persists. That is the same discipline to use with three or another segmental/word-form hypothesis.

A useful self-check therefore asks:

  • Did the same word fail more than once?
  • Did the surrounding words stay stable?
  • Did recording conditions stay roughly the same?
  • Can an independent pronunciation source confirm the feature you plan to practise?

If not, keep the result out of your error list.

Your screen reader’s pronunciation is not a phonetic verdict either

There is a second layer of uncertainty here: text-to-speech itself can pronounce written words unexpectedly. The W3C Web Accessibility Initiative documents pronunciation problems in screen readers and other text-to-speech systems, especially where spelling, context, abbreviations, names, or specialized terms leave room for ambiguity.

So if your screen reader reads a proper name or unfamiliar word in a surprising way, do not automatically copy that pronunciation as the model. The screen reader is reading text; it is not analyzing the acoustic quality of the speech you just produced.

See the W3C’s overview of text-to-speech pronunciation issues.

A real screen-reader Word walkthrough, if you want orientation

If you are newer to Microsoft Word with a screen reader, VisionAccess has a public walkthrough covering Word setup with JAWS, NVDA, and Narrator. It is useful as orientation to the environment; it is not evidence that Word can judge pronunciation accuracy, and you do not need the video to complete the method in this article.

Watch the VisionAccess Microsoft Word screen-reader walkthrough on YouTube.

A note for NVDA users

The current NVDA 2026.2 documentation describes its normal Windows keyboard navigation, browse/focus behavior, Microsoft Word support, and UI Automation support for Microsoft Word and Chromium-based browsers. Those mechanics can help you understand why web pages and editable application controls may behave differently.

This article does not add extra NVDA keystrokes that Microsoft does not require for the Word Dictate task. If focus or editing mode behaves differently in your setup, use the current NVDA and Microsoft 365 documentation rather than a guessed shortcut from a pronunciation tutorial.

Open the current NVDA user guide.

Privacy and service limits

Microsoft’s current Word accessibility page says Dictate sends what you say to Microsoft to provide text results and says the service does not store your audio data or transcribed text. It links to Microsoft’s Connected Experiences information for more detail. This guide does not infer anything beyond that published statement.

Also remember the access limits: Word dictation is tied to Microsoft 365, needs an internet connection, and feature availability can change as Microsoft rolls out updates. That makes this a verified workflow, not a universal free solution for every screen-reader user.

When should you actually create a pronunciation target?

Create a practice target only when the evidence becomes specific enough to act on:

  • the same word-recognition mismatch repeats;
  • the surrounding sentence is otherwise recognized consistently;
  • your recording setup is reasonably stable; and
  • a trusted pronunciation model, teacher, dictionary audio, or other qualified source independently confirms the sound or word form you want to change.

Then practise that independently verified feature. Do not keep re-running dictation until a favorable transcript appears and call that learning. The point is to identify a plausible target, verify it, practise it deliberately, and later check it again in fresh language.

What this guide deliberately does not recommend

  • No list of pronunciation apps merely because they have an accessibility statement.
  • No visual waveform or color score as required feedback.
  • No claim that a successful transcript proves a sound, stress pattern, or accent is correct.
  • No attempt to infer stress, rhythm, pitch, or intonation from a plain transcript.
  • No claim that NVDA, Narrator, or JAWS analyzes pronunciation.
  • No assumption that blind or low-vision learners need different pronunciation goals.
  • No claim that FunFluen has been screen-reader certified by this guide.

Bottom line

Good screen-reader pronunciation practice does not need a visual score and it does not need an accessibility guessing game. Use one workflow whose keyboard and screen-reader operation has actually been documented and tested, keep the speech-to-text result modest, and separate “the recognizer recovered my word” from “my pronunciation is correct.”

One sentence. One target. Two takes. Then classify the evidence instead of blaming yourself—or blindly trusting the software.

Sources

Explore more language-learning guides in Media-Based Language Learning.