FunFluenLearn

English Listening Difficulty Ladder: A Practical 12-Week Plan

Build a 12-week English listening plan that raises one difficulty variable at a time, uses observable pass criteria, and shows when to step back.

The short answer

Make English listening harder by changing one controllable variable at a time while keeping the other five stable.

You understood yesterday’s English video. Today you tried something “more advanced” and heard soup. The tempting conclusion is: my listening is worse than I thought. A better question is: what changed? Length? Language load? Delivery? Speaker overlap? Variety familiarity? Support or signal? If you know which dial moved, the result can actually teach you something.

HOLD FIVE -> RAISE ONE -> PROVE IT -> REPEAT OR STEP BACK

That is the whole listening ladder in one line. The six dials are language load, length and structure, delivery, speaker load, variety familiarity, and support and signal. One harder thing, five stable things.

This is a practical twelve-week schedule, not a CEFR placement test, a biological timeline, or a promise that twelve weeks will produce a particular proficiency level. The goal is simpler and more useful: make each harder listening session interpretable.

The six variables you can stage

“Hard English” is not one thing. A clip can be difficult because the words are unfamiliar, because it is long, because the speech is reduced, because people interrupt each other, because you have had less exposure to a speaker’s variety, or because the recording itself is a mess. Those are different problems. Treating them as one giant difficulty score throws away the most useful information.

  1. LANGUAGE LOAD. Topic familiarity, vocabulary coverage, idiom density, cultural references, and how much meaning depends on knowledge you do not yet have. Research on lexical coverage in listening supports treating vocabulary familiarity as a genuine variable, but it does not give you a universal percentage that means “ready.”
  2. LENGTH AND STRUCTURE. Clip duration, number of ideas, signposting, and memory load. A learner can decode individual sentences and still lose the sequence once the message stretches across several points. Research comparing utterance length and speech rate supports keeping those two dimensions separate rather than calling both “fast listening.”
  3. DELIVERY. Speech rate, pausing, reduction, and clarity. Delivery is broader than the playback-speed button. Research on speech-rate training and learner control shows that changing speed or pause control changes the listening task, but it does not support one magic playback rate for a level.
  4. SPEAKER LOAD. One speaker, orderly turn-taking, interruptions, and overlap. Two clear speakers can create a different tracking problem from one equally clear speaker, even when vocabulary and speed stay the same.
  5. VARIETY FAMILIARITY. How much exposure you have had to a speaker’s regional, social, or international English variety. This is not an accent ranking. A variety can feel difficult because it is less familiar to you, because the source is noisy, because the topic is unfamiliar, or because several factors interact.
  6. SUPPORT AND SIGNAL. Visual context, transcript or captions, replay control, recording clarity, music, background noise, and reverberation. Studies of captions, playback conditions, noise, and acoustics all make the same practical warning useful here: the listening task changes when the signal or support changes. Do not quietly count a bad recording as evidence that your English got worse.

Your 30-second baseline

Before Week 1, describe material you can currently follow without turning this into a score. Short phrases are enough.

Six-variable baseline







That card is your starting configuration. It is not your level. You are about to change one part of the configuration and see what survives.

Change one at a time

Suppose your current material is a 45-second explainer on a familiar topic, spoken clearly by one presenter, with clean audio and a transcript available after listening. A terrible way to “level up” is to jump to a seven-minute panel on an unfamiliar subject with three speakers, dense slang, no checking surface, less familiar varieties, interruptions, and café noise. That is not a ladder. It is six trapdoors opening together.

A useful progression is much less dramatic:

Hold five. Raise one. Listen blind. Prove what you still understand. Then either repeat the rung or step back on the dial you moved.

If the difficult part is simply choosing a suitable source for today’s session, use this separate guide to choose one clip that is difficult enough for the current session. This article owns the multi-week progression after you have material to work with.

Likewise, playback speed is only one part of delivery. If your specific goal is to train that variable by itself, use the playback-speed-only ladder rather than turning this twelve-week plan into a speed program.

Your first ten-minute session

A rough ten-minute shape is enough: spend about one minute choosing the target and noting the profile, four minutes on the blind listen and recall, three minutes checking the dropout, and two minutes logging or retrying. Treat those times as a practical container, not a proficiency rule.

  1. Choose a short, familiar, well-signposted one-speaker source with clear audio and a reliable transcript or captions available for checking later.
  2. Write the six-variable profile in short phrases. Choose one detail target before playback: a reason, a name, a contrast, a sequence step, or another concrete detail.
  3. Listen once without reading the transcript or captions.
  4. Say or write the main interaction or claim, the sequence, the speaker’s purpose, and your chosen detail from memory.
  5. Now use the checking surface. Compare it with the model you built in your head; do not turn the check into a second first pass.
  6. Record one dropout: what specifically disappeared?
  7. Retry once. Then decide whether the same rung needs another clip or whether one variable should be easier.

Use public, authorised listening sources and their normal playback tools. You do not need to download or redistribute audio, captions, transcripts, or test material to run this method. And save listening practice for situations where listening is safe; this is not an exercise to do while driving or doing something that needs your attention.

Advance criteria per rung

The ladder needs proof, but not fake precision. “I understood 82%” sounds scientific until you ask what the 82% measured. Words? Ideas? Names? Intent? A transcript comparison? A guess?

Instead, use four observable questions:

  • Main interaction or claim: Can you state what is happening or what the speaker is mainly saying?
  • Sequence: Can you put the important events, reasons, or points in the right order?
  • Speaker purpose: Can you tell what the speaker is trying to do: explain, disagree, persuade, request, warn, joke, clarify, or something else?
  • Chosen detail target: Can you recover the one detail you decided to listen for before playback?

For this guide, the default consistency rule is: meet the rung’s criterion on two different clips before advancing. That rule is an editorial guardrail, not a research-established proficiency threshold. It reduces the chance that you “pass” because you memorised one source or happened to know its topic unusually well.

The Council of Europe’s CEFR work is useful here for a narrower reason: it describes reception through observable “can do” performance across different listening and audio-visual situations. It does not make this ladder CEFR-validated, and one home-practice clip should not be used to assign yourself a CEFR level.

Ready to advance?

Quick check for the current rung

You do not need every word. If you miss one adjective but correctly understand that a tenant is asking a landlord to repair a leaking pipe before the weekend, you may have preserved the important message. If you catch every noun but reverse who promised to do what, your sequence or speaker-role model has failed. That distinction is much more useful than chasing a percentage.

When to drop back

A failed rung is not an instruction to retreat three months. It is an instruction to inspect the one variable you just raised.

The failed-rung recovery protocol

  1. Name the variable that was raised on this rung.
  2. Make the source easier along that variable only. If length rose, shorten it. If speaker load rose, restore clearer turns. If support fell, restore the last reliable support condition.
  3. Keep the other five variables as close as possible to the previous successful profile.
  4. Use one check to locate the dropout: transcript/captions after the blind pass, a short replay, or one very short transcription if the failure is specifically decoding.
  5. Retry a different clip at the easier setting.
  6. Only after repeated failure across different clips should you conclude that the rung needs more time.

The important move is small: step back on the dial you moved, not on your whole identity as a listener.

What should you change next?

Open the symptom that best matches your blind listen. These are practice decisions, not diagnoses.

I get the gist, but important details keep disappearing.

Do not make the clip longer yet. Keep length, speaker load, variety, and support stable. First inspect language load: was the missed detail carried by unfamiliar vocabulary, an idiom, or a reference you did not know? If yes, step back to a more familiar language-load profile and keep one detail target. If the words were familiar but merged acoustically, inspect delivery instead. Retry on a different clip before changing a second variable.

I can catch details in short clips, but longer clips collapse.

That points first to length and structure. Shorten the segment or reduce the number of ideas while keeping topic, delivery, speaker load, variety familiarity, and signal stable. Your target is not more word recognition; it is preserving sequence and purpose for longer. When that holds on two different clips, extend again.

I understand my familiar presenter, but another speaker throws me.

Inspect variety familiarity before deciding that the new speaker is inherently “harder.” Keep topic, length, delivery, speaker count, and signal unusually stable. Build exposure with bounded clips from the less familiar variety. If the source is also noisy or faster, fix those differences first so you are not blaming variety for a signal or delivery change.

I am fine with one speaker, but two speakers make me lose who said what.

Step back on speaker load. Use two speakers with orderly, visible turn-taking before adding interruption or overlap. Keep the vocabulary, length, delivery, variety familiarity, and signal close to the profile you already passed. Your next target is a turn map: who wants what, who responds, and what changes after the response?

With captions I feel fine; without them I cannot build the message.

Inspect support and signal. If captions are appropriate and available to you, keep them as a reliable checking surface—but try moving them after the blind first pass for practice sessions where that is accessible. Do not remove captions and simultaneously choose harder language or faster delivery. If you need stable captions for access, keep them; the ladder is allowed to adapt to you.

I understand only when playback is slowed down.

Inspect delivery. Keep the same length, language load, speaker profile, variety familiarity, and signal while you work with clearer or slightly less demanding delivery. Do not add a new variable just because one replay worked. Slower playback can be a temporary practice control; it is not a proficiency label, and there is no universal speed that proves readiness.

I understand in quiet conditions, then background noise destroys everything.

That is a support and signal change. Restore a cleaner signal first. Research on non-native listening in noise and classroom acoustics shows that noise and reverberation can materially change speech comprehension. Do not log this as “my English suddenly got worse.” Once clear-signal comprehension is stable, you can deliberately introduce a bounded signal challenge if it serves your real listening goals.

I understand familiar topics, but unfamiliar topics collapse.

Inspect language load. Keep the speaker, duration, delivery, support, and signal stable while you move into a less familiar topic. If the topic brings dense specialist vocabulary and references at the same time, narrow it: choose a source where only part of the language load is new. You are testing topic and vocabulary load, not every other listening skill at once.

After several replays I can match the transcript, but my first pass had no clear model.

Do not count transcript recognition as proof that the first-pass comprehension criterion held. Ask what vanished first. If familiar words merged together, step back delivery. If you heard the phrases but could not hold the sequence, step back length and structure. Choose one of those based on the observed dropout; do not change both. Then retry a different clip blind.

I still fail after checking the answer.

Return to the variable you raised. Make only that dimension easier and try a different clip. If the transcript contains many unknown expressions, reduce language load. If the transcript is easy but the audio remains opaque, reduce delivery difficulty. If the message is clear sentence by sentence but you lose the arc, reduce length. Repeated failure is a reason to spend more time on the rung—not a reason to assign yourself a lower level from one source.

I understand orderly dialogue, but interruptions and overlap break the scene.

That is speaker load. Keep two speakers but reduce overlap first; do not also add a less familiar variety or noisy signal. Your blind task should be modest: identify each speaker’s purpose and the point where the turn changes. Once that holds on two different clips, add a little more turn pressure.

I understand when the picture shows what is happening, but audio-only feels much harder.

That is a support and signal difference because visual context is one of your supports. Keep the language, length, delivery, speakers, and variety familiar while you reduce visual help. If audio-only listening is not important for your real goals, you do not have to remove useful visual context just to make practice look tougher.

Accessibility is part of the baseline, not an exception

Hearing loss, auditory-processing differences, and poor playback conditions can change what a listening task demands. Stable captions, better headphones or speakers, a quieter signal, or other accommodations may be the correct baseline rather than “help you must remove.” This ladder is not a hearing test and cannot diagnose hearing or auditory-processing conditions. If listening difficulties seem broader than language learning or interfere with everyday hearing, professional support may be appropriate.

A twelve-week ladder

Here is a default schedule. It deliberately begins with short, familiar, well-signposted, one-speaker material and ends with a bounded authentic segment whose complexity comes from settings you have already proved plus one newly raised variable.

Weeks 4 and 8 are consolidation weeks. Week 4 proves that the raised delivery setting transfers to new clips. Week 8 deliberately combines consolidation with a step-back: it restores the last reliable support condition and proves that stable comprehension again before you raise anything else.

Default advance rule: meet the named criterion on two different clips at the same rung. Again, that is a consistency rule for this plan, not a research-established threshold. If a week takes longer than a calendar week, keep working the rung. The calendar serves the method; the method does not owe the calendar a performance.

English listening difficulty ladder: twelve-week default progression
Week Material shape Single raised variable Blind-listen task Allowed check Pass criterion Step-back action
1 Short, familiar-topic, well-signposted, one speaker; clear signal; transcript or captions available after listening. Language load: slightly denser vocabulary or references inside a familiar topic. State the main claim, purpose, and one chosen detail. Transcript or captions after the blind pass; one focused replay. Main-point hold: recover main claim, purpose, and chosen detail on two different clips. Return to more familiar vocabulary/topic density only.
2 Same general topic, speaker profile, delivery, and support; somewhat longer and still clearly structured. Length and structure: more duration or ideas. Retell the main points in order and include the chosen detail. Transcript/captions after the blind pass; compare your sequence with the source. Sequence hold: preserve order, purpose, and chosen detail on two clips. Shorten duration or reduce idea count only.
3 Keep Week 2 length and topic; one speaker; clear signal; same support. Delivery: less paused, more reduced, or otherwise more demanding natural delivery. Recover main message, speaker purpose, and target detail without reading first. Checking surface after first pass; one replay if needed. Delivery hold: preserve the comprehension targets on two different clips. Choose clearer or more paused delivery only.
4 New clips with the same profile as Week 3. Delivery: consolidation at the raised setting; no second variable added. Repeat the same blind task on unfamiliar clips rather than rehearsing one source. Same checking surface as Week 3. Transfer hold: meet the Week 3 criterion on two new clips. Return to the previous successful delivery profile.
5 Same topic, length, delivery, variety familiarity, and support; two speakers with orderly turns. Speaker load: one speaker becomes two orderly speakers. Track who says or wants what, what changes, and one chosen detail. Transcript/captions after blind pass; replay one turn boundary. Turn-map hold: recover speaker roles, sequence, purpose, and target detail on two clips. Return to one speaker or make turn boundaries clearer only.
6 Keep language, length, delivery, speaker count, and signal at proven settings; use a speaker or English variety you have had less exposure to. Variety familiarity: less familiar to you, without ranking varieties. Recover the same main message, purpose, sequence, and detail target. Reliable transcript/captions after first pass; replay only after forming a model. Variety-transfer hold: meet the same comprehension targets on two clips from that exposure band. Use a more familiar speaker/variety while preserving the other five variables.
7 Keep source clear and the other five variables stable; retain a reliable checking surface but do not show it during the first pass. Support and signal: reduce one support, such as first-pass caption visibility. Build the main-message model before revealing the check. Reveal transcript/captions after first pass; replay one difficult line if needed. Blind-first-pass hold: recover main message, purpose, sequence, and detail before the check on two clips. Restore the last reliable support level; do not simultaneously simplify another variable.
8 Second consolidation week and deliberate step-back. Use new clips at the last stable profile before the Week 7 support challenge. Support and signal: this is the single variable raised in Week 7 and deliberately stepped back here; no different variable is introduced. Repeat the blind task with the restored support arrangement and prove the stable model transfers to new clips. The same reliable checking surface used at the prior successful rung. Recovery-and-consolidation hold: recover the previous stable criterion on two different new clips before re-raising anything. If recovery still fails, make support/signal one more step easier; keep the other five stable.
9 Recovered stable profile; bounded segment becomes longer or carries more ideas while delivery and speaker pressure stay proven. Length and structure: a second increase in memory/structure load. Retell the arc in order, state purpose, and recover the chosen detail. Outline/transcript/captions after blind pass; compare sequence, not word count. Memory-span hold: preserve sequence, purpose, and target detail across the longer segment on two clips. Reduce duration or idea count only.
10 Same language, length, delivery, variety familiarity, and signal; speakers now interrupt or overlap in a bounded way. Speaker load: more turn pressure. Track speaker roles, the disagreement/response sequence, and one target detail. Transcript/captions after first pass; replay one overlap point. Interaction hold: preserve roles, purpose, sequence, and detail on two clips despite turn pressure. Return to orderly turns or less overlap only.
11 Keep the other five variables at proven settings; choose another speaker/variety or context with lower familiarity for you. Variety familiarity: a second transfer challenge. Recover message, purpose, sequence, and chosen detail without treating unfamiliarity as an accent ranking. Reliable checking surface after the first model; focused replay if needed. Second-transfer hold: meet the same comprehension targets on two different clips. Increase familiarity/exposure only; preserve the other five variables.
12 One bounded authentic higher-complexity segment whose other five settings have already been demonstrated. Language load: less familiar topic, denser idiom load, or more references—one language-load increase, not a simultaneous signal disaster. State main interaction/claim, sequence, purpose, and one preselected detail. Reliable transcript/captions or authorised source notes after the blind pass; focused replay. Bounded-authenticity hold: recover all four comprehension targets on two different bounded segments. Restore a more familiar topic/vocabulary profile only.

Three learners can use the same ladder in a different order

The integrity of the ladder comes from the pass criteria and one-variable rule, not from pretending every learner has the same bottleneck.

  • Decoding-heavy path: keep clips short and speakers simple; bring delivery and then support and signal earlier. Work on hearing reduced or connected speech before adding much length or overlap. The pass criterion still requires main message, sequence, purpose, and the chosen detail on two different clips.
  • Memory-heavy path: keep speaker, delivery, and language load familiar; bring length and structure forward, consolidate it, then add speaker pressure. Do not “fix” memory collapse by slowing every source and changing the topic at the same time.
  • Variety-exposure-heavy path: hold topic, length, delivery, speaker load, and signal unusually steady while variety familiarity moves earlier. This is exposure management, not a claim that one accent is better, more correct, or naturally easier.

Your progress log

Do not write “good,” “bad,” or “too hard” and call it data. Record enough to reconstruct the experiment.

Listening ladder progress log









If you want other methods alongside this progression, explore more English listening exercises. Keep this page for the job it does best: deciding how difficulty should change across sessions and weeks.

Optional: control the check without changing the method

The ladder above works without FunFluen. If you use compatible subtitle-bearing video, FunFluen can be an optional listen-first/check-later practice layer: keep subtitles hidden for the blind pass, reveal them as a checking surface, and repeat a difficult line when needed.

Review the FunFluen extension listing before installing.

This requires compatible video and available subtitles. It supports deliberate practice; it does not assign a listening level, grade comprehension, verify transcript accuracy, repair source audio or captions, automatically run this twelve-week ladder, or promise support for every platform or title.

Combining with dictation

Dictation can be useful here, but only as a flashlight. Do not let it take over the building.

If your blind summary failed and you suspect a decoding problem—“I know those words on the page, but I could not hear where they were in the stream”—transcribe a very short problematic stretch. Then compare it with the authorised transcript or captions and ask one question: what exactly dropped out?

  1. Do the blind comprehension pass first.
  2. Choose only the short stretch that seems to contain the dropout.
  3. Transcribe what you genuinely hear before checking.
  4. Compare with the checking surface and identify the decoding issue: merged word boundary, reduced form, unfamiliar word, or something else you can actually observe.
  5. Replay the line in context, then return to the ladder criterion by stating the main message, sequence, purpose, or chosen detail again.

The boundary matters. Being able to reproduce a transcript after five replays is not the same thing as building a useful first-pass model of the message. The twelve-week ladder measures the latter. Short transcription simply helps locate one failure inside it.

Common sequencing errors

Most broken listening plans do not fail because the learner lacks discipline. They fail because the experiment becomes impossible to read. Watch for these patterns.

Changing several dials and calling the result “advanced”

A longer clip with denser language, faster delivery, more speakers, a less familiar variety, and worse signal may be authentically difficult—but it cannot tell you which ability changed. Earn complexity by combining settings you have already proved, then raise one new variable.

Letting the checking surface answer before you listen

If captions are an access need, keep them. If they are a practice support you are trying to reduce, make the blind pass genuinely blind before checking. The goal is not “captions bad.” The goal is knowing whether your first-pass mental model came from listening, reading, or both.

Making slow playback permanent because it feels safer

Playback control can be useful, and research does not support the lazy rule that slower is always better. If delivery is your active variable, use slower or clearer speech as a step-back condition and then re-test. Do not let the support quietly become a permanent level label.

Ranking accents instead of tracking familiarity

“I understand this variety more easily” is a useful observation. “This accent is easy and that accent is bad/hard” is not. Your exposure history, the speaker, topic, recording, noise, and task all matter. Log familiarity and conditions, not a hierarchy of English speakers.

Calling a noisy recording a language failure

If your comprehension collapses when the signal gets muddy, restore signal quality before changing vocabulary, speakers, or length. Noise and reverberation are task variables, not secret CEFR examiners hiding inside your laptop speakers.

Turning one ugly clip into a verdict

One clip can fail for too many reasons. That is why the default rule asks for evidence across two different clips at a rung and why a failed rung includes a retry on a different source. A bad session should produce a next action, not a press conference about your entire English level.

Writing vague progress notes

Your log is also useful English practice. Replace fuzzy judgments with language that tells you what happened.

More useful English for describing listening problems
What the learner wrote Classification What a listener may understand Likely intended meaning Natural alternative Context note
“I lost the conversation after the second speaker joined.” Unusual / non-idiomatic for comprehension You probably stopped following, but “lost the conversation” can sound as if the conversation, recording, connection, or thread itself disappeared. You could no longer follow who was saying what. “I lost track of the conversation after the second speaker joined.” For a listening log, “lose track of” or “couldn’t follow” is clearer. “Lost the conversation” can make more sense when referring literally to a missing recording, chat thread, or connection.
“I made a listening practice with a longer clip.” Unusual / non-idiomatic The meaning is guessable: you did some listening work with a longer clip. You practised listening by using a longer clip. “I did a listening exercise with a longer clip.” or “I practised listening with a longer clip.” English normally uses “do an exercise” or the verb “practise/practice” here rather than “make a practice.”

For register, compare these three ways to record the same kind of dropout:

  • Conversational: “I lost track when the speakers started talking over each other.”
  • Neutral: “I couldn’t follow the conversation once the turns began to overlap.”
  • Formal: “My comprehension dropped during overlapping speech.”

Useful listening collocations include catch the main point, follow the sequence, lose track, talk over each other, check against the transcript, and replay a line.

Say your next rung aloud

Complete both frames with your own session data:

  • “I could follow __________, but I lost track when __________.”
  • “Next time I’ll keep __________ stable and change __________.”

If your second sentence names two or three changes, rewrite it. The ladder is already telling you what went wrong.

Sources and what they do not prove

These sources support the variable boundaries and evidence cautions used above. None validates this exact twelve-week order as a universal curriculum.

  1. Council of Europe — CEFR Descriptors. The official CEFR material uses structured “can do” descriptors and distinguishes listening from audio-visual reception across situations. It supports observable task-specific performance. It does not validate this twelve-week ladder or justify assigning yourself a CEFR level from one clip.
  2. Hilde van Zeeland and Norbert Schmitt — “Lexical Coverage in L1 and L2 Listening Comprehension: The Same or Different from Reading Comprehension?” (Applied Linguistics, 2013). The study manipulated lexical coverage in four informal narratives with 36 native and 40 non-native listeners; comprehension generally rose with coverage, with substantial variation among non-native listeners at lower coverage. It supports treating vocabulary familiarity as a variable, not turning its percentages into a universal pass score.
  3. Kara McBride — “The effect of rate of speech and distributed practice on the development of listening comprehension” (Computer Assisted Language Learning, 2011). The study compared fast, slow, learner-choice, and pause-control training conditions with 141 native-Spanish EFL learners from six Chilean universities. Outcomes differed across measures. It does not show that slower is always better or establish one correct playback rate for a level.
  4. Odette Scharenborg and Marjolein van Os — “Why listening in background noise is harder in a non-native language than in a native language: A review” (Speech Communication, 2019). The review synthesises research on non-native spoken-word recognition in background noise and discusses language exposure and multiple processing effects. It supports separating noise/signal from language ability; it cannot diagnose an individual learner.
  5. Almitra Medina, Gilda Socarrás, and Sridhar Krishnamurti — “L2 Spanish Listening Comprehension: The Role of Speech Rate, Utterance Length, and L2 Oral Proficiency” (The Modern Language Journal, 2020). The study examined 31 native-English upper-level learners of L2 Spanish listening to 32 sentences that varied rate and length. It supports treating length and rate as separable task variables. It does not provide fixed English thresholds, and its Spanish sentence-level results should not be transplanted into a universal English progression.
  6. Laura Mahalingappa, Jiaxuan Zong, and Nihat Polat — “The impact of captioning and playback speed on listening comprehension of multilingual English learners at varying proficiency levels” (System, 2024). In a quasi-experimental TED-based study with 287 multilingual English learners in China and Turkey, outcomes varied with caption condition, playback speed, item difficulty, proficiency/listening measures, and background factors. It supports treating captions and speed as adjustable supports, not universal fixes.
  7. Zhao Ellen Peng and Lily M. Wang — “Effects of noise, reverberation and foreign accent on native and non-native listeners' performance of English speech comprehension” (Journal of the Acoustical Society of America, 2016). The study involved 115 adults across 15 classroom-acoustic conditions, with native-English or native-Mandarin-Chinese talkers and English ability considered. It supports controlling signal conditions and resisting accent-only explanations. It does not justify ranking accents or generalising beyond the study design.

You do not need the next clip to prove what level you are. You need it to answer a smaller question: which variable did I move, and what still held? Once you can answer that, a hard clip stops being a verdict. It becomes a useful rung.

One harder thing. Five stable things. Prove it, repeat it, or step back on the dial you actually changed.