Compare similar speaking or writing samples over time and track spontaneous retrieval, range across topics, delayed reuse, and appropriate use—not how many cards you saved.

Your flashcard app knows you have 3,842 cards. Your mouth, during a Tuesday meeting, appears to know “good,” “important,” and “thing.” The mismatch is not mysterious: saved and recognized vocabulary are not the same as vocabulary you can retrieve on demand.

Track four signals: RETRIEVAL → RANGE → STABILITY → PRECISION. Keep the task type and sample length similar before comparing.

Your deck measures storage. You need evidence of retrieval.

Vocabulary research distinguishes receptive knowledge from productive knowledge. In research comparing receptive and productive vocabulary size, receptive vocabulary was larger than productive vocabulary. Recognizing a word in a card, subtitle, or test is useful—but it is not the same task as producing the word when nobody has shown it to you first.

If you want to measure active vocabulary growth, ask: Which words now arrive when I need them?

This article uses “active vocabulary” in that practical sense: vocabulary you can retrieve and use in output. It is not offering a psychometric estimate of your total productive vocabulary size.

First rule: compare like with like

A two-minute story retell today and a 700-word essay next month are not useful competitors. Different tasks invite different vocabulary. Research on oral productive vocabulary has used comparable guided tasks to observe development over time, and recent methods work on spoken lexical diversity warns that some familiar indices are sensitive to text length or unstable across tasks.

For self-tracking:

  • Use the same task family: opinion with opinion, retell with retell.
  • Use roughly the same time or length.
  • Do not review your old answer immediately before the new one.
  • Use a different but comparable prompt so you are not simply memorizing wording.

A simple baseline could be three tasks: a two-minute retell, a picture description, and a two-minute opinion answer. Repeat parallel versions later.

The Active Vocabulary Evidence Card

Do not add these rows into a fake total. Each answers a different question.

1. Retrieval — did useful words arrive without being cued?

If a flashcard shows you significant and you remember its meaning, great. If you independently choose significant while explaining a change in results, that is different evidence.

2. Range — does the vocabulary survive a change of topic?

One fancy adjective does not receive tenure because it appeared once in a restaurant story. Range asks whether vocabulary is becoming available beyond the exact context in which you studied it.

3. Stability — can you retrieve it later?

Productive vocabulary growth is not a screenshot; it is a pattern.

4. Precision — did you actually use it well?

Do not reward rarity for its own sake. A common word used exactly right is stronger evidence than an exotic word wearing the wrong collocation like somebody else’s shoes.

Why “unique words used” is not enough

Counting unique word forms sounds scientific because a spreadsheet can do it. The problem is that sample length and task can distort the comparison. A 2024 methods study of lexical-diversity measures in L2 oral responses found that some commonly used indices were not reliably independent of text length or stable across tasks.

Raw unique-word count can be a note, but it should not be your verdict. If Sample B is twice as long as Sample A, “I used 18 more unique words!” may mostly mean “I spoke twice as long.”

Look for better evidence than rarity

Suppose your first opinion sample contains:

It is good because it is important for people. The idea is good, but there are bad things.

A month later, under a comparable prompt, you naturally produce:

The change is useful, but the main drawback is that it creates an extra burden for smaller teams.

The interesting improvement is not that drawback is “advanced.” It is that you retrieved more precise language for contrast and evaluation without being given those words.

Now test range and stability. Do useful items appear appropriately in later topics? If yes, your evidence gets stronger.

Track circumlocution differently

Circumlocution is not failure. If you cannot retrieve landlord and say “the person who owns the apartment,” you successfully kept communicating.

For measurement, notice whether a word you used to explain around becomes directly retrievable later. That transition—“the person who…”landlord—is a practical sign of lexical access improving.

A three-sample protocol you can repeat

  1. Choose three task families: retell, description, and opinion are enough.
  2. Set a fixed speaking time or comparable writing length.
  3. Save the output or transcript.
  4. Mark genuinely useful spontaneous vocabulary—not every different word.
  5. Label evidence under Retrieval, Range, Stability, and Precision.
  6. After an illustrative interval such as two to four weeks, use parallel prompts and repeat.
  7. Compare the evidence cards, not a giant “active vocabulary total.”

The interval is not a magic dosage. It simply creates a later retrieval opportunity rather than immediate memory of the previous answer.

What counts as growth?

  • a previously passive word appearing spontaneously;
  • the same useful expression surviving across several topics;
  • retrieval becoming quicker, so you need fewer vague placeholders;
  • more precise collocations and register choices;
  • less repeated dependence on one safe adjective or verb.

It may not look like a dramatic explosion in unique-word count. That is fine. You are measuring usable access, not vocabulary Pokémon.

If you are learning English, create another speaking sample

You can run this protocol with any source of output. If you are learning English and want another place to speak, choose a speaking-practice path in FunFluen and use a response as another practice sample when the conditions are comparable.

The rubric here remains manual. FunFluen is not being presented as an automatic active-vocabulary counter or as a scorer for Retrieval, Range, Stability, and Precision.

Do you need an exact active-vocabulary number?

For serious assessment, use a validated productive vocabulary instrument designed for that construct. A home speaking sample is better treated as longitudinal evidence: it tells you what changed in your actual output under defined conditions.

“My active vocabulary is 2,731 words” sounds satisfying. “I can now retrieve precise evaluation language across three topics without prompts” tells you what you can actually do.

Measure retrieval, not storage

Your flashcard count can still help manage review. It just should not impersonate your speaking ability.

Build comparable samples. Look for retrieval. Check range. Wait for stability. Verify precision. Keep an evidence trail instead of worshipping one giant number.

When vocabulary comes from shows, videos, or other real contexts, the same principle fits naturally with media-based language learning: collect less, retrieve more, and make the language do a job.