PTE Summarize Spoken Text
Learn PTE Summarize Spoken Text with current Pearson rules, a five-part note map, an original worked example, scoring checks, and practice.
For PTE Summarize Spoken Text, build a map rather than a transcript: capture the topic, controlling idea, major supports, relationships, and conclusion, then turn that map into a coherent 50–70-word summary.
The worst second in Summarize Spoken Text is often the quiet one after the audio stops. You have notes. Plenty of notes. Unfortunately, they look as if a small spider attended the lecture on your behalf. The fix is not faster handwriting. It is giving each useful note a job, so the page tells you what the speaker meant rather than merely proving that words happened.
Current PTE Academic Summarize Spoken Text snapshot
Official-source check: 26 August 2026. These rules are for the current PTE Academic and PTE Academic UKVI tests. Pearson’s 10 July 2025 update says the 2025 enhancements affected those two products, while PTE Core and PTE Home were not affected; it also says the original question types kept the same task, format, timing, and response type. Summarize Spoken Text is therefore treated here using the current post-2025 Academic/UKVI rules, not older prep material or PTE Core rules.
| What matters | Current rule and practical meaning |
|---|---|
| Prompt | Pearson describes a 60–90-second audio prompt. The audio starts automatically and plays once. |
| Response time | You have 10 minutes to listen and write. |
| Target length | 50–70 words earns full Form credit under the current Score Guide. |
| Near the word limit | 40–49 or 71–100 words receives reduced Form credit under the current Score Guide. |
| Below 40 or above 100 | The Score Guide assigns Form 0, and Pearson’s current task page goes further: fewer than 40 or more than 100 words receives zero points across all five factors, so the summary is scored zero. |
| Other Form-0 conditions | The current Score Guide also assigns Form 0 when a summary is written in capital letters, contains no punctuation, or consists only of bullet points or very short sentences. |
| Assessed skills | Listening and Writing. |
| Current question count | The July 2025 PTE Academic Score Guide currently lists 1 Summarize Spoken Text question. |
| Scoring traits | Content, Form, Grammar, Vocabulary, and Spelling. Partial credit applies. |
| Current trait scales | The current Score Guide gives Content a 0–4 scale and Form, Grammar, Vocabulary, and Spelling 0–2 scales. Do not turn those trait descriptors into an invented standalone SST band or score conversion. |
| Scoring review | Pearson says the traits are AI-scored and a human expert reviews Content before the final task score. |
That word-count distinction matters. A 49-word or 73-word response is not automatically a zero for the whole task under the current rubric: it loses Form credit. A 39-word or 101-word response is different; Pearson’s current task page says the whole summary is scored zero. Your practical target remains 50–70 words.
If you need the broader exam structure rather than this one task, use the PTE Listening format and study plan. Here, we stay with Summarize Spoken Text.
What SST actually measures: meaning under compression
Summarize Spoken Text is not a contest to see who can become a court stenographer for 90 seconds. Pearson’s current Content rubric focuses on whether you understand the lecture, identify and synthesize its main points, connect ideas coherently, and summarize them in your own words. The task assesses both Listening and Writing.
That changes what “good notes” means. A page packed with accurate nouns can still be poor preparation for the response if you cannot tell which noun is the main claim, which is evidence, and which was just a memorable example. Your final summary needs the structure of the message, not a souvenir from every sentence.
Pearson’s current score report shows an Overall score plus Reading, Listening, Speaking, and Writing scores on its reporting scale. It does not report one SST response as a separate “band,” so this guide discusses rubric alignment and response quality rather than invented task-level score conversions or guaranteed thresholds.
The five-part note map: T → CI → S → R → C
Instead of asking “What words can I catch?”, give your notes five possible jobs. You do not need every box to be equally large, and a lecture may not state an explicit conclusion. The point is to force hierarchy while the audio is still fresh.
| Role | Question to ask | Useful note style | What to resist |
|---|---|---|---|
| T — Topic | What is this lecture broadly about? | Two or three words: “urban trees + heat” | Writing a whole opening sentence while the speaker keeps talking. |
| CI — Controlling idea | What is the speaker’s main point about the topic? | A compressed claim: “trees = infrastructure, not planting count” | Confusing the most vivid example with the main point. |
| S — Major support | Which reasons or findings make the main point believable? | Short chunks: “shade + transpiration cool”; “mature canopy > young tree” | Collecting every example, number, name, or side comment. |
| R — Relationship | How do the ideas connect? | Symbols or words: “→”, “because”, “but”, “therefore”, “contrast” | Keeping two facts but losing the cause, contrast, or qualification between them. |
| C — Conclusion | What does the speaker finally want us to understand? | One line: “measure quality/survival/need, not totals” | Inventing a conclusion when the speaker did not actually give one. |
Pearson itself advises noting the main idea and supporting points. The five-role map is our practice method for making that advice usable: it adds explicit labels for the controlling idea, relationships, and conclusion so your notes can feed a summary instead of becoming a pile such as “solar / farms / birds / costs / govt??”. Accurate fragments are not enough if none of them knows what job it has.
From notes to 50–70 words: transform, do not transcribe
Once the audio ends, stop collecting. Your job changes from listening to compression. A reliable sequence is:
- Rank the map. Circle the controlling idea. Keep only the supports that explain, prove, qualify, contrast with, or lead to it.
- Restore relationships. If your notes say “heat ↑ / shade ↓ surface heat,” decide whether the speaker expressed cause, contrast, result, or simple addition before you write a connector.
- Draft the main meaning first. Write a sentence that carries the controlling idea instead of opening with three peripheral facts.
- Add major supports. Usually the best supports are the ones that explain the mechanism, evidence, contrast, or consequence behind the main claim.
- Remove unsupported detail. Do not “improve” the lecture with facts you know from elsewhere. SST rewards the message you heard, not the lecture you wish the speaker had given.
- Count and compress. Aim for 50–70 words. Cut repetition and low-value examples before cutting the main relationship.
- Proofread meaning before decoration. Fix grammar, spelling, and awkward wording without turning a cautious “suggests” into an absolute “proves.”
A flexible sentence shape can help you organize the result, but it is not a magic PTE Summarize Spoken Text template. For example: state the main point, connect two major supports, then finish with the speaker’s implication or conclusion. The container cannot choose the content for you. A template with the wrong cargo is still going to the wrong address.
One tiny transformation
Suppose your notes say: CI: remote work can widen hiring pool; S: location barrier ↓; BUT onboarding/mentoring harder; C: hybrid systems need deliberate support.
Before revealing a model, write one sentence that preserves the contrast. Do not add a claim about productivity unless it appears in the notes.
See one possible sentence
Remote work can widen an employer’s hiring pool by reducing location barriers, but organizations need deliberate onboarding and mentoring systems to offset the weaker informal support that can come with distance.
Language precision: accurate beats impressive
SST gives you very little room for decorative English. A simple accurate verb is safer than a grander synonym that changes the relationship. “The speaker suggests” is not automatically interchangeable with “the speaker proves.” “Because” and “however” are not interchangeable glue. They tell the reader what the lecture’s logic was.
| Original expression | Classification | What a listener or reader would understand | Likely learner intent | Natural alternative | Context note |
|---|---|---|---|---|---|
| “The lecture discussed about urban trees.” | Wrong | The meaning is understandable: the lecture’s topic was urban trees. | Introduce the subject of the lecture. | “The lecture discussed urban trees.” or “The speaker talked about urban trees.” | Discuss takes its object directly in this use; talk about uses about. |
| “Due to cities are hotter, trees are important.” | Wrong | The reader can infer a causal link between urban heat and the importance of trees. | Express a reason. | “Because cities are hotter, trees are important.” or “Due to higher urban temperatures, trees are important.” | Use because before a clause; use due to naturally before a noun phrase here. |
| “Researches show that mature canopies cool more.” | Unusual/non-idiomatic for the intended noun meaning | The intended claim is clear: research evidence supports the point. | Refer to research evidence or multiple studies. | “Research shows…” or “Studies show…” | Researches is normal as a verb in a sentence such as “She researches urban heat,” but it is not the natural noun choice for this intended general meaning. |
| “Trees make neighborhoods more healthier.” | Wrong | The reader understands that trees improve neighborhood conditions. | Use the comparative form of healthy. | “Trees make neighborhoods healthier.” | Do not combine more with the comparative healthier. |
Reporting verbs are meaning choices
- says / explains: safe when you are reporting what the speaker states or clarifies.
- argues: useful when the speaker is advancing a position, not merely describing facts.
- suggests / indicates: useful when the conclusion is cautious or evidence points in a direction.
- proves: much stronger. Do not upgrade a tentative or correlational claim just because the word sounds academic.
Production step: take one sentence from your next practice summary and replace a vague reporting verb only if the new verb preserves the speaker’s certainty and relationship. If it changes the claim, put the simpler verb back.
Worked example: from an original lecture to a 55-word summary
The practice material below was created for this article. It is not a real PTE question, transcript, or answer. The transcript is shown so you can inspect every decision even without audio.
Original practice transcript
City planners often treat trees as decoration, but urban heat research suggests their more important role is infrastructure. In dense neighborhoods, asphalt and concrete absorb solar energy during the day and release it slowly at night, keeping local temperatures high. Tree canopies interrupt that cycle in two ways: shade reduces the energy reaching hard surfaces, while transpiration moves heat as water evaporates from leaves. However, simply counting trees can be misleading. A young tree in a narrow pit offers much less cooling than a mature canopy with enough soil and water. This means cities should not focus only on planting targets; they also need long-term maintenance and thoughtful placement. Wealthier neighborhoods often have more mature canopy, so a citywide planting campaign can still leave the hottest districts behind if survival rates and location are ignored. The main lesson is that urban trees work best as long-term climate infrastructure when cities measure canopy quality, survival, and neighborhood need—not just the number planted.
Before reading our notes, cover or scroll past the transcript and answer two questions from memory: What is the controlling idea? And which two or three supports are essential to that idea?
Reveal the note map
| T | urban trees + city heat |
|---|---|
| CI | trees = long-term climate infrastructure, not simple planting count |
| S1 | hard surfaces store heat → shade + transpiration cool |
| S2 | mature/healthy canopy cools more; tree counts can mislead |
| S3 / R | canopy unevenly distributed → placement + survival matter |
| C | measure canopy quality + survival + neighborhood need, not totals alone |
A weak response
The lecture is about trees in cities and asphalt stays hot at night. Trees give shade and water evaporates from leaves. Young trees in narrow pits give less cooling, while richer neighborhoods have more mature trees. Cities should plant more trees and count them carefully, and they should also maintain them because trees can reduce heat.
Word count: 56. The response is inside the full-credit Form range, which is useful because it exposes an important lesson: a correct word count cannot rescue weak content selection.
The response contains several accurate details, but it behaves like a compressed list. More importantly, “cities should plant more trees and count them carefully” distorts the speaker’s conclusion. The lecture argues that simple planting totals can be misleading and that canopy quality, survival, placement, and need matter.
An improved response
Urban trees can reduce city heat by shading hard surfaces and cooling through transpiration, but tree counts alone do not show their real effect. Because mature, well-maintained canopies provide more cooling and are unevenly distributed, cities should treat trees as long-term climate infrastructure and prioritize canopy quality, survival, and neighborhood need rather than planting totals.
Word count: 55.
Why the improved response fits the current rubric better
| Area | Weak response | Improved response |
|---|---|---|
| Content | Includes relevant details but underplays the controlling idea and reverses part of the conclusion by recommending careful counting. | States the cooling mechanism, the limitation of raw tree counts, the importance of mature canopy and distribution, and the final infrastructure conclusion. |
| Form | 56 words: inside the current 50–70 full-credit Form range. | 55 words: inside the current 50–70 full-credit Form range. |
| Grammar | The sentences are mostly grammatically complete. The bigger weakness is not Grammar itself but how the response organizes and represents the lecture’s meaning. | Uses controlled clauses to preserve contrast and cause without becoming syntactically crowded. |
| Vocabulary | Generally understandable, but “count them carefully” is an inaccurate lexical choice for the speaker’s recommendation. | Uses precise task-relevant wording such as “tree counts alone,” “well-maintained canopies,” and “prioritize” without adding unsupported claims. |
| Spelling | No deliberate spelling errors in this article-created sample. | No deliberate spelling errors in this article-created sample. |
| Coherence | Relevant facts sit beside one another, but the main hierarchy is weak. | Cause, contrast, and conclusion are visible. Note that Pearson does not list “Coherence” as a separate SST trait here; it matters through how clearly the response communicates and synthesizes Content with controlled language. |
There is no honest reason to call the improved version a “guaranteed full-score answer.” It is simply a stronger response because its content hierarchy, relationships, word range, and language align more closely with the current rubric.
Eight common SST failure modes—and the repair for each
Transcript-like copying
Symptom: your notes are long, chronological, and full of phrases from the recording. Repair: stop asking “what came next?” and ask “what job does this idea have?” Promote the controlling idea and major supports; demote repeated wording and peripheral examples.
Disconnected notes
Symptom: every fragment is accurate, but you cannot tell whether two points support, contrast with, or cause one another. Repair: record relationships with tiny markers such as →, vs, +, because, but, result.
Missing the main idea
Symptom: your response contains details but could not answer “What was the speaker’s point?” in one sentence. Repair: identify CI before drafting. If you cannot, reread your practice notes and ask which claim the other points are serving.
Invented detail
Symptom: you fill a gap with outside knowledge or a plausible-sounding statistic. Repair: delete it. A summary is not the place to rescue incomplete notes with imagination.
Excessive examples
Symptom: your response preserves the easiest example to remember but squeezes out the speaker’s conclusion. Repair: keep an example only when it is necessary to explain a major support. An accurate detail can still be low-value summary content.
Sentence fragments
Symptom: note shorthand leaks into the response: “Higher heat in cities. Trees reducing temperature. Better maintenance.” Repair: turn shorthand back into complete grammatical relationships before submission.
An overlong response
Symptom: you have 73 coherent words and panic because a prep page told you that 71 means automatic zero. Repair: stay calm and cut repetition or a peripheral detail. Under the current Pearson Score Guide, 71–100 receives reduced Form credit, not automatic zero for the whole task. Your practical target remains 50–70 for full Form credit. Once you go above 100, the rule changes: Pearson’s current task page says the whole summary receives zero.
Proofreading that changes meaning
Symptom: you replace “suggests” with “proves” because it sounds stronger, or change “may reduce” to “reduces” without evidence. Repair: proofread for correctness first. Thesaurus sabotage is still sabotage, even when it arrives wearing a tie.
Decision practice: build the map before you reveal ours
This is an article-created hierarchy drill, not a real PTE item. Open the short transcript, read it once, then close it before answering. The point is not memory trivia; it is deciding what deserves space in an SST summary.
Open the one-exposure practice transcript
At one university, cafeteria managers initially assumed food waste came mostly from unpopular dishes. A two-week audit showed a different pattern: much of the discarded food was untouched side portions that students had not chosen. Managers then made side portions smaller by default while allowing anyone to request a free second serving. Waste fell during the trial without reducing menu choices. One tray in the audit contained three untouched bread rolls, a memorable example but not the main finding. The project suggests institutions should measure where waste comes from before choosing a solution, because a visible problem can have a less obvious cause.
Read once, then close this section before making your choices.
Reveal answer and why
Best answer: B. The audit, smaller portions, and bread-roll example all serve the broader conclusion that measurement should come before intervention. A is built from a vivid example and turns it into a policy the speaker never proposed. C mistakes one feature of the intervention for the lecture’s central point.
Reveal answer and why
Best choices: A and C. A identifies the measured cause; C shows the response and result. D can be useful if you have room for the contrast between assumption and evidence, but it is less essential than the cause and intervention. B is accurate but peripheral: the bread rolls illustrate the waste pattern rather than define it.
Reveal answer and why
Best answer: B. It is memorable, specific, and true—and that is exactly why it can hijack your notes. The lecture’s summary can survive perfectly well without the bread-roll count.
Reveal answer and why
Best answer: A. It preserves the sequence of evidence → intervention → result. B invents student dislike and misstates what the audit did. C reverses the logic and adds a claim the passage does not make.
If you chose the right controlling idea but one wrong support, that is useful information. Your next practice target is not “take more notes.” It is “separate central support from memorable detail faster.”
The 90-second hierarchy sprint
Use this as a training constraint after one unaided listen to a new 60–90-second academic-style practice clip. These 30-second subdivisions are not official Pearson exam timings; they are a drill for prioritization.
- Spend about 30 seconds writing only T + CI.
- Spend about 30 seconds choosing at most three S/R notes: major supports plus the relationship between them.
- Spend about 30 seconds writing C, or “no explicit conclusion” if the speaker did not state one.
Then ask one diagnostic question: Did my notes get shorter while becoming easier to draft from? On the next pass through your practice routine, write the full 50–70-word summary from that map.
PTE Summarize Spoken Text response checklist
Pearson recommends leaving one or two minutes to check grammar, spelling, and punctuation. A fixed checklist is better than random last-second editing.
A short practice plan that transfers to new speakers and topics
Do not repeat the urban-tree example until you can reproduce it from memory and call that “progress.” SST needs transfer. Use fresh academic-style audio and keep the map constant while the speaker, accent, topic, and structure change.
| Session | Main job | What to record |
|---|---|---|
| Hierarchy only | Listen once and build T / CI / S / R / C. Do not write the full summary yet. | Which map role was hardest to hear? |
| Map to prose | Use a fresh clip, build the map, then write 50–70 words. | Which note did you remove because it was peripheral? |
| Full timed response | Use the current 10-minute task window and reserve final checking time. | Did time pressure damage Content, Form, or language accuracy first? |
| Error audit | Use another new topic, then compare your response with the checklist rather than with a memorized template. | Choose one repair target for the next session: CI, supports, relationships, compression, grammar, vocabulary, or spelling. |
The manual routine above is complete on its own. If you want extra off-test listening practice on supported subtitle-based video, FunFluen can be useful after your first unaided listen: use repeat or a slightly slower replay to check which part of your note map you missed, then work back toward normal speed. When subtitles are available, a listen-first comparison pass can also help you verify what you actually heard.
Review the FunFluen extension for off-test listening practice. The first step is to review the extension listing before installing. This does not simulate PTE, score Summarize Spoken Text, or change test conditions; subtitle-dependent practice features require subtitles to be available.
When the audio stops, read the map
The silence after the recording does not need to be the moment your notes turn into a crime scene. Give them a hierarchy before you draft: topic, controlling idea, major supports, relationships, conclusion. Then compress the message, not the audio.
Keep the speaker’s meaning ahead of impressive vocabulary. Keep major supports ahead of vivid examples. Keep your response inside 50–70 words when you can, then check grammar, vocabulary, spelling, and meaning without rewriting the lecture into something it never said.
For your next practice response, choose one map role to watch closely. If you usually collect details, hunt for the controlling idea. If you hear the main point but lose the logic, mark relationships. The goal is not more ink. It is reaching the end of one-play audio and knowing exactly what your notes are for.
Sources
- Pearson PTE — PTE Academic & UKVI test format: Listening (current task page checked 26 August 2026).
- Pearson PTE — PTE Academic — Score Guide — Test Taker (July 2025 edition; current Pearson-linked score guide at the 26 August 2026 check).
- Pearson — Establishing score concordance between the enhanced PTE Academic and IELTS Academic — Background: Construct Comparison (July 2025; used here for task-construct background, not as the governing learner scoring rubric).
- Pearson PTE — PTE changes 2025: everything you need to know (10 July 2025).