How to get a transcript of a Panopto lecture
The most useful form of a lecture is the one nobody asks for. Where to find the transcript Panopto already made, and how to make your own when it has not.
Short answer
Check Panopto first, because it has probably already made one. It machine-transcribes recordings to produce closed captions, and institutions are documenting a change to having that on by default for new content. Open the recording, open the Captions or Transcript panel, and look for a download option - what you get is a subtitle file, which is plain text with timestamps.
Captions are usually a separate permission from the video. They are often available when downloading is not. If there is no transcript, because it is an older recording, the wrong language, or downloads are switched off, you can make one yourself on your own computer with an offline speech-recognition model, with no upload and no account. Expect roughly 2 to 24 minutes per hour of lecture depending on the model; the measured table is below.
On this page
- Panopto has probably already made one
- Getting it out of the player
- When Panopto's transcript is not there, or not enough
- Making your own, on your own computer
- How long it actually takes, measured
- Which model, and the advice that is wrong on a laptop
- How accurate any of this is
- What a transcript is actually good for
- The permission question
- Common questions
Panopto has probably already transcribed your lecture, because captioning a recording is no longer something a lecturer has to remember to switch on. It is worth thirty seconds to check. A transcript runs a few hundred kilobytes where the video runs a few hundred megabytes. You can search it. It opens in anything. NotebookLM will read a transcript and will not read a local video file.
The free route comes first here; making your own comes after.
Panopto has probably already made one
Panopto machine-transcribes recordings to produce closed captions, and captions are a transcript with timestamps on. What changed recently is that it is no longer something a lecturer has to remember to switch on.
Brown University's IT knowledge base records the rollout and the reason: "On or before May 11, 2026 automatic captions will be turned on for all new content uploaded to Panopto", to meet the US Section 504 ruling and WCAG 2.1 Level AA. That is one institution documenting one rollout, and tenants are configured independently, so read it as "check yours" rather than as a guarantee. The direction only goes one way, though, driven by accessibility law rather than by anybody's product roadmap. Captions are far more likely to be present on a 2026 recording than on a 2023 one.
Two consequences worth having:
- Captions are usually not the same permission as the video. A lecturer who has switched downloading off has very often left captions alone, because turning accessibility features off is a decision nobody wants to defend. Check the transcript before concluding a lecture is locked.
- They arrive on a delay. Panopto's own FAQ says captions "can be created in about one-quarter of the total video length", so a two-hour lecture is roughly half an hour behind the recording. If a lecture went up an hour ago and has no captions, wait rather than conclude.
Getting it out of the player
Inside the Panopto viewer the transcript is a panel beside the video, and it is worth opening even if you never download it: clicking a line jumps the recording to that moment, which is the fastest way to find the four minutes of a lecture you actually needed.
- Open the recording in the Panopto player, inside Canvas or on your institution's Panopto site.
- Open the Captions or Transcript panel. If there is no such panel, the recording has no captions yet - see the next section.
- Look for a download option on that panel or in the player's menu. Panopto documents this as downloading captions from the viewer, and what you get is a subtitle file.
- If you only want the words, open the file in any text editor and delete the numbers. A subtitle file is plain text with timestamps; there is nothing proprietary in it.
A subtitle file is more useful than a plain one, so keep it.
.srt and .vtt carry a timestamp per line, which
means the text stays anchored to the recording. Search the transcript,
find the timestamp, jump to that minute. A plain transcript loses that
for the sake of tidiness you will not care about in April.
When Panopto's transcript is not there, or not enough
Four situations where the free route runs out, all ordinary:
- Older recordings. The default change applies to new content. A 2024 course may have nothing, and nothing will appear on its own.
- Caption downloads switched off while viewing is allowed. You can read it and not keep it.
- The wrong language, or a mixed one. ASR picks a language and commits. A lecture delivered in Danish with English slides and English technical vocabulary is exactly the case machine captioning handles worst, and it is completely normal in Europe.
- You are about to lose access. A transcript you can only read inside Panopto disappears with your account, which is precisely when you want it - see what happens to your Canvas access after graduation.
Making your own, on your own computer
The alternative is to transcribe the audio yourself. This used to mean an upload to a paid service; it does not any more. Open speech-recognition models run on an ordinary laptop, and the whole job is local: the audio never leaves your machine, there is no account, no quota and no per-minute charge.
The shape of it is always the same, whatever tool you use:
- Get the audio. You do not need the video - a lecture's audio track is a fraction of the size and carries everything a transcript can use.
- Download a speech-recognition model once. They run from about 75 MB to 3 GB depending on how accurate you want to be.
- Run it. After the first download, nothing needs a network connection.
- Keep both outputs if the tool offers them: a plain
.txtto read and search, and a timestamped.srtto navigate with.
The catch is not accuracy. It is time. Almost nothing published about this is specific about how much.
How long it actually takes, measured
Here is a real measurement rather than an estimate. Five models on the same clip, on one machine, timed as multiples of realtime - so 6x means an hour of lecture takes ten minutes.
| Model | Download size | Speed on this CPU | One hour of lecture takes |
|---|---|---|---|
| Tiny | 75 MB | 25.5x realtime | about 2 minutes |
| Base | 145 MB | 17.5x realtime | about 3 minutes |
| Small | 484 MB | 6.2x realtime | about 10 minutes |
| Large v3 Turbo | 1.6 GB | 3.3x realtime | about 18 minutes |
| Medium | 1.5 GB | 2.5x realtime | about 24 minutes |
| Large v3 | 3.1 GB | not timed - wants a GPU | - |
The method, and how small the sample is. One machine, an Apple M4 with 10 CPU cores running on the CPU, over 180 seconds of real Danish lecture audio, at int8 quantisation, beam size 5, with voice-activity filtering on. That is a single clip on a single laptop, not a benchmark suite. It is published because the relative ordering is stable and useful and nobody else states it at all, not because four significant figures would mean anything. Your machine will differ; the shape of the table will not.
Two things fall out of it that matter more than the exact numbers. First, on a CPU the range across models is ten-fold, so the choice is not a detail at all. Second, a semester is not one lecture: at 2.5x realtime, thirty two-hour lectures is about a full day of your laptop running flat out, and at 25x it is under three hours.
Which model, and the advice that is wrong on a laptop
The standard advice is to use Large v3 Turbo, because it is nearly as accurate as Large v3 and several times faster. On a GPU that is correct. On a CPU it is wrong. The reason is structural rather than a matter of degree.
Turbo's speed comes from its decoder, which has four layers where Large v3 has thirty-two. Its encoder is Large v3's, unchanged. On a GPU the decoder dominates the time, so cutting it to an eighth shows up almost in full. On a CPU the encoder dominates. The encoder did not change, so almost none of that speed-up arrives. In the table above Turbo lands between Small and Medium, in the group that is too slow to be practical.
What that costs in real terms, from a real course: 36 recordings that take 40 to 60 minutes on Tiny would have taken about six and a half hours on the model a naive "pick the best one" rule selects.
So the honest rule is short:
- No GPU? Nothing above Small. Small is the floor rather than Tiny because the tier below it loses real accuracy on non-English lecture audio, which is what most of these recordings are.
- A GPU with 4 GB or more free? Turbo, and this is where its reputation comes from.
- A GPU with 6 GB or more? Large v3, if you want the last sliver of accuracy on a hard language.
How accurate any of this is
Both routes are machine transcription and both are wrong sometimes. Worth knowing where.
Panopto states that "ASR captions are typically 90-95% accurate depending on the audio quality in the recording", and institutions restate that with a warning attached: Brown's guidance says machine captions "may contain errors" and that creators are expected to review and edit them. Which is worth reading twice: 90 to 95% correct means roughly one word in fifteen is wrong, and they are not distributed evenly.
The open models sit in a similar band. They were trained at a scale that makes them fairly tolerant of accents and background noise: the current large model was built on 1 million hours of weakly labelled audio and 4 million hours of pseudo-labelled audio. Its own model card claims a "10% to 20% reduction of errors" over the version before it, across a wide range of languages.
Where both fail is the same place, and it is the worst possible place for a student: proper nouns and technical vocabulary. A lecturer's name, a theory named after somebody, a chemical, a case citation, a Danish word inside an English sentence. The general words come out fine and the exact term you were going to search for does not.
The practical answer is not a better model. Treat the transcript as a searchable index into the recording rather than as a document of record. That is what the timestamps are for, and it is why keeping the recording, or at least a shortcut to it, beside the transcript is worth the disk.
What a transcript is actually good for
- Finding the four minutes that mattered. Search the text, read the timestamp, jump there. This alone is worth the exercise and it is not possible with video.
- Revision that is not rewatching. Reading a ninety-minute lecture takes about fifteen minutes.
- AI study tools. A local video file is the one thing they will not take - see getting your Canvas files into NotebookLM, where the transcript is the whole point of this exercise.
- Quoting accurately. With the caveat above: check the exact words against the audio before a quotation goes into anything assessed.
- Keeping the lecture at all. Text is the format most likely to still open in ten years, on any machine, with no player and no login.
The permission question
Getting a transcript out of Panopto is subject to the same rules as getting the video, and those rules genuinely differ between institutions. How to download Panopto lecture recordings covers it properly; the short version is that a personal study copy is widely allowed, redistribution is prohibited essentially everywhere, and asking is a two-click favour rather than an imposition.
Transcribing audio you are entitled to have does not change that calculation in either direction. If you may keep the recording, you may keep a transcript of it. If you may not, a transcript is not a loophole. And if your lectures turn out not to be in Panopto at all, start with how to download lecture videos from Canvas, which covers all five systems.
A whole semester of lectures, as text you can search
Tick as many courses as you like. Canvas Downloader finds the Panopto lecture recordings in each one and saves them as a searchable transcript and a timestamped subtitle file, alongside the video or audio if you want it, and alongside the slides and readings from the same courses.
The transcription runs on your own computer with the models in the table above - nothing is uploaded, there is no account and no per-minute charge, and after the first model download it works with no network at all. Free and open source, Windows and macOS.
Common questions
Does Panopto give you a transcript automatically?
Usually, now. Panopto machine-transcribes recordings to produce closed captions, and institutions are documenting a change to having that on by default for new content. Brown University's IT knowledge base dates it to on or before 11 May 2026, driven by Section 504 and WCAG 2.1 Level AA accessibility requirements. Tenants are configured independently, so check yours. Older recordings are not covered either way.
How do I download the transcript from a Panopto lecture?
Open the recording in the Panopto player, open the Captions or Transcript panel, and look for a download option there or in the player's menu. Panopto documents this as downloading captions from the viewer. What you get is a subtitle file, which is plain text with timestamps and opens in any text editor.
Can I get the transcript if downloading the video is blocked?
Often, yes. Captions are usually a separate permission from the recording, and a lecturer who has switched downloads off has frequently left captions alone, because turning accessibility features off is harder to justify. It is worth checking before concluding a lecture is locked.
How long does it take to transcribe a lecture on my own laptop?
It depends entirely on the model. The range is about tenfold. Measured on a 10-core laptop CPU with no GPU: roughly 2 minutes per hour of lecture on the smallest model, about 10 minutes on a mid-sized one, and about 24 minutes on a large one. That is one machine on one clip rather than a benchmark, but the ordering holds.
Should I use Large v3 Turbo for transcription?
On a GPU, yes. On a CPU, no. The reason is structural. Turbo is fast because its decoder has four layers instead of thirty-two. Its encoder is unchanged from the full model. On a CPU the encoder is what dominates the time, so almost none of the speed-up arrives. On a laptop it lands in the group that is too slow to be practical.
Is machine transcription accurate enough to revise from?
For following the argument, yes. Panopto states its automatic captions are typically 90 to 95% accurate depending on audio quality. The open models sit in a similar band. The errors are not spread evenly, though: proper nouns and technical terms are where both fail, which is exactly what you were going to search for. Treat a transcript as a searchable index into the recording rather than as a document of record.
Does transcribing a lecture upload it anywhere?
It does not have to. Open speech-recognition models run on your own computer, so the audio never leaves it, and after the model is downloaded once the process needs no network at all. Paid online services do upload, which for lecture recordings is worth thinking about, since they usually contain other students.
What is the difference between a .txt transcript and a .srt subtitle file?
Only timestamps. A .srt carries a time for every line, so the text stays anchored to the recording and you can search the transcript, find the timestamp and jump to that minute. A .txt is tidier to read. Keep both if you can - the timestamps are what make a transcript a navigation tool rather than just a wall of words.