
We organized 668 Derek Moneyberg YouTube transcripts. If your videos are hard to find and reuse, start with a list. Here’s how we linked each transcript to that list so our team can turn it into an article.
We’re building an article workflow around Derek Moneyberg’s personal website. His YouTube video library gives us the source material. Our YouTube channel inventory guide explains how to catalog that material before assigning it to a writer.

Count the source videos before collecting transcripts
The channel snapshot contained 691 videos. We focused on uploads in the Videos tab, kept podcasts and personal videos together, and excluded videos under three minutes. Shorts and the separate livestream archive weren’t part of this pull.
That scope matters. “Every video” sounds simple until it mixes full episodes, short clips, and the same conversation on several platforms. We wanted a useful article queue with one record per YouTube video.
| Result | Videos |
|---|---|
| Auto-generated captions collected | 666 |
| Manual caption tracks collected | 2 |
| No caption tracks available | 8 |
| Under three minutes, intentionally skipped | 15 |
| Total inventoried | 691 |
The 668 reading copies contain roughly 2.6 million words. That’s the size of the source library, not a claim that we’ve reviewed every word or written 668 articles. Each transcript still needs editorial checks before we use it.

Let the script handle the repeated work
I used Claude to help write and troubleshoot a Python script, then ran it on my Mac. The script used the yt-dlp project’s Python library to collect video details and available English captions. It downloaded existing captions; it didn’t send each recording through a new speech-to-text model.
We began with a three-video test. Once the test produced three readable files, the script worked through the channel and recorded the outcome for each video. A saved progress log let it resume without downloading successful files again.
The script selected one English caption track per video, preferring a manual track when available. It saved the original timestamped caption file and created a separate Markdown reading copy. It removed caption markup and overlapping repeated lines while retaining timestamps that help an editor return to the recording.
Keeping both versions gives us a way to inspect the source if the cleaned copy looks wrong. Removing every repeated sentence would be a mistake: speakers sometimes repeat a point on purpose.
Fix empty results before trusting the batch
The first setup reported missing captions for videos that later downloaded successfully. We stopped, installed an updated yt-dlp in its own Python environment, added the JavaScript runtime used by the tool, and corrected the caption-selection logic. The small test then passed.
Because several things changed together, we can’t assign that recovery to one change alone. The useful lesson is to test the output before letting a tool label an entire channel “no captions.”
A network interruption also left three videos with errors. A later run recovered all three. The eight remaining videos still had no caption tracks, so they stayed marked unavailable. We didn’t invent text to fill those rows.
The first empty-result safeguard also stopped a retry of known captionless videos. We adjusted it to distinguish a fresh run with no successes from a cleanup pass through previously empty records. Recovery behavior needs testing too.
Keep the video ID with every transcript
We named the reading copies with this pattern:
YYYY-MM-DD-slug-VIDEOID.md
The date helps with sorting. The short title helps a person recognize the file. The exact YouTube video ID connects it to the right recording even if the title changes.
Each file also records its source URL, publication date, duration, caption type, and retrieval date. Auto-caption text carries a reminder to verify names and homophones against the recording before quoting it.

We kept the source index separate from the full text. An editor or agent can select a useful video from the index, then open that transcript. There’s no need to load millions of words just to find one episode.
Put the files where the next writer can find them
After extraction, we uploaded all 668 reading copies to a Raw Transcripts folder in Google Drive. We put the existing live Google Sheet in the client’s parent folder and added the direct transcript link to each matching video row. The original timestamped caption files remain in the source archive.
The tracker has three tabs:
- Read Me: how to use the sheet and where the working files live.
- Personal YouTube + Podcasts: videos on Derek’s channel, including podcast episodes.
- Featured In Videos: appearances on other people’s channels.
The whole tracker is the article queue. A podcast episode on Derek’s channel belongs in the second tab; it doesn’t need another copy in a separate podcast queue. The nine featured-video records were outside this channel extraction.
For each source, the sheet keeps the YouTube URL, transcript link, transcript status, article status, QA status, and eventual article link. “Transcript received” and “article published” are separate fields because they describe different work.

We matched the links by exact video ID. Then we checked the Drive folder against the expected filenames, file sizes, and uploaded file IDs. The final folder contained 668 unique expected reading copies. We also read back every written tracker link and compared one downloaded transcript’s text with its local source.
Those checks catch a problem that a successful upload message can miss: a file may exist while the sheet points to the wrong recording.
Turn the library into useful articles
This puts Derek’s existing recordings into the Process stage of the Content Factory’s four-stage workflow: Produce, Process, Post, and Promote. The recordings already exist. We’ve now organized the text and the handoff to writing.
Next, our team will select a source, watch the recording, and identify the question it answers. We’ll check names, quotes, and the meaning of the conversation, then write in Derek’s voice. The article writing and review guidelines govern that editorial pass.
We’ll also check whether his site already answers the same question. Several videos may support one strong article. A single interview may contain several distinct ideas. The source count shouldn’t become a publishing quota.
The SEO Tree content structure gives each article a place: connect a specific story to its broader topic and link related material where it helps the reader. Derek’s site becomes easier to explore when readers can move from an article to the original conversation, his biography, or another relevant lesson.
We’ve already documented another format in our Derek Moneyberg story-card example. That work used positive mentions to make visual assets. This transcript library supports longer articles. Both start with existing source material, and both need source checks before publication.
The purpose of the content matters too. Our Pete Hazzard website audit connects recommendations to relevant speaking inquiries. For Derek, this library gives writers source material for his personal-brand site. Choosing what to publish should start with the audience and the question being answered.
Separate the measured result from the next outcome
Claude helped with scripting and troubleshooting; I ran the extraction commands and set the scope. ChatGPT then handled the Drive organization, tracker updates, and reconciliation checks. The software did the repetitive collection and matching. Editorial judgment is still part of choosing and preparing each article.
The extraction handoff records roughly 7.5 elapsed hours from initial setup to the final retry, with much of the main run unattended. We didn’t capture a reliable hands-on time total or a complete billing record, so a dollars-saved comparison would be guesswork.
| Measurement | What we can report |
|---|---|
| Source inventory | 691 video records |
| Transcript files and tracker links | 668 of each |
| Extraction method | Existing captions, with no per-video language-model call |
| Total AI tokens and allocated cost | Unknown |
| Human labor saved | Not measured |
| New articles from this batch | Writing and review are the next step |
The benefit we can show now is a source library with a clear path into article production. Search traffic, leads, and ranking changes will need their own evidence after articles are published. A transcript count can’t prove those outcomes.
We’re recording the method using the BlitzMetrics meta-article process so the next run starts with the fixes: test a small batch, preserve source files, resume failed downloads, and match links by ID.
If you’re doing this for your own channel, start with the inventory guide above and test three videos before collecting the rest. If you want help turning an existing video library into a publishing plan, start with a marketing audit. Bring the channel and the website where the articles will live.
Originally published .

