How We Organized Derek Moneyberg’s YouTube Transcripts

Derek Moneyberg seated at a table wearing a black and gold Moneyberg shirt.

We organized 668 Derek Moneyberg YouTube transcripts. If your videos are hard to find and reuse, start with a list. Here’s how we linked each transcript to that list so our team can turn it into an article.

We’re building an article workflow around Derek Moneyberg’s personal website. His YouTube video library gives us the source material. Our YouTube channel inventory guide explains how to catalog that material before assigning it to a writer.

Derek Moneyberg seated at a table wearing a black and gold Moneyberg shirt.
Derek Moneyberg. Photo supplied for his personal-brand content.

Count the source videos before collecting transcripts

The channel snapshot contained 691 videos. We focused on uploads in the Videos tab, kept podcasts and personal videos together, and excluded videos under three minutes. Shorts and the separate livestream archive weren’t part of this pull.

That scope matters. “Every video” sounds simple until it mixes full episodes, short clips, and the same conversation on several platforms. We wanted a useful article queue with one record per YouTube video.

ResultVideos
Auto-generated captions collected666
Manual caption tracks collected2
No caption tracks available8
Under three minutes, intentionally skipped15
Total inventoried691

The 668 reading copies contain roughly 2.6 million words. That’s the size of the source library, not a claim that we’ve reviewed every word or written 668 articles. Each transcript still needs editorial checks before we use it.

Of 691 videos inventoried, 668 transcripts were collected, eight had no captions, and 15 were under three minutes.
September 24, 2026 snapshot: every inventoried video has a recorded outcome. Tap the graphic to enlarge.

Let the script handle the repeated work

I used Claude to help write and troubleshoot a Python script, then ran it on my Mac. The script used the yt-dlp project’s Python library to collect video details and available English captions. It downloaded existing captions; it didn’t send each recording through a new speech-to-text model.

We began with a three-video test. Once the test produced three readable files, the script worked through the channel and recorded the outcome for each video. A saved progress log let it resume without downloading successful files again.

The script selected one English caption track per video, preferring a manual track when available. It saved the original timestamped caption file and created a separate Markdown reading copy. It removed caption markup and overlapping repeated lines while retaining timestamps that help an editor return to the recording.

Keeping both versions gives us a way to inspect the source if the cleaned copy looks wrong. Removing every repeated sentence would be a mistake: speakers sometimes repeat a point on purpose.

Fix empty results before trusting the batch

The first setup reported missing captions for videos that later downloaded successfully. We stopped, installed an updated yt-dlp in its own Python environment, added the JavaScript runtime used by the tool, and corrected the caption-selection logic. The small test then passed.

Because several things changed together, we can’t assign that recovery to one change alone. The useful lesson is to test the output before letting a tool label an entire channel “no captions.”

A network interruption also left three videos with errors. A later run recovered all three. The eight remaining videos still had no caption tracks, so they stayed marked unavailable. We didn’t invent text to fill those rows.

The first empty-result safeguard also stopped a retry of known captionless videos. We adjusted it to distinguish a fresh run with no successes from a cleanup pass through previously empty records. Recovery behavior needs testing too.

Keep the video ID with every transcript

We named the reading copies with this pattern:

YYYY-MM-DD-slug-VIDEOID.md

The date helps with sorting. The short title helps a person recognize the file. The exact YouTube video ID connects it to the right recording even if the title changes.

Each file also records its source URL, publication date, duration, caption type, and retrieval date. Auto-caption text carries a reminder to verify names and homophones against the recording before quoting it.

One exact YouTube video ID links the original recording, its transcript file, and the corresponding article-tracker row.
Match by video ID, even when a title changes. This diagram illustrates the workflow without exposing client files.

We kept the source index separate from the full text. An editor or agent can select a useful video from the index, then open that transcript. There’s no need to load millions of words just to find one episode.

Put the files where the next writer can find them

After extraction, we uploaded all 668 reading copies to a Raw Transcripts folder in Google Drive. We put the existing live Google Sheet in the client’s parent folder and added the direct transcript link to each matching video row. The original timestamped caption files remain in the source archive.

The tracker has three tabs:

  • Read Me: how to use the sheet and where the working files live.
  • Personal YouTube + Podcasts: videos on Derek’s channel, including podcast episodes.
  • Featured In Videos: appearances on other people’s channels.

The whole tracker is the article queue. A podcast episode on Derek’s channel belongs in the second tab; it doesn’t need another copy in a separate podcast queue. The nine featured-video records were outside this channel extraction.

For each source, the sheet keeps the YouTube URL, transcript link, transcript status, article status, QA status, and eventual article link. “Transcript received” and “article published” are separate fields because they describe different work.

Four stages: collect caption reading copies, organize Drive files and Sheet links, review the recording and text, then publish approved articles.
The collection and organization are complete. Selecting, writing, and reviewing the articles come next.

We matched the links by exact video ID. Then we checked the Drive folder against the expected filenames, file sizes, and uploaded file IDs. The final folder contained 668 unique expected reading copies. We also read back every written tracker link and compared one downloaded transcript’s text with its local source.

Those checks catch a problem that a successful upload message can miss: a file may exist while the sheet points to the wrong recording.

Turn the library into useful articles

This puts Derek’s existing recordings into the Process stage of the Content Factory’s four-stage workflow: Produce, Process, Post, and Promote. The recordings already exist. We’ve now organized the text and the handoff to writing.

Next, our team will select a source, watch the recording, and identify the question it answers. We’ll check names, quotes, and the meaning of the conversation, then write in Derek’s voice. The article writing and review guidelines govern that editorial pass.

We’ll also check whether his site already answers the same question. Several videos may support one strong article. A single interview may contain several distinct ideas. The source count shouldn’t become a publishing quota.

The SEO Tree content structure gives each article a place: connect a specific story to its broader topic and link related material where it helps the reader. Derek’s site becomes easier to explore when readers can move from an article to the original conversation, his biography, or another relevant lesson.

We’ve already documented another format in our Derek Moneyberg story-card example. That work used positive mentions to make visual assets. This transcript library supports longer articles. Both start with existing source material, and both need source checks before publication.

The purpose of the content matters too. Our Pete Hazzard website audit connects recommendations to relevant speaking inquiries. For Derek, this library gives writers source material for his personal-brand site. Choosing what to publish should start with the audience and the question being answered.

Separate the measured result from the next outcome

Claude helped with scripting and troubleshooting; I ran the extraction commands and set the scope. ChatGPT then handled the Drive organization, tracker updates, and reconciliation checks. The software did the repetitive collection and matching. Editorial judgment is still part of choosing and preparing each article.

The extraction handoff records roughly 7.5 elapsed hours from initial setup to the final retry, with much of the main run unattended. We didn’t capture a reliable hands-on time total or a complete billing record, so a dollars-saved comparison would be guesswork.

MeasurementWhat we can report
Source inventory691 video records
Transcript files and tracker links668 of each
Extraction methodExisting captions, with no per-video language-model call
Total AI tokens and allocated costUnknown
Human labor savedNot measured
New articles from this batchWriting and review are the next step

The benefit we can show now is a source library with a clear path into article production. Search traffic, leads, and ranking changes will need their own evidence after articles are published. A transcript count can’t prove those outcomes.

We’re recording the method using the BlitzMetrics meta-article process so the next run starts with the fixes: test a small batch, preserve source files, resume failed downloads, and match links by ID.

If you’re doing this for your own channel, start with the inventory guide above and test three videos before collecting the rest. If you want help turning an existing video library into a publishing plan, start with a marketing audit. Bring the channel and the website where the articles will live.

Originally published .

Dylan Haugen
Dylan Haugen
Dylan Haugen is a professional dunker, content creator, and editor at the Content Factory, where he transforms podcasts and interviews into strategic brand assets. He collaborates with Dennis Yu to support young entrepreneurs and business owners in building their personal brands through education, transparency, and effective content marketing. As the host of the Dunk Talk podcast and a dedicated advocate for establishing dunking as a recognized sport, Dylan combines athletic expertise, storytelling, and digital strategy to help elevate the next generation of creators.