Podcast to Articles: One Episode Becomes a Whole Section
What it does: Turns one podcast episode — on YouTube, on Spotify, or on both — into a set of articles carrying what you actually said: your arguments, your figures, your guest’s own words, each embedding the episode, linking to the others, and ending with your subscribe link.
Antradus AI Pro v2.5.10 · Last updated 2026-08-15
Why an episode needs articles at all
A two-hour episode is the best-researched thing you make, and it is invisible. Search engines cannot read speech. AI assistants cannot quote a video. Your back catalogue is a library nobody can search.
The usual fixes don’t work. A raw transcript ranks for nothing — it’s unstructured, repetitive and unreadable. Show notes are too thin to rank. A “watch my video” landing page is worse than either: it holds nothing back for a reader who arrives cold, so it never earns the visit that would have converted them.
What works is a real article that happens to be built from your episode: complete enough to rank on its own, specific enough that only you could have written it, and pointed at the video for the people it convinces.
What makes this different from the SEO Cluster tab
They look similar and behave very differently.
| SEO Clusters | Podcast to Articles | |
|---|---|---|
| What the source is used for | Planning the titles only | Planning and writing |
| How much is read | The first ~12,000 characters | The whole transcript, in passes |
| What each article is written from | Its title, plus the model’s general knowledge | A brief of the episode’s own points |
| Quotes | None | Verified verbatim, or none at all |
| Video, sponsors, subscribe link | No | Yes |
The distinction matters most for the money question: a cluster article about your topic competes with your episode. A podcast article promotes it.
Building one
Go to Bulk Publishing → 🎙️ Podcast to Articles.
Step 1 — Say what you’re promoting
The first thing the screen asks is what are you promoting? — Spotify, YouTube, or both. Everything else follows from the answer, and you get one address box for the service you picked (two, in Both mode). Paste the address and click 🎬 Read this episode: the browser extension reads the transcript in your own browser, including multi-hour episodes, and hands it back without opening anything or taking you off this page.
| Mode | You give | What you get |
|---|---|---|
| 🎧 Spotify | The Spotify episode | The best result available here. Punctuated transcript, named speakers, exact timestamps, the episode’s own artwork |
| ▶️ YouTube | The video | Automatic captions — no punctuation, no speaker names. You are asked who said each quotation instead of being told |
| ⚖️ Both (advanced) | Both addresses | Spotify’s words with YouTube’s picture, and you choose what your readers see |
Pick Spotify if you have the choice. Spotify’s transcripts arrive punctuated and diarised — it marks where one person stops and the other starts. YouTube’s automatic captions do neither. That single difference is the whole of the “who said what” problem below, and everything in Who is speaking exists only when a Spotify link was given.
The address has to match the mode. Paste a YouTube link while Spotify is selected and you are told so immediately, by name — “That is a YouTube address, and you chose Spotify above” — rather than having it quietly re-sorted into the other box. Nothing outside YouTube and Spotify is read at all: a link to your own site, an RSS feed, an Apple Podcasts page are ignored rather than half-read, because a transcript is what this needs and only these two hand one over.
Both mode: the best of each
Each platform is better at one of the things this feature needs, and neither is better at both. In Both mode you give it both addresses and then say what happens with them.
The transcript always comes from Spotify. That one is shown as a ticked box you cannot untick, because it is the reason the mode is worth using — and because a person choosing which service to feature deserves to know which one the words came from. What the article is built from and what your readers see are separate questions, and the rest of the box answers the second one:
| Control | What it decides |
|---|---|
| Which player goes in each article | The embed at the top. YouTube by default. The other service is still offered as a plain link beside your subscribe link, so readers always get both |
| Which timestamps | Spotify (exact) or YouTube (measured — see below) |
| If the two clocks cannot be matched | What to do when YouTube’s timings can’t be trusted: fall back to Spotify’s, publish without timestamps, or stop and let you decide |
| Read the show notes from both | On by default. A link one platform truncated is often whole on the other. Unticked, the second platform’s page is not opened at all |
The picture comes from YouTube in this mode: its thumbnail is a 16:9 still made for this episode, which is the shape a featured image wants. Spotify’s is square and has to be filled out to fit — but on a show that commissions artwork per episode it is the guest’s face too, which is why Spotify mode uses it happily.
Two links to the same service are refused — that is two episodes, not one episode twice.
Putting Spotify’s words on YouTube’s clock
This is the part of Both mode worth understanding, because it is the one that could be silently wrong.
Every quotation carries the moment it was said, as a link that starts the episode there. The words were read from Spotify, so the seconds are Spotify’s. The two platforms do not agree on where zero is — a YouTube upload with a cold open runs ahead of the audio feed by however long that open lasts, and an ad break inserted into one and not the other moves everything after it again. Printing Spotify’s seconds under a YouTube link would send a reader to the wrong sentence, and a reader who follows a timestamp and hears something else concludes the quotation was invented.
So the offset is measured, not assumed. When you ask for YouTube timestamps, the episode is read twice — Spotify for the words, YouTube purely for its caption timings — and then passages from across the Spotify transcript are located in the YouTube captions. Each one that is found gives a real pair of moments, and the difference between them is the offset. It is done in dozens of places rather than one, so:
- a phrase that occurs twice in either recording is thrown away rather than believed, because that is the one thing that produces a confidently wrong answer;
- anchors that disagree with each other are dropped, and an inserted ad break shows up as a step the map bends around rather than a slope smeared across the minutes either side;
- the anchors have to cover the episode, not just its first ten minutes.
Then, once the episode has been read and the quotations are known, each quotation is looked for in the YouTube captions too — and every one that is found adds an anchor sitting exactly where a timestamp is about to be used. This matters because the passages above are spread evenly across the episode, which measures the offset somewhere; a quotation is where the offset is actually spent. Between two anchors either side of an ad break, a quotation can be given the offset from the wrong side of it. An anchor on the sentence itself cannot be.
A pin is authoritative only for the moment it sits on. Everywhere else the evenly-spread passages still do the work, and that is deliberate rather than tidy: measured against six moments checked by ear on a real episode, letting the quotation anchors into the general calculation made the un-anchored moments slightly worse. Where the two recordings genuinely run at a constant offset, the error in any one anchor is caption chunking — one service starting its caption a few words earlier — and that is noise. Spread-out passages average it away; an extra anchor right beside where you asked just hands you its own noise. So a quotation found on both gets the exact second its words start over there, and everything else keeps the behaviour that was already landing within a second or two.
It costs nothing — the words were already checked against the recording — and it is put through every test above rather than trusted for where it came from. The two services ran different speech recognition over the same audio, so a long quotation has many chances to disagree over a word; a short run of words from inside it has far fewer, and is still said only once in an episode. A quotation that occurs twice, or that cannot be found at all, is simply not anchored: it still gets its timestamp from the anchors around it, exactly as before. And an anchor that claims an offset no neighbour supports is dropped, however well its words matched — that is the only failure worth being afraid of here.
The plan screen then tells you what was found — “The YouTube version runs 37 seconds ahead of the Spotify one, measured across the episode”, and how many quotations were found on both recordings — and every printed timestamp moves with its link, so the number a reader sees is the moment on the recording they are being sent to.
The link itself opens two seconds before that number, on purpose. A measured offset is good to about a second either way, because its anchors are caption chunks and a caption chunk starts where the chunk starts rather than where the sentence does. A line Spotify puts at 0:05 is at 0:04 or 0:06 on YouTube and nothing here can tell which. Those two outcomes are not worth the same: a second late opens the player mid-sentence or past the line, and the one reader who checks a quotation concludes you invented it, while two seconds early costs them the tail of the previous sentence — which is what anyone scrubbing to a quotation does by hand anyway. So the link aims at the harmless side. The printed number stays where it is, because it is a claim about when the words were said and it is the true one. Timestamps that were not translated — any single-service episode — are read straight off the recording you are being sent to and get no run-up; there is nothing for them to be early for.
If it cannot be measured — a re-recorded intro, a heavily edited upload, a video with no captions to read — nothing is guessed. Your answer to “if the two clocks cannot be matched” decides what happens, and all three answers are honest ones. The warning also names which of the five ways matching can fail actually happened: passages never found in the other recording usually means the two links are not the same episode, while passages found more than once means a caption track that repeats itself.
If you are told an episode arrived more than once: that is YouTube leaving an already-opened transcript panel in the page, so the extension saw two copies of the same recording. The plugin now reads it once regardless, which is why the message is informational — but updating the browser extension fixes it at the source, and until you do, every YouTube episode is being downloaded twice even though it is only read once.
Reading twice takes about twice as long. That is the price, and it is charged only when you ask for it: choosing Spotify timestamps reads the YouTube page for its show notes alone.
You’ll see roughly how many minutes of speech came back, and how many requests the planning will take. That number matters: the episode is read in passes of about 20,000 characters, one request each, plus one to plan. A 30-minute episode is about 3 requests; a three-hour one is closer to 20. This is the honest cost of reading the whole thing rather than the first fifteen minutes.
Then click 🎧 Read the episode & plan. It works through the episode a part at a time, telling you where it is — “Reading the episode — part 7 of 19” — because a three-hour episode takes a few minutes and a silent spinner tells you nothing.
Each part is its own request, deliberately. Web servers kill a request after 30 to 60 seconds, so reading twenty parts inside one request doesn’t finish slowly — it fails outright with a server error. Splitting it also means progress is kept: if your connection drops or a part fails, clicking again carries on from where it stopped instead of re-reading (and re-paying for) the whole episode.
Leave the tab open while it reads.
Step 2 — Set your channel
Paste your channel address once — a YouTube channel or a Spotify show page. It’s remembered, and it’s what the subscribe link at the end of every article is built from. A YouTube channel opens YouTube’s own subscribe dialog rather than just landing on your channel page; a Spotify show page is where the Follow button lives, so the link there says Follow rather than Subscribe, because that is what the button says. If you leave it empty, articles still link to the episode; they just can’t offer one-click subscribing.
Step 3 — Choose how many articles
| Choice | Result |
|---|---|
| Let the AI decide (recommended) | The episode article plus however many spin-offs the conversation genuinely supports |
| Just the episode article | One article, nothing else |
| 4 / 8 / 12 spin-offs | Exactly that many |
More is not better. One strong article per episode outranks ten thin ones, and a back catalogue turned into hundreds of overlapping pages can drag a whole site down. If an episode only really covered three things, ask for three.
Style and tone: reported, but spoken
Leave these on Blog and Conversational. They are the defaults, and they are the two the articles are tuned around.
The register these articles want is not a choice between a reporter and a friend — it is both at once. A reporter’s discipline: everything attributed, nothing asserted that nobody said, the specifics kept rather than summarised away. A talker’s surface: short sentences, ordinary words, “you” rather than “one”, and the point in the first line instead of after a run-up. That combination is what makes a reader who arrived from a podcast keep reading, and it is written into the prompt as countable rules rather than left to the two words in these dropdowns — most sentences under 22 words, fewer than one in eight under eight words and never two of those side by side, no sentence that opens with It is worth noting, no adjective telling the reader that something is interesting.
Those counts are deliberately for the finished article, not for each section, and 2.5.8 rewrote several of them that way after a measured run showed why. A per-section allowance is an allowance a writer spends: one keyphrase per section across six sections produced fourteen, and at least one short sentence per section produced a third of the article in fragments. The same rule stated once, as a total, holds.
What actually makes a reader excited is never the tone setting. It is the specifics: the line your guest really said, the number, the name, the story they told. Those come out of the episode, and nearly everything in this feature exists to stop them being flattened into “start small and build up” on the way to the page.
How long an article is — and why you cannot set it
There is no word count anywhere in this feature, and that is the point. Until 2.5.8 there was one: a spin-off was asked for 900–1400 words and four to six sections. Handed a stretch of episode carrying one projection and three quotations, the only way to obey both numbers is to say the same things repeatedly — and that is exactly what came back. One measured article ran to 1,483 words in which the same population projection was restated under four of its six headings, its timestamp printed five times, and the keyphrase used fourteen times.
Nothing in it was disobedience. Every anti-padding rule was already in the prompt and every one of them lost to the number, because a length target on material of unknown size is an instruction to pad.
So the size is now counted from what the episode actually gave that article: its verified quotations, figures, the speaker’s own qualifications, and the objections put to them. Roughly one section per two of those, and the article ends when the material does. An episode that genuinely only said three things about a subject gets three sections and a short article — which is the honest result, and it reads far better than the same three things said six times.
Two consequences worth expecting:
- The same episode can produce articles of very different lengths, and a short one is not a failure. It means that stretch of the conversation was thin, which is information about the episode rather than about the plugin.
- A source with less material makes shorter articles. A Spotify-only run loses your show notes — so no sponsors and no links list — but it keeps the transcript, the timings and the speaker turns, so the articles themselves are unaffected.
The other choices narrow it. News tightens further and drops the second person — right for an episode built on an announcement or a set of figures. Listicle and How-To only work when the episode genuinely is a list or a method; forced onto an ordinary conversation they invent a structure that was never there. Formal and Authoritative read as a report and cost you the show’s own voice. Humorous is risky, because the jokes would be the article’s rather than your guest’s.
There is no Opinion style on this tab. Every other rule here exists to stop the article having a view of its own about a real person’s ideas — and a dropdown asking for one out-argues all of them, because it reaches the model as a single word at the very top of the instructions. It was removed in 2.5.7.
Step 4 — Check the plan
You get the episode article plus each proposed spin-off, and a line saying how the episode was divided. Below it, what was verified:
- how many parts the episode was read in
- how many quotes survived checking
- the guest’s name, if the episode has one
- any sponsors found, with their links and codes
Nothing is written yet. Untick anything, rename anything — and the sponsor and link tables are editable here too, so this is where you fix a misread brand or add a sponsor the description never mentioned.
The ⓘ icons
Almost everything on this screen has an explanation, and for a while all of them were on the page at once — fourteen paragraphs of good advice stacked between four controls, which is a page of reading rather than a screen you can use. They’re now folded into the small ⓘ beside whatever they explain. Hover it, or tab to it and it opens; click it to pin it open on a touch screen, and press Escape to close it.
What stays unfolded is anything that reports state: a warning, a count, something that just went wrong. A hint explains and can wait to be asked for. A warning has to be seen.
What every article gets
Assembled by the plugin, not written by the AI — so the links are always right and can never be duplicated or invented:
- The episode, embedded at the top, with a line naming the guest — the player being whichever service you’re promoting, or both if you asked for both.
- A watch-and-subscribe section at the end, using your own call to action (editable in Settings). The other service, when you gave one, is offered here as a plain link to the episode.
- The sponsor block, if the episode had sponsors and you left the option ticked.
- Links to the other articles in the set, added once they all exist.
The marks on the episode links
Each link on that closing row carries its service’s own logo — the red YouTube badge, the green Spotify circle — in the real colours, drawn identically on every device. They are two small files inside the plugin, so nothing is fetched from anybody else’s server and there is nothing extra for a reader to load.
The subscribe link takes the mark of wherever it actually points, which is not always the episode’s. You can quote the Spotify recording and grow a YouTube channel, and those two links then sit side by side pointing at different places — a green Spotify circle beside a link to youtube.com is the kind of small wrongness that makes a reader distrust everything around it. It is read from that link’s own address, so a URL that merely mentions a platform in a query string gets no icon rather than the wrong one, and a channel link to anything else entirely — an Apple Podcasts page, say — is published exactly as it is, with no icon and nothing broken.
The icons are also the reason to know one thing about WordPress: it deletes inline <svg> from post content. An icon written that way vanishes the moment the post is saved, leaving only the link text, and an image built from a data: address is mangled into a broken path for the same reason. That is why these are ordinary images pointing at files. Loosening that filter would be a genuine security hole, so it is not something to work around.
If the plugin is ever removed, those images 404 on old posts. The links themselves still read correctly — Watch this episode on YouTube — because the icon never carried the meaning. That is also why they have no alt text: a screen reader should not announce “YouTube” and then read “Watch this episode on YouTube”.
Quotes: nothing is ever put in your guest’s mouth
This is the part worth understanding, because it’s the risk that matters when a named public figure is involved.
Quotes are collected while the episode is being read, then checked word-for-word against the transcript. Punctuation and capitalisation may be fixed — a spoken sentence has neither. Changing, adding or reordering a word fails the check, and the quote is dropped.
Every quotation mark in the finished article is checked, not just the pull quotes. That distinction was learned the hard way: in testing, all six <blockquote> quotes in an article were exactly right, and all three quotations inside sentences were wrong — including a named public figure on record saying life is “short, brutal, and uncertain” when the recording says life “is short and can be brutal”. It read like a quote, it was punctuated like a quote, and nothing had ever checked it.
So the check now runs on both, twice: once on the plan, and once on the finished article.
- A pull quote that isn’t on the verified list is demoted to an ordinary paragraph, with the quotation marks and the attribution removed.
- A quotation inside a sentence that can’t be found in the recording keeps its words and loses its quotation marks, so the paragraph still reads. You are told which one, and why.
- An ellipsis is allowed, and is checked properly: both halves must have been said, in that order.
Nothing is credited twice
The credit line under each quotation — — Jordan B. Peterson (1:18:41) — is printed by the plugin from checked data, and the writer is told not to write one. It obeys, and then used to write the same thing again as an ordinary paragraph directly underneath:
“Think things are bad just because they’re bad?”
— Jordan B. Peterson (1:18:41)
Jordan B. Peterson says that at (1:18:41).
Every quotation credited twice, with the full name and the moment both printed twice. It slipped through because the rule only ever spoke about what goes inside the quotation, and that paragraph is outside it.
The instruction now covers the sentence after a quotation as well, and anything that still gets through is removed from the finished article — you are told how many, because deleting a paragraph from your article is not something to do quietly.
What is removed is deliberately narrow. The test is a whitelist, not a pattern: the sentence goes only when everything left after taking out the speaker’s name and that quotation’s moment is filler the plugin recognises by name. One word it does not recognise is content, so the paragraph stays.
| After the quotation | What happens |
|---|---|
| Jordan B. Peterson says that at (1:18:41). | removed — it says nothing the line above does not |
| He says this at (1:18:41). | removed |
| Peterson says that at 1:18:45. His point is that bitterness and revenge do not merely express suffering… | kept — a redundant opening on a paragraph that carries the argument |
| That is the heart of his warning. | kept |
| He says that at (1:18:41) the market had already turned. | kept |
Short quoted phrases — a term being introduced, a book title — are left alone, because below about five words a match proves nothing either way.
Speech is tidied by the plugin, not by the AI. People restart sentences: the recording really says “you’re the you’re your you’re the only control group”. An article that prints that has to apologise for it, and an AI asked to clean it up rewrites it. So stutters and restarts are trimmed here, by a rule that can only ever delete a repetition — no word is changed, added or reordered.
That rule used to be too eager, and it is worth knowing what it was doing, because nothing caught it. A sentence that begins the same way twice is normal speech, not a stutter — “if they knock the door will open, if they ask they’ll receive” — and the old rule read the second beginning as a restart and deleted everything before it. What was left was still word for word, so every check passed while the quote lost its first clause and its timestamp stayed with the deleted words. Measured across three real episodes it was cutting about one stretch of ordinary speech in forty. A trim now needs the restart to come straight back and to stumble over the same words, which is what an actual false start does. A trimmed quote is also timed from where the trimmed words start, so the link lands on the sentence you are reading.
A new chapter means the recording stopped saying
Spotify’s transcript is one flat list carrying three kinds of marker: a speaker label, a chapter heading, and the lines themselves. A label stays in force until the next label — that is how a diarised transcript works, and it is why the label is carried forward.
A chapter heading is not a label, and a chapter is very often exactly where the conversation changes hands, because that is what a chapter is for. Spotify writes no new label at that point. So the previous speaker used to run straight through the heading and into the next person’s words, silently, with nothing anywhere to say so.
It cost a real article. The host opened a new chapter by reading out Census Bureau population projections; the guest had been speaking before it; and the article credited every figure to “the Census Bureau figures Peterson cites”. Nothing else in this feature could have caught that — the numbers were checked, the words were checked, and the only thing wrong was the name on them.
Since 2.5.7 a chapter start means the recording has stopped telling us. Nobody is named until it says again.
- Nothing tries to work out who the new speaker is. A passage that opens a chapter is quoted with no name rather than under a guess — the same choice this feature makes everywhere else, and you can still set it yourself on the plan screen.
- A chapter that does open with its own speaker label loses nothing: the label wins.
- The read tells you how many chapter starts this affected, because it is a genuine trade rather than a free win.
- The searchable recording is untouched. The markers come back off the words before anything is searched, exactly like the speaker labels and the caption timings, so every quotation still verifies byte for byte.
Measured on a real 203-minute episode: not one of its 26 chapters carried a speaker label, including the first, where the host opens the show and is never named at all. This is how the format works, not an occasional glitch — Spotify has one title slot per marker, and where there is a chapter the chapter takes it.
Browser extension v1.0.11 or later does the other half, and it is worth updating for. Two things happen there. The extension stops carrying a speaker through a chapter, so the wrong name never reaches your website in the first place. And when Spotify names that speaker again a moment later, the extension now passes the label on instead of skipping it as a repeat — which is what returns attribution to the conversation as soon as the recording offers it. On that episode the unnamed stretch after a chapter went from about two and a half minutes down to forty seconds. With an older extension nothing breaks; those passages simply stay unattributed for longer.
One thing you may notice: if the episode is already open in one of your tabs, the plugin used to read that tab quietly, without clicking or scrolling. On screen a chapter heading and a speaker name look identical — both are just a bold line — so that read cannot tell where the chapters are. Rather than write an article on attribution it cannot stand behind, v1.0.11 opens its own tab for a few seconds instead whenever the episode names its speakers. Episodes with no speaker labels have nothing to get wrong and are still read straight from your tab.
When the recording says who is speaking, that is who it was. Spotify diarises its transcripts: the words arrive marked with which turn they belong to. Those labels are lifted off the text before anything is searched — the same treatment the caption timings get, and for the same reason, because “Speaker 2” left sitting inside a sentence would stop a genuine quotation matching the recording at all.
What the labels then buy is attribution that is read rather than guessed — and it gets its own step on the plan screen, above the quotations rather than inside them.
Who is speaking
On a Spotify episode the plan screen shows one small table with a row per voice:
| The recording calls them | We think this is | Why we think so | Hear this line |
|---|---|---|---|
| Speaker 1<br>“Jordan, that’s the part I wanted to ask you about…” | Chris Williamson (host) | Says “Jordan Peterson” far more than the other voices do, and is never called it — that is the person asking the questions. 38% of the talking, 6 quotations below. | 1:10 |
| Speaker 2<br>“Well, the thing about meaning is…” | Jordan Peterson (guest) | Named by the voice above and does most of the answering. 62% of the talking, 30 quotations below. | 0:56 |
That is one question about the whole episode, and answering it attributes every quotation in it. It used to sit as the first row of the quotations table, where it read as the beginning of a thirty-row chore rather than the thing that finishes the job in one click.
The names offered are a real suggestion with its reason attached. The old default was whoever talks most is the guest, which is true often enough to be useful and wrong in the case that matters — an episode that opens on a clip of the guest, or a host who does most of the talking. The stronger signal is who says whose name: a host says the guest’s name constantly, introducing them, addressing them, reading out their book, and almost nobody says their own. So the voice that names the guest far more often than it is named is read as the host, and that arithmetic settles both voices at once. Talking time is only the tie-break, and when it is the tie-break the reason says so out loud — “nothing in the recording confirms it”.
Two things keep that count honest, and both were learned from an episode it got backwards:
- It is counted per thousand words, not per episode. A voice holding three quarters of the recording picks up more of every term simply by being longer — the guest’s own name included, because a diariser that lumps the interviewer’s question into the guest’s block files that name under the guest.
- It is never allowed to name the majority voice the host. An interviewer who talks two thirds of the time is rare; that lumping is not. Where the two signals disagree the share wins, and the reason on screen says the count disagreed and marks the row as the one to check first.
Each row also carries the first substantial thing that voice says and a link to the moment it says it. One line of somebody’s own words identifies them faster than any percentage can; the timestamp is there for when it doesn’t. The time is that line’s own moment, not the first noise the voice makes — a voice opens with “wrong, beginner’s luck” far more often than with a sentence, and four seconds there identify nobody.
It is still a guess, and every part of the screen says so. Leave a voice on Not sure and there is no name to publish its lines under, so those rows arrive unticked in the table below — tick one and it is quoted as an unidentified speaker.
The quotations table below then reports that answer rather than asking again. On a labelled recording each row shows who said it as plain text, not as a dropdown — because the recording already answered, and thirty dropdowns invited somebody to overrule the evidence line by line. (One row is the exception, and it is the one the recording could not answer; see below.) Change a name in Who is speaking and every row that voice owns updates underneath. The per-row dropdowns are still there on a YouTube episode, where there are no labels and the question genuinely is open.
One case still comes back unattributed on a labelled recording, and it is the important one: a quotation that starts in one person’s turn and ends in the next. Every word of it is real, so nothing else here would ever catch it — and no single name is right for it, because it is two people’s words in one set of quotation marks. Those arrive unticked — there is no name to publish them under — and the plan says how many.
Those rows are the one place a labelled recording still asks you. The row says Not named — this line crosses a change of speaker, and underneath it there is a dropdown of the people in the episode. It starts on Leave it unnamed, so the row stays unticked and nothing changes if you ignore it — which is usually the right answer, because the recording genuinely does not say. But “the recording cannot say” is not “nobody can”: often the second turn is an interruption or a “yeah, exactly”, and the line really is one person’s. Click the timestamp, listen for four seconds, and if you know whose it is, name them — and edit the quotation down to their words while you are there, since anything you leave in still has to be in the recording word for word or the row is dropped. Naming somebody ticks the row, and your answer beats everything else in the system, exactly as it does on a YouTube episode. Naming a voice in Who is speaking never reaches these rows: they belong to no voice, so an answer you give here is yours alone.
Without labels, nobody is named where a guess is known to go wrong. YouTube’s automatic captions mark nothing, so at the moment one person stops and the other agrees, an inference goes wrong — the host’s sentence gets attributed to your guest. Where an affirmation (“exactly”, “right”, “absolutely”) follows a line, the quote arrives with no name attached rather than the wrong one — and therefore unticked — and the plan tells you how many.
And either way, you settle it. A label is evidence and an affirmation is a hint; neither is you having listened. So the plan screen shows a Quotations table: every checked quotation, the moment it was said as a link into the episode, and who said it — reported from the recording where there are labels, and offered as a dropdown of the people in it where there are none. Setting a row beats both.
- On a labelled recording the name is read straight off the recording, through whatever you set in Who is speaking. Change it there and every row that voice owns follows. The only rows carrying their own dropdown are the ones that crossed a change of speaker, because no voice owns those.
- On an unlabelled one, leaving a row alone gets you exactly what you got before: the guest, which is what the writer assumed on its own. The difference is that the assumption is now on screen where you can see it.
- Set a row and your answer beats everything else in the system, including the plugin’s own doubt about that line.
- Choose Not sure and the row unticks itself. Tick it back on and the quotation is used, just never attributed — not by a name, not by “the guest”, and not by a “he” in the sentence after it, which is how a careful non-attribution used to get undone one line later.
- A row nobody can be named for starts unticked, whichever way it got there. The words were said — that was checked — but publishing one prints — unidentified speaker under a pull quote, and that is a decision about your article rather than a fact about the recording. So it is yours to make: tick it and it goes out exactly as it reads on screen. The line above the table says how many are sitting there and why.
- Untick a row to drop that quote entirely.
Click the timestamp and the video opens at that second in a new tab. That is the whole point of the column: settling who spoke takes about four seconds of listening, and asking somebody to scrub a three-hour recording by hand is asking for the answer they can give without it.
Nothing on that screen is a machine’s opinion about who spoke. The read used to be asked which speaker a quote belonged to, and it answered — from a transcript with no speaker labels, where the answer is not available. On one episode three consecutive lines of the guest’s arrived pre-set to the host. It is not asked any more. A row starts on the guest if the recording can place it, on nobody if it cannot, and on your answer the moment you give one.
Each quotation also comes with a note on whether it stands up alone. Some lines are word-perfect and still useless out of context: “that’s what love should do” is a complete sentence in the room and a riddle on a page, because the thing it points at was said thirty seconds earlier. Those are labelled “checked and correct, but probably not worth an article”. It is a note and nothing else — it never unticks a row, and if you disagree, use it. Thin is a matter of taste; the words are not.
The tics come off, and the punctuation goes on. Automatic captions have no capitals and no full stops, and people end sentences by checking you are still listening — one real plan had twelve quotes out of twelve ending in “right”. So a quotation arrives with the scaffolding trimmed off its ends. Only ever by deleting: no word is changed, added or moved.
That trimming is careful about the words that can be the point of the sentence. “You have to get it right” keeps its “right”, and so does “the young adults are not okay” — an early version cut that to “the young adults are not”, which lost a word and, worse, moved where the quotation ended, so the plugin stopped noticing that somebody agrees right afterwards.
Then the commas go in. A checked quotation with no punctuation in it is word-perfect and unreadable:
this is who I am and if you don’t want me that’s a drag because I’m looking for a job but by the same token I’m not going to pretend to be someone other than who I am so I can work here what a stupid way of starting your career
A rule cannot place those commas. Knowing that “if you don’t want me” ends before “that’s” means reading the clause, and a rule that guesses the boundary wrong moves a qualifier from one clause into the other — a change of meaning wearing the costume of a typographic fix. So once the quotes are checked and trimmed, they go through one more pass whose only job is punctuation, and it lands as:
This is who I am, and if you don’t want me, that’s a drag because I’m looking for a job, but by the same token, I’m not going to pretend to be someone other than who I am so I can work here. What a stupid way of starting your career.
What makes that safe is not the instruction. It is the check afterwards. Every punctuated line is stripped back to bare words and compared with the bare words that went in, and anything that is not identical is thrown away and the original kept. A word added, dropped, replaced or reordered all fail that comparison — and so does a helpful one: where the captions misheard “Sapolsky” as “spolski”, the punctuation pass is not allowed to repair it, because a mishearing gets corrected on the Names the captions got wrong table below, by you, where you can see it happen. The plan tells you how many quotations were left alone for this reason. It is the guard working, not a failure.
The step runs after the plan appears, so nothing waits on it — the commas land a second or two later, and if the model is unreachable you simply get the quotations as they were. Rows you have already started editing are never overwritten.
The quotation itself is editable too, because automatic captions mishear — “Yung” for “Jung”, “naivity” for “naivety”, a swallowed word at the start of a sentence. Fix it in the box.
What you cannot do is write one. Whatever comes back is looked up in the recording again when you press the button: words that are not there do not become a quotation, and you are told how many were left out. An edited line is also re-timed from where the words actually are, not from the moment the original was found — a word or two of difference can put a quotation in the next caption, and carrying the old timestamp over would link a reader to a line that no longer says this.
This is also what finally puts your host in the articles. Seven test versions never named the interviewer once — and that was not just a missing credit. With only one person named, every quotable line had exactly one plausible owner, so the host’s own sentences were written up as the guest’s.
And the same rule pointed the other way. An invented quote puts words in your guest’s mouth; the opposite failure takes words out of it. An article carried the sentence “if you tell a sufficiently seductive lie, people do not fall for you” in its own voice — seven words straight off the recording, with no quotation marks on them. Every check above passed it, because the words really were his. What was missing was the credit: the guest’s own formulation read as the writer’s.
So the finished article is also checked the other way round. Any run of six or more words lifted from the recording and written as the article’s own is reported before you publish, with the phrase quoted so you can find it. Names are excused — there is no other way to write “the Harvard Graduate School of Education” — as are the figures and their labels, which the article was told to carry. What is left is phrasing that had alternatives, which is what makes reusing it unattributed a choice. It is reported, never edited: adding quotation marks needs a speaker and a timestamp to be honest, and rewriting someone’s sentence for them is your call.
All of this is enforced in code rather than asked for in the prompt, because in testing the model was told exactly once not to invent quotes and did it anyway.
Every pull quote is set by the plugin, not by the writer. A quotation is only checkable with three things on it — the marks, a name, and the moment — and across three test versions of one article the writer supplied a different two of them each time. So it is no longer asked. It writes the words; the plugin adds the credit line underneath:
“You’re only courageous if there’s a risk.”
— Jordan B. Peterson (1:22:13)
Anything the writer put there itself is taken off first, and never at the cost of a word of the quotation: each cut is tried, checked against the verified wording, and undone if it went too far.
Where nobody could settle the speaker, the credit line says “unidentified speaker” out loud. That looks like a small thing and it is the most important line on this page. A bare quotation sitting in a paragraph about one person has been attributed to that person — a reader has nowhere else to put it — and that is precisely how the host’s own line went out as the guest’s in four consecutive test articles, with no name anywhere near it. Saying it costs a little polish and closes the last route by which a reader can be misled about who spoke.
The check runs the other way as well: if an unattributed quotation ends up in a passage where exactly one person is named, you are told before you publish, because the paragraph is doing the attributing even though the quote is not.
Timestamps: a link to the moment it was said
Quotes and figures can carry the point in the recording where they happen, written as (12:34) and linked so a reader lands on that second of the video.
The times are worked out by the plugin from the transcript’s own caption timings. The AI is never asked for one, because a model asked for a timestamp invents it — and any time it writes that we did not give it is removed before the article is saved.
Three details worth knowing:
- Which recording they point at is normally the one the words came from — the only place they are certainly true. Both mode can move them onto the other service, and only ever by measuring the difference between the two clocks; see Putting Spotify’s words on YouTube’s clock above. When they move, the printed number moves with them, and the link opens two seconds before it so a measurement that is a second out still lands you in front of the line rather than past it.
- A figure is timed by where it was actually said, not by the first place those digits appear. “20 plus years”, “every 20 minutes” and “a 20% discount” all contain a 20, and an earlier version linked all three to the same wrong moment. Where the recording can’t say which occurrence is meant, the figure keeps its place in the article and simply gets no link.
- This needs the browser extension at v1.0.3 or later, which sends the caption timings along with the words. With an older one everything else works exactly as described, with no times.
The same column runs down the Names the captions got wrong table, where it earns its place twice over. That table asks you to approve a rewrite of something a person said, and on the page isai bin → Isaiah Berlin is a guess you are being asked to take on trust. Click the time and you hear it, and it stops being a guess. It costs nothing to work out — the check that proved the mangled spelling is in the recording already knows where it is.
If you retype the left-hand box, the moment beside it disappears. It belonged to the words that were there before, and a link that opens the video on something else is worse than no link at all: somebody clicks it, hears a different sentence, and has no way to tell whether the correction or the plugin is the thing that is wrong.
Each article quotes its own part of the episode
This is the subtlest failure the feature has had, and the one worth understanding, because everything about it looked correct.
Quotes used to be shared out between the articles by position in a list. That takes no account of subject. An article about young adults and mental health — a topic running from 1:22 to 1:29 in the recording — was handed quotes from 1:01, 1:04, 1:08 and 1:18, and told to use at least two of them. Every one was word-perfect. Every one had the right speaker and the right timestamp. Every one was about something else: a media pile-on, a newsroom, the Book of Job.
No check could see it, because nothing about the words was wrong. A made-up quote fails a text comparison in a fraction of a second. This one passes every gate in this document and is still misleading, because it implies the speaker said that thing about this subject — and a reader who opens the video at that moment lands in a different conversation.
So each article’s subject is now located in the recording first. The plan writes each article a brief; the plugin finds where in the episode that brief’s material actually sits, and the article is only offered quotes and figures from that stretch.
Some articles have more than one. An episode is not a list of separate topics, and a brief that pulls a theme together — truth and lying and image and persona — is drawing on four passages spread over three hours, not one. So the placing keeps every stretch that genuinely rivals the strongest, not just the tallest, and the article’s quotes are dealt one from each in turn.
Getting this wrong was invisible for the same reason as the original bug. An article covering four passages was placed on the opening four minutes — which won by a single point — and every quote it received came from there. All four were word-perfect, correctly attributed, correctly timestamped, and in scope by the rule as it then stood. Three of the article’s five sections carried a quotation about a different argument. The head of a recording wins this contest by default, because the first minutes of an episode state every theme at once, so any article whose subject was spread quietly ended up quoting the introduction.
Four consequences worth knowing:
- An article with nothing of its own gets nothing. No importing from elsewhere to fill a gap. It is told plainly not to use quotation marks at all, which is the same instruction an episode with no verified quotes produces.
- No line is ever published in two articles. See below — this took a second attempt to get right.
- Several stretches is not “anywhere you like”. At most six are kept, and on a real three-hour episode a four-passage article claims about a ninth of the running time. A single-subject brief still gets exactly one stretch, unchanged.
- Anything that still slips through is flagged, at the top of the note, before you publish — including an article that covers several passages but spends all its quotes on one.
If the transcript has no caption timings — an older browser extension, or a pasted transcript — none of this can be worked out, and the sharing falls back to dealing one quote to each article in turn until they run out. That is also the only situation where the article has no timestamps to mislead anybody with.
Whose line is it
Placing each article on its own stretch left one thing unasked. Stretches are worked out one article at a time, from each brief’s own words, so nothing makes them disjoint — two neighbouring subjects overlap, and inside the overlap both articles genuinely reach the same line.
The first version dealt with that by dealing quotes one article at a time and letting an already-used line sink to the bottom of the next article’s list, on the grounds that a repeat beats a quotation from a different subject. True as far as it goes, and it left the real case unhandled: where an article’s stretch holds exactly one quote, demoting the only candidate changes nothing. It is dealt twice.
Measured on a real 3h13m episode, thirteen articles: five quotations went out in two articles each. And they went to the wrong ones, because the first article to reach a line got it — “I think they’re probably still getting worse” (1:38:05) went to the article about casual sex, whose stretch runs 1:31–1:42, simply because it came first in the plan. The article about population decline, held to 1:36–1:40 and about nothing else, got the same line as its only quote. First past the post is not an answer to “whose line is this”, and it was never asked.
So the whole plan is now dealt at once, every claim on a quotation is compared, the best one wins it, and the losers do not get it at all. In order:
- The article whose own brief talks about that line most. Where the words decide, they decide — though a seven-word quotation usually shares no distinctive word with anybody’s prose, so usually they do not.
- The article that passage is, over one that merely touches it. An episode’s opening states every theme it will cover, so almost every article gets a weak secondary claim on the first few minutes. Without this rule, “the veils have fallen from your eyes” (3:56) went to the article about exposure therapy — which reaches back into the opening on a hill scoring 62% of its own best minute — instead of to the article about cynicism, whose subject those minutes are.
- The tighter passage, then the article with less else of its own, then the plan’s order, so the same episode always deals the same way twice.
A line the winner has no room for is offered to whoever wanted it next, rather than going to waste while a sibling covering the same passage publishes with nothing.
Re-dealt over that same real batch: the five repeats become none, no quotation is lost, and both contested lines land on the article a person reading them would have chosen.
The cost is honest and worth stating: an article whose only in-stretch quotation belongs to a sibling now publishes without one. That is the same trade this whole section is built on — a missing pull quote is a smaller failure than a misleading one, and it is the one you can see.
Two things changed alongside it, because restricting each article to its own stretch means that stretch has to be well covered. More quotes and figures are kept from every reading pass, so a five-minute subject has something in it rather than nothing. And a spin-off now writes four to six sections rather than nine — one subtopic of an episode does not carry nine separately-evidenced points, and splitting thin material across more headings is what produced sections with nothing in them in the first place.
The main episode article works differently, because its honest range is the whole recording: it takes one quote from each stretch of the running time. Before, it took the first few on the list — which is the order they were said — so a survey of a three-hour conversation quoted nothing after the first half hour.
What else each article is handed from its own stretch
Quotes and figures were the first two. Two more travel the same way — dealt from the article’s own part of the recording, or not at all.
The names spoken in it. Every section has to carry something specific from the episode, and a named person, place, institution or work has always counted — but only if the writer happened to reach for one. A section narrating a passage about an academic job market came back saying “he ignored warnings”, for minutes in which the university, the colleague and the length of the working relationship are all said out loud. Those names are now handed over with the article, spelled correctly, each with the moment it was said. A name’s timestamp is a link like any other, which gives a section with no pull quote and no figure something a reader can still go and check.
The places the speaker qualified themselves. An article is not allowed to invent an objection — if nobody challenged a point on the recording, it stands unchallenged. What it is allowed to carry is the speaker’s own limit on their own claim, and people do that constantly: “that doesn’t mean…”, “I’m not saying…”, “I could be wrong about this”. Those lines are now found in the article’s stretch and offered to the writer with their timestamps, to sit next to the claim they qualify.
One case is deliberately left out. Where a hedge falls exactly at the point the conversation changes speaker, it is not somebody qualifying themselves — it is the other person disagreeing, and an undiarised recording cannot tell you which. Presenting a rebuttal as a caveat would misattribute it twice over, so those are dropped.
And the moments the conversation pushed back. That was the half of this that looked impossible: naming who challenged whom needs speaker labels, and there are none. But noticing that a challenge happened does not. The recording still contains “can you not think about yourself in a positive way”, and still contains the answer in the caption after it.
So the objection and the reply are both handed over, with the moment, and the writer is told plainly that it does not know who asked and may not say. “The objection put to that is…” and “asked whether…, the answer given was…” are allowed; “the host challenged him” and “he replied” are not, because inventing the roles is the same failure as inventing a quotation, aimed one step to the left.
The markers for this are deliberately narrow, and the reason is worth stating: a loose one does not produce a weak finding here, it produces an argument that never happened. An early version read the caption fragment “but doesn’t matter you can also retreat into” as an objection — and a writer told that two lines disagree will build a paragraph explaining how. On a three-hour episode the tightened version finds one genuine exchange, which is one more than eight previous versions carried between them.
Percentages travel in packs. Where a study is read out, the read pass reliably comes back with three of the four numbers — and a different three each time. So the transcript is checked directly: for any verified figure with a moment, the caption it sits in and the one after it are scanned for a percentage nobody extracted, and it is added with the words that follow it. This only ever fills gaps in a run already found, so it can never introduce a number from a part of the episode nothing was taken from.
Who brought the statistic
On an interview the host brings the research at least as often as the guest. One real episode had the host introducing a named report, reading out four percentages and a comparison, with the guest’s entire contribution being “yeah, I believe that” — and the article opened by crediting the guest, then repeated the credit three more times.
An automatic transcript has no speaker labels, so the plugin cannot say who read the numbers out. It does not guess. A figure is credited to the report, survey or institution that published it, never to a person — “a report from X found…”, not “your guest points to figures from X”. The only exception is a verified quote in which the speaker is plainly doing the citing.
Where the recording shows a number being read out to somebody — an introduction like “there’s a report I wanted to bring this to you”, or an agreement landing right after the numbers — the figure is marked, and the article is told explicitly to name the study and no person at all.
The specifics have to survive into the writing
The single most common failure in this feature has nothing to do with invented material. It is the opposite: the article compresses everything concrete out of the episode and publishes the lesson without the evidence. “Seven women over 70 outswam him” becomes “start small and build up” — and the second one is what a reader forgets and an AI assistant ignores.
Three things now push against that.
Every article gets figures. Each spin-off is handed the numbers that belong to its own subject. The broad episode article is handed the episode’s figures whatever its brief happens to mention, ordered by what it is most about — because it covers the whole conversation, so everything in it is on topic. That last part was a real bug: the plan is told to put each number with the article covering its subject, which sends them all to the spin-offs, and the main article — the one most people land on — was arriving with none and then being judged for having no specifics.
The opening has to touch the ground. The first paragraph may answer the question in the abstract. The one after it has to carry something real: a figure, a person, a study, an age, a place, or a story somebody told. Two abstract paragraphs at the top is the shape of an article written from general knowledge rather than from your episode, and it is obvious within about ten seconds of reading.
Every section is checked on its own. A section with no quotation and no moment a reader can go and hear is reported to you by name, alongside the section that carries nothing specific at all. Whole-article totals hide this: an article can average well and still have one long heading written entirely from the model’s own sense of the subject.
Nothing here is repaired automatically. A missing figure belongs in a sentence, credited to whoever said it, and a machine pasting it back in would produce exactly the kind of forced line this feature exists to avoid — so you are told what went missing and where, and you decide.
Reading the note in the job list
Every article in a podcast batch gets one, and it always ends “Worth a read before publishing.” It is a checklist, not a list of errors — the article is finished and the quotations in it are already verified. Long ones fold to three lines with a Show more beside them; nothing is truncated, and the CSV export always carries the whole note.
The Queue tab refreshes itself every few seconds so progress from WP-Cron shows up without a page reload. Until 2.5.7 that refresh rebuilt the table, so a note you had just opened closed itself again a moment later — Show more with an invisible Show less on a timer. What is open is now remembered across refreshes, and a refresh that finds nothing changed leaves the table completely alone, so text you are part-way through selecting stays selected.
What each part is asking you to do:
| The note says | What it means |
|---|---|
| CHECK THIS FIRST — n numbers… | Numbers reach the writing through the plan brief as well as through the verified figures, and only the figures are checked. Grouped by sentence, so one sentence is one thing to confirm however many numbers are in it. |
| Phrases taken word for word… with no quotation marks | Your guest’s own phrasing, written in the article’s voice. Either mark it as a quotation or reword it. |
| Sections have neither a quotation nor a moment | That heading has nothing a reader can go and hear. Worth a look; it is the shape of a section written from general knowledge. |
| Sections carry no quote, no figure and nothing named | Stronger version of the same thing — nothing specific in it at all. |
| Reported-voice check | A passage reads as the article’s opinion rather than as something that was said. |
| All n quotations come from one stretch | The article covers several passages but quotes only one of them, so the rest carries no voice. |
| CHECK THIS FIRST — the table has n empty cells | A column the article could not fill on every row. Delete the column or fill the gaps from the recording. |
| CHECK THIS FIRST — n timestamps in a column headed “…” | A time has been printed as though it were a statistic. Head that column When, or take the times out. |
| The table names more than one person in a “Who said it” cell | Each line was said by one person. Name that one, or drop the row. |
| The exact phrase “…” appears n times | The keyphrase repeated past the point a reader stops noticing the subject and starts noticing the phrase. Stuffing also lowers the chance of an AI assistant quoting the page. |
| The guest’s full name is written out n times | Twice is plenty. The surname alone reads as writing; the full name on every mention reads as a machine filling a field. |
Tables, and when there shouldn’t be one
A table is the most quotable thing an AI assistant can find on a page, which is why the general SEO rules ask for one wherever a subject compares options along the same few dimensions. A conversation has no options to compare — so on this tab a table is optional and usually wrong, and saying so is now part of the instructions rather than left implied.
If an article writes one, it has a fixed shape: Topic | What was said | Who said it | When, four columns and no others, every cell filled on every row, one name per speaker cell, and every time taken from the list of moments the recording actually gave. Three rows minimum. If it cannot be filled that way the article is told to write no table at all, because no table beats a table of the article’s own summary.
That is narrower than it used to be, and it is narrower because of a real one. A published spin-off produced a five-row table with entirely legal column headings — and a column headed Figure cited holding nothing but timestamps, two empty cells, a speaker cell naming two people, and every row a shorter version of the paragraph above it. Nothing could see it: the old rule checked the headings, and the headings were fine.
So the finished tables are now read as markup, cell by cell, and anything that does not add up is reported in the job note above. It is a report and never a repair — deleting a column or rewriting a row is an editorial decision, and the two-second look you give the table is worth more than a guess.
The focus keyphrase never bends a sentence
Every article gets a focus keyphrase, and the general SEO rules ask for it in the first sentence and then at a steady density down the page. On an ordinary article that is harmless, because the keyphrase is a phrase. On a podcast spin-off the title is often a question, the plan turns it into two words, and those two words are not English on their own.
One real spin-off was titled “Is cynicism helpful after naivety breaks?” and given the keyphrase cynicism helpful. It opened: “Cynicism helpful is the question people ask right after something breaks.” Then four more times, each sentence bent around the same two words. It reads as the title having fallen into the body, which is exactly what it is.
Podcast articles now override those rules outright:
- A keyphrase that cannot be spoken aloud as an ordinary phrase never appears in that form. It gets the words that make it a sentence, or an inflection, or a question heading carries it.
- The exact phrase appears at most twice in the whole body, and never twice in one section. Its individual words appear as often as the subject needs them — it is the fixed block that has to be rare.
- The article never opens with it, and never opens with a paragraph that reads as the title in disguise.
- No H2 repeats the title or rewords it. The title is answered in the opening paragraph; the first section moves on.
None of this costs you the keyphrase. Yoast and Rank Math read variants and word forms, and a sentence nobody could say out loud was never earning a ranking anyway.
Sponsor reads stay out of your articles
Every episode carries ad reads, and they are full of the kind of number this feature hunts for: 20% off, a 60-day trial, free shipping on your first box. One of them reached a published test article as though it were something the guest argued.
Numbers spoken inside an ad read are now left out of the article bodies. They are not lost — sponsors are published in their own block, with a disclosure line, further down the same page.
Sponsors and links: your reads keep earning
Tick 💰 Include the episode’s sponsors and links and the episode description is read. Sponsors, discount codes and links all live in that description, so getting hold of it is the whole game.
The description comes from YouTube, or from you. Since 2.5.7 there are exactly two sources, in this order:
- What you paste into the Episode description box. This always wins, because it is the one you can see.
- YouTube, read by the browser extension from the same tab it opens for the transcript, and dropped straight into that box so you can check the sponsor block is really in there before you spend a request. If the extension cannot get it, your website asks YouTube itself as a last resort.
Spotify’s description is not read at all — not by the extension, not by your website, not in any mode. What Spotify serves is a share preview: the first couple of hundred characters of your notes with its own “Listen to this episode from …” line in front of them, and the sponsor block below the cut. That would be a manageable half-measure on its own; what made it a bug is that it was winning. On a Spotify or a Both-mode run it took the description box — the box that outranks everything else — and the same episode’s complete YouTube notes were displaced by it and no longer visible anywhere on the plan screen. A truncated blurb standing in for your show notes costs you the sponsor reads it cut off, silently, and the plugin has no way to tell that it happened.
So the shapes are:
| Mode | Where the notes come from |
|---|---|
| ▶️ YouTube | YouTube’s description, automatically. Nothing to do |
| ⚖️ Both | YouTube’s description, automatically — this is one of the three things the second address is for, alongside the picture and the clock |
| 🎧 Spotify | Nothing automatic. Copy the description from the episode and paste it in, or the articles publish with no sponsor block |
That last row is the price of Spotify mode, and the screen says so where you would look for it rather than leaving you to find out on the plan screen. Pasting also works while the episode is still being read, so it costs no extra time.
Your website’s own request is the last resort behind YouTube for a reason worth knowing. YouTube treats a request from your home connection and a request from a web server very differently: from a shared host it commonly answers with a rate-limit or a bot check instead of the page. So the same episode can produce three sponsors on a test install and none on the live site — same plugin, same episode, different address. If that happens you’ll be told, and the fix is to paste the notes in.
When there are two sets of notes, they are joined rather than chosen between — that happens when you paste a description and YouTube supplies one. The check that matters asks whether a URL or a discount code is really in the notes, and a link one source truncated is often whole in the other. Nothing is loosened by it: a code that appears in neither still fails. Each sponsor then carries a small badge saying where it was found — in both, YouTube only or pasted only — and one source on its own is the useful fact, because that is the one that cut your notes short. With a single set of notes there is nothing to compare, so no badges are shown at all rather than every row reading “YouTube only”, which is true and reads as a warning.
The description is split by its own headings, and what happens next depends on whether it labels its sponsors.
When the description says “Sponsors:”
Everything under that heading becomes the Episode sponsors block: brand, offer, discount code. Every URL and code is checked against the description character for character, and anything that doesn’t match is dropped rather than guessed — a wrong affiliate link sends your audience somewhere neither of you chose, and costs you the commission.
Everything else the description linked to — the guest’s book, their social accounts, your own newsletter — goes into a separate Links from this episode list. Those are not called sponsors, because the description didn’t say they were.
When it doesn’t
Plenty of shows don’t label anything. In that case nothing is claimed to be a sponsor at all: every link goes into Links from this episode with the note “Some of these may be affiliate or sponsored links.”
That’s the honest answer. Guessing which links are paid is the one judgement that must not be made on your behalf — calling an ordinary link a sponsor is a false claim, and missing a real one is a compliance problem.
Editing them
Both lists are editable on the plan screen, and what is on screen when you press Write these articles is what gets published. Change a brand the AI misread, fix an offer, drop a row with ✕, or add one with ➕ Add a sponsor.
Adding is the one worth knowing about. Everything read out of the description has to be found in the description to be published — that check exists so the AI can never invent a link. A row you type is not checked that way, because it does not need to be: you are the source. That makes the common case possible at last — a sponsor read that only ran in the audio and was never written in the description anywhere.
A row needs a brand and a real http(s) link to be published. The offer and the code are optional.
What’s always true
- A disclosure line above the links, not below them. Below, it isn’t a disclosure — it’s a footnote.
rel="sponsored"on every outbound link, sponsor or not. The asymmetry is deliberate: marking an ordinary link as sponsored costs you nothing — you were never passing ranking credit to someone’s Amazon page — while leaving one paid link unmarked exposes the whole site to a manual action.- Links back to your own site are skipped. Those are internal links, not episode resources.
Edit the disclosure wording in Antradus AI → Settings → Podcast to Articles.
Images: your guests’ faces on every article
Tick 🎬 Build those images from the episode thumbnail and each article’s featured image is built from the episode’s own picture rather than invented from nothing.
Which picture that is follows the mode. YouTube and Both use the video’s thumbnail — a 16:9 still made for this episode, which is the shape a featured image wants. Spotify uses the episode’s own artwork, taken from Spotify’s public oEmbed endpoint (which answers from any host, including the shared ones where nothing else about Spotify works) and falling back to the picture the extension saw on the episode page. On a show that commissions artwork per episode that is the guest’s face; on one that doesn’t, it is the show’s cover, and every article in the set will share it.
What stays: the people. Same faces, same identities, same clothing, same positions and scale. What changes: everything else — a new scene, built for that article’s subject, with its own objects, depth and lighting.
Everyone on the thumbnail stays on the picture. A panel episode with four faces keeps all four, in the same arrangement, with the new scene built into the space around them. An earlier version kept “the most prominent” one and deleted the rest, which meant the plugin was deciding which of your guests mattered — not a judgement it has any business making about your episode.
Two things are stripped on the way. The thumbnail’s big promotional headline is deleted, because a featured image with “STOP WASTING YOUR LIFE” burned into it is unusable on an article about something else. And nothing is written in its place unless you ask for it.
The result is a set that reads as one section of a site: the same recognisable people across a dozen articles, each picture obviously about its own subject.
Words on the picture
Off by default. Tick 🔠 Put a short headline on those images and each picture carries two to four words about that article — “Online Dating Burnout”, “The Cost of Status”, “Finding Real Meaning”.
This part has been got wrong twice, and both failures shaped how it works now.
The first version asked for a short question. An image model spells worse with every extra word, so it wants to be brief — and being brief, it deletes the words that make a sentence work. Out went “How Find Meaning?”, and later “What Says About Female Attraction?”. Neither is a sentence, and both were burned into a picture where nothing can be edited afterwards.
The second version stopped generating anything and printed the article’s focus keyword instead. That fixed the grammar and broke the meaning: a one-word keyword produced a picture labelled “Meaning”. Meaning of what? Nobody has ever seen a picture called that and known what it was about.
So it is written as a phrase, never a question. A phrase has no auxiliary verb and no subject to drop, which is precisely how the first two failed. Then it is checked before it is used:
- At least two words, at most four. One word is never published.
- No punctuation at all — no question marks, colons or dashes. Image models render punctuation worse than they render letters.
- It has to be about your article. At least one substantial word must be shared with the title or the keyword. A phrase that wandered off-subject is readable and about the wrong thing, which is worse than a dull one.
Anything that fails those checks is thrown away, and the article’s own focus keyword is used — then, failing that, a phrase cut from its title at a natural break. Both of those are human-readable by construction, so there is always something sensible to fall back to.
Three more things are deliberate:
- Shorter is set larger. Two words get big, heavy type that reads across a room; four are sized down so they still fit. If the size and the margin ever disagree, the size gives way.
- It goes where the original headline went. The words are placed in the same part of the frame the thumbnail’s own headline occupied — across the top if that’s where theirs ran, beside the guest if that’s where it sat — so the picture looks like it belongs to your channel rather than like type dropped on a photo.
- It cannot be cropped. Wherever it lands, it is pulled inward until there’s a wide empty margin on all four sides and it is clear of anybody’s face, because your theme will crop the picture a little on every edge. Not being cut off outranks matching the position.
Nothing else in the picture carries writing — not the thumbnail’s original headline, not a caption, not a label on an object. Leave the box unticked and the images have no text at all.
What lands in your Media Library
All three fields are filled, and none of them with anything you would have to delete.
- Caption and description: one plain sentence describing the actual picture — “A moody layered scene of glowing phones, lonely interiors and hazy nightlife.” It’s written in the same request that plans the scene, so it describes the picture that was actually made. The description adds which article the image belongs to.
- Alt text: the words printed on the picture, then that same sentence. Someone using a screen reader gets the headline first and the scene second, which is the order they’d want.
This used to be worse: the caption and description carried the instructions written for the image model — “On the open side, build a layered interior montage…” — which is meaningless to a reader and shows up publicly wherever a theme prints captions. Anything that still reads like an instruction is now refused and replaced with a plain line naming the article.
What you need
An image model that can work from an existing picture. Most cannot — they only write images from scratch, and handed a source they will happily return a photograph of somebody else entirely. So the podcast tab checks your current model before you start and says plainly whether it can do this.
| Provider | Works | Doesn’t |
|---|---|---|
| OpenAI | gpt-image-1, gpt-image-1-mini, gpt-image-2 |
dall-e-3, dall-e-2 |
| Gemini | gemini-2.5-flash-image and other gemini-*-image models |
imagen-* |
| OpenRouter | google/gemini-2.5-flash-image, openai/gpt-image-1, flux-kontext, qwen-image-edit |
plain flux, most others |
| Anthropic, DeepSeek | — | no image generation at all |
If a particular article’s image fails anyway, it is written from scratch instead and the queue tells you that happened. You never end up with one article missing a picture because of this.
Details worth knowing
- The thumbnail is downloaded once, for the whole episode, and saved to your Media Library. Every article works from that one file — which is what makes the set look consistent rather than like a dozen separate attempts.
- Images are always landscape here, whatever the shape set in Settings. A thumbnail is 16:9 and the guest is composed inside that frame; asking for a square crops the sides off and shoves them against an edge.
- No browser extension needed for this part. Thumbnails come from YouTube’s image CDN, which answers a normal server request from anywhere — unlike the description, this works on shared hosting. A Spotify episode uses the artwork from Spotify’s own public oEmbed endpoint instead, which also answers from anywhere. That artwork is often the show’s, identical on every episode of the season, which is why giving the YouTube link as well is worth doing: whenever there is one, the picture comes from there.
- It costs one extra request per article on top of the image itself — a short one that plans the scene, writes the words for the picture and writes the Media Library caption, all in the same answer. Turning the words on adds nothing to that. Image editing is slower than plain generation — budget roughly 30-60 seconds per article.
What it costs
Per episode: one request per ~20,000 characters of transcript, one to plan, then one per article written (plus one per article for the linking pass, and one per image if you turn images on — two if those images are built from the thumbnail, whether or not you put words on them). Reading is the cheap part — those passes use short outputs. The articles are the same cost as any other article.
A reading session is kept for 12 hours, so a failed run is never charged twice: whatever was already read is reused.
Known limits
- Timestamps need the current extension. Deep links to a moment in the episode come from the transcript’s caption timings, which only reach the plugin from extension v1.0.3 or later. Older versions, or a pasted transcript, simply produce articles with no times in them.
- Attribution around chapters wants v1.0.11. Nothing is wrong with an older extension — a passage that opens a chapter is never given the wrong name on any version from 2.5.7 onwards. What v1.0.11 buys is how quickly the recording gets to name somebody again afterwards: roughly forty seconds instead of two and a half minutes, on the episode this was measured against. Until you update, more of your quotations arrive for you to attribute by hand on the plan screen.
- YouTube transcripts have no speaker labels. Only Spotify’s do. On a YouTube-only episode the plugin refuses to name anyone at the points where inference is known to go wrong, publishes the quote unattributed, and asks you row by row. If the episode is also on Spotify, give that link too and the question is answered from the recording instead.
- A label is a number, not a name. Spotify usually says “Speaker 2”, not “Jordan Peterson” — so Who is speaking offers a name for each label and asks you to confirm it. Where Spotify does supply real names, they are used as they are.
- The suggested names are a reading, not a fact. Who-says-whose-name is a strong signal and it is still a guess: a host who never uses their guest’s first name, or an episode with three voices in it, will beat it. Check the row rather than trusting it — that is why the reason and a link to that voice’s own words are printed beside it.
- Moving the timestamps needs both recordings to be readable. In Both mode, pointing them at YouTube means reading the YouTube captions as well and measuring the two clocks against each other. A video with no caption track, or an upload edited enough that its passages can’t be matched, can’t be measured — and then nothing is guessed: your chosen fallback applies. Outside Both mode the times always belong to the one recording that was read.
- A measured offset is exact between anchors, not through a break. Where one version has an ad break the other doesn’t, the step is found and the map bends around it — but the exact second the break starts is only known to within the distance between two anchors. A quotation landing in that narrow window can be a few seconds out; everywhere else is exact. A quotation that was found on both recordings is not affected, because its own anchor sits on it — which is most of them, and which is why they are looked for.
- One episode at a time. Processing a whole back catalogue in one run isn’t available yet.
- Sponsors need a description. If your episode’s notes list no links at all, there is nothing to read, and both blocks are skipped rather than invented. If they do list them and they don’t appear, the notes didn’t reach us — paste them into the Episode description box and plan again.
- Spotify’s own copy of the notes can be locked out for hours. The allowance is per network and a signed-out browser spends it quickly, so a run can fall back to reading the page. You will be told when that happens, and told what fixes it: sign in to Spotify in that browser. Until then the paste box is the reliable route.
- The show name is the channel name. YouTube gives the channel (“Chris Williamson”), not the series (“Modern Wisdom”), so the subscribe link uses the channel name and the closing heading stays generic.