Why AI Podcasts Sound Natural but Fall Apart Under Real Conversation
Google's NotebookLM can transform any document into a lively two-host podcast that sounds remarkably like a real radio show, but new research reveals a fundamental gap: AI-generated dialogues operate by completely different rules than human conversation. A study from the University of Aberdeen and University of Edinburgh compared NotebookLM-generated podcasts with human-produced ones on identical topics, uncovering striking differences in how speakers take turns, overlap, and adapt to one another.
What Makes Human Conversation So Hard to Replicate?
Human conversation is governed by an intricate system of predictive coordination that researchers have studied for decades. Listeners don't simply wait for speakers to finish; they anticipate the end of a turn by picking up on grammatical, semantic, and vocal cues, preparing their response while the other person is still talking. This allows smooth transitions averaging just 200 to 250 milliseconds between speakers, far faster than the roughly 600 milliseconds humans actually need to plan and produce speech. This means real conversation requires constant prediction, not mere reaction.
Beyond timing, human podcasters follow genre-specific social scripts. Hosts introduce topics and step back into a novice role to give guests a platform, while guests predominantly adopt expert stances. These role expectations, rooted in broadcasting tradition, make conversational behavior predictable despite frequent shifts between expert and novice positions. Overlapping speech, brief vocal acknowledgments like "mm-hm," and backchannels add dynamic, expressive quality that signals engagement rather than disruption.
How Do AI Podcasts Differ From Human Ones?
The researchers annotated approximately ten-minute excerpts from two human podcasts and two NotebookLM-generated podcasts, using acoustic analysis software to code speaker contributions, role shifts, turn durations, overlaps, and transitions. On the surface, the AI podcasts behaved similarly to their human counterparts: questions structured the conversation, topics expanded through follow-ups, and both hosts and guests provided feedback through backchannels.
But the deeper measurements told a strikingly different story. Human turns were dramatically longer, averaging 22.3 seconds for hosts and 49.4 seconds for guests, compared with just 11.3 and 14.8 seconds for the AI agents. Human distributions of turn length were skewed toward long, developed contributions, while AI turns were short and symmetrically distributed. The most striking divergence appeared in overlapping speech: human guests overlapped for an average of more than four seconds and human hosts for over two seconds, whereas AI overlaps lasted mere milliseconds.
AI backchannels rarely overlapped at all. Instead, the current speaker paused, waited for the backchannel to be uttered, and then resumed, producing a mechanical politeness that no human interlocutor would display. The researchers interpret human overlaps as signals of engagement and emotional alignment, made possible by the fact that human listeners continuously predict where turns will end. NotebookLM sidesteps the problem of endpoint prediction entirely by pre-planning the entire dialogue script before converting it to speech, strictly enforcing the avoidance of gaps and overlaps in a way that humans never do.
What Happens When Humans Join the Conversation?
The study also tested a newer NotebookLM feature that allows a human to join the AI conversation live, and the results were revealing. Every human intervention disrupted the pre-planned script: transitions slowed dramatically, overlaps vanished, audio glitches caused sudden voice switches, and the agents could only acknowledge the human's contribution after long, silence-based delays before returning to their prepared narrative. Because spontaneous intervention requires resource-intensive, real-time context integration that pre-planning was designed to avoid, the hybrid dialogues lost the fluidity that made the fully scripted versions sound convincing.
"NotebookLM clearly hits its limits when responding to unpredictable input," the researchers concluded, noting that pre-determining role behavior and turn-taking only goes so far.
Yasmin A. Carruthers and Johannes M. Heim, University of Aberdeen and University of Edinburgh
Key Limitations of AI-Generated Dialogue
- Turn Length Asymmetry: AI agents produce uniformly short turns averaging 11-15 seconds, while human speakers vary dramatically, with guests speaking for nearly 50 seconds on average, allowing for deeper exploration of ideas.
- Absence of Meaningful Overlap: Human overlaps lasting multiple seconds signal engagement and emotional connection, but AI overlaps last milliseconds, creating mechanical interactions that lack the warmth of real conversation.
- Rigid Role Adherence: AI hosts and guests fail to break genre conventions dynamically, instead maintaining predetermined expert-novice positions without the fluid adaptation that makes human podcasts feel natural and responsive.
- Inability to Handle Real-Time Disruption: When humans interrupt the pre-planned script, AI systems cannot integrate new information smoothly, causing delays, glitches, and loss of conversational flow.
The researchers argue that AI-generated dialogue can mimic individual features of natural conversation in isolation but fails to synthesize them into the dynamic negotiation of information and social relation that defines human talk. Human conversation is not optimized for seamless information exchange but for building common ground, an emergent and collaborative process shaped by mutual attention, verbal feedback, and shared, growing context.
Variable turn-taking timings and adaptive conversational role behavior remain unmatched human skills. Achieving genuine naturalness will require AI systems that can interpret social contexts and respond dynamically to the fluid negotiation of interlocutor relations, not merely imitate their surface elements. For now, NotebookLM's Audio Overview feature excels at creating engaging, scripted dialogues that sound polished and professional, but they operate in a fundamentally different conversational universe than the messy, unpredictable, deeply human art of real dialogue.