Avoiding wrong answers in browser automation: reliability checks for a NotebookLM integration
This article describes kof-notebooklm-mcp’s web integration. Google Cloud documents enterprise notebook-management APIs, checked September 21, 2026. The earlier blanket claim that NotebookLM has no public API was too broad. Replacement depends on the account, licensing, and operations required.
So I built an MCP server that drives the real NotebookLM page with Playwright, wraps "add source", "ask", and "extract answer" into callable tools, and shipped it to PyPI.
Version one worked after a weekend. Then I spent the next two releases fixing a single category of problem: it would return things that looked entirely plausible and were wrong.
Why browser automation errors are uniquely poisonous
APIs can expose 4xx, 5xx, and timeouts, but can also return semantically wrong success payloads. Browser integrations add another boundary: whether the DOM node actually represents the result of the current operation.
Browser automation is different. You are scraping a dynamic page designed for humans, and its failure mode is "grabbed something — just not the thing you wanted":
| Failure | Symptom | Downstream damage |
|---|---|---|
| Stale answer | Ask question two, receive the reply to question one | Perfectly formatted non sequitur; very hard to notice |
| Partial answer | Harvested while the stream was still typing | Conclusions cut mid-sentence; citations missing |
| The UI itself | Feature-button labels — "AI slides", "mind map" — stitched into a paragraph | A completely meaningless "answer" |
All three can be wrapped in a successful tool result. Whether the transport is HTTP or MCP, a success flag alone is insufficient. Downstream agents can propagate the error into later steps.
Failure one: an answer only counts when it is consecutively stable
AI answers stream; the DOM node keeps growing. When is it "done"?
Waiting a fixed number of seconds is the intuitive answer and the wrong one: short answers overwait, long answers underwait. The correct signal is not time — it is stability:
# Poll the answer node: content must be identical
# across 3 consecutive checks to count as complete
if text == last_text:
stable_answer_checks += 1
else:
stable_answer_checks = 0 # still changing — reset
last_text = text
if has_changed_answer and stable_answer_checks >= 3:
return text # new, and stableFreshness and short-term stability reject some stale or premature extractions, but do not prove completion. A stream may pause and resume, and two questions may produce identical text. The three-check heuristic needs a current reply identity, generation state, citation checks, and an overall timeout. Without a completion signal, report uncertainty.
Failure two: detecting UI pollution
The third class is the most insidious. NotebookLM's workspace is full of feature entry points — AI slides, flashcards, mind maps, Audio Overview — and when a selector starts matching too wide a node after a UI update, those button labels get stitched into an "answer".
It is not garbage bytes. It is a string of real words that fails to mean anything. A human spots it instantly; a program cannot.
My fix is a signature library: list the workspace's control-element strings and feature names as markers, and count hits in the extracted content:
def detect_answer_ui_pollution(text: str) -> list[str]:
"""Detect workspace UI captured as if it were answer text."""
normalized = re.sub(r"\s+", " ", text).strip().lower()
control_hits = [m for m in _ANSWER_UI_CONTROL_MARKERS if m in normalized]
feature_hits = [m for m in _ANSWER_UI_FEATURE_MARKERS if m in normalized]
# >=2 control hits, or >=4 feature hits: this is interface, not an answer
if len(control_hits) >= 2 or len(feature_hits) >= 4:
return [*control_hits, *feature_hits]
return []Contamination thresholds can reject valid answers, especially when a question asks about NotebookLM features. Multiple feature names should trigger inspection, not prove that text came from controls. Preserve the matched rule and target-node context so adjustments can be based on evidence.
What happens on detection? This was the most important decision in the project:
Return an explicit ANSWER_EXTRACTION_FAILED error instead of passing along whatever was scraped. Better the caller knows "nothing this time" than receives wrong content it will treat as true. For tools consumed by AI agents this principle cannot be overstated — agents have zero immunity against well-formatted wrong content.
Failure three: layered selector defense
With no API contract, the DOM is the contract — and the other side may change it at will. Selectors are not a "write them correctly" problem; they are a "how do they age" problem.
Every target element gets a fallback array, ordered most-specific to most-generic:
"sources_panel": [
"section.source-panel", # current real structure: precise
'[data-testid="sources-panel"]', # test attribute: fairly stable
'[aria-label*="Sources"]', # a11y label: survives redesigns well
'[aria-label*="來源"]', # same layer, Chinese UI
'[class*="sources"]', # widest net: last resort
],The ordering is the strategy. The top entries are precise but brittle — a redesign snaps them. The bottom ones live long but match wide — which is precisely one source of UI pollution. So layered selectors and pollution detection are a pair: wide selectors keep basic function alive across redesigns, and pollution detection catches what the wide selectors wrongly grab.
Two supporting rules: every optional element read gets a short timeout (a removed element must not stall the whole flow), and "page still loading" gets an explicit state — a loading notebook page naively reads as "an unnamed notebook with zero sources", which is again plausible-looking wrong data.
Separate offline tests from authenticated checks
Testing such a project has one peculiar difficulty: you cannot mock NotebookLM. Unit tests cover the parsing logic — pollution detection, stability judgment, error classification — but "do the selectors still match the real page" has exactly one verification method: a real account, real notebooks, real long streaming answers.
The earlier article mentioned 133 tests without a versioned output attached here. That count is not evidence of current service availability. Parsing can be tested offline; authentication, selectors, long answers, and citations need live checks. This revision did not sign in or rerun NotebookLM.
If you are wrapping a product with no API
Five principles, in order of importance:
One — distrust extracted content by default. Everything passes integrity checks (stable? new? not UI?) before leaving the tool.
Two — prefer explicit failure. An error code always beats plausible-looking wrong content, doubly so when the caller is an AI.
Three — wait on stability signals, never fixed durations.
Four — layer your selectors, and accept that wide selectors must be balanced by pollution detection.
Five — real-environment verification is part of the release process, not a post-incident remedy.
When an official API covers the required operation, compare its permissions and maintenance cost. For remaining browser work, document why it is needed and how a UI change will be detected.
Related: From zero to PyPI — making NotebookLM programmableUse a small fixture to expose failure classes
Create a test source containing only “Project: Bluebird; delivery date: 2026-10-15.” Ask for the date, then the project name. The second reply should not reuse the first, and its citation should lead to the fixture. Ask for an owner the document never names to check unsupported attribution. This is a proposed acceptance case, not a completed live test.
Also test an interrupted stream, a source still processing, and an expired session. Record waiting, explicit failure, or an acceptable result rather than just whether text returned. Use non-sensitive fixture data so repeated checks remain easy to compare.