All posts
Guide

What should you remove from pages to improve AI extraction?

By Janis Plume, Founder, Outbound Pros · 9 min read · 2026-08-27

Quick answer

Remove anything that makes facts harder to fetch, isolate, and quote. The usual offenders are JavaScript-only content, repetitive marketing filler, contradictory claims across pages, oversized comparison grids, and sections where key facts are buried inside tabs, accordions, or image assets. AI extraction improves when one page states one thing clearly in plain text, near the top, with stable wording and clean structure.

What should you remove first?

Start with anything that prevents a crawler from seeing the page as text on first fetch. That is the highest leverage cleanup because if the model pipeline cannot reliably access the underlying facts, every downstream tactic is cosmetic.

The most important verified point here is simple. AI crawlers do not execute JavaScript. In the Vercel and MERJ server log study from late 2024, they fetched JavaScript files and never ran them. So if your page depends on client-side rendering to reveal core claims, remove that dependency for the facts that matter.

  • Remove JavaScript-only rendering for product facts, definitions, author bios, company descriptions, and eligibility criteria
  • Remove tabs or accordions that hide the only copy version of a critical claim
  • Remove image-only tables when the same facts can be written as text
  • Remove hero sections that push the real answer far below the first viewport
  • Remove animated text swaps that change wording before a crawler can anchor on one stable phrase

A lot of teams ask what they should add for GEO. Fair question, but the first win is usually subtractive. If your page forces extraction systems to reconstruct meaning from scripts, click states, and design flourishes, you are creating work for a machine that is not committed to doing it.

If you need the background on crawler behavior, read this breakdown of AI crawlers and JavaScript.

Which copy patterns should you remove?

Next, remove copy that sounds persuasive to a human skimmer but gives a model nothing quotable. AI assistants are much better at lifting stable facts than interpreting fluffy positioning language. When every paragraph says some version of leading, seamless, next-generation, powerful, or end-to-end, you have burned space that could have carried extractable detail.

  • Remove slogan-heavy intros that delay the actual answer
  • Remove vague category labels without definitions
  • Remove repeated claims written three different ways on one page
  • Remove unsupported superlatives unless you are willing to source and defend them
  • Remove paragraphs that mix narrative, positioning, and factual statements into one block

Operator view, if a sentence cannot survive being copied into an answer box without your sales rep standing next to it to explain the context, rewrite or remove it. AI extraction rewards plain claims with clear subjects and objects. It struggles when your copy relies on brand tone to imply the meaning.

This is also where many teams overuse recency theater. They keep updating headline phrasing and swapping terminology because they think freshness alone will help. Fresh pages can matter in some systems, but unstable wording can make entity matching worse if your core description keeps drifting.

What page structure makes extraction worse?

Bad structure is usually not dramatic. It is death by small obstacles. The facts exist, but they are split across decorative sections, buried under soft headers, or surrounded by so many side notes that the main point loses salience.

Remove or reduceWhy it hurts extractionWhat to keep instead
Long hero slogansDelays the page answer and weakens topical focusOne direct summary sentence near the top
Nested accordionsHides important claims behind interaction statesExpanded plain text for core facts
Duplicate FAQ blocksCreates multiple slightly different answer versionsOne canonical answer in body copy
Mixed-purpose paragraphsBlends facts, opinion, and CTA languageShort single-purpose paragraphs
Image-based proof pointsPrevents clean text extractionText equivalents beside visual assets
Overgrown comparison matricesMakes row and column relationships ambiguousSmaller tables with obvious labels

Notice what is happening in each case. The issue is not that AI is stupid. The issue is that extraction systems are optimized to move fast across imperfect pages. If your structure forces them to infer relationships instead of reading them directly, you lose.

This is why I usually tell operators to reduce page ambition. A page trying to be a brand manifesto, sales pitch, thought leadership essay, feature index, and documentation article all at once often gets cited for none of those jobs.

Should you remove FAQ schema and llms.txt expectations?

Yes, remove the belief that either is a shortcut. That mindset causes teams to avoid the harder content cleanup work.

FAQ rich results are fully deprecated, stopped appearing on 2026-05-07. That does not mean question and answer formatting is useless for extraction. It means you should stop treating FAQ markup as a visibility hack. Keep useful Q and A content when it clarifies the page. Remove bloated FAQ sections created only for old search result decoration.

Same story with llms.txt. Google states it is not used by Search. The SE Ranking study across around 300,000 domains found 10.13% adoption, zero adoption among the top 1,000 sites, and no citation lift after controls. So remove the expectation that publishing llms.txt will rescue weak pages. If you publish it at all, do it for internal organization and experimentation, not as a substitute for extractable page copy.

We covered that trade off more directly in this llms.txt guide.

What content should you consolidate or delete?

Remove conflicting versions of the same fact across your site. This is a quiet citation killer. Teams often have one statement on the homepage, another in a blog post, and a third in documentation. Humans can tolerate that mess. AI assistants often cannot. They either pick the wrong version or avoid citing you at all.

  • Delete outdated category pages that describe the company differently from the current positioning
  • Consolidate near-duplicate blog posts answering the same question with different wording
  • Remove old screenshots if the labels in them contradict present terminology
  • Delete thin glossary pages that restate another page without adding a cleaner definition
  • Consolidate fact ownership so each important claim has one canonical home

This is where content teams resist because deletion feels like losing inventory. In practice, the site often becomes more citable after you cut overlap. Extraction gets easier when one page owns one answer and supporting pages point back to it conceptually instead of competing with it.

If your contradiction problem is broader than one page, start with entity consistency work first. Clean names, descriptions, role labels, and service definitions matter more than publishing more content into the confusion.

For that problem specifically, see our guide to brand and entity consistency.

Where does this advice fail?

It fails when the page has no real authority, no original information, or no reason to be cited in the first place. Clean extraction does not create demand or trust by itself. It only reduces friction between your facts and the systems trying to read them.

It also fails if your business case depends on persuasion more than factual retrieval. Some pages need emotional weight, visual storytelling, or interactive demos to convert a human buyer. You should not flatten those pages into sterile fact sheets just to please a crawler. The better pattern is to keep persuasive experiences for people and create a parallel page that states the facts plainly.

And this advice is not for every team. If you are working on outbound execution, list building, or campaign operations, that belongs under Outbound Pros rather than here. We run managed outbound there, so we are not neutral about where that work should live, but the separation is useful because extraction work and outbound execution solve different problems.

My practical rule is simple. Remove obstacles from pages that are supposed to answer a question. Keep richer design on pages whose job is to persuade, demo, or sell. Confusing those jobs is what creates extraction issues in the first place.

Common questions

Should I remove all accordions from every page?

No. Remove them only when the accordion contains the only copy of an important fact. If the content exists clearly in visible plain text elsewhere, the accordion is less risky.

Does better AI extraction mean writing shorter pages?

Not always. It means making the important facts easier to isolate. Some pages can stay long if the answer appears early, the structure is clean, and each section has one job.

Should I publish llms.txt before doing page cleanup?

No. Treat page cleanup as the primary work. llms.txt is not used by Google Search, and current adoption evidence does not show citation lift after controls.

Do I need schema to fix extraction problems?

Schema can help clarify page meaning, but it will not rescue hidden, contradictory, or vague copy. Start with visible plain text and stable page structure first.

What is the single worst thing to leave on a page?

Critical facts that only appear after client-side JavaScript runs. If the crawler never executes the code, those facts are effectively absent.

Last updated: 2026-08-27

Talk through your AI visibility with people who measure it

30 minutes. We will look at what assistants can actually retrieve from your site and tell you plainly what is worth fixing first.

Book a strategy call

30 minutes, no obligation. The calendar shows real availability.

Or start with the free GTM audit from Outbound Pros