Log inGet Started

Site Crawl Agent

A real multi-page crawl, not a single-page check dressed up as one

Crawls a handful of your real pages, renders the ones that need JavaScript to show their content, and checks what it finds for near-duplicate text — no invented metrics, no pretending it crawled a thousand pages it didn't touch.

Free, self-hosted · PageSpeed key optional

Without vs. with Marlo

Why teams switch to a real multi-page crawl agent

Without Marlo

  • Manually checking each page for duplicate content
  • No idea if your SPA routes even render for crawlers
  • Audits that stop at the homepage
  • Full-site crawler tools that cost hundreds/mo
  • PageSpeed calls burning rate limits on every single page

With Marlo Site Crawl Agent

  • Crawls up to 8 pages beyond your homepage automatically
  • Real headless-browser fallback catches SPA content
  • Cross-page duplicate detection via Jaccard similarity
  • PageSpeed scoring is opt-in, per page — no wasted calls
  • Self-hosted — no monthly retainer, your own API keys

Definition

What the Site Crawl Agent actually does

A real crawl runs against up to 8 of your live pages beyond the homepage — bounded, capped at an 8-second timeout per page, and automatically skipping auth-flow paths like login and signup since there's nothing there worth indexing. When a page comes back with too little content — the telltale sign of a client-rendered SPA — it re-fetches through a real headless-browser render instead of giving up.

Every crawled page is then compared against every other page using word-set Jaccard similarity at a 0.85 threshold, flagging real near-duplicate content — not just byte-identical pages. PageSpeed scoring is available per page, but only when you deliberately ask for it.

How it works

How the Site Crawl Agent Maps Your Site

Set it up once. The agent crawls on demand, every time.

Crawl queue
/CRAWLED
/pricingCRAWLED
/blogCRAWLED
/loginSKIPPED
Step 01

Crawl up to 8 real pages

Real HTTP fetches, 8-second timeout per page, auth-flow paths auto-skipped. Example shown above, illustrative not live data.

/products/a vs /products/a-print93%
/blog/post-1 vs /blog/post-241%
Step 02

Detect real cross-page duplicates

Word-set Jaccard similarity between every page pair, flagged at a 0.85 threshold — near-duplicates, not just byte-identical text.

PageSpeed — opt-in
/pricingRun check →
Step 03

Score performance only when you ask

Real PageSpeed Insights scoring per page — a deliberate, separate action, not burned automatically on every crawl.

Capabilities

What it actually does

A single crawl run walks real pages over HTTP — no invented page counts, no pretending it crawled a site it didn't touch.

Multi-page crawl, not a single-page check

Crawls up to 8 pages beyond your homepage with real HTTP fetches — bounded and capped at an 8-second timeout per page so a slow or hung page can't stall the whole run. It automatically skips auth-flow paths like login, signup, sign-in, sign-up, and logout URLs, since there's nothing to index there.

Falls back to a real headless-browser render for SPAs

The same JS-render fallback as the SEO agent: when a plain crawl of a page comes back with too little content — the telltale sign of a client-rendered app that needs JavaScript to fill in the page — it re-fetches that page through a real headless-browser render via Jina AI Reader instead of giving up on it.

Real cross-page duplicate content detection

Compares every crawled page against every other page using word-set Jaccard similarity, not exact-match text comparison, at a 0.85 threshold. It flags pages that are near-duplicates of each other, not just byte-identical ones — catching the real-world duplicate-content problem, like a product page and its printable variant sharing 90% of the same text, that exact-match checks miss entirely.

Opt-in PageSpeed scoring, per page

Real PageSpeed Insights / Lighthouse scoring is available per crawled page, but it's not run automatically as part of a crawl. PSI has rate limits and each call takes 10–20 seconds, so it's a deliberate, separate action you trigger only when you actually want performance numbers for a specific page.

Auth-path skip

Login, signup, sign-in, sign-up, and logout URLs are automatically excluded from the crawl — nothing to index there, so no time wasted on them.

8-second per-page timeout

Every page fetch is bounded at 8 seconds, so a single slow or hung page can never stall the whole crawl run.

0.85 similarity threshold

Word-set Jaccard similarity above 0.85 flags a near-duplicate — tuned to catch real duplicate-content problems, not coincidental overlap.

Site Crawl Agent vs. the alternatives

Marlo vs. a generic crawler tool and an agency retainer

Self-hosted and open-source vs. $49–$199/mo tools or a $1,500–$5,000+/mo agency retainer.

Generic crawler toolSEO agencyMarlo
Crawl depthHomepage onlyManual site reviewUp to 8 pages, real HTTP
SPA contentOften missedRarely checkedHeadless-render fallback
Duplicate detectionExact-match onlySpot-checked0.85 Jaccard, cross-page
Cost$49–$199/mo$1,500–$5,000+/moSelf-hosted, your own API keys

Why Site Crawl Agent?

One Page Never Tells the Whole Story.

Duplicate content and unrendered SPA routes hide across your site, not just on the homepage. The Site Crawl Agent checks multiple real pages so problems don't stay invisible until rankings drop.

8

Pages Crawled Per Run

8s

Timeout Cap Per Page

0.85

Duplicate Similarity Threshold

$0

Monthly Cost, Self-Hosted

Connecting a Google PageSpeed key is optional

The crawl itself never depends on it — pages get crawled, rendered, and checked for duplicates regardless. A PageSpeed key only unlocks the opt-in, per-page performance score when you deliberately ask for one, so you never burn rate limits you didn't mean to spend.

ILLUSTRATIVE EXAMPLES

What a marketing system like this looks like in practice

Marlo is early — these are illustrative scenarios showing what lean teams could achieve, not real customer results (yet).

SG
Example: SaaS growth teamIllustrative

A 2-person SaaS team could run SEO, GEO, and content on autopilot

40%
Typical target for organic traffic growth
3x
Content output with the same team size
DB
Example: DTC brandIllustrative

A lean DTC brand could launch a full quarter of campaigns without a marketing hire

12
Campaigns in a single quarter, illustratively
28%
Realistic email conversion lift range

Pricing

Three ways to run Marlo

Self-host it for free, let us host it and cover the AI cost, or host it with us and bring your own model key. Same 10 agents, every tier.

Self-Hosted

Run it yourself, own everything.

$0/mo, forever
  • All 10 agents — Articles, SEO, GEO, Site Crawl, Reddit, X, LinkedIn, GitHub, Hacker News, Leads
  • Bring your own LLM key — OpenAI, Anthropic, or any OpenAI-compatible endpoint
  • Bring your own integration keys — GA4, GSC, Gmail, GitHub, PageSpeed, Tavily, Places
  • Your data stays in your own database, never on a server we control
  • AGPL-licensed — audit the code, modify it, run it forever
Get Marlo on GitHub
Most convenient

Hosted, Built-In Models

We host it, we cover the AI model cost.

$29/mo
  • All 10 agents, fully hosted — nothing to install or maintain
  • Built-in AI models included — no LLM API key required
  • Bring your own integration keys — GA4, GSC, Gmail, GitHub, PageSpeed, Tavily, Places
  • Automatic updates, managed infrastructure
  • Cancel anytime

No credit card required

Hosted, Your Own Key

We host it, you bring your own model key.

$9/mo
  • All 10 agents, fully hosted — nothing to install or maintain
  • Bring your own LLM key — OpenAI, Anthropic, or any OpenAI-compatible endpoint
  • You pay your model provider directly — this fee covers hosting only
  • Bring your own integration keys — GA4, GSC, Gmail, GitHub, PageSpeed, Tavily, Places
  • Cancel anytime

No credit card required

Run your first real crawl

Enter your website — the Site Crawl Agent walks your real pages and shows you exactly what it found, in seconds.

Self-host free, or hosted from $9/mo