Directory search

Find the right SEO tool.

Start typing a product name, category or capability.
Practical guide / Original research

One faulty site, three open-source SEO tools: our test

Compare real SiteOne, lychee and Sitemap Cohort Auditor results on one controlled sample. Download the fixture, commands and JSON reports to repeat the test.

One small site with defects we can inspect

Which tool finds a broken page, and which one catches a link to a missing heading? We ran SiteOne Crawler 2.5.1, lychee 0.24.2 and Sitemap Cohort Auditor 0.2.2 on a controlled local sample on 29 September 2026. Each covers a different part of an audit. The comparison is about the evidence they produced under these settings.

The fixture serves eight HTML pages on 127.0.0.1:4319, one redirect and two missing resources. It also contains a missing title, duplicate titles, a missing description, a noindex page and a link to an absent fragment. Its sitemap deliberately repeats a URL, includes an invalid lastmod, and lists a missing page and a noindex page.

We created these defects for the test. They are not findings about a customer website. Both HTTP tools used the same local pages; the sitemap auditor received saved XML from that sample. Runs used macOS arm64, initial HTML and no paid APIs. No JavaScript-rendered content, large-site load or speed benchmark was included.

Download the Node.js fixture, before sitemap, after sitemap and reproduction notes with artifact checksums.

The reports answer different questions

Known conditionSiteOne 2.5.1lychee 0.24.2Sitemap auditor 0.2.2
Missing page and image2 HTTP 404 responses2 broken resource linksDoes not fetch pages
Missing fragment on a 200 pageNot reported in this runReported with fragment checks onOutside this check
Missing / duplicate titlesReportedOutside link checkingOutside sitemap checking
Missing description / noindexReportedOutside link checkingDoes not inspect page directives
Duplicate sitemap URL / invalid dateNot assessed in this runNot assessed in this run1 duplicated URL / 1 invalid date
Sitemap changesNo snapshot comparison runNo snapshot comparison run3 added / 1 removed

“Not reported” describes this configuration, not a claim that the tool can never detect the condition. SiteOne followed links from the root; lychee received eight explicit input URLs. The sitemap check read two files without crawling their URLs. We did not calculate a combined accuracy score across these different roles.

SiteOne Crawler

Our test · SiteOne Crawler v2.5.1

Input
Local fixture, root URL; one worker, 2 requests/second, 30-URL cap.
Observed result
Visited 11 resources: 8 successful responses, 1 redirect and 2 missing resources. Reported a missing title, duplicate titles, a missing description and a noindex page.
Limit
The missing fragment was not reported in this run. HTTP-only loopback hosting also triggers findings that are irrelevant to a public HTTPS site.

Download result JSON · Method and reproduction steps →

Install: Native CLI; macOS arm64 tested. License: MIT. Exact release and project documentation ↗

Sitemap Cohort Auditor

Our test · Sitemap Cohort Auditor v0.2.2

Input
before.xml and after.xml from the same fixture, with JSON output and URL-set comparison.
Observed result
Read 6 entries representing 5 unique URLs. Found 1 duplicated URL and 1 invalid lastmod; the comparison showed 3 additions and 1 removal.
Limit
It does not fetch listed pages, so the declared 404 and noindex URLs need a separate crawl. The test is pinned to v0.2.2; later interfaces may differ.

Download result JSON · Method and reproduction steps →

Install: Node.js 20+ CLI; local XML files tested. License: MIT. Exact release and project documentation ↗

Repeat the test with the same versions

  1. Download the three fixture files linked above into one directory. Use Node.js 20 or newer. Start node fixture.mjs in one terminal; it binds only to loopback. Keep it running during the HTTP checks.
  2. Download SiteOne v2.5.1 and lychee v0.24.2 from the linked official releases for your operating system. Extract them and either put the executables on your PATH or use their full paths below.
  3. Download the Sitemap Cohort Auditor v0.2.2 package from its release and extract it so that package/bin/sitemap-cohort-auditor.mjs exists. This tested package needs no additional runtime dependencies.

SiteOne: follow links from the root

siteone-crawler --url=http://127.0.0.1:4319/ --workers=1 --max-reqs-per-sec=2 --max-visited-urls=30 --no-color --output-json-file=siteone.json --output-html-report=siteone.html --output-text-file=siteone.txt

The run completed with exit code 0. It visited 11 resources. JSON, HTML and text reports were generated; we publish the JSON with workstation paths replaced by neutral placeholders.

lychee: check the eight HTML inputs and their fragments

lychee --format json --output lychee.json --include-fragments=anchor-only --max-concurrency 2 --max-retries 0 http://127.0.0.1:4319/ http://127.0.0.1:4319/healthy http://127.0.0.1:4319/missing-title http://127.0.0.1:4319/duplicate-a http://127.0.0.1:4319/duplicate-b http://127.0.0.1:4319/noindex http://127.0.0.1:4319/missing-description http://127.0.0.1:4319/final

Exit code 2 is expected for this fixture: lychee found errors. Its 27 checks count link occurrences, not 27 unique pages. The result is three errors: two missing resources and one missing anchor.

Sitemap Cohort Auditor: compare saved declarations

node ./package/bin/sitemap-cohort-auditor.mjs after.xml --compare before.xml --json > sitemap-cohort.json

The report has schemaVersion 1. The run exited 0 despite the findings; no failure policy was supplied. It reads six entries representing five unique URLs, compared with three unique URLs before. Later project versions may change the interface, so use the linked release when reproducing this result.

Stop the local fixture with Ctrl+C when finished. Timings, paths and incidental environment checks may differ on another machine. Compare the observed defects rather than byte-for-byte report identity.

A separate live-page check with Mydentify

We also tested a different task on our own public guide. This is not a fourth crawler in the local comparison.

Our test · Mydentify CLI v1.1.4

Input
https://seotechlist.com/guides/ai-crawler-access, checked on 29 September 2026 at 01:43 UTC.
Observed result
The page and robots.txt returned HTTP 200. All 6 tested OpenAI and Anthropic tokens were permitted; all 6 user-agent probes returned HTTP 200.
Limit
This was a separate live-page check, not the local fixture comparison. Requests used our connection, not verified provider IPs; no indexing or citation outcome was measured.

Download result JSON · Method and reproduction steps →

node ./src/cli.js https://seotechlist.com/guides/ai-crawler-access --json

The command ran from the Mydentify v1.1.4 source release, using Node.js 20+ and no additional dependencies. Its MIT license covers the companion CLI. A user-agent probe from our machine cannot authenticate an official crawler. Read the access diagnosis guide before changing bot rules.

Use this as a reproducible starting point

The sample is intentionally small and uses HTTP on loopback. SiteOne therefore also raises environment-related findings such as HTTPS and hostname checks; those are not evidence of a production defect. We did not measure throughput, memory use, JavaScript rendering, authenticated pages, international sites, every parser edge case or Google indexing.

For a first audit, use the crawler to build a page inventory, the link checker for broken links and fragments, and the sitemap comparison around releases. Keep each report’s scope visible. There was no paid placement or vendor review of these results.

Choose the next step: turn findings into a fix queue, choose an open-source tool by task, or compare newer directory additions.