One small site with defects we can inspect
Which tool finds a broken page, and which one catches a link to a missing heading? We ran SiteOne Crawler 2.5.1, lychee 0.24.2 and Sitemap Cohort Auditor 0.2.2 on a controlled local sample on 29 September 2026. Each covers a different part of an audit. The comparison is about the evidence they produced under these settings.
The fixture serves eight HTML pages on 127.0.0.1:4319, one redirect and two missing resources. It also contains a missing title, duplicate titles, a missing description, a noindex page and a link to an absent fragment. Its sitemap deliberately repeats a URL, includes an invalid lastmod, and lists a missing page and a noindex page.
We created these defects for the test. They are not findings about a customer website. Both HTTP tools used the same local pages; the sitemap auditor received saved XML from that sample. Runs used macOS arm64, initial HTML and no paid APIs. No JavaScript-rendered content, large-site load or speed benchmark was included.
Download the Node.js fixture, before sitemap, after sitemap and reproduction notes with artifact checksums.
The reports answer different questions
| Known condition | SiteOne 2.5.1 | lychee 0.24.2 | Sitemap auditor 0.2.2 |
|---|---|---|---|
| Missing page and image | 2 HTTP 404 responses | 2 broken resource links | Does not fetch pages |
| Missing fragment on a 200 page | Not reported in this run | Reported with fragment checks on | Outside this check |
| Missing / duplicate titles | Reported | Outside link checking | Outside sitemap checking |
| Missing description / noindex | Reported | Outside link checking | Does not inspect page directives |
| Duplicate sitemap URL / invalid date | Not assessed in this run | Not assessed in this run | 1 duplicated URL / 1 invalid date |
| Sitemap changes | No snapshot comparison run | No snapshot comparison run | 3 added / 1 removed |
“Not reported” describes this configuration, not a claim that the tool can never detect the condition. SiteOne followed links from the root; lychee received eight explicit input URLs. The sitemap check read two files without crawling their URLs. We did not calculate a combined accuracy score across these different roles.
SiteOne Crawler
Our test · SiteOne Crawler v2.5.1
- Input
- Local fixture, root URL; one worker, 2 requests/second, 30-URL cap.
- Observed result
- Visited 11 resources: 8 successful responses, 1 redirect and 2 missing resources. Reported a missing title, duplicate titles, a missing description and a noindex page.
- Limit
- The missing fragment was not reported in this run. HTTP-only loopback hosting also triggers findings that are irrelevant to a public HTTPS site.
Install: Native CLI; macOS arm64 tested. License: MIT. Exact release and project documentation ↗
lychee
Our test · lychee v0.24.2
- Input
- The same fixture, with all 8 HTML pages supplied explicitly; fragment checks enabled.
- Observed result
- Checked 27 link occurrences and reported 3 errors: the missing page, missing image and a link to an absent fragment.
- Limit
- This invocation checks links from supplied pages; it is not a full recursive SEO audit or a metadata check.
Install: Native CLI; macOS arm64 tested. License: MIT / Apache-2.0. Exact release and project documentation ↗
Sitemap Cohort Auditor
Our test · Sitemap Cohort Auditor v0.2.2
- Input
- before.xml and after.xml from the same fixture, with JSON output and URL-set comparison.
- Observed result
- Read 6 entries representing 5 unique URLs. Found 1 duplicated URL and 1 invalid lastmod; the comparison showed 3 additions and 1 removal.
- Limit
- It does not fetch listed pages, so the declared 404 and noindex URLs need a separate crawl. The test is pinned to v0.2.2; later interfaces may differ.
Install: Node.js 20+ CLI; local XML files tested. License: MIT. Exact release and project documentation ↗
Repeat the test with the same versions
- Download the three fixture files linked above into one directory. Use Node.js 20 or newer. Start
node fixture.mjsin one terminal; it binds only to loopback. Keep it running during the HTTP checks. - Download SiteOne v2.5.1 and lychee v0.24.2 from the linked official releases for your operating system. Extract them and either put the executables on your PATH or use their full paths below.
- Download the Sitemap Cohort Auditor v0.2.2 package from its release and extract it so that
package/bin/sitemap-cohort-auditor.mjsexists. This tested package needs no additional runtime dependencies.
SiteOne: follow links from the root
siteone-crawler --url=http://127.0.0.1:4319/ --workers=1 --max-reqs-per-sec=2 --max-visited-urls=30 --no-color --output-json-file=siteone.json --output-html-report=siteone.html --output-text-file=siteone.txtThe run completed with exit code 0. It visited 11 resources. JSON, HTML and text reports were generated; we publish the JSON with workstation paths replaced by neutral placeholders.
lychee: check the eight HTML inputs and their fragments
lychee --format json --output lychee.json --include-fragments=anchor-only --max-concurrency 2 --max-retries 0 http://127.0.0.1:4319/ http://127.0.0.1:4319/healthy http://127.0.0.1:4319/missing-title http://127.0.0.1:4319/duplicate-a http://127.0.0.1:4319/duplicate-b http://127.0.0.1:4319/noindex http://127.0.0.1:4319/missing-description http://127.0.0.1:4319/finalExit code 2 is expected for this fixture: lychee found errors. Its 27 checks count link occurrences, not 27 unique pages. The result is three errors: two missing resources and one missing anchor.
Sitemap Cohort Auditor: compare saved declarations
node ./package/bin/sitemap-cohort-auditor.mjs after.xml --compare before.xml --json > sitemap-cohort.jsonThe report has schemaVersion 1. The run exited 0 despite the findings; no failure policy was supplied. It reads six entries representing five unique URLs, compared with three unique URLs before. Later project versions may change the interface, so use the linked release when reproducing this result.
Stop the local fixture with Ctrl+C when finished. Timings, paths and incidental environment checks may differ on another machine. Compare the observed defects rather than byte-for-byte report identity.
A separate live-page check with Mydentify
We also tested a different task on our own public guide. This is not a fourth crawler in the local comparison.
Our test · Mydentify CLI v1.1.4
- Input
- https://seotechlist.com/guides/ai-crawler-access, checked on 29 September 2026 at 01:43 UTC.
- Observed result
- The page and robots.txt returned HTTP 200. All 6 tested OpenAI and Anthropic tokens were permitted; all 6 user-agent probes returned HTTP 200.
- Limit
- This was a separate live-page check, not the local fixture comparison. Requests used our connection, not verified provider IPs; no indexing or citation outcome was measured.
node ./src/cli.js https://seotechlist.com/guides/ai-crawler-access --jsonThe command ran from the Mydentify v1.1.4 source release, using Node.js 20+ and no additional dependencies. Its MIT license covers the companion CLI. A user-agent probe from our machine cannot authenticate an official crawler. Read the access diagnosis guide before changing bot rules.
Use this as a reproducible starting point
The sample is intentionally small and uses HTTP on loopback. SiteOne therefore also raises environment-related findings such as HTTPS and hostname checks; those are not evidence of a production defect. We did not measure throughput, memory use, JavaScript rendering, authenticated pages, international sites, every parser edge case or Google indexing.
For a first audit, use the crawler to build a page inventory, the link checker for broken links and fragments, and the sitemap comparison around releases. Keep each report’s scope visible. There was no paid placement or vendor review of these results.
Choose the next step: turn findings into a fix queue, choose an open-source tool by task, or compare newer directory additions.