Site audit P0 (1/3): issue engine, incremental persistence, block detection (#362)

* Site audit P0 (1/3): issue engine, incremental persistence, block detection

Server-side foundation of the P0 feature set from docs/site-audit-pm-research.md:

- Issue engine: shared registry of issue types (severity, explanation,
  how-to-fix). Per-page reporters run inside crawl steps; cross-page checks
  (duplicate titles/descriptions/content, broken internal links, redirect
  chains/loops, orphan pages) run at finalize as SQL over the persisted crawl.
- New audit_links + audit_issues tables, audit_pages columns (depth, content
  hash, header signals, fetch class, sitemap flag); audit tables moved to
  src/db/{,pg/}audit.schema.ts; migrations 0029 (D1) / 0006 (PG).
- Incremental persistence: pages/links/issues written inside each crawl-batch
  step with deterministic row ids + upserts (retry idempotent); slim step
  state; robots.txt checkpointed as step state; merged progress steps keep a
  10k-page crawl within the Workflows step budget.
- Crawler: manual redirect handling with inline follow of normalization-
  equivalent redirects, response header capture (X-Robots-Tag, Link
  rel=canonical), BFS depth, sitemap-last seeding, SSRF check on discovered
  links, honest 'we were blocked' classification (403/429/cf-mitigated/
  challenge).
- MCP: run_site_audit, get_audit_status, get_audit_issues, get_audit_pages;
  limitTier resolved via shared AuditService.resolveAuditLimitTier.
- Lighthouse strategies reduced to auto/none (legacy all/manual map on read).
- Self-healing: getStatus reconciles audits whose workflow instance errored/
  terminated without reaching mark-failed.

The Issues UI and the badseo.dev e2e fixture site stack on top of this PR.

Deploy notes: run db:migrate:prod (additive); terminate running audits before
deploying — the workflow step structure changed and in-flight instances cannot
replay under the new code (a finalize guard fails them loudly instead of
completing empty).

* Store only internal link edges in audit_links

Both consumers (broken-internal-link and orphan checks) filter on
isInternal; per-page external counts already live on audit_pages.
Dropping external rows cuts stored edges on outbound-heavy sites.
Column stays so P1 external-link checks can re-add rows without a
migration.

* Review fixes: failAudit CAS guard, dedupe hash helpers, cheaper checks

- failAudit only transitions running audits, so the getStatus reconciler
  can't flip a just-completed audit to failed when it races finalize
- collapse the duplicate SHA-256 helper into audit/ids.ts
- finalize integrity guard uses a limit-1 existence probe instead of
  fetching every page row
- get_audit_status MCP tool no longer reads the audit row twice when an
  explicit auditId is given
This commit is contained in:
Ben Senescu 2026-07-07 12:29:32 -04:00 committed by Ben Senescu
parent 7caaebbbac
commit 1c74fded7b
36 changed files with 9512 additions and 662 deletions

View File

@ -0,0 +1,142 @@
# OpenSEO Site Audit: 80/20 Feature Set & Crawl Architecture Proposal (rev. 3, founder decision applied)
**Author:** PM research synthesis | **Date:** 2026-06-09 | **Inputs:** SEOnaut + LibreCrawl code reviews, OpenSEO baseline review, competitor/market research, architecture research, adversarial panel review. Revision notes: panel critiques accepted are folded in silently; where we overrode the panel, an explicit **Panel pushback** note marks the disagreement. **Rev. 3:** founder decision — build on the in-house Workers crawler, do not adopt DataForSEO OnPage; "we were blocked" flagging accepted as the blocking posture. §4 now weighs that decision's tradeoffs instead of recommending an engine.
---
## 1. How site audits actually work (primer)
Every site audit tool is the same three-stage pipeline:
1. **Crawl.** Start from a URL, fetch robots.txt and sitemaps, fetch pages, extract links, enqueue new same-site URLs (BFS), repeat until a page cap. Two engine choices matter: _plain HTTP fetch_ (fast, cheap, blind to JavaScript) vs _headless rendering_ (sees what Google sees, ~10x cost). JS rendering is table stakes in commercial tools; differentiation is fidelity and whether it's paywalled.
2. **Check.** Per-page rules (missing title, noindex, broken canonical) plus cross-crawl rules (duplicate titles, redirect chains, orphan pages). The cross-page checks carry most of the value, and they're set-queries over crawl data, not parsing.
3. **Report.** Severity grouping, plain-English explanations, drill-down to affected URLs, **export**, and diffs against previous crawls.
**Desktop vs cloud.** Screaming Frog ($259/yr) and Sitebulb run on the user's laptop: residential IP, no bot-blocking, but no scheduling, RAM-limited crawls (the #1 SF complaint), and the machine must be on. Cloud tools (Ahrefs, Semrush, Moz) crawl datacenter IPs on a schedule — and **getting blocked is their #1 recurring complaint**: WAFs 403/challenge their bots, producing phantom 404s and false positives.
**Why blocking happens.** WAFs score IP reputation (datacenter = suspicious) and fingerprint (no JS = bot). It hits every cloud auditor, including ours. The industry playbook: documented UAs, published allowlistable IPs, Cloudflare's Verified Bots program, and ownership verification to justify higher crawl rates (Ahrefs advertises dramatically higher rates for verified owners; the oft-quoted 30→30,000 URLs/min figure is unsourced in our fact base and Ahrefs' verification is its own, not GSC — treat as directional only).
---
## 2. Where OpenSEO's audit stands today
Honest summary: **we have a crawler and a Lighthouse integration, but not an audit product.** We extract per-page fields and show filterable tables, but no rule engine names issues; the only "issues" users see are Lighthouse failures. SEOnaut ships ~88 issue types; LibreCrawl ~26; Screaming Frog ~300. We effectively ship ~2 as table styling.
Structural problems, in priority order:
- **The link graph is thrown away.** Link lists are extracted but only counts persist to D1, so broken links, orphans, redirect chains, and inlink analysis are impossible without schema change.
- **Persistence is finalize-only and Workflow-shaped.** `runCrawlBatch` returns full page arrays (links, images) as durable step state (~1MiB/step cap — already borderline on link-dense sites), `allPages` accumulates everything in 128MB Worker memory, and `batchWriteResults` writes all rows inside the single finalize step (DB_BATCH_SIZE=100 → a 1M-row link graph would mean ~10,000 subrequests in one step, ~10x the limit). **Any link-graph feature requires restructuring persistence first.** This was deferred to P1 in draft 1; the panel correctly identified that as backwards — it is now the P0 prerequisite.
- **Step budget caps crawl size — ledger, not vibes.** Workflows allows ~1,024 steps/instance. Today: 3 steps per 25-page batch → ~341 batches ≈ **~8.5k pages** fits (minus discovery/finalize overhead), so the existing 10,000-page clamp slightly exceeds the budget. The sharper constraint is Lighthouse strategy `all`: 2 steps per 10-URL batch ≈ 1,000 steps at 5k pages — over budget on its own. Politeness delays must live _inside_ steps (wall-clock waits are fine), never as smaller batches. See §4 for the post-refactor ledger and caps.
- **No audit MCP tools** — keywords, SERP, backlinks, rank-tracker, GSC all have them; audit has zero (verified in `src/server/mcp/tools`). Ahrefs, meanwhile, **already ships site-audit MCP tools** (`site-audit-issues`, `site-audit-page-explorer`, etc.); Semrush has an MCP server too. We are behind here, not ahead — see §3.4 for what our actual wedge is.
- **SSRF gap:** `url-policy.ts` validates only the start URL; discovered links are fetched unvalidated. Must be fixed in P0 for as long as the in-house crawler runs.
Condensed check inventory: status/redirects, title/meta, canonicals, indexability, headings, images = **partial** (extracted, unflagged; redirects followed silently; header variants like `X-Robots-Tag` unchecked). Duplicates (3035% of sites, Semrush study¹), broken links (52% of domains¹), redirect chains (~12%¹), orphans, security headers, hreflang, JS rendering, scheduling/diffs, issue UX, MCP, block resilience = **missing**. Lighthouse/CWV = **strong**, better than both OSS tools. _(¹ Prevalence figures from market research §3; primary sources to be attached as a footnoted appendix before this circulates externally.)_
(Folklore correction stands: the 100k-page lifetime cap settles to actual crawled counts; only failed/stuck audits permanently pay their reservation.)
---
## 3. The 80/20 feature set
**Pricing posture (decides everything downstream).** Audit is the most commoditized feature in SEO — Ahrefs Webmaster Tools and Screaming Frog ≤500 URLs are free. Per our SEO strategy (goal = signups), **audit is acquisition, not monetization**: default 50-page audits are free per site; credits are charged for large crawls, scheduling, and history. The free posture also reinforces the engine decision (§4): with the in-house crawler, a free audit's marginal cost is D1 writes and Workers compute largely inside quotas we already pay for — no per-audit vendor bill, no metered cost on failed runs (DataForSEO charges even for failed tasks).
The prevalence evidence says a credible audit needs ~6 check families, not 300. P0 below runs on the **in-house crawler (founder decision, §4)**; crawler-side hardening is therefore mandatory P0 work, not contingency. Everything except the crawler-side items remains **engine-independent by design** (schema, SQL checks, UX, MCP) — that portability is the cheap hedge that keeps a future vendor or local-agent engine swap on the table.
### P0 — the audit becomes real (~68 weeks)
**0. Incremental persistence (prerequisite, was P1).** Write `audit_pages`/`audit_links` rows to D1 _inside each crawl-batch step_ (one extra `db.batch` per batch keeps subrequests bounded), return only slim summaries (id, url, status) as step state, drop the `allPages` accumulator. Finalize becomes: multipage SQL step → completion. Bonus: failed audits keep partial results. Also fix the SSRF gap (run discovered links through url-policy).
**1. Issue engine + schema.** `audit_links` edge table (source, target, anchor, rel, internal) and `audit_issues` table. **Reporter contract (decided now, per panel):** reporters consume a normalized page record — `(record: AuditPageRecord) => boolean` where the record includes response headers captured at crawl time — _not_ SEOnaut's `(page, $, headers)` cheerio signature. Cheerio runs once in `crawlPage`; reporters are DOM-free. This is what makes the engine portable across our crawler, a future local agent, or a vendor feed if the §4 revisit triggers ever fire. Cross-page checks are D1 SQL passes (duplicates = `GROUP BY hash HAVING count > 1`, per SEOnaut's multipage reporters). **Effort: 23 weeks including #0's restructure** (was "1 week" — panel was right).
**2. High-prevalence checks** (post-#1; thresholds cribbed from SEOnaut/LibreCrawl):
- _Broken internal links + redirect chains/loops_: switch to `redirect: "manual"`, record hops, enqueue targets. Decision: redirect hops do **not** count against maxPages or billing (they're cheap HEAD-sized fetches; metering/progress semantics budgeted at +1 day). Broken-link checks fire **only for targets we actually fetched or HEAD-checked** — never inferred from absence on partial crawls. ~4 days.
- _Duplicates_: body-hash group-by catches mirrored/parameterized duplicates only (any dynamic token defeats exact hashing); **duplicate titles/descriptions group-by is the workhorse** and carries the prevalence claim. Avoid LibreCrawl's O(n²) similarity scan. ~12 days.
- _Title/meta length, missing/multiple H1, heading order_ (SEOnaut thresholds). ~12 days.
- _Canonical + indexability with header fallbacks_ (`X-Robots-Tag`, `Link: rel=canonical` — SEOnaut's standout trick; headers are in the normalized record by design). ~2 days.
- _Orphan pages_: reorder seeding to **sitemap-last** (we currently seed sitemap URLs _first_, the opposite of the SEOnaut pattern we cite, so sitemap pages eat the page budget and link discovery never exhausts), and **emit orphan findings only when the crawl completed** (queue exhausted below maxPages) — on truncated crawls "no observed inlinks" is true of nearly everything and would mass-false-positive, the exact Semrush health-score noise we mock. ~34 days including reordering and gating.
- _Thin content + crawl depth_: explicit disposition — **build in P0** (word count already stored; BFS depth is one column). _Page size_: P1. ~1 day.
**3. Issues UX.** Critical/Warning/Info grouping, one-paragraph plain-English explanation per type (adapt SEOnaut's YAML copy), affected-URL drill-down, LibreCrawl's ~100-glob default exclusions (wp-admin, checkout, feeds), and **CSV export** (previously unaccounted; it's one of the three core report capabilities from §1). ~45 days.
**4. MCP audit tools — reframed.** Draft 1 claimed "a demo no competitor can match." **That was wrong: Ahrefs ships site-audit MCP tools today and Semrush has an MCP server.** Our wedge is not MCP existence; it's **credit-priced, zero-setup, agent-actionable audits**: audit any URL on demand for cents with no project setup or $129+/mo subscription, and return per-issue _fix instructions an agent can execute_ (the issue record carries remediation steps, not just labels). Tool shapes designed against Ahrefs' actual, inspectable MCP audit tools: `run_site_audit`, `get_audit_status`, `get_audit_issues(severity, type)`, `get_audit_pages(filter)`, each issue carrying `how_to_fix`. ~4 days.
**5. Block detection — moved to weeks 12 (was week 5).** The panel is right that for our segment (indie founders on Vercel/Cloudflare/Netlify — bot protection there is a one-toggle setting and IP-reputation challenges hit datacenter egress even without it; the week-1 baseline will establish the real rate) blocking is a **first-run activation killer**, not an ops metric — and our own onboarding audit shows activation leaks are the existential problem. Classify 403s/challenge bodies/`cf-mitigated` headers as "we were blocked," say so honestly instead of recording error pages, and show owner-targeted allowlist instructions inline (GSC-verified flow, §4). This layer is response-classification, hence engine-independent, and instruments block-rate telemetry from day one. ~3 days.
### P1 — weeks 814
- **Scheduled audits + diffing**: extend the existing cron; store per-run issue counts, show deltas ("12 new broken links since last week"). The retention feature and the reason cloud beats desktop. **Data retention policy (new):** full link graphs kept for the latest 2 runs only; older runs keep issue counts/deltas (a 10k-page graph is ~100150MB against D1's 10GB cap — weekly large audits would eat gigabytes/year otherwise). Large-audit edge lists optionally archived as compressed R2 blobs.
- **AI-crawler accessibility audit — promoted from P2.** Which AI bots (GPTBot, ClaudeBot) can reach your content, robots/Cloudflare AI-block detection, is-key-content-in-static-HTML. The panel is right that our own words ("open greenfield, perfectly on-brand") contradicted P2 placement: it's mostly computable from data we already collect, and "is ChatGPT blocked from your site?" is the launch/marketing asset for the whole relaunch — nobody switches auditors for HSTS checks. (Skip llms.txt as a headline; no major LLM vendor consumes it.)
- **JS rendering — deferred (consequence of the no-vendor decision, §4).** Plain fetch can't render SPAs. P1 ships the honest mitigation instead: detect client-rendered shells (near-empty body text with a large script payload / root-div-only HTML) and flag "this site renders client-side; HTML-level checks are incomplete" rather than reporting false missing-title/thin-content issues. If demand materializes, the revisit path is Cloudflare Browser Rendering as an opt-in, credit-priced add-on for GSC-verified own-site audits (where the owner can allowlist, neutralizing the fingerprint problem) — not a vendor API.
- **Security/hygiene checks** (HTTPS, mixed content, HSTS/CSP, insecure forms) and **external broken links** (HEAD checks, per-crawl cache) — demoted below AI-crawler work.
- **Cloudflare Verified Bots**: reframed as a **lottery ticket, not a mitigation**. Eligibility requires established, meaningful crawl volume; a new low-volume crawler routinely doesn't qualify, and Web Bot Auth is beta on Cloudflare's side. File opportunistically for our own UA (`OpenSEOBot`); plan zero dependency on it. The owner-targeted allowlist flow is the only mitigation we count on.
### P2 — later
- **`npx openseo-audit` local agent** (see §4 option D) — gated on block-rate telemetry, which now exists from week 1.
- **Crawl snapshots in R2** ("what we saw" evidence; lightweight SEOnaut WACZ analog).
- **Hreflang validation, structured-data validation, page-size checks, link-placement classification.**
**Skip entirely (explicit dispositions):** 300-check parity, crawl visualizations, E-E-A-T scores, multi-language UI.
---
## 4. Architecture: where should the crawl run?
| | **A. In-house Workers crawl + mitigations** | **B. DataForSEO OnPage offload** | **C. Desktop app** | **D. Hybrid: cloud + `npx` local agent** |
| -------------------- | -------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------- | ----------------------------------------------------------- |
| Blocking reliability | Low-medium: no stable IPs, no JS-challenge survival | Medium: still datacenter, but published IPs + named UA are _allowlistable_ | Highest (residential IP, real browser) | High when used; manual trigger only |
| Scheduled audits | Yes (≤ a few k pages — step budget) | Yes | **No** | Yes (cloud path); **no for the blocked sites D exists for** |
| JS rendering | No | Yes (priced) | Yes | Yes (local Playwright) |
| Eng effort | Ongoing hardening forever | Low (vendor/SDK/billing integrated) | **Months** + signing/notarization/3-platform QA | ~35 weeks (see below) |
| Cost per 1k pages | **Not $0**: D1 link writes ≈ $0.30/1k (with indexes) + Workers/Workflows | ~$0.1250.50 base, ~$1.25 JS _(list prices from architecture research — re-verify if the §4 revisit triggers ever fire)_ | $0 marginal | $0 marginal + ingest |
| Max crawl size | ~8.5k today, ~1225k post-refactor (Lighthouse `all` binds first; instance chaining removes the ceiling) | Vendor-side, no Workflow ceiling | RAM-limited | Local machine-limited |
| MCP / billing fit | Good / fine | Good / per-page cost maps to credits | **Bad** / awkward | Excellent for developers / fine |
**Excluded options, explicitly (panel asked):** _Cloudflare Browser Rendering_ — solves JS but not blocking (egress is still Cloudflare datacenter IPs), adds per-browser-hour cost, and headless-Chrome fingerprints start their own anti-bot treadmill. With the vendor path now declined (see decision below), this becomes the only _cloud_ JS-rendering route (the P2 `npx` agent's local Playwright is the other) — scoped as a P2 opt-in for GSC-verified own-site audits where the owner can allowlist us first. _Residential/ISP proxy egress_ — rejected on principle and practice: it disguises crawler identity, which contradicts our documented-UA/allowlist posture and many proxy networks' sourcing is ethically murky; cost (~$18/GB) is also worst-in-class.
**Decision: A — in-house Workers crawler (founder call: avoid DataForSEO vendor costs).** The original recommendation was B, contingent on a week-1 spike; the founder has decided against adding vendor spend, and "we were blocked" honest flagging is accepted as the blocking posture. That is a defensible position, but the proposal's job is to price it accurately. What the decision buys, and what it costs:
**What going in-house buys:**
- **No per-audit vendor bill on a free acquisition feature.** Every OnPage call is marginal cash out (~$0.0060.025 per free 50-page audit, $1.255 per 10k-page crawl at list prices), and DataForSEO meters even failed tasks — a real footgun we've already hit on the SERP side. Crucially, the comparison must not double-count: D1 storage writes are common to **both** engines (vendor results would land in the same `audit_pages`/`audit_links` tables to run the same SQL checks), so the vendor fee is purely additive. In-house wins on marginal cost at every crawl size.
- **No new vendor dependency or pricing-change exposure** on the product's centerpiece free feature; full control of crawl behavior, UA identity, and politeness.
- **One less integration**: no task-queue/webhook polling model, no mapping a vendor's page schema into the normalized record.
**What it costs — priced honestly:**
1. **"Basically free" is true only until the free tier works.** D1 Paid includes 50M row-writes/month _shared with the entire product_. A 50-page free audit writes ~815k rows (pages + links + indexes), so **~35k free audits/month consumes the whole included allotment**; beyond it, $1/M rows — at 10k free audits/month, ~$50100/month in-house vs ~$60250/month vendor-side. A ~2x cost win, not zero, and it eats quota headroom from every other D1 feature. Required companion: a **TTL on free-tier audit data** (e.g., purge one-off free audits' link graphs after 30 days) — the P1 two-run retention policy only covers scheduled audits and doesn't protect the 10GB D1 cap from thousands of one-off free crawls.
2. **JS rendering is off the roadmap** until/unless we build it on Cloudflare Browser Rendering. SPA-heavy sites (a real slice of indie-founder stacks: Next.js CSR pages, Vite SPAs) get degraded audits. Mitigation: detect-and-disclose (§3 P1), never silently mis-report — and count CSR-shell prevalence in the week-1 telemetry so the JS-rendering revisit trigger has a denominator too.
3. **A step-budget ceiling — real, but it's Lighthouse that binds, not the crawl.** The ledger: today's 3 steps per 25-page batch supports ~8.5k pages in the 1,024-step budget; after the P0 #0 refactor (12 steps/batch), the crawl alone supports **~1225k pages** — the 10k clamp can stand. The binding constraint is Lighthouse strategy `all` (2 steps per 10-URL batch ≈ 1,000 steps at 5k pages, over budget by itself): cap `all` at ~2k pages or switch to sampling, and document page caps _coupled to Lighthouse strategy_. Nor is the ceiling permanent: Workflow instances can chain (the P0 #0 refactor already puts frontier state in D1; a continuation instance resumes it) — roughly a week of cross-instance dedupe work, not an engine swap. The agency/large-site segment is deferred, not foreclosed.
4. **Weaker blocking mitigation — and the allowlist fix is unverified on the exact tiers our ICP uses.** Workers egress comes from shared Cloudflare IP ranges — we cannot publish a dedicated allowlistable IP, so the GSC-verified fix flow must be **User-Agent-based** (`OpenSEOBot/1.0` via a WAF custom rule; spoofable, tolerable when the owner allowlists their _own_ site). But: on **Cloudflare Free, Bot Fight Mode supports no exceptions** — a skip-by-UA rule does not bypass it; the owner's only fix there is turning BFM off entirely, a much heavier ask than "add this one rule." Vercel's challenge bypass is plan-gated and may be unavailable on Hobby. **Week 1 therefore includes a 12 day validation**: crawl test zones on Cloudflare Free (BFM on), Cloudflare Pro (Super BFM), and Vercel Hobby/Pro (Attack Challenge Mode), and write per-platform fix copy that is honest where the real answer is "disable bot protection." Verified Bots filing (P1) stays a lottery ticket.
5. **The ~12 weeks of crawler work inside P0 that the vendor path would have eliminated, plus permanent ownership.** Manual-redirect handling, per-host politeness inside steps, sitemap-last seeding, SSRF fix, step-budget restructuring, timeout/retry edges — already counted inside the 68-week P0 estimate (not additional to it), but under B most of it would have been deleted, and the hardening treadmill (new WAF behaviors, new edge cases) is ours indefinitely.
6. **No fallback engine for blocked sites.** Under B, a site blocking us might still be crawlable via the vendor's allowlisted infrastructure. In-house, a blocked site stays blocked until the owner acts. The honest-flagging posture makes this acceptable for owned sites: "we were blocked, here's the fix" instead of pretending.
7. **Third-party-URL audits have no remedy at all.** The MCP wedge includes "audit any URL on demand" (competitor audits, agent workflows) — but the allowlist flow only works for sites the user _owns_. For non-owned URLs, "we were blocked" is a terminal state, exactly where vendor infrastructure would have helped. Disposition: scope the claim — on-demand MCP audits are first-class for owned/verified sites, best-effort for everything else, and the tool response says which one it was.
**Revisit triggers (set now, so the decision is cheap to reverse).** Define "audit blocked" first: start URL blocked, or ≥20% of fetched URLs classified as challenge/`cf-mitigated`. Then: (a) >1520% of first-run audits blocked _and_ the fix funnel (block shown → instructions viewed → re-run unblocked, instrumented in week 1) shows owners not converting; (b) sustained demand for >10k-page crawls that chaining doesn't satisfy, or for JS rendering (CSR-shell prevalence from week-1 telemetry is the denominator); (c) DataForSEO/competitor pricing shifts. The reporter contract and schema are engine-agnostic by design, so adopting OnPage later — or the `npx` local agent (D) — is an adapter, not a rewrite.
**Implementation notes for A (from §2's constraints):** politeness delays serialized _inside_ steps (wall-clock waits are fine; more steps are not), progress steps collapsed to stay under the 1,024-step budget, incremental per-batch D1 writes (P0 #0), block-rate baseline pulled from existing audit data (status 0/403 rates + PostHog) in week 1 so the revisit triggers have a denominator.
One honest limitation, stated plainly: **for sites that block all datacenter crawlers, every recommended option loses scheduled audits** — D is a manual local run. The cloud-scheduling advantage evaporates for exactly that segment; our remedy there is the GSC-verified allowlist-fix flow, which restores cloud crawling rather than working around it.
**Why not C (the desktop question, directly).** Desktop genuinely solves blocking — it's Screaming Frog's model. But scheduled audits die, server-side MCP can't wake a sleeping laptop, and it's a second product (signing, notarization, auto-update, months of work) whose own failure modes — RAM crashes, machine tie-up — are that model's loudest complaints. Draft 1's "~80% of the benefit for ~5% of the cost" for D was invented precision; the honest version: D recovers desktop's _crawl-origin_ benefit (residential IP, optional local Playwright) but not unattended operation, requires a **new** authenticated bulk-ingest endpoint feeding `audit_pages`/`audit_links` (no ingest path exists today; D depends on the P0 persistence refactor) plus a local crawler — the cheerio extraction logic is plain TS and ports, the Workflow orchestration doesn't. Estimate **~35 weeks**. **Panel pushback:** the panel suggested that if D balloons to 68 weeks the C-vs-D calculus changes. It doesn't: C strictly contains D's costs (local crawler, Playwright, sync) _plus_ installers/signing/3-platform QA, and still loses scheduling and MCP. Worst-case D remains cheaper than best-case C. Scope honesty: D serves the developer half of the ICP; for SEO freelancers (not terminal-native, and whose clients' sites are exactly the WAF'd ones) the realistic remedy is the GSC-verified allowlist flow — which is why that flow ranks above the CLI.
**The GSC idea, honestly.** GSC verification **does not bypass any WAF** — Cloudflare neither knows nor cares about Google's ownership records. What it buys: (1) **policy cover** to raise crawl _rate_ for verified owners (industry-standard move; the Ahrefs precedent is directional, see §1) — but robots.txt respect stays the default with a per-project owner toggle, not auto-ignored: crawl traps and server load are real, and "verified" doesn't repeal politeness; (2) **the trustworthy block-fix loop** — on a detected block for a GSC-verified site, show the _owner_ the exact Cloudflare/Vercel allowlist rule for our UA (`OpenSEOBot` — UA-based since Workers egress IPs aren't publishable, see decision above). Highest-leverage mitigation available; OAuth already built; now surfaced in onboarding, not post-P0; (3) later, GSC-enriched audits (true orphans = GSC impressions, zero inlinks).
---
## 5. Recommended roadmap
**Week 1:** `audit_pages`/`audit_links`/`audit_issues` schema + engine-agnostic reporter input contract; ship block detection + honest "we were blocked" messaging + telemetry — audit-level blocked definition, fix-funnel events (block shown → instructions viewed → re-run unblocked), CSR-shell prevalence counting, baseline from existing audit data (status 0/403 + PostHog; note raw 403s overcount — classify challenges, don't just count status codes); set the documented UA (`OpenSEOBot/1.0` with a /bot info page); **12 day allowlist validation** on Cloudflare Free (BFM), Cloudflare Pro (SBFM), Vercel Hobby/Pro (Attack Challenge) → per-platform fix copy.
**Weeks 24:** incremental persistence refactor + SSRF fix; crawler hardening (manual redirects, per-host politeness inside steps, sitemap-last seeding, step collapse; page caps coupled to Lighthouse strategy — `all` capped ~2k or sampled, crawl clamp stays 10k); issue engine + per-page reporters on the normalized record; multipage SQL checks (duplicates, broken links, chains); orphans (sitemap-last + completeness gating).
**Weeks 46:** Issues UX (severity, copy, drill-down, exclusions, CSV export); MCP audit tools with `how_to_fix` payloads; GSC-verified allowlist-fix flow in onboarding (UA-based WAF rule instructions).
**Weeks 68:** client-side-rendering detection + disclosure; pricing posture shipped (free 50-page audits, credits beyond) with the free-tier TTL/retention policy from §4 cost #1; buffer for the hardening tail that always exists when you own the crawler.
**Weeks 814 (P1):** scheduled audits + deltas with the 2-run link-graph retention policy; **AI-crawler accessibility report as the relaunch marketing asset**; thin-content/depth checks if not landed; security + external-link checks; opportunistic Verified Bots filing for `OpenSEOBot`.
**Later (P2):** `npx` local agent (new ingest endpoint; gate on block telemetry vs the §4 triggers); Cloudflare Browser Rendering JS audits for GSC-verified sites if demand shows; R2 crawl snapshots; hreflang/structured-data/page-size.
**The one-sentence version:** keep audits in the cloud on our own crawler (founder call — no vendor spend), give the small ones away to drive signups, spend engineering on a portable issue engine plus agent-actionable MCP output, answer bot-blocking with day-one honest "we were blocked" flagging and owner-targeted allowlist fixes validated per-platform in week 1, accept no JS rendering for now and Lighthouse-coupled page caps — with pre-set, measurable telemetry triggers that tell us if the no-vendor bet needs revisiting, a local CLI later, and a desktop app never.

View File

@ -0,0 +1,36 @@
CREATE TABLE "audit_issues" (
"id" text PRIMARY KEY NOT NULL,
"audit_id" text NOT NULL,
"page_id" text,
"page_url" text NOT NULL,
"issue_type" text NOT NULL,
"severity" text DEFAULT 'info' NOT NULL,
"details_json" text
);
--> statement-breakpoint
CREATE TABLE "audit_links" (
"id" text PRIMARY KEY NOT NULL,
"audit_id" text NOT NULL,
"source_page_id" text NOT NULL,
"source_url" text NOT NULL,
"target_url" text NOT NULL,
"anchor" text,
"is_internal" boolean DEFAULT true NOT NULL,
"is_nofollow" boolean DEFAULT false NOT NULL
);
--> statement-breakpoint
ALTER TABLE "audit_pages" ADD COLUMN "x_robots_tag" text;--> statement-breakpoint
ALTER TABLE "audit_pages" ADD COLUMN "header_canonical_url" text;--> statement-breakpoint
ALTER TABLE "audit_pages" ADD COLUMN "crawl_depth" integer;--> statement-breakpoint
ALTER TABLE "audit_pages" ADD COLUMN "in_sitemap" boolean DEFAULT false NOT NULL;--> statement-breakpoint
ALTER TABLE "audit_pages" ADD COLUMN "content_hash" text;--> statement-breakpoint
ALTER TABLE "audit_pages" ADD COLUMN "fetch_class" text DEFAULT 'ok' NOT NULL;--> statement-breakpoint
ALTER TABLE "audit_issues" ADD CONSTRAINT "audit_issues_audit_id_audits_id_fk" FOREIGN KEY ("audit_id") REFERENCES "public"."audits"("id") ON DELETE cascade ON UPDATE no action;--> statement-breakpoint
ALTER TABLE "audit_issues" ADD CONSTRAINT "audit_issues_page_id_audit_pages_id_fk" FOREIGN KEY ("page_id") REFERENCES "public"."audit_pages"("id") ON DELETE cascade ON UPDATE no action;--> statement-breakpoint
ALTER TABLE "audit_links" ADD CONSTRAINT "audit_links_audit_id_audits_id_fk" FOREIGN KEY ("audit_id") REFERENCES "public"."audits"("id") ON DELETE cascade ON UPDATE no action;--> statement-breakpoint
ALTER TABLE "audit_links" ADD CONSTRAINT "audit_links_source_page_id_audit_pages_id_fk" FOREIGN KEY ("source_page_id") REFERENCES "public"."audit_pages"("id") ON DELETE cascade ON UPDATE no action;--> statement-breakpoint
CREATE INDEX "audit_issues_audit_id_idx" ON "audit_issues" USING btree ("audit_id");--> statement-breakpoint
CREATE INDEX "audit_issues_audit_type_idx" ON "audit_issues" USING btree ("audit_id","issue_type");--> statement-breakpoint
CREATE INDEX "audit_links_audit_id_idx" ON "audit_links" USING btree ("audit_id");--> statement-breakpoint
CREATE INDEX "audit_links_audit_target_idx" ON "audit_links" USING btree ("audit_id","target_url");--> statement-breakpoint
CREATE INDEX "audit_pages_audit_url_idx" ON "audit_pages" USING btree ("audit_id","url");

File diff suppressed because it is too large Load Diff

View File

@ -50,6 +50,13 @@
"when": 1783438455708,
"tag": "0006_location_name",
"breakpoints": true
},
{
"idx": 7,
"version": "7",
"when": 1783306049167,
"tag": "0007_same_marvel_zombies",
"breakpoints": true
}
]
}

View File

@ -0,0 +1,36 @@
CREATE TABLE `audit_issues` (
`id` text PRIMARY KEY NOT NULL,
`audit_id` text NOT NULL,
`page_id` text,
`page_url` text NOT NULL,
`issue_type` text NOT NULL,
`severity` text DEFAULT 'info' NOT NULL,
`details_json` text,
FOREIGN KEY (`audit_id`) REFERENCES `audits`(`id`) ON UPDATE no action ON DELETE cascade,
FOREIGN KEY (`page_id`) REFERENCES `audit_pages`(`id`) ON UPDATE no action ON DELETE cascade
);
--> statement-breakpoint
CREATE INDEX `audit_issues_audit_id_idx` ON `audit_issues` (`audit_id`);--> statement-breakpoint
CREATE INDEX `audit_issues_audit_type_idx` ON `audit_issues` (`audit_id`,`issue_type`);--> statement-breakpoint
CREATE TABLE `audit_links` (
`id` text PRIMARY KEY NOT NULL,
`audit_id` text NOT NULL,
`source_page_id` text NOT NULL,
`source_url` text NOT NULL,
`target_url` text NOT NULL,
`anchor` text,
`is_internal` integer DEFAULT true NOT NULL,
`is_nofollow` integer DEFAULT false NOT NULL,
FOREIGN KEY (`audit_id`) REFERENCES `audits`(`id`) ON UPDATE no action ON DELETE cascade,
FOREIGN KEY (`source_page_id`) REFERENCES `audit_pages`(`id`) ON UPDATE no action ON DELETE cascade
);
--> statement-breakpoint
CREATE INDEX `audit_links_audit_id_idx` ON `audit_links` (`audit_id`);--> statement-breakpoint
CREATE INDEX `audit_links_audit_target_idx` ON `audit_links` (`audit_id`,`target_url`);--> statement-breakpoint
ALTER TABLE `audit_pages` ADD `x_robots_tag` text;--> statement-breakpoint
ALTER TABLE `audit_pages` ADD `header_canonical_url` text;--> statement-breakpoint
ALTER TABLE `audit_pages` ADD `crawl_depth` integer;--> statement-breakpoint
ALTER TABLE `audit_pages` ADD `in_sitemap` integer DEFAULT false NOT NULL;--> statement-breakpoint
ALTER TABLE `audit_pages` ADD `content_hash` text;--> statement-breakpoint
ALTER TABLE `audit_pages` ADD `fetch_class` text DEFAULT 'ok' NOT NULL;--> statement-breakpoint
CREATE INDEX `audit_pages_audit_url_idx` ON `audit_pages` (`audit_id`,`url`);

File diff suppressed because it is too large Load Diff

View File

@ -211,6 +211,13 @@
"when": 1783438453855,
"tag": "0029_location_name",
"breakpoints": true
},
{
"idx": 30,
"version": "6",
"when": 1783306047190,
"tag": "0030_legal_reaper",
"breakpoints": true
}
]
}

View File

@ -345,123 +345,3 @@ export const rankSnapshots = sqliteTable(
),
],
);
// ============================================================================
// Site Audit tables
// ============================================================================
// One row per audit run
export const audits = sqliteTable(
"audits",
{
id: text("id").primaryKey(),
projectId: text("project_id")
.notNull()
.references(() => projects.id, { onDelete: "cascade" }),
startedByUserId: text("started_by_user_id").notNull(),
startUrl: text("start_url").notNull(),
status: text("status", {
enum: ["running", "completed", "failed"],
})
.notNull()
.default("running"),
workflowInstanceId: text("workflow_instance_id"),
// JSON config: { maxPages, lighthouseStrategy }
config: text("config").notNull().default("{}"),
// Progress & summary
pagesCrawled: integer("pages_crawled").notNull().default(0),
pagesTotal: integer("pages_total").notNull().default(0),
lighthouseTotal: integer("lighthouse_total").notNull().default(0),
lighthouseCompleted: integer("lighthouse_completed").notNull().default(0),
lighthouseFailed: integer("lighthouse_failed").notNull().default(0),
currentPhase: text("current_phase").default("discovery"),
startedAt: text("started_at")
.notNull()
.default(sql`(current_timestamp)`),
completedAt: text("completed_at"),
},
(table) => [
index("audits_project_id_idx").on(table.projectId),
index("audits_started_by_user_id_idx").on(table.startedByUserId),
],
);
// One row per crawled page
export const auditPages = sqliteTable(
"audit_pages",
{
id: text("id").primaryKey(),
auditId: text("audit_id")
.notNull()
.references(() => audits.id, { onDelete: "cascade" }),
url: text("url").notNull(),
statusCode: integer("status_code"),
redirectUrl: text("redirect_url"),
// Metadata
title: text("title"),
metaDescription: text("meta_description"),
canonicalUrl: text("canonical_url"),
robotsMeta: text("robots_meta"),
// Open Graph
ogTitle: text("og_title"),
ogDescription: text("og_description"),
ogImage: text("og_image"),
// Headings
h1Count: integer("h1_count").notNull().default(0),
h2Count: integer("h2_count").notNull().default(0),
h3Count: integer("h3_count").notNull().default(0),
h4Count: integer("h4_count").notNull().default(0),
h5Count: integer("h5_count").notNull().default(0),
h6Count: integer("h6_count").notNull().default(0),
headingOrderJson: text("heading_order_json"),
// Content
wordCount: integer("word_count").notNull().default(0),
// Images
imagesTotal: integer("images_total").notNull().default(0),
imagesMissingAlt: integer("images_missing_alt").notNull().default(0),
imagesJson: text("images_json"),
// Links
internalLinkCount: integer("internal_link_count").notNull().default(0),
externalLinkCount: integer("external_link_count").notNull().default(0),
// Structured data
hasStructuredData: integer("has_structured_data", { mode: "boolean" })
.notNull()
.default(false),
// Hreflang
hreflangTagsJson: text("hreflang_tags_json"),
// Indexability
isIndexable: integer("is_indexable", { mode: "boolean" })
.notNull()
.default(true),
// Performance
responseTimeMs: integer("response_time_ms"),
},
(table) => [index("audit_pages_audit_id_idx").on(table.auditId)],
);
// One row per Lighthouse test (mobile + desktop per page).
export const auditLighthouseResults = sqliteTable(
"audit_lighthouse_results",
{
id: text("id").primaryKey(),
auditId: text("audit_id")
.notNull()
.references(() => audits.id, { onDelete: "cascade" }),
pageId: text("page_id")
.notNull()
.references(() => auditPages.id, { onDelete: "cascade" }),
strategy: text("strategy", { enum: ["mobile", "desktop"] }).notNull(),
performanceScore: integer("performance_score"),
accessibilityScore: integer("accessibility_score"),
bestPracticesScore: integer("best_practices_score"),
seoScore: integer("seo_score"),
lcpMs: real("lcp_ms"),
cls: real("cls"),
inpMs: real("inp_ms"),
ttfbMs: real("ttfb_ms"),
errorMessage: text("error_message"),
r2Key: text("r2_key"),
payloadSizeBytes: integer("payload_size_bytes"),
},
(table) => [index("audit_lighthouse_results_audit_id_idx").on(table.auditId)],
);

201
src/db/audit.schema.ts Normal file
View File

@ -0,0 +1,201 @@
import {
sqliteTable,
text,
integer,
real,
index,
} from "drizzle-orm/sqlite-core";
import { sql } from "drizzle-orm";
import { projects } from "./app.schema";
// ============================================================================
// Site Audit tables
// ============================================================================
// One row per audit run
export const audits = sqliteTable(
"audits",
{
id: text("id").primaryKey(),
projectId: text("project_id")
.notNull()
.references(() => projects.id, { onDelete: "cascade" }),
startedByUserId: text("started_by_user_id").notNull(),
startUrl: text("start_url").notNull(),
status: text("status", {
enum: ["running", "completed", "failed"],
})
.notNull()
.default("running"),
workflowInstanceId: text("workflow_instance_id"),
// JSON config: { maxPages, lighthouseStrategy }
config: text("config").notNull().default("{}"),
// Progress & summary
pagesCrawled: integer("pages_crawled").notNull().default(0),
pagesTotal: integer("pages_total").notNull().default(0),
lighthouseTotal: integer("lighthouse_total").notNull().default(0),
lighthouseCompleted: integer("lighthouse_completed").notNull().default(0),
lighthouseFailed: integer("lighthouse_failed").notNull().default(0),
currentPhase: text("current_phase").default("discovery"),
startedAt: text("started_at")
.notNull()
.default(sql`(current_timestamp)`),
completedAt: text("completed_at"),
},
(table) => [
index("audits_project_id_idx").on(table.projectId),
index("audits_started_by_user_id_idx").on(table.startedByUserId),
],
);
// One row per crawled page
export const auditPages = sqliteTable(
"audit_pages",
{
id: text("id").primaryKey(),
auditId: text("audit_id")
.notNull()
.references(() => audits.id, { onDelete: "cascade" }),
url: text("url").notNull(),
statusCode: integer("status_code"),
redirectUrl: text("redirect_url"),
// Metadata
title: text("title"),
metaDescription: text("meta_description"),
canonicalUrl: text("canonical_url"),
robotsMeta: text("robots_meta"),
// Open Graph
ogTitle: text("og_title"),
ogDescription: text("og_description"),
ogImage: text("og_image"),
// Headings
h1Count: integer("h1_count").notNull().default(0),
h2Count: integer("h2_count").notNull().default(0),
h3Count: integer("h3_count").notNull().default(0),
h4Count: integer("h4_count").notNull().default(0),
h5Count: integer("h5_count").notNull().default(0),
h6Count: integer("h6_count").notNull().default(0),
headingOrderJson: text("heading_order_json"),
// Content
wordCount: integer("word_count").notNull().default(0),
// Images
imagesTotal: integer("images_total").notNull().default(0),
imagesMissingAlt: integer("images_missing_alt").notNull().default(0),
imagesJson: text("images_json"),
// Links
internalLinkCount: integer("internal_link_count").notNull().default(0),
externalLinkCount: integer("external_link_count").notNull().default(0),
// Structured data
hasStructuredData: integer("has_structured_data", { mode: "boolean" })
.notNull()
.default(false),
// Hreflang
hreflangTagsJson: text("hreflang_tags_json"),
// Indexability
isIndexable: integer("is_indexable", { mode: "boolean" })
.notNull()
.default(true),
// Indexability/canonical signals from response headers
xRobotsTag: text("x_robots_tag"),
headerCanonicalUrl: text("header_canonical_url"),
// Crawl metadata
// null depth = not reached via links (e.g. sitemap-seeded)
crawlDepth: integer("crawl_depth"),
inSitemap: integer("in_sitemap", { mode: "boolean" })
.notNull()
.default(false),
// SHA-256 of the visible body text, for duplicate-content grouping
contentHash: text("content_hash"),
// How the fetch resolved: ok | blocked (WAF/bot challenge) | error
fetchClass: text("fetch_class", { enum: ["ok", "blocked", "error"] })
.notNull()
.default("ok"),
// Performance
responseTimeMs: integer("response_time_ms"),
},
(table) => [
index("audit_pages_audit_id_idx").on(table.auditId),
index("audit_pages_audit_url_idx").on(table.auditId, table.url),
],
);
// One row per unique (source page, target URL) link edge. Currently only
// internal edges are stored (see AuditRepository); isInternal stays so
// external-link checks can start writing rows without a migration.
export const auditLinks = sqliteTable(
"audit_links",
{
id: text("id").primaryKey(),
auditId: text("audit_id")
.notNull()
.references(() => audits.id, { onDelete: "cascade" }),
sourcePageId: text("source_page_id")
.notNull()
.references(() => auditPages.id, { onDelete: "cascade" }),
sourceUrl: text("source_url").notNull(),
targetUrl: text("target_url").notNull(),
anchor: text("anchor"),
isInternal: integer("is_internal", { mode: "boolean" })
.notNull()
.default(true),
isNofollow: integer("is_nofollow", { mode: "boolean" })
.notNull()
.default(false),
},
(table) => [
index("audit_links_audit_id_idx").on(table.auditId),
index("audit_links_audit_target_idx").on(table.auditId, table.targetUrl),
],
);
// One row per (issue type, affected page)
export const auditIssues = sqliteTable(
"audit_issues",
{
id: text("id").primaryKey(),
auditId: text("audit_id")
.notNull()
.references(() => audits.id, { onDelete: "cascade" }),
pageId: text("page_id").references(() => auditPages.id, {
onDelete: "cascade",
}),
pageUrl: text("page_url").notNull(),
issueType: text("issue_type").notNull(),
severity: text("severity", { enum: ["critical", "warning", "info"] })
.notNull()
.default("info"),
// JSON details specific to the issue type (e.g. broken link target)
detailsJson: text("details_json"),
},
(table) => [
index("audit_issues_audit_id_idx").on(table.auditId),
index("audit_issues_audit_type_idx").on(table.auditId, table.issueType),
],
);
// One row per Lighthouse test (mobile + desktop per page).
export const auditLighthouseResults = sqliteTable(
"audit_lighthouse_results",
{
id: text("id").primaryKey(),
auditId: text("audit_id")
.notNull()
.references(() => audits.id, { onDelete: "cascade" }),
pageId: text("page_id")
.notNull()
.references(() => auditPages.id, { onDelete: "cascade" }),
strategy: text("strategy", { enum: ["mobile", "desktop"] }).notNull(),
performanceScore: integer("performance_score"),
accessibilityScore: integer("accessibility_score"),
bestPracticesScore: integer("best_practices_score"),
seoScore: integer("seo_score"),
lcpMs: real("lcp_ms"),
cls: real("cls"),
inpMs: real("inp_ms"),
ttfbMs: real("ttfb_ms"),
errorMessage: text("error_message"),
r2Key: text("r2_key"),
payloadSizeBytes: integer("payload_size_bytes"),
},
(table) => [index("audit_lighthouse_results_audit_id_idx").on(table.auditId)],
);

View File

@ -2,6 +2,7 @@
// which is the provider-aware barrel) so the D1 client always binds to the
// SQLite tables regardless of DATABASE_PROVIDER.
export * from "../app.schema";
export * from "../audit.schema";
export * from "../sam.schema";
export * from "../better-auth-schema";
export * from "../billing.schema";

View File

@ -338,112 +338,3 @@ export const rankSnapshots = pgTable(
),
],
);
// ============================================================================
// Site Audit tables
// ============================================================================
// One row per audit run
export const audits = pgTable(
"audits",
{
id: text("id").primaryKey(),
projectId: text("project_id")
.notNull()
.references(() => projects.id, { onDelete: "cascade" }),
startedByUserId: text("started_by_user_id").notNull(),
startUrl: text("start_url").notNull(),
status: text("status", {
enum: ["running", "completed", "failed"],
})
.notNull()
.default("running"),
workflowInstanceId: text("workflow_instance_id"),
// JSON config: { maxPages, lighthouseStrategy }
config: text("config").notNull().default("{}"),
// Progress & summary
pagesCrawled: integer("pages_crawled").notNull().default(0),
pagesTotal: integer("pages_total").notNull().default(0),
lighthouseTotal: integer("lighthouse_total").notNull().default(0),
lighthouseCompleted: integer("lighthouse_completed").notNull().default(0),
lighthouseFailed: integer("lighthouse_failed").notNull().default(0),
currentPhase: text("current_phase").default("discovery"),
startedAt: timestampColumn("started_at").notNull().default(isoNow),
completedAt: timestampColumn("completed_at"),
},
(table) => [
index("audits_project_id_idx").on(table.projectId),
index("audits_started_by_user_id_idx").on(table.startedByUserId),
],
);
// One row per crawled page
export const auditPages = pgTable(
"audit_pages",
{
id: text("id").primaryKey(),
auditId: text("audit_id")
.notNull()
.references(() => audits.id, { onDelete: "cascade" }),
url: text("url").notNull(),
statusCode: integer("status_code"),
redirectUrl: text("redirect_url"),
// Metadata
title: text("title"),
metaDescription: text("meta_description"),
canonicalUrl: text("canonical_url"),
robotsMeta: text("robots_meta"),
// Open Graph
ogTitle: text("og_title"),
ogDescription: text("og_description"),
ogImage: text("og_image"),
// Headings
h1Count: integer("h1_count").notNull().default(0),
h2Count: integer("h2_count").notNull().default(0),
h3Count: integer("h3_count").notNull().default(0),
h4Count: integer("h4_count").notNull().default(0),
h5Count: integer("h5_count").notNull().default(0),
h6Count: integer("h6_count").notNull().default(0),
headingOrderJson: text("heading_order_json"),
// Content
wordCount: integer("word_count").notNull().default(0),
// Images
imagesTotal: integer("images_total").notNull().default(0),
imagesMissingAlt: integer("images_missing_alt").notNull().default(0),
imagesJson: text("images_json"),
internalLinkCount: integer("internal_link_count").notNull().default(0),
externalLinkCount: integer("external_link_count").notNull().default(0),
hasStructuredData: boolean("has_structured_data").notNull().default(false),
hreflangTagsJson: text("hreflang_tags_json"),
isIndexable: boolean("is_indexable").notNull().default(true),
responseTimeMs: integer("response_time_ms"),
},
(table) => [index("audit_pages_audit_id_idx").on(table.auditId)],
);
// One row per Lighthouse test (mobile + desktop per page).
export const auditLighthouseResults = pgTable(
"audit_lighthouse_results",
{
id: text("id").primaryKey(),
auditId: text("audit_id")
.notNull()
.references(() => audits.id, { onDelete: "cascade" }),
pageId: text("page_id")
.notNull()
.references(() => auditPages.id, { onDelete: "cascade" }),
strategy: text("strategy", { enum: ["mobile", "desktop"] }).notNull(),
performanceScore: integer("performance_score"),
accessibilityScore: integer("accessibility_score"),
bestPracticesScore: integer("best_practices_score"),
seoScore: integer("seo_score"),
lcpMs: real("lcp_ms"),
cls: real("cls"),
inpMs: real("inp_ms"),
ttfbMs: real("ttfb_ms"),
errorMessage: text("error_message"),
r2Key: text("r2_key"),
payloadSizeBytes: integer("payload_size_bytes"),
},
(table) => [index("audit_lighthouse_results_audit_id_idx").on(table.auditId)],
);

196
src/db/pg/audit.schema.ts Normal file
View File

@ -0,0 +1,196 @@
import { sql } from "drizzle-orm";
import {
boolean,
index,
integer,
pgTable,
real,
text,
} from "drizzle-orm/pg-core";
import { projects } from "./app.schema";
// Timestamps are stored as *text* (same column shape as the SQLite schema); see
// the note in pg/app.schema.ts. `isoNow` matches `new Date().toISOString()` so
// DB-defaulted and app-written values sort together lexicographically.
const isoNow = sql`to_char(now() AT TIME ZONE 'utc', 'YYYY-MM-DD"T"HH24:MI:SS.MS"Z"')`;
const timestampColumn = (name: string) => text(name);
// ============================================================================
// Site Audit tables
// ============================================================================
// One row per audit run
export const audits = pgTable(
"audits",
{
id: text("id").primaryKey(),
projectId: text("project_id")
.notNull()
.references(() => projects.id, { onDelete: "cascade" }),
startedByUserId: text("started_by_user_id").notNull(),
startUrl: text("start_url").notNull(),
status: text("status", {
enum: ["running", "completed", "failed"],
})
.notNull()
.default("running"),
workflowInstanceId: text("workflow_instance_id"),
// JSON config: { maxPages, lighthouseStrategy }
config: text("config").notNull().default("{}"),
// Progress & summary
pagesCrawled: integer("pages_crawled").notNull().default(0),
pagesTotal: integer("pages_total").notNull().default(0),
lighthouseTotal: integer("lighthouse_total").notNull().default(0),
lighthouseCompleted: integer("lighthouse_completed").notNull().default(0),
lighthouseFailed: integer("lighthouse_failed").notNull().default(0),
currentPhase: text("current_phase").default("discovery"),
startedAt: timestampColumn("started_at").notNull().default(isoNow),
completedAt: timestampColumn("completed_at"),
},
(table) => [
index("audits_project_id_idx").on(table.projectId),
index("audits_started_by_user_id_idx").on(table.startedByUserId),
],
);
// One row per crawled page
export const auditPages = pgTable(
"audit_pages",
{
id: text("id").primaryKey(),
auditId: text("audit_id")
.notNull()
.references(() => audits.id, { onDelete: "cascade" }),
url: text("url").notNull(),
statusCode: integer("status_code"),
redirectUrl: text("redirect_url"),
// Metadata
title: text("title"),
metaDescription: text("meta_description"),
canonicalUrl: text("canonical_url"),
robotsMeta: text("robots_meta"),
// Open Graph
ogTitle: text("og_title"),
ogDescription: text("og_description"),
ogImage: text("og_image"),
// Headings
h1Count: integer("h1_count").notNull().default(0),
h2Count: integer("h2_count").notNull().default(0),
h3Count: integer("h3_count").notNull().default(0),
h4Count: integer("h4_count").notNull().default(0),
h5Count: integer("h5_count").notNull().default(0),
h6Count: integer("h6_count").notNull().default(0),
headingOrderJson: text("heading_order_json"),
// Content
wordCount: integer("word_count").notNull().default(0),
// Images
imagesTotal: integer("images_total").notNull().default(0),
imagesMissingAlt: integer("images_missing_alt").notNull().default(0),
imagesJson: text("images_json"),
// Links
internalLinkCount: integer("internal_link_count").notNull().default(0),
externalLinkCount: integer("external_link_count").notNull().default(0),
// Structured data
hasStructuredData: boolean("has_structured_data").notNull().default(false),
// Hreflang
hreflangTagsJson: text("hreflang_tags_json"),
// Indexability
isIndexable: boolean("is_indexable").notNull().default(true),
// Indexability/canonical signals from response headers
xRobotsTag: text("x_robots_tag"),
headerCanonicalUrl: text("header_canonical_url"),
// Crawl metadata
// null depth = not reached via links (e.g. sitemap-seeded)
crawlDepth: integer("crawl_depth"),
inSitemap: boolean("in_sitemap").notNull().default(false),
// SHA-256 of the visible body text, for duplicate-content grouping
contentHash: text("content_hash"),
// How the fetch resolved: ok | blocked (WAF/bot challenge) | error
fetchClass: text("fetch_class", { enum: ["ok", "blocked", "error"] })
.notNull()
.default("ok"),
// Performance
responseTimeMs: integer("response_time_ms"),
},
(table) => [
index("audit_pages_audit_id_idx").on(table.auditId),
index("audit_pages_audit_url_idx").on(table.auditId, table.url),
],
);
// One row per unique (source page, target URL) link edge. Currently only
// internal edges are stored (see AuditRepository); isInternal stays so
// external-link checks can start writing rows without a migration.
export const auditLinks = pgTable(
"audit_links",
{
id: text("id").primaryKey(),
auditId: text("audit_id")
.notNull()
.references(() => audits.id, { onDelete: "cascade" }),
sourcePageId: text("source_page_id")
.notNull()
.references(() => auditPages.id, { onDelete: "cascade" }),
sourceUrl: text("source_url").notNull(),
targetUrl: text("target_url").notNull(),
anchor: text("anchor"),
isInternal: boolean("is_internal").notNull().default(true),
isNofollow: boolean("is_nofollow").notNull().default(false),
},
(table) => [
index("audit_links_audit_id_idx").on(table.auditId),
index("audit_links_audit_target_idx").on(table.auditId, table.targetUrl),
],
);
// One row per (issue type, affected page)
export const auditIssues = pgTable(
"audit_issues",
{
id: text("id").primaryKey(),
auditId: text("audit_id")
.notNull()
.references(() => audits.id, { onDelete: "cascade" }),
pageId: text("page_id").references(() => auditPages.id, {
onDelete: "cascade",
}),
pageUrl: text("page_url").notNull(),
issueType: text("issue_type").notNull(),
severity: text("severity", { enum: ["critical", "warning", "info"] })
.notNull()
.default("info"),
// JSON details specific to the issue type (e.g. broken link target)
detailsJson: text("details_json"),
},
(table) => [
index("audit_issues_audit_id_idx").on(table.auditId),
index("audit_issues_audit_type_idx").on(table.auditId, table.issueType),
],
);
// One row per Lighthouse test (mobile + desktop per page).
export const auditLighthouseResults = pgTable(
"audit_lighthouse_results",
{
id: text("id").primaryKey(),
auditId: text("audit_id")
.notNull()
.references(() => audits.id, { onDelete: "cascade" }),
pageId: text("page_id")
.notNull()
.references(() => auditPages.id, { onDelete: "cascade" }),
strategy: text("strategy", { enum: ["mobile", "desktop"] }).notNull(),
performanceScore: integer("performance_score"),
accessibilityScore: integer("accessibility_score"),
bestPracticesScore: integer("best_practices_score"),
seoScore: integer("seo_score"),
lcpMs: real("lcp_ms"),
cls: real("cls"),
inpMs: real("inp_ms"),
ttfbMs: real("ttfb_ms"),
errorMessage: text("error_message"),
r2Key: text("r2_key"),
payloadSizeBytes: integer("payload_size_bytes"),
},
(table) => [index("audit_lighthouse_results_audit_id_idx").on(table.auditId)],
);

View File

@ -1,4 +1,5 @@
export * from "./app.schema";
export * from "./audit.schema";
export * from "./sam.schema";
export * from "./better-auth-schema";
export * from "./billing.schema";

View File

@ -1,11 +1,13 @@
import { getDatabaseProvider } from "./provider";
import * as sqliteApp from "./app.schema";
import * as sqliteAudit from "./audit.schema";
import * as sqliteSam from "./sam.schema";
import * as sqliteAuth from "./better-auth-schema";
import * as sqliteBilling from "./billing.schema";
import * as sqliteGsc from "./gsc.schema";
import * as sqliteReddit from "./reddit-attribution.schema";
import * as pgApp from "./pg/app.schema";
import * as pgAudit from "./pg/audit.schema";
import * as pgSam from "./pg/sam.schema";
import * as pgAuth from "./pg/better-auth-schema";
import * as pgBilling from "./pg/billing.schema";
@ -23,6 +25,7 @@ import * as pgReddit from "./pg/reddit-attribution.schema";
// schema is the one structural artifact NOT regenerated by `db:generate`, so the
// parity test is its drift guard.
type AppSchema = typeof sqliteApp &
typeof sqliteAudit &
typeof sqliteSam &
typeof sqliteAuth &
typeof sqliteBilling &
@ -31,9 +34,18 @@ type AppSchema = typeof sqliteApp &
const runtimeSchema =
getDatabaseProvider() === "postgres"
? { ...pgApp, ...pgSam, ...pgAuth, ...pgBilling, ...pgGsc, ...pgReddit }
? {
...pgApp,
...pgAudit,
...pgSam,
...pgAuth,
...pgBilling,
...pgGsc,
...pgReddit,
}
: {
...sqliteApp,
...sqliteAudit,
...sqliteSam,
...sqliteAuth,
...sqliteBilling,
@ -57,6 +69,8 @@ export const {
rankSnapshots,
audits,
auditPages,
auditLinks,
auditIssues,
auditLighthouseResults,
samSessions,
samProjectMemory,

View File

@ -1,17 +1,34 @@
/**
* Data access layer for site audit tables.
* Provider-aware (D1 or Postgres) via the `@/db` handle.
* Provider-aware (D1 or Postgres) via the `@/db` handle. Covers audits,
* audit_pages, audit_links, audit_issues, and stored Lighthouse results.
*/
import { and, desc, eq } from "drizzle-orm";
import { db } from "@/db";
import { audits, auditLighthouseResults, auditPages } from "@/db/schema";
import {
audits,
auditIssues,
auditLighthouseResults,
auditLinks,
auditPages,
} from "@/db/schema";
import { executeInBatches } from "@/db/runBatch";
import { AUDIT_ISSUE_TYPES } from "@/shared/audit-issues";
import { deterministicAuditRowId } from "@/server/lib/audit/ids";
import type { DetectedIssue } from "@/server/lib/audit/issues/page-reporters";
import type {
AuditConfig,
CrawledPageResult,
LighthouseResult,
StepPageResult,
} from "@/server/lib/audit/types";
// Only internal links are stored: both consumers (broken-internal-link and
// orphan checks) filter on isInternal, and per-page external counts already
// live on audit_pages. External rows come back when P1 adds external-link
// checks. Mega-menu/footer-heavy sites can carry 1000+ links per page; cap
// what we store so a 10k-page crawl can't write tens of millions of link rows.
const MAX_STORED_LINKS_PER_PAGE = 500;
async function createAudit(data: {
id: string;
projectId: string;
@ -84,6 +101,9 @@ async function completeAudit(
}
async function failAudit(auditId: string, workflowInstanceId: string) {
// Only a running audit can transition to failed: the getStatus reconciler
// races the workflow's own finalize, and without this guard it could flip
// a just-completed audit to failed.
await db
.update(audits)
.set({
@ -95,6 +115,7 @@ async function failAudit(auditId: string, workflowInstanceId: string) {
and(
eq(audits.id, auditId),
eq(audits.workflowInstanceId, workflowInstanceId),
eq(audits.status, "running"),
),
);
}
@ -111,25 +132,24 @@ async function getAuditForWorkflow(
});
}
async function batchWriteResults(
/**
* Persist one crawl batch (pages + link edges + per-page issues).
* Called inside the crawl-batch Workflow step so results land in D1
* incrementally instead of accumulating in memory until finalize.
*
* Idempotent on step retry: callers assign deterministic page ids
* (deterministicAuditRowId) and link/issue ids are derived from stable
* content. Page rows upsert (a retried fetch may legitimately differ last
* attempt wins, matching what the step returns); links and issues are
* insert-or-ignore.
*/
async function insertCrawledBatch(
auditId: string,
pages: StepPageResult[],
lighthouseResults: LighthouseResult[],
pages: CrawledPageResult[],
issues: DetectedIssue[],
) {
// The `finalize` workflow step can retry after a partial write (multi-chunk
// inserts aren't atomic, and steps after the insert can throw). Clear any
// rows from a prior attempt first so the re-run is idempotent — otherwise
// stable page ids collide on the PK and lighthouse rows silently duplicate.
// audit_lighthouse_results.page_id FKs audit_pages, so delete it first.
await db
.delete(auditLighthouseResults)
.where(eq(auditLighthouseResults.auditId, auditId));
await db.delete(auditPages).where(eq(auditPages.auditId, auditId));
await executeInBatches(pages, (tx, page) =>
tx.insert(auditPages).values({
id: page.id,
auditId,
await executeInBatches(pages, (tx, page) => {
const dataColumns = {
url: page.url,
statusCode: page.statusCode,
redirectUrl: page.redirectUrl,
@ -137,6 +157,8 @@ async function batchWriteResults(
metaDescription: page.metaDescription,
canonicalUrl: page.canonicalUrl,
robotsMeta: page.robotsMeta,
xRobotsTag: page.xRobotsTag,
headerCanonicalUrl: page.headerCanonicalUrl,
ogTitle: page.ogTitle,
ogDescription: page.ogDescription,
ogImage: page.ogImage,
@ -148,25 +170,83 @@ async function batchWriteResults(
h6Count: page.h6Count,
headingOrderJson: JSON.stringify(page.headingOrder),
wordCount: page.wordCount,
contentHash: page.contentHash,
imagesTotal: page.imagesTotal,
imagesMissingAlt: page.imagesMissingAlt,
imagesJson: JSON.stringify(page.images),
internalLinkCount: page.internalLinks.length,
externalLinkCount: page.externalLinks.length,
internalLinkCount: page.links.filter((l) => l.isInternal).length,
externalLinkCount: page.links.filter((l) => !l.isInternal).length,
hasStructuredData: page.hasStructuredData,
hreflangTagsJson: JSON.stringify(page.hreflangTags),
isIndexable: page.isIndexable,
fetchClass: page.fetchClass,
crawlDepth: page.crawlDepth,
inSitemap: page.inSitemap,
responseTimeMs: page.responseTimeMs,
}),
};
return tx
.insert(auditPages)
.values({ id: page.id, auditId, ...dataColumns })
.onConflictDoUpdate({ target: auditPages.id, set: dataColumns });
});
const linkRows = await Promise.all(
pages.flatMap((page) =>
page.links
.filter((link) => link.isInternal)
.slice(0, MAX_STORED_LINKS_PER_PAGE)
.map(async (link) => ({
id: await deterministicAuditRowId(auditId, page.url, link.targetUrl),
auditId,
sourcePageId: page.id,
sourceUrl: page.url,
targetUrl: link.targetUrl,
anchor: link.anchor,
isInternal: link.isInternal,
isNofollow: link.isNofollow,
})),
),
);
await executeInBatches(linkRows, (tx, row) =>
tx.insert(auditLinks).values(row).onConflictDoNothing(),
);
if (lighthouseResults.length === 0) {
return;
}
await insertIssues(auditId, issues);
}
await executeInBatches(lighthouseResults, (tx, result) =>
tx.insert(auditLighthouseResults).values({
id: crypto.randomUUID(),
async function insertIssues(auditId: string, issues: DetectedIssue[]) {
const issueRows = await Promise.all(
issues.map(async (issue) => ({
id: await deterministicAuditRowId(
auditId,
issue.pageUrl,
issue.issueType,
issue.dedupeKey ?? "",
),
auditId,
pageId: issue.pageId,
pageUrl: issue.pageUrl,
issueType: issue.issueType,
severity: AUDIT_ISSUE_TYPES[issue.issueType].severity,
detailsJson: issue.details ? JSON.stringify(issue.details) : null,
})),
);
await executeInBatches(issueRows, (tx, row) =>
tx.insert(auditIssues).values(row).onConflictDoNothing(),
);
}
async function insertLighthouseResults(
auditId: string,
lighthouseResults: LighthouseResult[],
) {
const rows = await Promise.all(
lighthouseResults.map(async (result) => ({
id: await deterministicAuditRowId(
auditId,
result.pageId,
result.strategy,
),
auditId,
pageId: result.pageId,
strategy: result.strategy,
@ -181,8 +261,17 @@ async function batchWriteResults(
errorMessage: result.errorMessage ?? null,
r2Key: result.r2Key ?? null,
payloadSizeBytes: result.payloadSizeBytes ?? null,
}),
})),
);
// Upsert: a step retry can charge a second DataForSEO call whose result
// must not be silently dropped in favor of a failed first attempt.
await executeInBatches(rows, (tx, row) => {
const { id: _id, auditId: _auditId, ...dataColumns } = row;
return tx.insert(auditLighthouseResults).values(row).onConflictDoUpdate({
target: auditLighthouseResults.id,
set: dataColumns,
});
});
}
async function getAuditForProject(auditId: string, projectId: string) {
@ -191,6 +280,58 @@ async function getAuditForProject(auditId: string, projectId: string) {
});
}
async function getLatestAuditForProject(projectId: string) {
return db.query.audits.findFirst({
where: eq(audits.projectId, projectId),
orderBy: desc(audits.startedAt),
});
}
async function getIssuesForAudit(
auditId: string,
filters: { severity?: "critical" | "warning" | "info"; issueType?: string },
) {
return db.query.auditIssues.findMany({
where: and(
eq(auditIssues.auditId, auditId),
filters.severity ? eq(auditIssues.severity, filters.severity) : undefined,
filters.issueType
? eq(auditIssues.issueType, filters.issueType)
: undefined,
),
});
}
async function getPagesForAudit(auditId: string) {
return db
.select({
id: auditPages.id,
url: auditPages.url,
statusCode: auditPages.statusCode,
fetchClass: auditPages.fetchClass,
redirectUrl: auditPages.redirectUrl,
title: auditPages.title,
metaDescription: auditPages.metaDescription,
wordCount: auditPages.wordCount,
isIndexable: auditPages.isIndexable,
crawlDepth: auditPages.crawlDepth,
inSitemap: auditPages.inSitemap,
internalLinkCount: auditPages.internalLinkCount,
responseTimeMs: auditPages.responseTimeMs,
})
.from(auditPages)
.where(eq(auditPages.auditId, auditId));
}
async function hasPagesForAudit(auditId: string): Promise<boolean> {
const rows = await db
.select({ id: auditPages.id })
.from(auditPages)
.where(eq(auditPages.auditId, auditId))
.limit(1);
return rows.length > 0;
}
async function getAuditsByProject(projectId: string) {
const rows = await db
.select({ audit: audits })
@ -223,19 +364,22 @@ async function getAuditUsageForUser(userId: string) {
async function getAuditResultsForProject(auditId: string, projectId: string) {
const audit = await getAuditForProject(auditId, projectId);
if (!audit) {
return { audit: null, pages: [], lighthouse: [] };
return { audit: null, pages: [], lighthouse: [], issues: [] };
}
const [pages, lighthouse] = await Promise.all([
const [pages, lighthouse, issues] = await Promise.all([
db.query.auditPages.findMany({
where: eq(auditPages.auditId, auditId),
}),
db.query.auditLighthouseResults.findMany({
where: eq(auditLighthouseResults.auditId, auditId),
}),
db.query.auditIssues.findMany({
where: eq(auditIssues.auditId, auditId),
}),
]);
return { audit, pages, lighthouse };
return { audit, pages, lighthouse, issues };
}
async function getLighthouseResultById(input: {
@ -285,8 +429,14 @@ export const AuditRepository = {
completeAudit,
failAudit,
getAuditForWorkflow,
batchWriteResults,
insertCrawledBatch,
insertIssues,
insertLighthouseResults,
getAuditForProject,
getLatestAuditForProject,
getIssuesForAudit,
getPagesForAudit,
hasPagesForAudit,
getAuditsByProject,
getAuditUsageForUser,
getAuditResultsForProject,

View File

@ -1,5 +1,9 @@
import { env } from "cloudflare:workers";
import type { BillingCustomerContext } from "@/server/billing/subscription";
import {
customerHasManagedAccess,
customerHasPaidPlan,
type BillingCustomerContext,
} from "@/server/billing/subscription";
import { AuditRepository } from "@/server/features/audit/repositories/AuditRepository";
import {
AUDIT_LIMITS,
@ -15,6 +19,24 @@ import {
type LighthouseStrategy,
} from "@/server/lib/audit/types";
import { normalizeAndValidateStartUrl } from "@/server/lib/audit/url-policy";
import { isHostedServerAuthMode } from "@/server/lib/runtime-env";
// Plan-tier limits are the abuse bound in hosted mode: free accounts get one
// small audit at a time, paid keeps the full limits, and customers with no
// Autumn product at all are turned away. Self-hosted isn't gated.
async function resolveAuditLimitTier(
organizationId: string,
): Promise<AuditLimitTier> {
if (!(await isHostedServerAuthMode())) return "paid";
const [hasManagedAccess, hasPaidPlan] = await Promise.all([
customerHasManagedAccess(organizationId),
customerHasPaidPlan(organizationId),
]);
if (!hasManagedAccess) {
throw new AppError("PAYMENT_REQUIRED", "Subscribe to run site audits");
}
return hasPaidPlan ? "paid" : "free";
}
async function startAudit(input: {
actorUserId: string;
@ -98,8 +120,29 @@ async function startAudit(input: {
}
async function getStatus(auditId: string, projectId: string) {
const audit = await AuditRepository.getAuditForProject(auditId, projectId);
if (!audit) throw new AppError("NOT_FOUND");
let audit = await AuditRepository.getAuditForProject(auditId, projectId);
if (!audit)
throw new AppError("NOT_FOUND", "Audit not found in this project.");
// Self-heal audits whose workflow died without reaching the mark-failed
// step (instance terminated, mark-failed itself failed, deploys, ...).
// Without this they stay "running" forever and hold capacity.
if (audit.status === "running" && audit.workflowInstanceId) {
try {
const instance = await env.SITE_AUDIT_WORKFLOW.get(
audit.workflowInstanceId,
);
const { status } = await instance.status();
if (status === "errored" || status === "terminated") {
await AuditRepository.failAudit(audit.id, audit.workflowInstanceId);
audit =
(await AuditRepository.getAuditForProject(auditId, projectId)) ??
audit;
}
} catch {
// Instance not found or status unavailable — leave the audit as-is.
}
}
return {
id: audit.id,
@ -117,7 +160,7 @@ async function getStatus(auditId: string, projectId: string) {
}
async function getResults(auditId: string, projectId: string) {
const { audit, pages, lighthouse } =
const { audit, pages, lighthouse, issues } =
await AuditRepository.getAuditResultsForProject(auditId, projectId);
if (!audit) throw new AppError("NOT_FOUND");
@ -140,6 +183,7 @@ async function getResults(auditId: string, projectId: string) {
},
pages,
lighthouse,
issues,
};
}
@ -216,6 +260,7 @@ async function remove(auditId: string, projectId: string) {
}
export const AuditService = {
resolveAuditLimitTier,
startAudit,
getStatus,
getCrawlProgress,

View File

@ -21,16 +21,6 @@ describe("audit capacity helpers", () => {
lighthouseTotal: 0,
total: 100,
});
expect(
getEstimatedAuditCapacity({
maxPages: 100,
lighthouseStrategy: "manual",
}),
).toEqual({
pagesTotal: 100,
lighthouseTotal: 0,
total: 100,
});
expect(
getEstimatedAuditCapacity({ maxPages: 100, lighthouseStrategy: "auto" }),
).toEqual({

View File

@ -22,39 +22,39 @@ export interface RobotsResult {
}
/**
* Fetch and parse robots.txt for a given origin.
* Returns a helper to check if URLs are allowed + discovered sitemap URLs.
* Fetch the raw robots.txt body (null = missing/unreachable). Kept separate
* from parsing so Workflows can checkpoint the text as durable step state and
* re-derive the parsed result deterministically on replay.
*/
export async function fetchRobotsTxt(origin: string): Promise<RobotsResult> {
const robotsUrl = `${origin}/robots.txt`;
async function fetchRobotsTxtText(origin: string): Promise<string | null> {
try {
const response = await fetch(robotsUrl, {
const response = await fetch(`${origin}/robots.txt`, {
headers: { "User-Agent": "OpenSEO-Audit/1.0" },
signal: AbortSignal.timeout(10_000),
});
if (!response.ok) {
// No robots.txt = everything allowed
return {
isAllowed: () => true,
sitemapUrls: [],
};
if (!response.ok) return null;
return await response.text();
} catch (error) {
console.warn("Failed to fetch robots.txt:", error);
return null;
}
}
/** Deterministic: same text in, same result out. Null = everything allowed. */
export function parseRobotsTxt(
origin: string,
text: string | null,
): RobotsResult {
if (text === null) {
return { isAllowed: () => true, sitemapUrls: [] };
}
const text = await response.text();
const robots = robotsParser(robotsUrl, text);
const robots = robotsParser(`${origin}/robots.txt`, text);
return {
isAllowed: (url: string) => robots.isAllowed(url) ?? true,
sitemapUrls: robots.getSitemaps(),
};
} catch (error) {
console.warn("Failed to fetch robots.txt:", error);
return {
isAllowed: () => true,
sitemapUrls: [],
};
}
}
/**
@ -184,8 +184,9 @@ async function fetchSitemapDocumentWithRetry(sitemapUrl: string): Promise<{
export async function discoverUrls(
origin: string,
maxPages = 50,
): Promise<{ urls: string[]; robots: RobotsResult; sitemapUrls: Set<string> }> {
const robots = await fetchRobotsTxt(origin);
): Promise<{ urls: string[]; robotsText: string | null }> {
const robotsText = await fetchRobotsTxtText(origin);
const robots = parseRobotsTxt(origin, robotsText);
// Collect sitemap URLs: from robots.txt + default location
const sitemapSources = new Set(robots.sitemapUrls);
@ -262,9 +263,10 @@ export async function discoverUrls(
);
}
// Cap at the crawl's page budget: these are seeds, the crawl can never use
// more — and an uncapped list can blow the ~1MiB Workflow step-state limit.
return {
urls: Array.from(allUrls),
robots,
sitemapUrls: allUrls,
urls: Array.from(allUrls).slice(0, maxPages),
robotsText,
};
}

View File

@ -0,0 +1,24 @@
/**
* Deterministic row ids for audit data.
*
* Crawl/lighthouse persistence happens inside Workflow steps, which retry on
* failure. Deriving ids from stable content (audit id + URL + ...) combined
* with `onConflictDoNothing` makes those writes idempotent across retries
* a partially-written batch is simply completed on the next attempt instead
* of duplicated under fresh random ids.
*/
export async function sha256Hex(text: string): Promise<string> {
const digest = await crypto.subtle.digest(
"SHA-256",
new TextEncoder().encode(text),
);
return Array.from(new Uint8Array(digest))
.map((byte) => byte.toString(16).padStart(2, "0"))
.join("");
}
export async function deterministicAuditRowId(
...parts: string[]
): Promise<string> {
return (await sha256Hex(parts.join("|"))).slice(0, 36);
}

View File

@ -0,0 +1,180 @@
/**
* Pure cross-page checks (no database access): duplicate grouping and
* redirect chain/loop detection. The D1-backed checks (broken links,
* orphans) live in multipage.ts.
*/
import type { DetectedIssue } from "@/server/lib/audit/issues/page-reporters";
const DUPLICATE_GROUP_SAMPLE = 3;
export interface SlimPage {
id: string;
url: string;
statusCode: number | null;
fetchClass: "ok" | "blocked" | "error";
title: string | null;
metaDescription: string | null;
contentHash: string | null;
redirectUrl: string | null;
wordCount: number;
isIndexable: boolean;
canonicalUrl: string | null;
headerCanonicalUrl: string | null;
}
function isOkHtmlPage(page: SlimPage): boolean {
return (
page.fetchClass === "ok" &&
page.statusCode !== null &&
page.statusCode >= 200 &&
page.statusCode < 300
);
}
/**
* Pages the owner already de-duplicated (noindex, or canonicalized to
* another URL) don't belong in duplicate groups flagging them tells the
* user to fix something they already fixed.
*/
function isDuplicateCandidate(page: SlimPage): boolean {
if (!isOkHtmlPage(page) || !page.isIndexable) return false;
const effectiveCanonical = page.canonicalUrl ?? page.headerCanonicalUrl;
return !effectiveCanonical || effectiveCanonical === page.url;
}
export function findDuplicates(pages: SlimPage[]): DetectedIssue[] {
const okPages = pages.filter(isDuplicateCandidate);
const groupBy = (
keyOf: (page: SlimPage) => string | null,
): Map<string, SlimPage[]> => {
const groups = new Map<string, SlimPage[]>();
for (const page of okPages) {
const key = keyOf(page);
if (!key) continue;
const group = groups.get(key);
if (group) group.push(page);
else groups.set(key, [page]);
}
return groups;
};
const issues: DetectedIssue[] = [];
const emitGroups = (
groups: Map<string, SlimPage[]>,
issueType: DetectedIssue["issueType"],
) => {
for (const group of groups.values()) {
if (group.length < 2) continue;
for (const page of group) {
issues.push({
issueType,
pageId: page.id,
pageUrl: page.url,
details: {
groupSize: group.length,
otherUrls: group
.filter((other) => other.id !== page.id)
.slice(0, DUPLICATE_GROUP_SAMPLE)
.map((other) => other.url),
},
});
}
}
};
emitGroups(
groupBy((page) => page.title || null),
"duplicate-title",
);
emitGroups(
groupBy((page) => page.metaDescription || null),
"duplicate-meta-description",
);
emitGroups(
groupBy((page) => (page.wordCount > 0 ? page.contentHash : null)),
"duplicate-content",
);
return issues;
}
export function findRedirectChainsAndLoops(pages: SlimPage[]): DetectedIssue[] {
const redirects = new Map<string, SlimPage>();
for (const page of pages) {
const isRedirect =
page.statusCode !== null &&
page.statusCode >= 300 &&
page.statusCode < 400 &&
page.redirectUrl;
if (isRedirect) redirects.set(page.url, page);
}
const redirectTargets = new Set(
Array.from(redirects.values(), (page) => page.redirectUrl!),
);
const issues: DetectedIssue[] = [];
const walked = new Set<string>();
// Walk from chain heads (redirects nothing else redirects to), so a 5-hop
// chain yields one issue, not five.
for (const [url, head] of redirects) {
if (redirectTargets.has(url)) continue;
const hops: string[] = [url];
const seen = new Set(hops);
walked.add(url);
let current = head.redirectUrl;
let isLoop = false;
while (current) {
if (seen.has(current)) {
isLoop = true;
hops.push(current);
break;
}
hops.push(current);
seen.add(current);
if (redirects.has(current)) walked.add(current);
current = redirects.get(current)?.redirectUrl ?? null;
}
if (isLoop) {
issues.push({
issueType: "redirect-loop",
pageId: head.id,
pageUrl: url,
details: { hops },
});
} else if (hops.length > 2) {
// url -> a -> b: two redirects before content = a chain
issues.push({
issueType: "redirect-chain",
pageId: head.id,
pageUrl: url,
details: { hops, finalUrl: hops[hops.length - 1] },
});
}
}
// Headless cycles (every member is also a target — e.g. a↔b, or a→a) are
// never reached from a head; emit one loop issue per cycle.
for (const [url, page] of redirects) {
if (walked.has(url)) continue;
const cycle: string[] = [];
let current: string | null = url;
while (current && !walked.has(current)) {
walked.add(current);
cycle.push(current);
current = redirects.get(current)?.redirectUrl ?? null;
}
issues.push({
issueType: "redirect-loop",
pageId: page.id,
pageUrl: url,
details: { hops: [...cycle, url] },
});
}
return issues;
}

View File

@ -0,0 +1,154 @@
/**
* Cross-page (multipage) issue checks.
*
* These run once after the crawl, against the rows the crawl wrote to D1:
* duplicates, broken internal links, redirect chains/loops, orphan pages.
* Pure set-queries over crawl data no fetching, no DOM.
*/
import { and, eq, gte, lt, ne, notExists, sql } from "drizzle-orm";
import { alias } from "drizzle-orm/sqlite-core";
import { db } from "@/db";
import { auditLinks, auditPages } from "@/db/schema";
import { normalizeUrl } from "@/server/lib/audit/url-utils";
import {
findDuplicates,
findRedirectChainsAndLoops,
type SlimPage,
} from "@/server/lib/audit/issues/multipage-checks";
import type { DetectedIssue } from "@/server/lib/audit/issues/page-reporters";
const BROKEN_LINK_ISSUE_CAP = 2_000;
export async function runMultipageChecks(input: {
auditId: string;
startUrl: string;
/** Orphan detection only makes sense when the crawl wasn't truncated. */
crawlCompleted: boolean;
}): Promise<DetectedIssue[]> {
const pages: SlimPage[] = await db
.select({
id: auditPages.id,
url: auditPages.url,
statusCode: auditPages.statusCode,
fetchClass: auditPages.fetchClass,
title: auditPages.title,
metaDescription: auditPages.metaDescription,
contentHash: auditPages.contentHash,
redirectUrl: auditPages.redirectUrl,
wordCount: auditPages.wordCount,
isIndexable: auditPages.isIndexable,
canonicalUrl: auditPages.canonicalUrl,
headerCanonicalUrl: auditPages.headerCanonicalUrl,
})
.from(auditPages)
.where(eq(auditPages.auditId, input.auditId));
const issues: DetectedIssue[] = [
...findDuplicates(pages),
...findRedirectChainsAndLoops(pages),
...(await findBrokenInternalLinks(input.auditId)),
];
if (input.crawlCompleted) {
// Page rows store normalized URLs; normalize the start URL the same way
// so the orphan exclusion matches.
const normalizedStart = normalizeUrl(input.startUrl) ?? input.startUrl;
issues.push(...(await findOrphanPages(input.auditId, normalizedStart)));
}
return issues;
}
async function findBrokenInternalLinks(
auditId: string,
): Promise<DetectedIssue[]> {
// Only flag targets we actually crawled and saw fail — never inferred from
// absence. Blocked targets (WAF challenges) are excluded: a 403 from bot
// protection is not evidence of a broken link.
const rows = await db
.select({
sourcePageId: auditLinks.sourcePageId,
sourceUrl: auditLinks.sourceUrl,
targetUrl: auditLinks.targetUrl,
targetStatus: auditPages.statusCode,
})
.from(auditLinks)
.innerJoin(
auditPages,
and(
eq(auditPages.auditId, auditLinks.auditId),
eq(auditPages.url, auditLinks.targetUrl),
),
)
.where(
and(
eq(auditLinks.auditId, auditId),
eq(auditLinks.isInternal, true),
gte(auditPages.statusCode, 400),
eq(auditPages.fetchClass, "ok"),
),
)
.limit(BROKEN_LINK_ISSUE_CAP);
return rows.map((row) => ({
issueType: "broken-internal-link" as const,
pageId: row.sourcePageId,
pageUrl: row.sourceUrl,
dedupeKey: row.targetUrl,
details: { targetUrl: row.targetUrl, targetStatus: row.targetStatus },
}));
}
async function findOrphanPages(
auditId: string,
startUrl: string,
): Promise<DetectedIssue[]> {
// A live 2xx page is an orphan when no OTHER crawled page links to it and
// nothing redirects to it. Only meaningful on a completed crawl: on a
// truncated one, "no observed inlinks" is true of nearly everything.
// Error/redirect rows aren't orphans — they already get their own issues.
const inlink = db
.select({ one: sql`1` })
.from(auditLinks)
.where(
and(
eq(auditLinks.auditId, auditId),
eq(auditLinks.isInternal, true),
eq(auditLinks.targetUrl, auditPages.url),
// Self-links (breadcrumbs, anchors) don't make a page reachable.
ne(auditLinks.sourcePageId, auditPages.id),
),
);
const redirectSourcePages = alias(auditPages, "redirect_sources");
const redirectSources = db
.select({ one: sql`1` })
.from(redirectSourcePages)
.where(
and(
eq(redirectSourcePages.auditId, auditId),
eq(redirectSourcePages.redirectUrl, auditPages.url),
),
);
const rows = await db
.select({ id: auditPages.id, url: auditPages.url })
.from(auditPages)
.where(
and(
eq(auditPages.auditId, auditId),
ne(auditPages.url, startUrl),
eq(auditPages.fetchClass, "ok"),
gte(auditPages.statusCode, 200),
lt(auditPages.statusCode, 300),
notExists(inlink),
notExists(redirectSources),
),
);
return rows.map((row) => ({
issueType: "orphan-page" as const,
pageId: row.id,
pageUrl: row.url,
}));
}

View File

@ -0,0 +1,308 @@
import { describe, expect, it } from "vitest";
import { runPageReporters } from "@/server/lib/audit/issues/page-reporters";
import {
findDuplicates,
findRedirectChainsAndLoops,
type SlimPage,
} from "@/server/lib/audit/issues/multipage-checks";
import type { CrawledPageResult } from "@/server/lib/audit/types";
function makePage(overrides: Partial<CrawledPageResult>): CrawledPageResult {
return {
id: "page-1",
url: "https://example.com/a",
statusCode: 200,
fetchClass: "ok",
redirectUrl: null,
title: "A perfectly reasonable page title",
metaDescription:
"A reasonable meta description that says something useful about the page.",
canonicalUrl: null,
robotsMeta: null,
xRobotsTag: null,
headerCanonicalUrl: null,
ogTitle: null,
ogDescription: null,
ogImage: null,
h1Count: 1,
h2Count: 0,
h3Count: 0,
h4Count: 0,
h5Count: 0,
h6Count: 0,
headingOrder: [1, 2, 3],
wordCount: 500,
contentHash: "abc123",
isHtml: true,
imagesTotal: 0,
imagesMissingAlt: 0,
images: [],
links: [],
hasStructuredData: false,
hreflangTags: [],
isIndexable: true,
responseTimeMs: 200,
crawlDepth: 1,
inSitemap: true,
...overrides,
};
}
function issueTypes(page: CrawledPageResult): string[] {
return runPageReporters(page).map((issue) => issue.issueType);
}
describe("runPageReporters", () => {
it("reports nothing for a healthy page", () => {
expect(issueTypes(makePage({}))).toEqual([]);
});
it("reports only blocked-page for a blocked fetch", () => {
expect(
issueTypes(makePage({ fetchClass: "blocked", statusCode: 403 })),
).toEqual(["blocked-page"]);
});
it("reports nothing for a fetch error", () => {
expect(
issueTypes(makePage({ fetchClass: "error", statusCode: 0 })),
).toEqual([]);
});
it("classifies error statuses by range", () => {
expect(issueTypes(makePage({ statusCode: 500 }))).toEqual(["server-error"]);
expect(issueTypes(makePage({ statusCode: 404 }))).toEqual(["broken-page"]);
expect(
issueTypes(
makePage({
statusCode: 301,
redirectUrl: "https://example.com/b",
}),
),
).toEqual([]);
});
it("checks titles and meta descriptions", () => {
expect(issueTypes(makePage({ title: "" }))).toContain("missing-title");
expect(issueTypes(makePage({ title: "x".repeat(70) }))).toContain(
"title-too-long",
);
expect(issueTypes(makePage({ title: "Tiny" }))).toContain(
"title-too-short",
);
expect(issueTypes(makePage({ metaDescription: "" }))).toContain(
"missing-meta-description",
);
expect(
issueTypes(makePage({ metaDescription: "x".repeat(200) })),
).toContain("meta-description-too-long");
});
it("checks headings", () => {
expect(issueTypes(makePage({ h1Count: 0 }))).toContain("missing-h1");
expect(issueTypes(makePage({ h1Count: 3 }))).toContain("multiple-h1");
expect(issueTypes(makePage({ headingOrder: [1, 2, 4] }))).toContain(
"heading-order-skip",
);
});
it("skips content checks for non-HTML responses", () => {
const nonHtml = makePage({
isHtml: false,
title: "",
metaDescription: "",
h1Count: 0,
headingOrder: [],
wordCount: 0,
contentHash: null,
});
expect(issueTypes(nonHtml)).toEqual([]);
});
it("still checks empty-shell HTML pages", () => {
const shell = makePage({
isHtml: true,
title: "",
metaDescription: "",
h1Count: 0,
headingOrder: [],
wordCount: 0,
contentHash: null,
});
const types = issueTypes(shell);
expect(types).toContain("missing-title");
expect(types).toContain("missing-h1");
expect(types).toContain("thin-content");
});
it("flags indexability and canonical signals", () => {
expect(
issueTypes(makePage({ isIndexable: false, robotsMeta: "noindex" })),
).toContain("noindex-page");
const conflicted = issueTypes(
makePage({
canonicalUrl: "https://example.com/canonical-a",
headerCanonicalUrl: "https://example.com/canonical-b",
}),
);
expect(conflicted).toContain("canonical-conflict");
expect(conflicted).toContain("canonicalized-page");
expect(
issueTypes(makePage({ canonicalUrl: "https://example.com/a" })),
).not.toContain("canonicalized-page");
});
it("flags thin content only on indexable pages", () => {
expect(issueTypes(makePage({ wordCount: 50 }))).toContain("thin-content");
expect(
issueTypes(
makePage({ wordCount: 50, isIndexable: false, robotsMeta: "noindex" }),
),
).not.toContain("thin-content");
});
it("flags slow responses and deep pages", () => {
expect(issueTypes(makePage({ responseTimeMs: 3000 }))).toContain(
"slow-response",
);
expect(issueTypes(makePage({ crawlDepth: 6 }))).toContain("deep-page");
expect(issueTypes(makePage({ crawlDepth: null }))).not.toContain(
"deep-page",
);
});
});
function makeSlimPage(overrides: Partial<SlimPage>): SlimPage {
return {
id: overrides.url ?? "page",
url: "https://example.com/a",
statusCode: 200,
fetchClass: "ok",
title: null,
metaDescription: null,
contentHash: null,
redirectUrl: null,
wordCount: 100,
isIndexable: true,
canonicalUrl: null,
headerCanonicalUrl: null,
...overrides,
};
}
describe("findDuplicates", () => {
it("flags duplicate titles across pages and includes the other URLs", () => {
const issues = findDuplicates([
makeSlimPage({ url: "https://example.com/a", title: "Same" }),
makeSlimPage({ url: "https://example.com/b", title: "Same" }),
makeSlimPage({ url: "https://example.com/c", title: "Different" }),
]);
const duplicateTitles = issues.filter(
(issue) => issue.issueType === "duplicate-title",
);
expect(duplicateTitles).toHaveLength(2);
expect(duplicateTitles[0].details?.otherUrls).toEqual([
"https://example.com/b",
]);
});
it("excludes noindexed and canonicalized pages from duplicate groups", () => {
const issues = findDuplicates([
makeSlimPage({ url: "https://example.com/a", title: "Same" }),
makeSlimPage({
url: "https://example.com/b",
title: "Same",
canonicalUrl: "https://example.com/a",
}),
makeSlimPage({
url: "https://example.com/c",
title: "Same",
isIndexable: false,
}),
]);
expect(issues).toHaveLength(0);
});
it("ignores non-2xx and blocked pages", () => {
const issues = findDuplicates([
makeSlimPage({ url: "https://example.com/a", title: "Same" }),
makeSlimPage({
url: "https://example.com/b",
title: "Same",
fetchClass: "blocked",
statusCode: 403,
}),
]);
expect(issues).toHaveLength(0);
});
it("groups duplicate content by hash only when there is text", () => {
const issues = findDuplicates([
makeSlimPage({ url: "https://example.com/a", contentHash: "h1" }),
makeSlimPage({ url: "https://example.com/b", contentHash: "h1" }),
makeSlimPage({
url: "https://example.com/empty-1",
contentHash: "h2",
wordCount: 0,
}),
makeSlimPage({
url: "https://example.com/empty-2",
contentHash: "h2",
wordCount: 0,
}),
]);
expect(
issues.filter((issue) => issue.issueType === "duplicate-content"),
).toHaveLength(2);
});
});
describe("findRedirectChainsAndLoops", () => {
const redirect = (url: string, target: string) =>
makeSlimPage({ url, statusCode: 301, redirectUrl: target });
it("ignores single redirects", () => {
expect(
findRedirectChainsAndLoops([
redirect("https://example.com/a", "https://example.com/b"),
makeSlimPage({ url: "https://example.com/b" }),
]),
).toHaveLength(0);
});
it("flags a chain once, on its head", () => {
const issues = findRedirectChainsAndLoops([
redirect("https://example.com/a", "https://example.com/b"),
redirect("https://example.com/b", "https://example.com/c"),
makeSlimPage({ url: "https://example.com/c" }),
]);
expect(issues).toHaveLength(1);
expect(issues[0].issueType).toBe("redirect-chain");
expect(issues[0].pageUrl).toBe("https://example.com/a");
expect(issues[0].details?.hops).toEqual([
"https://example.com/a",
"https://example.com/b",
"https://example.com/c",
]);
});
it("flags loops", () => {
const issues = findRedirectChainsAndLoops([
redirect("https://example.com/a", "https://example.com/b"),
redirect("https://example.com/b", "https://example.com/a"),
]);
expect(
issues.filter((issue) => issue.issueType === "redirect-loop").length,
).toBeGreaterThan(0);
});
it("flags self-loops", () => {
const issues = findRedirectChainsAndLoops([
redirect("https://example.com/a", "https://example.com/a"),
]);
expect(issues).toHaveLength(1);
expect(issues[0].issueType).toBe("redirect-loop");
});
});

View File

@ -0,0 +1,146 @@
/**
* Per-page issue reporters.
*
* Each reporter is a pure function over a single crawled page record
* DOM-free by design (cheerio runs once in crawlPage), so the engine works
* over any crawl source that can produce a CrawledPageResult.
*
* Cross-page checks (duplicates, broken links, orphans, redirect chains)
* live in multipage.ts and run over D1 after the crawl.
*/
import type { AuditIssueType } from "@/shared/audit-issues";
import type { CrawledPageResult } from "@/server/lib/audit/types";
export interface DetectedIssue {
issueType: AuditIssueType;
pageId: string | null;
pageUrl: string;
details?: Record<string, unknown>;
/**
* Distinguishes multiple issues of the same type on the same page
* (e.g. one broken-internal-link issue per target). Part of the
* deterministic row id, so step retries don't duplicate issues.
*/
dedupeKey?: string;
}
const TITLE_MAX_CHARS = 60;
const TITLE_MIN_CHARS = 10;
const META_DESCRIPTION_MAX_CHARS = 160;
const THIN_CONTENT_WORDS = 150;
const SLOW_RESPONSE_MS = 1500;
const DEEP_PAGE_DEPTH = 5;
function hasHeadingLevelSkip(headingOrder: number[]): boolean {
for (let i = 1; i < headingOrder.length; i++) {
if (headingOrder[i] > headingOrder[i - 1] + 1) return true;
}
return false;
}
export function runPageReporters(page: CrawledPageResult): DetectedIssue[] {
const issues: DetectedIssue[] = [];
const report = (
issueType: AuditIssueType,
details?: Record<string, unknown>,
) => issues.push({ issueType, pageId: page.id, pageUrl: page.url, details });
if (page.fetchClass === "blocked") {
report("blocked-page", { statusCode: page.statusCode });
return issues;
}
if (page.fetchClass === "error") {
return issues;
}
if (page.statusCode >= 500) {
report("server-error", { statusCode: page.statusCode });
return issues;
}
if (page.statusCode >= 400) {
report("broken-page", { statusCode: page.statusCode });
return issues;
}
// Redirects are normal on their own; chains/loops are flagged in multipage.
if (page.statusCode >= 300) {
return issues;
}
if (page.responseTimeMs > SLOW_RESPONSE_MS) {
report("slow-response", { responseTimeMs: page.responseTimeMs });
}
// Content checks only make sense for analyzed HTML documents (a PDF has no
// title tag to miss; an empty-shell HTML page very much does).
if (!page.isHtml) {
return issues;
}
// Titles
if (!page.title) {
report("missing-title");
} else if (page.title.length > TITLE_MAX_CHARS) {
report("title-too-long", { length: page.title.length });
} else if (page.title.length < TITLE_MIN_CHARS) {
report("title-too-short", { length: page.title.length });
}
// Meta description
if (!page.metaDescription) {
report("missing-meta-description");
} else if (page.metaDescription.length > META_DESCRIPTION_MAX_CHARS) {
report("meta-description-too-long", {
length: page.metaDescription.length,
});
}
// Headings
if (page.h1Count === 0) {
report("missing-h1");
} else if (page.h1Count > 1) {
report("multiple-h1", { h1Count: page.h1Count });
}
if (hasHeadingLevelSkip(page.headingOrder)) {
report("heading-order-skip");
}
// Indexability + canonical signals
if (!page.isIndexable) {
report("noindex-page", {
robotsMeta: page.robotsMeta,
xRobotsTag: page.xRobotsTag,
});
}
if (
page.canonicalUrl &&
page.headerCanonicalUrl &&
page.canonicalUrl !== page.headerCanonicalUrl
) {
report("canonical-conflict", {
htmlCanonical: page.canonicalUrl,
headerCanonical: page.headerCanonicalUrl,
});
}
const effectiveCanonical = page.canonicalUrl ?? page.headerCanonicalUrl;
if (effectiveCanonical && effectiveCanonical !== page.url) {
report("canonicalized-page", { canonicalUrl: effectiveCanonical });
}
// Content quality
if (page.isIndexable && page.wordCount < THIN_CONTENT_WORDS) {
report("thin-content", { wordCount: page.wordCount });
}
if (page.imagesMissingAlt > 0) {
report("images-missing-alt", {
imagesMissingAlt: page.imagesMissingAlt,
imagesTotal: page.imagesTotal,
});
}
// Structure
if (page.crawlDepth !== null && page.crawlDepth >= DEEP_PAGE_DEPTH) {
report("deep-page", { crawlDepth: page.crawlDepth });
}
return issues;
}

View File

@ -1,4 +1,4 @@
import { detectUrlTemplate } from "./url-utils";
import { detectUrlTemplate, normalizeUrl } from "./url-utils";
import type { BillingCustomerContext } from "@/server/billing/subscription";
import { createDataforseoClient } from "@/server/lib/dataforseo";
import type { LighthouseResult, LighthouseStrategy } from "./types";
@ -127,16 +127,13 @@ export function selectLighthouseSample(
(p) => p.statusCode >= 200 && p.statusCode < 300,
);
if (strategy === "manual") {
// manual = user picks after crawl; for now return empty
return [];
}
// strategy === "auto": homepage + 1 per URL pattern, capped at 10
const selected = new Set<string>();
// Always include the start URL / homepage
const startPage = validPages.find((p) => p.url === startUrl);
// Always include the start URL / homepage. Page URLs are normalized;
// normalize the start URL the same way or the comparison silently misses.
const normalizedStart = normalizeUrl(startUrl) ?? startUrl;
const startPage = validPages.find((p) => p.url === normalizedStart);
if (startPage) selected.add(startPage.url);
// Group by URL template pattern

View File

@ -7,7 +7,7 @@
*/
import * as cheerio from "cheerio";
import { normalizeUrl, isSameOrigin } from "./url-utils";
import type { PageAnalysis } from "./types";
import type { PageAnalysis, PageLink } from "./types";
/**
* Analyze an HTML string and extract all SEO-relevant data.
@ -71,9 +71,8 @@ export function analyzeHtml(
});
});
// --- Links ---
const internalLinks: string[] = [];
const externalLinks: string[] = [];
// --- Links (deduped by target URL; first anchor wins) ---
const linksByTarget = new Map<string, PageLink>();
$("a[href]").each((_, el) => {
const href = $(el).attr("href");
@ -84,13 +83,18 @@ export function analyzeHtml(
const resolved = normalizeUrl(href, pageUrl);
if (!resolved) return;
if (linksByTarget.has(resolved)) return;
if (isSameOrigin(resolved, pageUrl)) {
internalLinks.push(resolved);
} else {
externalLinks.push(resolved);
}
const anchor = $(el).text().replace(/\s+/g, " ").trim().slice(0, 200);
const rel = $(el).attr("rel")?.toLowerCase() ?? "";
linksByTarget.set(resolved, {
targetUrl: resolved,
anchor: anchor || null,
isInternal: isSameOrigin(resolved, pageUrl),
isNofollow: rel.split(/\s+/).includes("nofollow"),
});
});
const links = Array.from(linksByTarget.values());
// --- Structured data (JSON-LD) ---
let hasStructuredData = false;
@ -119,9 +123,9 @@ export function analyzeHtml(
h1s,
headingOrder,
wordCount,
bodyText,
images,
internalLinks,
externalLinks,
links,
hasStructuredData,
hreflangTags,
};

View File

@ -6,7 +6,7 @@ import { z } from "zod";
import { MIN_AUDIT_PAGES, PAID_MAX_AUDIT_PAGES } from "@/shared/audit-limits";
import { jsonCodec } from "@/shared/json";
export type LighthouseStrategy = "auto" | "manual" | "none";
export type LighthouseStrategy = "auto" | "none";
export interface AuditConfig {
maxPages: number;
@ -14,11 +14,21 @@ export interface AuditConfig {
}
// Read-side only (writes stringify a typed AuditConfig). Stored rows may hold
// retired strategies (e.g. "all"); fall back to "auto" instead of failing the
// whole config parse and making the audit's results unviewable.
// retired strategies ("all", "manual") from older audits; map them onto the
// closest surviving strategy — and fall back to "auto" on anything unknown —
// instead of failing the whole config parse and making the audit's results
// unviewable.
const lighthouseStrategySchema = z
.enum(["auto", "all", "manual", "none"])
.transform(
(value): LighthouseStrategy =>
value === "all" ? "auto" : value === "manual" ? "none" : value,
)
.catch("auto");
const auditConfigSchema = z.object({
maxPages: z.number().int().min(MIN_AUDIT_PAGES).max(PAID_MAX_AUDIT_PAGES),
lighthouseStrategy: z.enum(["auto", "manual", "none"]).catch("auto"),
lighthouseStrategy: lighthouseStrategySchema,
});
const auditConfigCodec = jsonCodec(auditConfigSchema);
@ -29,6 +39,17 @@ export function parseAuditConfig(configRaw: string | null): AuditConfig | null {
return result.success ? result.data : null;
}
/** How a page fetch resolved. "blocked" = WAF/bot challenge stood in the way. */
export type PageFetchClass = "ok" | "blocked" | "error";
/** One outgoing link edge, deduped by target URL within a page. */
export interface PageLink {
targetUrl: string;
anchor: string | null;
isInternal: boolean;
isNofollow: boolean;
}
/** Data extracted from a single page via cheerio. */
export interface PageAnalysis {
url: string;
@ -51,13 +72,13 @@ export interface PageAnalysis {
// Content
wordCount: number;
bodyText: string;
// Images
images: Array<{ src: string | null; alt: string | null }>;
// Links (raw href values from the HTML)
internalLinks: string[];
externalLinks: string[];
// Links (normalized, deduped by target)
links: PageLink[];
// Structured data
hasStructuredData: boolean;
@ -84,15 +105,22 @@ export interface LighthouseResult {
payloadSizeBytes?: number | null;
}
export interface StepPageResult {
/**
* Full result of crawling one page. Persisted to D1 inside the crawl-batch
* step; never accumulated in memory or returned as durable step state.
*/
export interface CrawledPageResult {
id: string;
url: string;
statusCode: number;
fetchClass: PageFetchClass;
redirectUrl: string | null;
title: string;
metaDescription: string;
canonicalUrl: string | null;
robotsMeta: string | null;
xRobotsTag: string | null;
headerCanonicalUrl: string | null;
ogTitle: string | null;
ogDescription: string | null;
ogImage: string | null;
@ -104,13 +132,37 @@ export interface StepPageResult {
h6Count: number;
headingOrder: number[];
wordCount: number;
contentHash: string | null;
/**
* True when an HTML document was fetched and analyzed. Gates the content
* checks in page reporters (an empty-shell HTML page must still be
* checked; a PDF must not). Transient not persisted.
*/
isHtml: boolean;
imagesTotal: number;
imagesMissingAlt: number;
images: Array<{ src: string | null; alt: string | null }>;
internalLinks: string[];
externalLinks: string[];
links: PageLink[];
hasStructuredData: boolean;
hreflangTags: string[];
isIndexable: boolean;
responseTimeMs: number;
/** null = not reached via links (e.g. sitemap-seeded). */
crawlDepth: number | null;
inSitemap: boolean;
}
/**
* Slim per-page summary returned as durable step state from a crawl batch.
* Keep this small: full page data lives in D1, not in Workflow step state.
*/
export interface StepPageSummary {
id: string;
url: string;
statusCode: number;
fetchClass: PageFetchClass;
redirectUrl: string | null;
title: string;
/** Normalized same-origin link targets, for frontier expansion. */
internalLinks: string[];
}

View File

@ -186,6 +186,26 @@ async function hostnameResolvesToBlockedAddress(
}
}
/**
* Synchronous SSRF check for URLs discovered mid-crawl (links, redirect
* targets, sitemap entries). Blocks non-http(s) schemes, private/loopback IP
* literals, and internal hostnames. DNS resolution is only performed for the
* start URL (see normalizeAndValidateStartUrl); per-link DoH lookups would be
* prohibitively slow.
*/
export function isCrawlableUrl(url: string): boolean {
let parsed: URL;
try {
parsed = new URL(url);
} catch {
return false;
}
if (parsed.protocol !== "http:" && parsed.protocol !== "https:") {
return false;
}
return !isBlockedHost(parsed.hostname);
}
export async function normalizeAndValidateStartUrl(
input: string,
): Promise<string> {

View File

@ -22,6 +22,12 @@ import {
getSearchConsolePerformanceTool,
inspectUrlsTool,
} from "@/server/mcp/tools/search-console-tools";
import {
getAuditIssuesTool,
getAuditPagesTool,
getAuditStatusTool,
runSiteAuditTool,
} from "@/server/mcp/tools/site-audit-tools";
import { whoamiTool } from "@/server/mcp/tools/whoami";
// Each handler is wrapped with instrumentMcpToolHandler so failures reach
@ -201,4 +207,24 @@ export function registerOpenSeoMcpTools(server: McpServer) {
inspectUrlsTool.handler,
),
);
server.registerTool(
runSiteAuditTool.name,
runSiteAuditTool.config,
runSiteAuditTool.handler,
);
server.registerTool(
getAuditStatusTool.name,
getAuditStatusTool.config,
getAuditStatusTool.handler,
);
server.registerTool(
getAuditIssuesTool.name,
getAuditIssuesTool.config,
getAuditIssuesTool.handler,
);
server.registerTool(
getAuditPagesTool.name,
getAuditPagesTool.config,
getAuditPagesTool.handler,
);
}

View File

@ -0,0 +1,399 @@
import { z } from "zod";
import { AuditRepository } from "@/server/features/audit/repositories/AuditRepository";
import { AuditService } from "@/server/features/audit/services/AuditService";
import { AppError } from "@/server/lib/errors";
import { captureServerEvent } from "@/server/lib/posthog";
import {
AUDIT_ISSUE_TYPES,
getIssueDescriptor,
ISSUE_SEVERITY_ORDER,
} from "@/shared/audit-issues";
import { mcpResponse } from "@/server/mcp/formatters";
import { buildProjectMeta } from "@/server/mcp/context";
import {
looseObjectOutputSchema,
optionalMetaOutputSchema,
} from "@/server/mcp/output-schemas";
import { withMcpProjectAuth } from "@/server/mcp/project-auth";
import { projectIdSchema } from "@/server/mcp/schemas";
const auditIdSchema = z
.string()
.optional()
.describe("Audit ID. If omitted, uses the project's most recent audit.");
async function resolveAudit(projectId: string, auditId?: string) {
const audit = auditId
? await AuditRepository.getAuditForProject(auditId, projectId)
: await AuditRepository.getLatestAuditForProject(projectId);
if (!audit) {
throw new AppError(
"NOT_FOUND",
auditId
? `Audit ${auditId} not found in this project.`
: "No audits exist for this project yet. Start one with run_site_audit.",
);
}
return audit;
}
function auditPath(projectId: string, auditId: string) {
return `/p/${projectId}/audit?auditId=${auditId}`;
}
// ─── run_site_audit ──────────────────────────────────────────────────────────
const runInputSchema = {
projectId: projectIdSchema,
url: z.string().min(1).max(2048).describe("Start URL to crawl."),
maxPages: z
.number()
.int()
.min(10)
.max(10_000)
.optional()
.describe("Page budget for the crawl (default 50)."),
runLighthouse: z
.boolean()
.optional()
.describe(
"Run Lighthouse on a sample of up to 10 representative pages (default true).",
),
} as const;
type RunArgs = z.infer<z.ZodObject<typeof runInputSchema>>;
export const runSiteAuditTool = {
name: "run_site_audit",
config: {
title: "Run site audit",
description:
"Start a site audit: crawls the site (robots.txt-aware, same-origin), checks every page for SEO issues (broken links, duplicate/missing titles and descriptions, redirect chains, orphan pages, canonical problems, thin content, and more), and optionally runs Lighthouse on a sample of pages. Runs in the background — poll get_audit_status, then read get_audit_issues. If the site blocks our crawler, pages are honestly flagged as blocked rather than misreported.",
inputSchema: runInputSchema,
outputSchema: z
.object({
auditId: z.string(),
...optionalMetaOutputSchema,
})
.passthrough(),
annotations: {
readOnlyHint: false,
openWorldHint: true,
destructiveHint: false,
},
},
handler: withMcpProjectAuth(async (args: RunArgs, context) => {
const lighthouseStrategy = (args.runLighthouse ?? true) ? "auto" : "none";
const limitTier = await AuditService.resolveAuditLimitTier(
context.auth.organizationId,
);
let auditId: string;
try {
({ auditId } = await AuditService.startAudit({
actorUserId: context.auth.userId,
billingCustomer: context.billing,
projectId: args.projectId,
startUrl: args.url,
maxPages: args.maxPages,
lighthouseStrategy,
limitTier,
}));
} catch (error) {
if (
error instanceof AppError &&
error.code === "AUDIT_CAPACITY_REACHED"
) {
return mcpResponse({
text: "Audit capacity reached for this account — delete old audits in the dashboard to free capacity, then try again.",
meta: buildProjectMeta(
context,
args.projectId,
`/p/${args.projectId}/audit`,
),
});
}
throw error;
}
await captureServerEvent({
distinctId: context.auth.userId,
event: "site_audit:start",
organizationId: context.auth.organizationId,
properties: {
project_id: args.projectId,
max_pages: args.maxPages ?? 50,
run_lighthouse: lighthouseStrategy !== "none",
source: "mcp",
},
});
return mcpResponse({
text: `Audit ${auditId} started for ${args.url}. Poll get_audit_status until it completes, then call get_audit_issues for the prioritized issue report.`,
meta: buildProjectMeta(
context,
args.projectId,
auditPath(args.projectId, auditId),
),
structuredContent: { auditId },
});
}),
};
// ─── get_audit_status ────────────────────────────────────────────────────────
const statusInputSchema = {
projectId: projectIdSchema,
auditId: auditIdSchema,
} as const;
type StatusArgs = z.infer<z.ZodObject<typeof statusInputSchema>>;
export const getAuditStatusTool = {
name: "get_audit_status",
config: {
title: "Get site audit status",
description:
"Check the progress of a site audit (phase, pages crawled, Lighthouse progress). Free — reads OpenSEO state. Omit auditId for the most recent audit.",
inputSchema: statusInputSchema,
outputSchema: z
.object({
status: looseObjectOutputSchema,
...optionalMetaOutputSchema,
})
.passthrough(),
annotations: {
readOnlyHint: true,
openWorldHint: false,
destructiveHint: false,
},
},
handler: withMcpProjectAuth(async (args: StatusArgs, context) => {
// getStatus fetches (and self-heals) the audit row itself; only hit the
// DB here when we need to default to the most recent audit.
const auditId = args.auditId ?? (await resolveAudit(args.projectId)).id;
const status = await AuditService.getStatus(auditId, args.projectId);
const lighthouseNote =
status.lighthouseTotal > 0
? `, lighthouse ${status.lighthouseCompleted + status.lighthouseFailed}/${status.lighthouseTotal}`
: "";
return mcpResponse({
text: `Audit ${status.id} (${status.startUrl}): ${status.status} — phase ${status.currentPhase}, ${status.pagesCrawled}/${status.pagesTotal} pages${lighthouseNote}.${status.status === "completed" ? " Call get_audit_issues for the issue report." : ""}`,
meta: buildProjectMeta(
context,
args.projectId,
auditPath(args.projectId, status.id),
),
structuredContent: { status },
});
}),
};
// ─── get_audit_issues ────────────────────────────────────────────────────────
const issuesInputSchema = {
projectId: projectIdSchema,
auditId: auditIdSchema,
severity: z
.enum(["critical", "warning", "info"])
.optional()
.describe("Only return issues of this severity."),
issueType: z
.string()
.optional()
.describe(
`Only return issues of this type. One of: ${Object.keys(AUDIT_ISSUE_TYPES).join(", ")}`,
),
limit: z
.number()
.int()
.min(1)
.max(1_000)
.optional()
.describe("Max issues to return (default 200)."),
} as const;
type IssuesArgs = z.infer<z.ZodObject<typeof issuesInputSchema>>;
export const getAuditIssuesTool = {
name: "get_audit_issues",
config: {
title: "Get site audit issues",
description:
"Read the prioritized issue report from a completed site audit. Every issue carries a how_to_fix with concrete remediation steps an agent can act on. Free — reads OpenSEO state. Omit auditId for the most recent audit.",
inputSchema: issuesInputSchema,
outputSchema: z
.object({
summary: z.array(looseObjectOutputSchema),
issues: z.array(looseObjectOutputSchema),
...optionalMetaOutputSchema,
})
.passthrough(),
annotations: {
readOnlyHint: true,
openWorldHint: false,
destructiveHint: false,
},
},
handler: withMcpProjectAuth(async (args: IssuesArgs, context) => {
const audit = await resolveAudit(args.projectId, args.auditId);
const unsorted = await AuditRepository.getIssuesForAudit(audit.id, {
severity: args.severity,
issueType: args.issueType,
});
// Severity-first so truncation drops info rows, never critical ones.
const rows = unsorted.toSorted(
(a, b) =>
ISSUE_SEVERITY_ORDER[a.severity] - ISSUE_SEVERITY_ORDER[b.severity] ||
a.issueType.localeCompare(b.issueType),
);
const counts = new Map<string, number>();
for (const row of rows) {
counts.set(row.issueType, (counts.get(row.issueType) ?? 0) + 1);
}
const summary = Array.from(counts.entries())
.map(([issueType, count]) => {
const descriptor = getIssueDescriptor(issueType);
return {
issueType,
title: descriptor?.title ?? issueType,
severity: descriptor?.severity ?? "info",
count,
};
})
.toSorted(
(a, b) =>
ISSUE_SEVERITY_ORDER[a.severity] - ISSUE_SEVERITY_ORDER[b.severity] ||
b.count - a.count,
);
const limit = args.limit ?? 200;
const issues = rows.slice(0, limit).map((row) => {
const descriptor = getIssueDescriptor(row.issueType);
return {
severity: row.severity,
issueType: row.issueType,
title: descriptor?.title ?? row.issueType,
url: row.pageUrl,
details: row.detailsJson
? (JSON.parse(row.detailsJson) as unknown)
: null,
howToFix: descriptor?.howToFix ?? null,
};
});
const text =
rows.length === 0
? args.severity || args.issueType
? `No issues found for audit ${audit.id} matching the given filters.`
: `No issues recorded for audit ${audit.id}. Note: audits run before issue checks existed have no issue data — re-run the audit with run_site_audit to get a real report.`
: [
`Audit ${audit.id} (${audit.startUrl}): ${rows.length} issues${rows.length > limit ? ` (showing ${limit})` : ""}.`,
"By type:",
...summary.map(
(entry) =>
`- [${entry.severity}] ${entry.title} (${entry.issueType}): ${entry.count}`,
),
"Full issue rows with how_to_fix instructions are in structuredContent.issues.",
].join("\n");
return mcpResponse({
text,
meta: buildProjectMeta(
context,
args.projectId,
auditPath(args.projectId, audit.id),
),
structuredContent: { summary, issues },
});
}),
};
// ─── get_audit_pages ─────────────────────────────────────────────────────────
const pagesInputSchema = {
projectId: projectIdSchema,
auditId: auditIdSchema,
fetchClass: z
.enum(["ok", "blocked", "error"])
.optional()
.describe(
'Filter by fetch outcome ("blocked" = the site\'s bot protection challenged the crawler).',
),
statusCode: z
.number()
.int()
.optional()
.describe("Filter by exact HTTP status code."),
urlContains: z
.string()
.optional()
.describe("Filter to URLs containing this substring."),
limit: z
.number()
.int()
.min(1)
.max(1_000)
.optional()
.describe("Max pages to return (default 100)."),
} as const;
type PagesArgs = z.infer<z.ZodObject<typeof pagesInputSchema>>;
export const getAuditPagesTool = {
name: "get_audit_pages",
config: {
title: "Get site audit pages",
description:
"List crawled pages from a site audit with per-page SEO data (status, title, description, word count, indexability, crawl depth, link counts). Free — reads OpenSEO state. Omit auditId for the most recent audit.",
inputSchema: pagesInputSchema,
outputSchema: z
.object({
pages: z.array(looseObjectOutputSchema),
total: z.number(),
...optionalMetaOutputSchema,
})
.passthrough(),
annotations: {
readOnlyHint: true,
openWorldHint: false,
destructiveHint: false,
},
},
handler: withMcpProjectAuth(async (args: PagesArgs, context) => {
const audit = await resolveAudit(args.projectId, args.auditId);
const allPages = await AuditRepository.getPagesForAudit(audit.id);
const filtered = allPages.filter(
(page) =>
(!args.fetchClass || page.fetchClass === args.fetchClass) &&
(args.statusCode === undefined ||
page.statusCode === args.statusCode) &&
(!args.urlContains || page.url.includes(args.urlContains)),
);
const limit = args.limit ?? 100;
const pages = filtered.slice(0, limit);
const text = [
`Audit ${audit.id}: ${filtered.length} pages${filtered.length > limit ? ` (showing ${limit})` : ""}.`,
...pages
.slice(0, 25)
.map(
(page) =>
`- ${page.statusCode} ${page.url}${page.fetchClass !== "ok" ? ` [${page.fetchClass}]` : ""} "${page.title ?? ""}"`,
),
"Full rows are in structuredContent.pages.",
].join("\n");
return mcpResponse({
text,
meta: buildProjectMeta(
context,
args.projectId,
auditPath(args.projectId, audit.id),
),
structuredContent: { pages, total: filtered.length },
});
}),
};

View File

@ -40,6 +40,11 @@ export class SiteAuditWorkflow extends WorkflowEntrypoint<Env, AuditParams> {
const { auditId, billingCustomer, projectId, startUrl, config } =
event.payload;
try {
// Inside a step so the D1 read is retried and replay-cached; a bare
// read here would re-execute on every replay and a transient failure
// would kill the instance before the catch below exists.
await pgStep(step, "validate-context", undefined, async () => {
const audit = await AuditRepository.getAuditForWorkflow(
auditId,
event.instanceId,
@ -52,8 +57,8 @@ export class SiteAuditWorkflow extends WorkflowEntrypoint<Env, AuditParams> {
if (audit.projectId !== projectId) {
throw new Error("Audit workflow project mismatch");
}
});
try {
await runAuditPhases(step, {
auditId,
workflowInstanceId: event.instanceId,

View File

@ -1,110 +1,258 @@
import type { StepPageResult } from "@/server/lib/audit/types";
import { isSameOrigin, normalizeUrl } from "@/server/lib/audit/url-utils";
import type {
CrawledPageResult,
PageFetchClass,
} from "@/server/lib/audit/types";
import { sha256Hex } from "@/server/lib/audit/ids";
import { normalizeUrl } from "@/server/lib/audit/url-utils";
const CRAWL_USER_AGENT = "OpenSEO-Audit/1.0";
/**
* Markers of a bot-mitigation challenge page. We classify these honestly as
* "blocked" instead of recording the challenge HTML as if it were the page.
*/
const CHALLENGE_BODY_MARKERS = [
"just a moment...",
"challenge-platform",
"cf-browser-verification",
"attention required! | cloudflare",
"verifying you are human",
];
function classifyFetch(
statusCode: number,
headers: Headers,
bodySnippet: string,
): PageFetchClass {
if (statusCode === 0) return "error";
if (headers.get("cf-mitigated")) return "blocked";
if (statusCode === 401 || statusCode === 403 || statusCode === 429) {
return "blocked";
}
if (statusCode === 503) {
const snippet = bodySnippet.toLowerCase();
if (CHALLENGE_BODY_MARKERS.some((marker) => snippet.includes(marker))) {
return "blocked";
}
}
return "ok";
}
/** Resolve a Location header against its base without normalizing. */
function resolveRawUrl(location: string, base: string): string | null {
try {
return new URL(location, base).toString();
} catch {
return null;
}
}
/** Parse `Link: <url>; rel="canonical"` response headers. */
function parseLinkHeaderCanonical(
linkHeader: string | null,
pageUrl: string,
): string | null {
if (!linkHeader) return null;
for (const part of linkHeader.split(",")) {
const match = part.match(/<([^>]+)>\s*;([^]*)/);
if (!match) continue;
if (/rel\s*=\s*"?canonical"?/i.test(match[2])) {
return normalizeUrl(match[1].trim(), pageUrl);
}
}
return null;
}
export async function crawlPage(
url: string,
crawlOrigin: string,
): Promise<StepPageResult | null> {
crawlDepth: number | null,
inSitemap: boolean,
): Promise<CrawledPageResult> {
const startTime = Date.now();
try {
const response = await fetch(url, {
// Manual redirect handling: each hop is recorded as its own page row and
// the target is enqueued by the frontier, so chains/loops are detectable.
// Exception: redirects whose target normalizes to this same URL (e.g.
// /docs -> /docs/ on slash-canonical sites — our normalizer strips the
// slash) are followed inline; recording them would create self-redirect
// rows the frontier can never resolve.
let fetchUrl = url;
let response: Response;
let hops = 0;
for (;;) {
response = await fetch(fetchUrl, {
headers: {
"User-Agent": "OpenSEO-Audit/1.0",
"User-Agent": CRAWL_USER_AGENT,
Accept: "text/html,application/xhtml+xml",
},
redirect: "follow",
redirect: "manual",
signal: AbortSignal.timeout(15_000),
});
const responseTimeMs = Date.now() - startTime;
const statusCode = response.status;
const finalUrl = normalizeUrl(response.url) ?? response.url;
if (!isSameOrigin(finalUrl, crawlOrigin)) return null;
if (response.status < 300 || response.status >= 400) break;
const redirectUrl =
response.redirected && response.url !== url ? response.url : null;
const contentType = response.headers.get("content-type") ?? "";
if (!contentType.includes("text/html")) {
return emptyPageResult(finalUrl, statusCode, redirectUrl, responseTimeMs);
const location = response.headers.get("location");
const rawTarget = location ? resolveRawUrl(location, fetchUrl) : null;
const normalizedTarget = location
? normalizeUrl(location, fetchUrl)
: null;
const isSelfAfterNormalization =
normalizedTarget === url &&
rawTarget !== null &&
rawTarget !== fetchUrl;
if (!isSelfAfterNormalization || hops >= 3) break;
fetchUrl = rawTarget;
hops += 1;
}
const html = await response.text();
const responseTimeMs = Date.now() - startTime;
const statusCode = response.status;
const xRobotsTag = response.headers.get("x-robots-tag");
const headerCanonicalUrl = parseLinkHeaderCanonical(
response.headers.get("link"),
fetchUrl,
);
if (statusCode >= 300 && statusCode < 400) {
const location = response.headers.get("location");
const redirectUrl = location ? normalizeUrl(location, fetchUrl) : null;
return emptyPageResult({
url,
statusCode,
fetchClass: "ok",
redirectUrl,
responseTimeMs,
xRobotsTag,
headerCanonicalUrl,
crawlDepth,
inSitemap,
});
}
const contentType = response.headers.get("content-type") ?? "";
const isHtml = contentType.includes("text/html");
const body = isHtml ? await response.text() : "";
const fetchClass = classifyFetch(
statusCode,
response.headers,
body.slice(0, 4_000),
);
if (!isHtml || fetchClass !== "ok" || statusCode >= 400) {
return emptyPageResult({
url,
statusCode,
fetchClass,
redirectUrl: null,
responseTimeMs,
xRobotsTag,
headerCanonicalUrl,
crawlDepth,
inSitemap,
});
}
// Resolve links/canonical against the URL that actually served the
// content (it may carry a trailing slash the recorded URL doesn't).
// Dynamic import keeps cheerio (page-analyzer's HTML parser) out of the
// worker's startup module graph: SiteAuditWorkflow is re-exported from
// src/server.ts, so a static import would evaluate cheerio in every
// isolate's baseline heap, not just when an audit actually crawls.
const { analyzeHtml } = await import("@/server/lib/audit/page-analyzer");
const analysis = analyzeHtml(
html,
finalUrl,
statusCode,
responseTimeMs,
redirectUrl,
);
const isIndexable = !(
analysis.robotsMeta?.toLowerCase().includes("noindex") ?? false
);
const h2Count = analysis.headingOrder.filter((h) => h === 2).length;
const h3Count = analysis.headingOrder.filter((h) => h === 3).length;
const h4Count = analysis.headingOrder.filter((h) => h === 4).length;
const h5Count = analysis.headingOrder.filter((h) => h === 5).length;
const h6Count = analysis.headingOrder.filter((h) => h === 6).length;
const analysis = analyzeHtml(body, fetchUrl, statusCode, responseTimeMs);
const robotsDirectives = [analysis.robotsMeta, xRobotsTag]
.filter(Boolean)
.join(",")
.toLowerCase();
const isIndexable = !robotsDirectives.includes("noindex");
const headingCount = (level: number) =>
analysis.headingOrder.filter((h) => h === level).length;
return {
id: crypto.randomUUID(),
url: finalUrl,
statusCode,
redirectUrl,
title: analysis.title,
metaDescription: analysis.metaDescription,
canonicalUrl: analysis.canonical,
robotsMeta: analysis.robotsMeta,
ogTitle: analysis.ogTitle,
ogDescription: analysis.ogDescription,
ogImage: analysis.ogImage,
h1Count: analysis.h1s.length,
h2Count,
h3Count,
h4Count,
h5Count,
h6Count,
headingOrder: analysis.headingOrder,
wordCount: analysis.wordCount,
imagesTotal: analysis.images.length,
imagesMissingAlt: analysis.images.filter(
(img) => !img.alt || img.alt === "",
).length,
images: analysis.images,
internalLinks: analysis.internalLinks,
externalLinks: analysis.externalLinks,
hasStructuredData: analysis.hasStructuredData,
hreflangTags: analysis.hreflangTags,
isIndexable,
responseTimeMs,
};
} catch (error) {
const responseTimeMs = Date.now() - startTime;
console.warn(`Failed to crawl ${url}:`, error);
return emptyPageResult(url, 0, null, responseTimeMs);
}
}
function emptyPageResult(
url: string,
statusCode: number,
redirectUrl: string | null,
responseTimeMs: number,
): StepPageResult {
return {
id: crypto.randomUUID(),
url,
statusCode,
redirectUrl,
fetchClass,
redirectUrl: null,
title: analysis.title,
metaDescription: analysis.metaDescription,
canonicalUrl: analysis.canonical
? (normalizeUrl(analysis.canonical, fetchUrl) ?? analysis.canonical)
: null,
robotsMeta: analysis.robotsMeta,
xRobotsTag,
headerCanonicalUrl,
ogTitle: analysis.ogTitle,
ogDescription: analysis.ogDescription,
ogImage: analysis.ogImage,
h1Count: analysis.h1s.length,
h2Count: headingCount(2),
h3Count: headingCount(3),
h4Count: headingCount(4),
h5Count: headingCount(5),
h6Count: headingCount(6),
headingOrder: analysis.headingOrder,
wordCount: analysis.wordCount,
contentHash: analysis.bodyText
? await sha256Hex(analysis.bodyText)
: null,
isHtml: true,
imagesTotal: analysis.images.length,
// Only a truly absent alt attribute counts: alt="" is the correct
// markup for decorative images.
imagesMissingAlt: analysis.images.filter((img) => img.alt === null)
.length,
images: analysis.images,
links: analysis.links,
hasStructuredData: analysis.hasStructuredData,
hreflangTags: analysis.hreflangTags,
isIndexable,
responseTimeMs,
crawlDepth,
inSitemap,
};
} catch (error) {
const responseTimeMs = Date.now() - startTime;
console.warn(`Failed to crawl ${url}:`, error);
return emptyPageResult({
url,
statusCode: 0,
fetchClass: "error",
redirectUrl: null,
responseTimeMs,
xRobotsTag: null,
headerCanonicalUrl: null,
crawlDepth,
inSitemap,
});
}
}
function emptyPageResult(input: {
url: string;
statusCode: number;
fetchClass: PageFetchClass;
redirectUrl: string | null;
responseTimeMs: number;
xRobotsTag: string | null;
headerCanonicalUrl: string | null;
crawlDepth: number | null;
inSitemap: boolean;
}): CrawledPageResult {
return {
id: crypto.randomUUID(),
url: input.url,
statusCode: input.statusCode,
fetchClass: input.fetchClass,
redirectUrl: input.redirectUrl,
title: "",
metaDescription: "",
canonicalUrl: null,
robotsMeta: null,
xRobotsTag: input.xRobotsTag,
headerCanonicalUrl: input.headerCanonicalUrl,
ogTitle: null,
ogDescription: null,
ogImage: null,
@ -116,14 +264,17 @@ function emptyPageResult(
h6Count: 0,
headingOrder: [],
wordCount: 0,
contentHash: null,
isHtml: false,
imagesTotal: 0,
imagesMissingAlt: 0,
images: [],
internalLinks: [],
externalLinks: [],
links: [],
hasStructuredData: false,
hreflangTags: [],
isIndexable: false,
responseTimeMs,
responseTimeMs: input.responseTimeMs,
crawlDepth: input.crawlDepth,
inSitemap: input.inSitemap,
};
}

View File

@ -1,13 +1,21 @@
import type { WorkflowStep } from "cloudflare:workers";
import type { RobotsResult } from "@/server/lib/audit/discovery";
import type { StepPageResult } from "@/server/lib/audit/types";
import type { StepPageSummary } from "@/server/lib/audit/types";
import { isSameOrigin, normalizeUrl } from "@/server/lib/audit/url-utils";
import { isCrawlableUrl } from "@/server/lib/audit/url-policy";
import { deterministicAuditRowId } from "@/server/lib/audit/ids";
import { runPageReporters } from "@/server/lib/audit/issues/page-reporters";
import { AuditRepository } from "@/server/features/audit/repositories/AuditRepository";
import { AuditProgressKV } from "@/server/lib/audit/progress-kv";
import { crawlPage } from "@/server/workflows/site-audit-workflow-helpers";
import { pgStep } from "@/server/workflows/pgStep";
const CRAWL_CONCURRENCY = 25;
// Keep durable step state under the ~1MiB limit: full link lists live in D1;
// the step return only carries new-to-the-frontier targets, deduped across
// the batch and capped.
const MAX_FRONTIER_LINKS_PER_BATCH = 2_000;
const MAX_SUMMARY_TITLE_CHARS = 300;
function shouldQueueCrawlLink(
link: string,
@ -18,12 +26,19 @@ function shouldQueueCrawlLink(
): boolean {
return (
isSameOrigin(link, origin) &&
isCrawlableUrl(link) &&
robots.isAllowed(link) &&
!visited.has(link) &&
!queued.has(link)
);
}
interface QueueEntry {
url: string;
/** Clicks from the start URL; null when only reachable via sitemap. */
depth: number | null;
}
type CrawlPhaseParams = {
auditId: string;
workflowInstanceId: string;
@ -34,10 +49,19 @@ type CrawlPhaseParams = {
sitemapUrls: string[];
};
/** What later phases need per page — no link lists (those stay in D1). */
export type CrawledPageSummary = Omit<StepPageSummary, "internalLinks">;
export type CrawlPhaseResult = {
pages: CrawledPageSummary[];
/** True when the frontier was exhausted before hitting maxPages. */
completed: boolean;
};
export async function runCrawlPhase(
step: WorkflowStep,
params: CrawlPhaseParams,
): Promise<StepPageResult[]> {
): Promise<CrawlPhaseResult> {
const {
auditId,
workflowInstanceId,
@ -48,43 +72,69 @@ export async function runCrawlPhase(
sitemapUrls,
} = params;
const visited = new Set<string>();
const queue: string[] = [];
const queued = new Set<string>();
const allPages: StepPageResult[] = [];
// Link-discovered URLs crawl first (BFS from the start URL); sitemap-only
// URLs drain last so link discovery isn't starved of page budget and
// orphan detection stays meaningful.
const linkQueue: QueueEntry[] = [];
const sitemapQueue: QueueEntry[] = [];
const sitemapSet = new Set<string>();
const summaries: CrawledPageSummary[] = [];
seedCrawlQueue({
startUrl,
origin,
robots,
sitemapUrls,
visited,
queued,
queue,
});
const normalizedStart = normalizeUrl(startUrl) ?? startUrl;
if (
robots.isAllowed(normalizedStart) &&
isSameOrigin(normalizedStart, origin)
) {
linkQueue.push({ url: normalizedStart, depth: 0 });
queued.add(normalizedStart);
}
for (const sitemapUrl of sitemapUrls) {
const normalized = normalizeUrl(sitemapUrl);
if (!normalized) continue;
sitemapSet.add(normalized);
if (!shouldQueueCrawlLink(normalized, origin, robots, visited, queued)) {
continue;
}
sitemapQueue.push({ url: normalized, depth: null });
queued.add(normalized);
}
let crawlBatchIndex = 0;
while (queue.length > 0 && allPages.length < maxPages) {
const urlsToCrawl = selectNextCrawlBatch(
queue,
while (
(linkQueue.length > 0 || sitemapQueue.length > 0) &&
summaries.length < maxPages
) {
const batchEntries = selectNextCrawlBatch({
linkQueue,
sitemapQueue,
queued,
visited,
robots,
maxPages - allPages.length,
);
if (urlsToCrawl.length === 0) continue;
remaining: maxPages - summaries.length,
});
if (batchEntries.length === 0) continue;
crawlBatchIndex += 1;
const crawledBatch = await runCrawlBatch(
step,
const crawledBatch = await runCrawlBatch(step, {
crawlBatchIndex,
urlsToCrawl,
origin,
auditId,
batchEntries,
sitemapSet,
visited,
queued,
});
// Keep only the slim summary in memory: at 10k pages, retaining link
// lists for the whole crawl would not fit in the 128MB Worker heap.
summaries.push(
...crawledBatch.map(({ internalLinks: _links, ...summary }) => summary),
);
allPages.push(...crawledBatch);
enqueueDiscoveredLinks({
crawledBatch,
queue,
batchEntries,
linkQueue,
queued,
visited,
origin,
@ -96,110 +146,151 @@ export async function runCrawlPhase(
auditId,
workflowInstanceId,
crawledBatch,
pagesCrawled: allPages.length,
pagesCrawled: summaries.length,
visitedCount: visited.size,
queueLength: queue.length,
queueLength: linkQueue.length + sitemapQueue.length,
maxPages,
});
}
return allPages;
return {
pages: summaries,
completed: linkQueue.length === 0 && sitemapQueue.length === 0,
};
}
function seedCrawlQueue({
startUrl,
origin,
robots,
sitemapUrls,
visited,
queued,
queue,
}: {
startUrl: string;
origin: string;
robots: RobotsResult;
sitemapUrls: string[];
visited: Set<string>;
function selectNextCrawlBatch(params: {
linkQueue: QueueEntry[];
sitemapQueue: QueueEntry[];
queued: Set<string>;
queue: string[];
visited: Set<string>;
robots: RobotsResult;
remaining: number;
}) {
const normalizedStart = normalizeUrl(startUrl) ?? startUrl;
if (
robots.isAllowed(normalizedStart) &&
isSameOrigin(normalizedStart, origin)
) {
queue.push(normalizedStart);
queued.add(normalizedStart);
}
for (const sitemapUrl of sitemapUrls) {
const normalized = normalizeUrl(sitemapUrl);
if (!normalized) continue;
if (!shouldQueueCrawlLink(normalized, origin, robots, visited, queued)) {
continue;
}
queue.push(normalized);
queued.add(normalized);
}
}
function selectNextCrawlBatch(
queue: string[],
queued: Set<string>,
visited: Set<string>,
robots: RobotsResult,
remaining: number,
) {
const { linkQueue, sitemapQueue, queued, visited, robots, remaining } =
params;
const batchSize = Math.min(CRAWL_CONCURRENCY, remaining);
const urlsToCrawl: string[] = [];
const batchEntries: QueueEntry[] = [];
while (queue.length > 0 && urlsToCrawl.length < batchSize) {
const url = queue.shift()!;
queued.delete(url);
if (visited.has(url)) continue;
if (!robots.isAllowed(url)) continue;
visited.add(url);
urlsToCrawl.push(url);
while (
(linkQueue.length > 0 || sitemapQueue.length > 0) &&
batchEntries.length < batchSize
) {
const entry = (linkQueue.length > 0 ? linkQueue : sitemapQueue).shift()!;
queued.delete(entry.url);
if (visited.has(entry.url)) continue;
if (!robots.isAllowed(entry.url)) continue;
visited.add(entry.url);
batchEntries.push(entry);
}
return urlsToCrawl;
return batchEntries;
}
async function runCrawlBatch(
step: WorkflowStep,
crawlBatchIndex: number,
urlsToCrawl: string[],
origin: string,
): Promise<StepPageResult[]> {
return step.do(`crawl-batch-${crawlBatchIndex}`, async () => {
const settled = await Promise.allSettled(
urlsToCrawl.map((url) => crawlPage(url, origin)),
input: {
crawlBatchIndex: number;
auditId: string;
batchEntries: QueueEntry[];
sitemapSet: Set<string>;
visited: Set<string>;
queued: Set<string>;
},
): Promise<StepPageSummary[]> {
const {
crawlBatchIndex,
auditId,
batchEntries,
sitemapSet,
visited,
queued,
} = input;
return pgStep(step, `crawl-batch-${crawlBatchIndex}`, undefined, async () => {
const pages = await Promise.all(
batchEntries.map((entry) =>
crawlPage(entry.url, entry.depth, sitemapSet.has(entry.url)),
),
);
return settled.flatMap((result) => {
if (result.status === "fulfilled" && result.value) {
return [result.value];
// Deterministic ids keep the D1 writes idempotent across step retries.
for (const page of pages) {
page.id = await deterministicAuditRowId(auditId, page.url);
}
return [];
const issues = pages.flatMap((page) => runPageReporters(page));
await AuditRepository.insertCrawledBatch(auditId, pages, issues);
// Frontier candidates only: drop targets already visited/queued and
// dedupe across the batch, so the step return stays far under the
// ~1MiB durable-state limit even on mega-menu sites.
const seenTargets = new Set<string>();
return pages.map((page) => {
const internalLinks: string[] = [];
for (const link of page.links) {
if (!link.isInternal) continue;
if (seenTargets.size >= MAX_FRONTIER_LINKS_PER_BATCH) break;
if (visited.has(link.targetUrl) || queued.has(link.targetUrl)) continue;
if (seenTargets.has(link.targetUrl)) continue;
seenTargets.add(link.targetUrl);
internalLinks.push(link.targetUrl);
}
return {
id: page.id,
url: page.url,
statusCode: page.statusCode,
fetchClass: page.fetchClass,
redirectUrl: page.redirectUrl,
title: page.title.slice(0, MAX_SUMMARY_TITLE_CHARS),
internalLinks,
};
});
});
}
function enqueueDiscoveredLinks(params: {
crawledBatch: StepPageResult[];
queue: string[];
crawledBatch: StepPageSummary[];
batchEntries: QueueEntry[];
linkQueue: QueueEntry[];
queued: Set<string>;
visited: Set<string>;
origin: string;
robots: RobotsResult;
}) {
const { crawledBatch, queue, queued, visited, origin, robots } = params;
const {
crawledBatch,
batchEntries,
linkQueue,
queued,
visited,
origin,
robots,
} = params;
const depthByUrl = new Map(
batchEntries.map((entry) => [entry.url, entry.depth]),
);
for (const pageResult of crawledBatch) {
for (const link of pageResult.internalLinks.filter((candidate) =>
shouldQueueCrawlLink(candidate, origin, robots, visited, queued),
)) {
queue.push(link);
const pageDepth = depthByUrl.get(pageResult.url) ?? null;
const childDepth = pageDepth === null ? null : pageDepth + 1;
for (const link of pageResult.internalLinks) {
if (!shouldQueueCrawlLink(link, origin, robots, visited, queued)) {
continue;
}
linkQueue.push({ url: link, depth: childDepth });
queued.add(link);
}
// Redirect targets continue the same navigation path: same depth.
const redirectTarget = pageResult.redirectUrl;
if (
redirectTarget &&
shouldQueueCrawlLink(redirectTarget, origin, robots, visited, queued)
) {
linkQueue.push({ url: redirectTarget, depth: pageDepth });
queued.add(redirectTarget);
}
}
}
@ -208,7 +299,7 @@ async function persistCrawlProgress(params: {
crawlBatchIndex: number;
auditId: string;
workflowInstanceId: string;
crawledBatch: StepPageResult[];
crawledBatch: StepPageSummary[];
pagesCrawled: number;
visitedCount: number;
queueLength: number;
@ -225,7 +316,15 @@ async function persistCrawlProgress(params: {
queueLength,
maxPages,
} = params;
await step.do(`kv-progress-batch-${crawlBatchIndex}`, async () => {
// KV push + D1 progress in one step — merging them halves the per-batch
// step count against the ~1k step budget. The D1 update is idempotent; the
// KV push can duplicate entries on a partial retry, which is acceptable for
// an ephemeral progress feed (capped list, short TTL).
await pgStep(
step,
`progress-batch-${crawlBatchIndex}`,
undefined,
async () => {
await AuditProgressKV.pushCrawledUrls(
auditId,
crawledBatch.map((pageResult) => ({
@ -235,13 +334,6 @@ async function persistCrawlProgress(params: {
crawledAt: Date.now(),
})),
);
});
await pgStep(
step,
`progress-batch-${crawlBatchIndex}`,
undefined,
async () => {
await AuditRepository.updateAuditProgress(auditId, workflowInstanceId, {
pagesCrawled,
pagesTotal: Math.min(visitedCount + queueLength, maxPages),

View File

@ -1,6 +1,6 @@
import type { WorkflowStep } from "cloudflare:workers";
import type { BillingCustomerContext } from "@/server/billing/subscription";
import { discoverUrls, fetchRobotsTxt } from "@/server/lib/audit/discovery";
import { discoverUrls, parseRobotsTxt } from "@/server/lib/audit/discovery";
import {
fetchAndStoreLighthouseResult,
selectLighthouseSample,
@ -8,13 +8,14 @@ import {
import { getOrigin } from "@/server/lib/audit/url-utils";
import { AuditRepository } from "@/server/features/audit/repositories/AuditRepository";
import { AuditProgressKV } from "@/server/lib/audit/progress-kv";
import type {
AuditConfig,
LighthouseResult,
StepPageResult,
} from "@/server/lib/audit/types";
import { runMultipageChecks } from "@/server/lib/audit/issues/multipage";
import type { AuditConfig } from "@/server/lib/audit/types";
import { captureServerEvent } from "@/server/lib/posthog";
import { runCrawlPhase } from "@/server/workflows/siteAuditWorkflowCrawl";
import {
runCrawlPhase,
type CrawledPageSummary,
type CrawlPhaseResult,
} from "@/server/workflows/siteAuditWorkflowCrawl";
import { pgStep } from "@/server/workflows/pgStep";
const LIGHTHOUSE_URL_BATCH_SIZE = 10;
@ -35,22 +36,6 @@ function capSitemapSeeds(urls: string[], maxPages: number): string[] {
return seeds;
}
function countLighthouseBatchResults(results: LighthouseResult[]): {
completed: number;
failed: number;
} {
let completed = 0;
let failed = 0;
for (const result of results) {
if (result.errorMessage) {
failed += 1;
continue;
}
completed += 1;
}
return { completed, failed };
}
type AuditPhasesParams = {
auditId: string;
workflowInstanceId: string;
@ -82,8 +67,11 @@ export async function runAuditPhases(
origin,
maxPages,
);
const robots = await fetchRobotsTxt(origin);
const allPages = await runCrawlPhase(step, {
// Parsed outside the step from checkpointed text, so replays see the exact
// robots rules the original run used (a live re-fetch could differ and
// desync the frontier from already-persisted crawl batches).
const robots = parseRobotsTxt(origin, discovery.robotsText);
const crawl = await runCrawlPhase(step, {
auditId,
workflowInstanceId,
origin,
@ -92,14 +80,14 @@ export async function runAuditPhases(
robots,
sitemapUrls: discovery.sitemapUrls,
});
const lighthouseResults = await runLighthousePhase(step, {
await runLighthousePhase(step, {
auditId,
workflowInstanceId,
billingCustomer,
projectId,
startUrl,
config,
allPages,
pages: crawl.pages,
});
await finalizeAudit({
step,
@ -107,9 +95,9 @@ export async function runAuditPhases(
workflowInstanceId,
billingCustomer,
projectId,
startUrl,
config,
allPages,
lighthouseResults,
crawl,
});
}
@ -126,7 +114,10 @@ async function runDiscoveryPhase(
pagesTotal: Math.min(result.urls.length + 1, maxPages),
currentPhase: "crawling",
});
return { sitemapUrls: capSitemapSeeds(result.urls, maxPages) };
return {
sitemapUrls: capSitemapSeeds(result.urls, maxPages),
robotsText: result.robotsText,
};
});
}
@ -137,13 +128,13 @@ type LighthousePhaseParams = {
projectId: string;
startUrl: string;
config: AuditConfig;
allPages: StepPageResult[];
pages: CrawledPageSummary[];
};
async function runLighthousePhase(
step: WorkflowStep,
params: LighthousePhaseParams,
): Promise<LighthouseResult[]> {
) {
const {
auditId,
workflowInstanceId,
@ -151,20 +142,19 @@ async function runLighthousePhase(
projectId,
startUrl,
config,
allPages,
pages,
} = params;
if (config.lighthouseStrategy === "none") return [];
if (config.lighthouseStrategy === "none") return;
const lighthouseWork = await selectLighthousePages({
step,
auditId,
workflowInstanceId,
allPages,
pages,
startUrl,
strategy: config.lighthouseStrategy,
});
const lighthouseResults: LighthouseResult[] = [];
let completedChecks = 0;
let failedChecks = 0;
let lighthouseBatchIndex = 0;
@ -172,78 +162,16 @@ async function runLighthousePhase(
for (let i = 0; i < lighthouseWork.length; i += LIGHTHOUSE_URL_BATCH_SIZE) {
const batch = lighthouseWork.slice(i, i + LIGHTHOUSE_URL_BATCH_SIZE);
lighthouseBatchIndex += 1;
const lighthouseBatchResults = await runLighthouseBatch({
step,
lighthouseBatchIndex,
batch,
billingCustomer,
projectId,
auditId,
});
const priorCompleted = completedChecks;
const priorFailed = failedChecks;
lighthouseResults.push(...lighthouseBatchResults);
const counts = countLighthouseBatchResults(lighthouseBatchResults);
failedChecks += counts.failed;
completedChecks += counts.completed;
await pgStep(
// Fetch, store (R2 + D1) and update progress inside one step. The step
// returns only counts; full results live in D1.
const counts = await pgStep(
step,
`lighthouse-progress-batch-${lighthouseBatchIndex}`,
`lighthouse-batch-${lighthouseBatchIndex}`,
undefined,
async () => {
await AuditRepository.updateAuditProgress(auditId, workflowInstanceId, {
lighthouseCompleted: completedChecks,
lighthouseFailed: failedChecks,
});
},
);
}
return lighthouseResults;
}
async function selectLighthousePages(params: {
step: WorkflowStep;
auditId: string;
workflowInstanceId: string;
allPages: StepPageResult[];
startUrl: string;
strategy: AuditConfig["lighthouseStrategy"];
}) {
const { step, auditId, workflowInstanceId, allPages, startUrl, strategy } =
params;
return pgStep(step, "select-lighthouse-sample", undefined, async () => {
const sample = selectLighthouseSample(allPages, startUrl, strategy);
const selectedUrls = new Set(sample);
await AuditRepository.updateAuditProgress(auditId, workflowInstanceId, {
currentPhase: "lighthouse",
lighthouseTotal: sample.length * 2,
lighthouseCompleted: 0,
lighthouseFailed: 0,
});
return allPages.flatMap((page) =>
selectedUrls.has(page.url) ? [{ url: page.url, pageId: page.id }] : [],
);
});
}
async function runLighthouseBatch(params: {
step: WorkflowStep;
lighthouseBatchIndex: number;
batch: Array<{ url: string; pageId: string }>;
billingCustomer: BillingCustomerContext;
projectId: string;
auditId: string;
}) {
const {
step,
lighthouseBatchIndex,
batch,
billingCustomer,
projectId,
auditId,
} = params;
return step.do(`lighthouse-batch-${lighthouseBatchIndex}`, async () => {
const perUrlResults = await Promise.all(
batch.map(async ({ url, pageId }) => {
const [mobileResult, desktopResult] = await Promise.all([
@ -267,8 +195,47 @@ async function runLighthouseBatch(params: {
return [mobileResult, desktopResult];
}),
);
const results = perUrlResults.flat();
await AuditRepository.insertLighthouseResults(auditId, results);
return perUrlResults.flat();
const failed = results.filter((result) => result.errorMessage).length;
const completed = results.length - failed;
await AuditRepository.updateAuditProgress(auditId, workflowInstanceId, {
lighthouseCompleted: priorCompleted + completed,
lighthouseFailed: priorFailed + failed,
});
return { completed, failed };
},
);
completedChecks += counts.completed;
failedChecks += counts.failed;
}
}
async function selectLighthousePages(params: {
step: WorkflowStep;
auditId: string;
workflowInstanceId: string;
pages: CrawledPageSummary[];
startUrl: string;
strategy: AuditConfig["lighthouseStrategy"];
}) {
const { step, auditId, workflowInstanceId, pages, startUrl, strategy } =
params;
return pgStep(step, "select-lighthouse-sample", undefined, async () => {
const sample = selectLighthouseSample(pages, startUrl, strategy);
const selectedUrls = new Set(sample);
await AuditRepository.updateAuditProgress(auditId, workflowInstanceId, {
currentPhase: "lighthouse",
lighthouseTotal: sample.length * 2,
lighthouseCompleted: 0,
lighthouseFailed: 0,
});
return pages.flatMap((page) =>
selectedUrls.has(page.url) ? [{ url: page.url, pageId: page.id }] : [],
);
});
}
@ -278,9 +245,9 @@ async function finalizeAudit(args: {
workflowInstanceId: string;
billingCustomer: BillingCustomerContext;
projectId: string;
startUrl: string;
config: AuditConfig;
allPages: StepPageResult[];
lighthouseResults: LighthouseResult[];
crawl: CrawlPhaseResult;
}) {
const {
step,
@ -288,23 +255,42 @@ async function finalizeAudit(args: {
workflowInstanceId,
billingCustomer,
projectId,
startUrl,
config,
allPages,
lighthouseResults,
crawl,
} = args;
await pgStep(step, "finalize", undefined, async () => {
await pgStep(step, "multipage-checks", undefined, async () => {
await AuditRepository.updateAuditProgress(auditId, workflowInstanceId, {
currentPhase: "finalizing",
});
await AuditRepository.batchWriteResults(
auditId,
allPages,
lighthouseResults,
// Integrity guard: pages are persisted inside crawl-batch steps. If the
// crawl claims pages but D1 has none (e.g. an instance started under the
// pre-incremental-persistence code was replayed under this code), fail
// loudly instead of completing with an empty audit.
if (
crawl.pages.length > 0 &&
!(await AuditRepository.hasPagesForAudit(auditId))
) {
throw new Error(
`Audit ${auditId}: crawl reported ${crawl.pages.length} pages but none were persisted`,
);
}
const issues = await runMultipageChecks({
auditId,
startUrl,
crawlCompleted: crawl.completed,
});
await AuditRepository.insertIssues(auditId, issues);
return { issueCount: issues.length };
});
await pgStep(step, "finalize", undefined, async () => {
await AuditRepository.completeAudit(auditId, workflowInstanceId, {
pagesCrawled: allPages.length,
pagesTotal: allPages.length,
pagesCrawled: crawl.pages.length,
pagesTotal: crawl.pages.length,
});
await captureServerEvent({
distinctId: billingCustomer.userId,
@ -313,8 +299,12 @@ async function finalizeAudit(args: {
properties: {
project_id: projectId,
status: "completed",
pages_crawled: allPages.length,
pages_total: allPages.length,
pages_crawled: crawl.pages.length,
pages_total: crawl.pages.length,
crawl_completed: crawl.completed,
pages_blocked: crawl.pages.filter(
(page) => page.fetchClass === "blocked",
).length,
run_lighthouse: config.lighthouseStrategy !== "none",
},
});

View File

@ -1,14 +1,7 @@
import { createServerFn } from "@tanstack/react-start";
import { waitUntil } from "cloudflare:workers";
import { AuditService } from "@/server/features/audit/services/AuditService";
import type { AuditLimitTier } from "@/server/features/audit/services/audit-capacity";
import {
customerHasManagedAccess,
customerHasPaidPlan,
} from "@/server/billing/subscription";
import { AppError } from "@/server/lib/errors";
import { captureServerEvent } from "@/server/lib/posthog";
import { isHostedServerAuthMode } from "@/server/lib/runtime-env";
import { requireProjectContext } from "@/serverFunctions/middleware";
import {
deleteAuditSchema,
@ -23,21 +16,9 @@ export const startAudit = createServerFn({ method: "POST" })
.middleware(requireProjectContext)
.validator(startAuditSchema)
.handler(async ({ data, context }) => {
// The crawler runs on our Workers compute and isn't credit-metered, so
// plan-tier limits are the abuse bound in hosted mode: free accounts get
// one small audit at a time, paid keeps the full limits, and customers
// with no Autumn product at all are turned away. Self-hosted isn't gated.
let limitTier: AuditLimitTier = "paid";
if (await isHostedServerAuthMode()) {
const [hasManagedAccess, hasPaidPlan] = await Promise.all([
customerHasManagedAccess(context.organizationId),
customerHasPaidPlan(context.organizationId),
]);
if (!hasManagedAccess) {
throw new AppError("PAYMENT_REQUIRED", "Subscribe to run site audits");
}
limitTier = hasPaidPlan ? "paid" : "free";
}
const limitTier = await AuditService.resolveAuditLimitTier(
context.organizationId,
);
const result = await AuditService.startAudit({
actorUserId: context.userId,

235
src/shared/audit-issues.ts Normal file
View File

@ -0,0 +1,235 @@
/**
* Registry of site-audit issue types.
*
* Shared between the server (issue engine, MCP tools) and the client
* (issues UI, CSV export). Each issue row in `audit_issues` references one
* of these types by id.
*/
type IssueSeverity = "critical" | "warning" | "info";
interface AuditIssueDescriptor {
severity: IssueSeverity;
title: string;
explanation: string;
howToFix: string;
}
export const AUDIT_ISSUE_TYPES = {
"blocked-page": {
severity: "critical",
title: "Crawler was blocked",
explanation:
"The site returned a bot challenge or access denial (e.g. a Cloudflare challenge, 403, or 429) instead of the page. We report this honestly rather than pretending the page is broken — but it means this page could not be audited, and other crawlers like search engines may face similar friction.",
howToFix:
'If you own this site, allowlist the "OpenSEO-Audit" user agent in your WAF/bot-protection settings (on Cloudflare: a WAF custom rule that skips bot protection when the user agent contains "OpenSEO-Audit"; on some free tiers you may need to relax bot protection). Then re-run the audit.',
},
"server-error": {
severity: "critical",
title: "Server error (5xx)",
explanation:
"The page returned a 5xx server error. Search engines that repeatedly see server errors will crawl the site less and may drop the page from the index.",
howToFix:
"Check the server logs for this URL and fix the underlying error. If the page is gone, return a 404/410 or redirect it to a relevant page instead of erroring.",
},
"broken-internal-link": {
severity: "critical",
title: "Broken internal link",
explanation:
"This page links to an internal URL that returns an error status (4xx/5xx). Broken links waste crawl budget, leak link equity, and frustrate users — they are among the most common and most damaging technical SEO issues.",
howToFix:
"Update the link to point at the correct live URL, or remove it. If the target was moved, prefer linking directly to the new URL rather than relying on a redirect.",
},
"missing-title": {
severity: "critical",
title: "Missing title tag",
explanation:
"The page has no <title>. The title is the strongest on-page relevance signal and the headline shown in search results; without it search engines generate one themselves, usually badly.",
howToFix:
"Add a unique, descriptive <title> of roughly 5060 characters that includes the page's primary topic.",
},
"broken-page": {
severity: "warning",
title: "Page returns an error (4xx)",
explanation:
"This crawled URL returned a client error (e.g. 404). If it is referenced from your sitemap or other pages, crawlers keep wasting requests on it.",
howToFix:
"If the page should exist, restore it. If it is intentionally gone, remove it from the sitemap and internal links, and consider a 301 redirect to the closest live page.",
},
"duplicate-title": {
severity: "warning",
title: "Duplicate title",
explanation:
"Multiple pages share the same title tag. Search engines use titles to differentiate pages; duplicates make pages compete with each other and depress click-through rates.",
howToFix:
"Write a unique title for each page describing its specific content. For templated pages, include the distinguishing attribute (name, category, location) in the template.",
},
"duplicate-meta-description": {
severity: "warning",
title: "Duplicate meta description",
explanation:
"Multiple pages share the same meta description, so search results show identical snippets and users cannot tell the pages apart.",
howToFix:
"Write a unique meta description per page, or remove the duplicated one entirely — search engines will generate a snippet from page content, which beats a wrong duplicate.",
},
"duplicate-content": {
severity: "warning",
title: "Duplicate page content",
explanation:
"Two or more URLs serve byte-identical visible text. Search engines pick one version to index and ignore the rest, and ranking signals get split across the duplicates.",
howToFix:
"Consolidate duplicates: pick the canonical URL, add rel=canonical from the others, and 301-redirect duplicate URLs where possible (common causes: trailing-slash variants, URL parameters, http/https or www variants).",
},
"missing-meta-description": {
severity: "warning",
title: "Missing meta description",
explanation:
"The page has no meta description. Search engines will assemble a snippet from page text, which is often less compelling and hurts click-through rate.",
howToFix:
"Add a meta description of roughly 70160 characters that summarizes the page and gives a reason to click.",
},
"missing-h1": {
severity: "warning",
title: "Missing H1 heading",
explanation:
"The page has no H1. The H1 tells users and search engines what the page is about; pages without one tend to have weaker topical clarity.",
howToFix:
"Add a single H1 that states the page's main topic, consistent with the title tag.",
},
"multiple-h1": {
severity: "warning",
title: "Multiple H1 headings",
explanation:
"The page has more than one H1, which dilutes the main-topic signal and usually indicates a templating mistake (e.g. a logo and a headline both marked up as H1).",
howToFix:
"Keep one H1 for the page's main heading and demote the others to H2/H3 (or unstyled elements for non-headings like logos).",
},
"redirect-chain": {
severity: "warning",
title: "Redirect chain",
explanation:
"Reaching the final page requires two or more consecutive redirects. Each hop adds latency, leaks link equity, and burns crawl budget; long chains may not be followed at all.",
howToFix:
"Point the first URL (and any internal links) directly at the final destination so there is at most one redirect.",
},
"redirect-loop": {
severity: "warning",
title: "Redirect loop",
explanation:
"This redirect eventually points back to itself, so the URL never resolves. Browsers and crawlers give up with an error.",
howToFix:
"Trace the redirect rules for this URL and break the cycle so the chain terminates at a real 200 page.",
},
"canonical-conflict": {
severity: "warning",
title: "Conflicting canonical signals",
explanation:
"The page declares different canonical URLs in its HTML <link rel=canonical> and its HTTP Link header. When signals conflict, search engines ignore both and choose their own canonical.",
howToFix:
"Pick one canonical URL and declare it in exactly one place (HTML head is the most common); remove or align the other declaration.",
},
"thin-content": {
severity: "warning",
title: "Thin content",
explanation:
"The page has very little visible text. Thin pages rarely rank, can drag down sitewide quality assessments, and (if the site renders client-side) may indicate content invisible to plain-HTML crawlers.",
howToFix:
"Either expand the page with genuinely useful content, noindex it, or consolidate it into a stronger page. If the content exists but is rendered by JavaScript, ensure it is server-rendered or pre-rendered.",
},
"images-missing-alt": {
severity: "warning",
title: "Images missing alt text",
explanation:
"One or more images on the page lack alt attributes. Alt text is an accessibility requirement and the main way search engines understand images.",
howToFix:
'Add descriptive alt text to meaningful images; use an empty alt (alt="") only for purely decorative ones.',
},
"orphan-page": {
severity: "warning",
title: "Orphan page",
explanation:
"No crawled page links to this URL — it was only discoverable via the sitemap. Pages without internal links receive little crawl attention and no internal link equity, and users can't find them by browsing.",
howToFix:
"Link to this page from relevant pages (navigation, related content, hub pages), or remove it from the sitemap if it shouldn't be indexed.",
},
"title-too-long": {
severity: "info",
title: "Title too long",
explanation:
"The title exceeds ~60 characters, so search results will truncate it and the ending may be cut off mid-phrase.",
howToFix:
"Shorten the title to roughly 5060 characters, front-loading the most important words.",
},
"title-too-short": {
severity: "info",
title: "Title too short",
explanation:
"The title is under ~10 characters, which is usually too generic to describe the page or attract clicks.",
howToFix:
"Expand the title into a descriptive phrase (roughly 3060 characters) that states what the page offers.",
},
"meta-description-too-long": {
severity: "info",
title: "Meta description too long",
explanation:
"The meta description exceeds ~160 characters, so search engines will truncate the snippet.",
howToFix:
"Trim the description to roughly 70160 characters while keeping the core message and call to action.",
},
"heading-order-skip": {
severity: "info",
title: "Heading levels skip",
explanation:
"The heading hierarchy skips levels (e.g. an H4 directly after an H2). This weakens document structure for accessibility tools and content parsing.",
howToFix:
"Adjust heading levels so they descend one step at a time (H1 → H2 → H3) without skipping.",
},
"slow-response": {
severity: "info",
title: "Slow server response",
explanation:
"The HTML response took over 1.5 seconds. Slow time-to-first-byte drags down every downstream performance metric and reduces crawl rate on large sites.",
howToFix:
"Investigate server/database time and caching for this route; serving cached or statically generated HTML usually fixes it.",
},
"noindex-page": {
severity: "info",
title: "Page is noindex",
explanation:
"The page asks search engines not to index it (via robots meta tag or X-Robots-Tag header). That's often intentional — this is a heads-up, not an error.",
howToFix:
"If this page should rank, remove the noindex directive. If it's intentional (admin, thank-you, filter pages), no action is needed.",
},
"canonicalized-page": {
severity: "info",
title: "Canonicalized to another URL",
explanation:
"The page declares a different URL as its canonical, telling search engines to index that URL instead. Fine when intentional (parameter pages, syndication) — a problem if this page was meant to rank.",
howToFix:
"If this page should rank on its own, set its canonical to itself. Otherwise no action is needed.",
},
"deep-page": {
severity: "info",
title: "Page is deep in the site structure",
explanation:
"The page is 5+ clicks from the homepage. Deep pages get crawled less often and receive less link equity.",
howToFix:
"Add links from higher-level pages (hubs, category pages, navigation) to flatten the path to this page.",
},
} as const satisfies Record<string, AuditIssueDescriptor>;
export type AuditIssueType = keyof typeof AUDIT_ISSUE_TYPES;
export const ISSUE_SEVERITY_ORDER: Record<IssueSeverity, number> = {
critical: 0,
warning: 1,
info: 2,
};
const issueRegistry: Record<string, AuditIssueDescriptor> = AUDIT_ISSUE_TYPES;
export function getIssueDescriptor(
issueType: string,
): AuditIssueDescriptor | null {
return issueRegistry[issueType] ?? null;
}

View File

@ -17,10 +17,7 @@ export const startAuditSchema = z.object({
.max(PAID_MAX_AUDIT_PAGES)
.optional()
.default(DEFAULT_AUDIT_PAGES),
lighthouseStrategy: z
.enum(["auto", "manual", "none"])
.optional()
.default("auto"),
lighthouseStrategy: z.enum(["auto", "none"]).optional().default("auto"),
});
export const getAuditStatusSchema = z.object({