* Site audit P0: issue engine, incremental persistence, block detection
Implements the P0 feature set from docs/site-audit-pm-research.md:
- Issue engine: 24 issue types (shared registry with severity,
explanation, how-to-fix). Per-page reporters run inside crawl steps;
cross-page checks (duplicate titles/descriptions/content, broken
internal links, redirect chains/loops, orphan pages) run at finalize
as SQL over the persisted crawl.
- New audit_links + audit_issues tables, audit_pages columns (depth,
content hash, header signals, fetch class, sitemap flag); audit
tables moved to src/db/audit.schema.ts.
- Incremental persistence: pages/links/issues written to D1 inside
each crawl-batch step with deterministic row ids + upserts (retry
idempotent); slim step state; robots.txt checkpointed as step state
for deterministic replay; merged progress steps keep a 10k-page
crawl within the Workflows step budget.
- Crawler: manual redirect handling with inline follow of
normalization-equivalent redirects (slash-canonical sites), response
header capture (X-Robots-Tag, Link rel=canonical), BFS depth,
sitemap-last seeding, SSRF check on discovered links, honest
"we were blocked" classification (403/429/cf-mitigated/challenge).
- UI: Issues tab (default) with severity grouping, per-type
explanations, drill-down, CSV/JSON/Sheets export, blocked banner.
- MCP: run_site_audit, get_audit_status, get_audit_issues (severity-
sorted, how_to_fix per issue), get_audit_pages.
- Lighthouse strategies reduced to auto/none (legacy all/manual map on
read); auto stays 10 URLs x 2 = 20 checks.
- Self-healing: getStatus reconciles audits whose workflow instance
errored/terminated without reaching mark-failed.
Deploy notes: run db:migrate:prod (additive migration 0022); terminate
running audits before deploying - the workflow step structure changed
and in-flight instances cannot replay under the new code (a finalize
guard fails them loudly instead of completing empty).
* feat(onboarding): hide agent chat step; subscribe after intro steps (#312)
* feat(onboarding): hide agent chat step; subscribe after intro steps
Remove the hosted-only strategy-chat diversion from the onboarding
sequence. After the three intro questions, hosted users now hit the
subscribe paywall directly, then return to the GSC and MCP connect
steps. The chat route and components stay in place but unlinked, to be
revisited later. Preserve the post-payment 'You're in!' interstitial by
carrying checkout=success through validateSearch.
* fix(onboarding): set checkout=success from subscribe route, not speculatively
The previous redirect baked checkout=success into the onboarding return
URL at the point needsSubscription is true — i.e. before the user had
paid. It only worked because the subscribe route gates its redirect on
actual access. Move the marker to the subscribe route's redirect-to-app
path, where checkoutCompleted reflects a real returned-from-Stripe
payment, so the 'You're in!' screen can never show pre-payment.
* website: change link
* fix(rank-tracking): unarchive config when re-adding an archived domain (#313)
* Unify dual-backend DB layer (D1 default + Postgres opt-in) (#238)
* D1 → Postgres data migration (ETL + runbook) (#274)
* Fix Postgres-only rank-tracking & site-audit workflow failures (#317)
* rank-tracking: raise per-project config limit from 20 to 100 (#318)
The cap was only a soft guard against runaway scheduled DataForSEO
workload, not a hard product constraint. Bump it to 100 so projects
tracking many domain/location combos aren't blocked.
Co-authored-by: Claude <noreply@anthropic.com>
* fix(db): add missing indexes and drop redundant ones (#319)
Postgres advisor flagged seq-scans and redundant indexes across both
backends (D1 + Postgres):
- add projects(organization_id) — org-scoped project listings seq-scanned
- add account(account_id, provider_id) — better-auth sign-in lookup
- add verification(expires_at) — expired-token cleanup range scan
- drop saved_keyword_tag_assignments_keyword_idx — covered by unique
(saved_keyword_id, tag_id) prefix
- drop rank_snapshots_run_idx — covered by unique
(run_id, tracking_keyword_id, device) prefix
Mirrored in both schema dialects + parity-test required-index guard.
* refactor(keywords): unify keyword-metric fetching behind one helper (#320)
* Fix production errors: onboarding crash hardening + DataForSEO spend/noise cleanup (#282)
* fix(ai-search): use valid Claude model_name and fail fast on unknown ones (#323)
DataForSEO dropped the Claude Sonnet 4.0 family from its llm_responses
catalog, so model_name=claude-sonnet-4-0 was rejected with 'Invalid
Field: model_name' while still billing the failed task. Point Claude at
claude-sonnet-4-5 and validate every model_name against DataForSEO's
accepted catalog before dispatching the paid call.
* fix(mcp): 405 the standalone GET SSE stream to stop /mcp OOM (#325)
The stateless MCP server returns JSON on POST (enableJsonResponse) and
pushes no server-initiated messages, so the optional standalone GET SSE
stream serves no purpose. Left enabled, each GET holds an SSE stream open
indefinitely (25s keepalive, no eventStore) and pins a fresh per-request
McpServer (~5MB of tools + Zod schemas); a few dozen concurrent connected
clients exceed the 128MB isolate limit. This was 100% of the /mcp
exceededMemory OOMs (GET only; POST never OOMed).
Return 405 (spec-compliant 'no standalone stream') before building the
server, so GET allocates nothing. Also removes the bulk of the elevated
GET canceled / responseStreamDisconnected outcomes.
* Re-add free plan as the floor; remove subscribe gate (#321)
* Pin production to Postgres via committed Hyperdrive binding (#329)
* Add Cloudflare Turnstile captcha on email signup (#326)
* Triage production log errors: audit crash, Autumn webhook FK, PostHog capture, auth rate-limit IP, log noise (#327)
* Add badseo.dev: a test site of deliberate SEO mistakes
An open-source Cloudflare Worker that serves ~27 pages, each breaking one
common technical-SEO rule (missing title, redirect loop, orphan page, thin
content, and so on). It doubles as the end-to-end fixture for the OpenSEO
site audit: every page declares the audit issues it should trigger, and
scripts/run-audit.ts drives the real audit engine against a running copy to
check that it does (36/36 checks, 25/25 issue types).
Styled to match the OpenSEO marketing site (web/). Maintained-by-OpenSEO
badge links back to openseo.so.
* badseo.dev: logo in pill, footer/hover polish, SEO-optimized titles
- Use the OpenSEO pine-tree logo (downscaled, base64-embedded, served at
/openseo-logo.png) in a light chip inside the badge, replacing the ◎ glyph.
- Footer band now fills to the bottom of the page (dropped the mismatched
body padding strip) with room for the floating badge.
- Index rows: remove the stray full-row underline and the stark white hover
box; hover is now a soft cream tint with the name underlined.
- Drop the "Maintained by OpenSEO" hero eyebrow; new H1 "A website
demonstrating common technical SEO problems" and a cleaner subtitle.
- Optimize homepage + catalog <title>/meta around real keywords from OpenSEO
keyword research (technical seo issues KD25/vol170; technical seo checklist
KD16/vol390), keeping meta lengths within limits.
* Site audit P0 (1/3): issue engine, incremental persistence, block detection
Server-side foundation of the P0 feature set from docs/site-audit-pm-research.md:
- Issue engine: shared registry of issue types (severity, explanation,
how-to-fix). Per-page reporters run inside crawl steps; cross-page checks
(duplicate titles/descriptions/content, broken internal links, redirect
chains/loops, orphan pages) run at finalize as SQL over the persisted crawl.
- New audit_links + audit_issues tables, audit_pages columns (depth, content
hash, header signals, fetch class, sitemap flag); audit tables moved to
src/db/{,pg/}audit.schema.ts; migrations 0029 (D1) / 0006 (PG).
- Incremental persistence: pages/links/issues written inside each crawl-batch
step with deterministic row ids + upserts (retry idempotent); slim step
state; robots.txt checkpointed as step state; merged progress steps keep a
10k-page crawl within the Workflows step budget.
- Crawler: manual redirect handling with inline follow of normalization-
equivalent redirects, response header capture (X-Robots-Tag, Link
rel=canonical), BFS depth, sitemap-last seeding, SSRF check on discovered
links, honest 'we were blocked' classification (403/429/cf-mitigated/
challenge).
- MCP: run_site_audit, get_audit_status, get_audit_issues, get_audit_pages;
limitTier resolved via shared AuditService.resolveAuditLimitTier.
- Lighthouse strategies reduced to auto/none (legacy all/manual map on read).
- Self-healing: getStatus reconciles audits whose workflow instance errored/
terminated without reaching mark-failed.
The Issues UI and the badseo.dev e2e fixture site stack on top of this PR.
Deploy notes: run db:migrate:prod (additive); terminate running audits before
deploying — the workflow step structure changed and in-flight instances cannot
replay under the new code (a finalize guard fails them loudly instead of
completing empty).
* Site audit P0 (2/3): Issues tab UI
- Issues tab (new default) with severity grouping, per-type explanations and
how-to-fix, drill-down to affected pages, CSV/JSON/Sheets export, and the
'we were blocked' banner when the crawl was challenged.
- Tabs always render (Issues/Pages, Performance when Lighthouse ran);
audit route search schema gains the issues tab and defaults to it.
Stacks on claude/audit-p0-server (issue engine + persistence).
* badseo.dev: render the badge logo as a white tree, no chip
The silver source logo was invisible on the dark pill, so it sat in a white
chip. Render it white via a CSS filter instead, so the tree fills the pill
with no backing background.
* badseo.dev: add build (typecheck) step before deploy
- Add 'build'/'typecheck' scripts (tsc --noEmit); 'deploy' now runs the build
before wrangler deploy.
- Scope the tsconfig typecheck to the Worker source (src/); the e2e harness in
scripts/ imports the main app and is run with tsx from the repo root.
- Document the deploy flow and first-time custom-domain setup in the README.
* badseo.dev: add trailing-slash redirect-cycle fixture + regression test
Reproduces the 508 "Loop Detected" class of bug from every-app/open-seo#61: a
CMS-style page whose canonical URL ends in a trailing slash, with the non-slash
form 301-redirecting to it. A crawler that strips trailing slashes turns the
canonical /foo/ back into /foo, follows the 301 to /foo/, strips it again, and
loops.
- New fixture at /redirect/trailing-slash: the non-slash form (intercepted in
index.ts on the raw path) 301s to the slash form, which is served as the
canonical 200.
- Harness asserts the page is crawled exactly once as a 200 with NO redirect
loop, plus a dedicated "Trailing-slash cycle -> 200, no loop" guard.
Verified the guard bites: temporarily disabling crawlPage's slash-canonical
inline-follow makes both checks fail (redirect-loop, status 301); with it in
place the harness is 38/38, 25/25 issue types.
* Add webapp-testing skill (installed via /reload-skills)
Vendors the anthropics/skills webapp-testing toolkit: real files under
.agents/skills/webapp-testing, a symlink from .claude/skills/, and skills-lock.json
pinning the source + hash. Matches how the other project skills are tracked.
* Site audit: redesign issues tab as grouped table + calmer page header
- Issues: single bordered table with severity sections (Critical/Warning/Info
headers carry the counts), dot indicators instead of filled pills, plain
right-aligned page counts, all rows collapsed by default; expanded rows get
a severity-colored left rule
- Removed the dead severity-count chips (they looked like filters but were
inert spans)
- Header: audited hostname is now the H1 with the status badge inline
- Blocked banner: compact tinted panel instead of a full-size alert
- Stats: hairline strip instead of four separate cards; issues stat shows a
severity breakdown, Lighthouse tile hidden when no tests ran, dropped the
orange issues-count coloring
* audit: fix trailing-slash redirect cycle at the root (preserve slashes)
Replaces the crawlPage inline-follow workaround with the root-cause fix, so we
don't carry two fixes for the same bug (every-app/open-seo#61).
- normalizeUrl: stop stripping trailing slashes. A trailing slash is the
canonical form on most CMSes, which 301 the non-slash version to it. Stripping
rewrote the canonical URL into its own redirect source and looped (508). Now
/path and /path/ are distinct and the redirect resolves normally.
- crawlPage: remove the isSelfAfterNormalization inline-follow (+ now-unused
resolveRawUrl). With slashes preserved it's dead code; a trailing-slash
redirect is recorded as an ordinary hop.
- add canonicalUrlKey (www/http/https-tolerant) and use it for the Lighthouse
homepage match, which had the same redirect-mismatch vulnerability.
- tests: preserve-trailing-slash + canonicalUrlKey unit tests; badseo harness
guard is now fix-agnostic (canonical resolves to 200, no loop/error).
Verified: 36 audit unit tests pass, tsc clean, badseo e2e 38/38. Reintroducing
stripping makes the trailing-slash guard fail (redirect-loop), confirming the
regression guard bites.
* Audit: add no-outgoing-links + meta-description-too-short checks, catch empty H1s
Two checks Ahrefs covers that we didn't, plus a fix: <h1></h1> now counts
as missing. badseo.dev gains fixtures for all three (41 checks, 27/27
issue types covered).
* Audit pages table: honest redirect/non-HTML rows, wrapped titles
- 3xx rows show their redirect target (dim →) instead of a red 'missing'
title, and dash out H1/Words/Images since nothing was analyzed
- red 'missing' only when the engine actually flagged missing-title, so
200 non-HTML files (security.txt) read as blank, not broken
- URL cells include the host when it differs from the audited site's, so
apex→www redirect sources no longer render identically to their target
- titles wrap to two lines (line-clamp) in a wider column instead of
truncating at 220px; PagesTable moved to its own file (lint max-lines)
* Audit pages table: canonical-host display, URL default sort, full title wrap
- host prefix now compares against the site's predominant 2xx host, not
the typed start URL — auditing apex 12port.com no longer prefixes every
www row with the host
- default sort by URL so the table opens as a site inventory instead of
leading with redirects on error-free sites
- titles wrap fully instead of clamping at two lines; long titles are the
thing being audited, so their tails shouldn't be hidden
* ci: exclude vendored skills from prettier; format test file
---------
Co-authored-by: Claude <noreply@anthropic.com>
* feat(rank-tracking): allow explicit language selection
* fix(rank-tracking): remove duplicate country picker and fix LocationSelect handler
The modal rendered two country pickers: a stale native <select> with a
broken empty onChange and the new LocationSelect combobox. The LocationSelect
handler also treated its numeric argument as a DOM event (Number(e.target.value)
-> NaN). Remove the dead <select> and pass the location code through directly.
* feat(rank-tracking): offer the full DataForSEO SERP language list
The language picker only listed the ~43 languages that happened to be a
country's default, so users could not track e.g. Hindi, Chinese (Simplified),
Tamil, or Urdu. Expand LANGUAGE_OPTIONS to the full set of languages the
DataForSEO SERP (Google) API accepts (sourced live from
/v3/serp/google/languages), dropping the deprecated 'iw' Hebrew duplicate and
the redundant 'no' (Norway uses 'nb', which both SERP and Labs accept). Mismatched
location+language pairs are accepted by SERP and invalid codes fail at zero cost,
so the wider list adds no charged-but-failed risk.
Also drop the now-unused LOCATION_OPTIONS re-export from the client shim.
* docs(rank-tracking): cite DataForSEO source for the language list
Note where LANGUAGE_OPTIONS comes from (/v3/serp/google/languages) and how it
relates to the country list (the Labs locations_and_languages endpoint), plus
the intentional deviations from the raw endpoint.
* chore(knip): treat scripts/** as entry points
The standalone CLI/dev scripts under scripts/ are invoked via package.json
scripts (tsx/node), not imported, so knip flagged all six as unused files.
Add scripts/** to knip entry so ci:check passes. (Pre-existing on main;
bundled here to keep the branch's ci:check green.)
* feat(rank-tracking): filter language picker to the country's supported languages
Showing all 128 languages for every country was noise. Restrict the picker to
the languages DataForSEO actually supports for the selected country via a new
getLanguageOptions() helper, backed by a per-country map from the Labs
locations_and_languages endpoint. Most countries expose only their default
language; the ~20 genuinely multilingual ones (US en/es, Canada en/fr,
Switzerland de/fr/it, India en/hi, etc.) list their real set. googleAdsOnly
countries have no per-country language data, so they show their default only.
The Language select is disabled when a country offers a single language.
LANGUAGE_OPTIONS becomes the internal master list (no longer exported).
---------
Co-authored-by: Ben Senescu <44480372+bensenescu@users.noreply.github.com>
Co-authored-by: Ben Senescu <bensenescu@gmail.com>
* feat: add personal access tokens
* feat: replace MCP tokens with OAuth foundation
* fix: keep OAuth constants private in auth foundation
* fix: clean up mcp oauth branch scope
* fix: expose oauth metadata endpoints
* fix: trim mcp oauth config to non-default options
Drop OIDC scopes, the org-id JWT claim, and the openid-configuration
metadata endpoint since the MCP integration is OAuth-only and the org
gets resolved server-side. Also remove options that just duplicated
better-auth defaults.
* fix: drop redundant oauth metadata helpers
Remove `session.storeSessionInDatabase: true` since better-auth only
enforces it when secondaryStorage is configured. Inline the
`getHostedBaseUrlForOAuthMetadata` alias and skip the async
`getOAuthServerConfig()` call in the protected-resource metadata handler
— the issuer is just `baseURL` without a custom jwt.issuer override.
* docs: explain cache headers on mcp metadata response
* Use escaped file routes for OAuth metadata
* save
* refactor: rename delegated auth user table
* feat: scaffold hosted better auth setup
* feat: add hosted auth flows
* refactor: scope project access to organizations
* fix: harden hosted auth entry points
* fix: stabilize org backfills and auth state
* refactor: simplify hosted organization setup
* fix: restore hosted auth signup flow
* fix: preserve hosted workspace access
* fix: preserve hosted auth redirects
* Improve hosted auth UX: auto-redirect to sign-up, hide header on auth pages, add form placeholders, and trust portless dev origins
- Auto-redirect unauthenticated users to /sign-up in hosted mode
- Hide top nav on /sign-in and /sign-up for a cleaner auth experience
- Add input placeholders across sign-in and sign-up forms
- Make name field optional on sign-up (falls back to email username)
- Update copy: remove 'hosted' from user-facing text, rename link to 'Create account'
- Trust *.open-seo.localhost:1355 in dev mode to fix Better Auth origin rejection with portless worktrees
* Simplify hosted auth flow and remove standalone PSI
Use TanStack Form for sign-in and sign-up, make hosted unauthenticated handling redirect-focused, and inline auth route errors. Remove the leftover standalone PSI route, services, and table so PSI only exists within site audits.
* Align project auth with Better Auth organizations
* Make server function auth middleware global
* Reduce auth server function boilerplate
* delete migrations
* fix regenerated migration data backfills
* Simplify hosted auth flow and project audit scoping
* Use active project context for audit actions
* Allow hosted session project updates
* Let agent dev server inherit auth mode
* Match hosted header to gateway account menu
* Scope project session updates to active project
* Inline authenticated server function setup
* Polish header project and account controls
* restore auth generate script
* Use explicit project access in server functions
Make project-scoped server functions take projectId input and enforce ownership through shared middleware instead of session-backed current project state. Document the tradeoffs in an ADR so future changes can follow the same boundary.
* fix ci dependency detection for auth tooling
* Harden project auth in server middleware
Authorize projectId automatically in authenticated server middleware and add a requireProject guard for project-scoped handlers. This makes the auth boundary harder to bypass and removes ad hoc non-null assertions from server functions.
* Inline project id input schemas
Remove tiny shared projectId schema helpers where they were adding indirection without reducing real complexity. Keep project-scoped validation explicit at each server function boundary.
* Skip hosted backlinks access checks
* Simplify auth mode helpers
* Avoid rerunning auth server middleware
* Simplify server function scoping ADR
* Fix backlinks project scoping in hosted auth
* Refine auth route foundations
* Simplify ensure user auth resolution
Split auth-mode context resolvers into focused modules so the middleware reads as request orchestration instead of implementation details. Reuse a shared ensured-user context type across server middleware.
* Simplify hosted organization bootstrap
Use Better Auth to own hosted organization creation and membership so hosted auth only needs to resolve a default active organization. Keep delegated-mode compatibility records isolated in a separate helper.
* Clarify hosted auth and backlinks behavior
Document the hosted AUTH_MODE deploy contract and explain why hosted deployments skip manual backlinks verification. This makes the platform-managed behavior explicit in the code paths that differ from self-serve mode.
* Document hosted org creation callback
Explain why auth.ts injects createOrganization into the hosted org helper. This makes the dependency direction explicit and avoids future import cycles while keeping the helper reusable.
* Fix CI check failures
* Fix nav link prop forwarding
* save
* Enforce safe TypeScript assertions and validate runtime payloads
* Fix CI knip config and floating promise lint
* Validate DataForSEO payloads with Zod schemas
Replace weak object guards with endpoint-level schema parsing so invalid API shapes fail fast instead of being silently filtered. Align downstream keyword mapping with the stricter validated payload contracts.
* remove Every App SDK and add auth modes for Cloudflare Access and local_noauth
* align local dev auth defaults and normalize Access team domain
* Apply suggestions from code review
* restore local drizzle D1 URL helper
* save
* fix auth error mapping and document self-hosting setup
* Tweak readme
* improve auth config error UI and remove manifest link
* fix team domain config validation and docs anchor