461 Commits

Author SHA1 Message Date
A Sivasubramanian Manoj
65f7baf8a1
Add Rank Tracking metric filters and fix null-last sort ordering (#75) 2026-07-13 13:25:14 -04:00
bookingseo
172126aec7
fix: register cache writes with waitUntil so workerd persists them (#73)
void setCached() leaves the R2 put unregistered, and workerd cancels
unregistered pending I/O once the response is sent — the write never
lands, so the caches behind domain overview, domain keyword/page pages,
and SERP research are silently re-fetched (and re-charged) on every
call. Register the writes with waitUntil, matching the pattern already
used in instrumentation.ts and brandLookup.ts.
2026-07-13 12:55:52 -04:00
Ben Senescu
47d638632a
release: v0.0.26 (#384) 2026-07-10 18:36:21 -04:00
Ben Senescu
d2287c2e47
Show app version in settings for self-hosted mode (#383) 2026-07-10 18:33:54 -04:00
Ben Senescu
8349bca43c
Migrate badseo.dev to TanStack Start (#382)
* Migrate badseo.dev to TanStack Start

* Fix badseo audit command
2026-07-10 17:53:42 -04:00
Ben Senescu
e264ee3472 Refine audit and rank tracking indexes (#378) 2026-07-10 15:45:02 -04:00
Ben Senescu
d7b52f0876
code factory: add papercuts skill, greptile rules + greptile rules skill (#77)
* chore: add versioned Greptile review policy

* chore: preserve review learnings and papercuts

* docs: align Claude and agent guidance

* docs: add shared engineering principles

* docs: trim engineering principles
2026-07-10 15:43:48 -04:00
Ben Senescu
3f2b4872ca
release: v0.0.25 (#377)
* release: v0.0.25

* chore: add release publishing command
2026-07-10 11:49:19 -04:00
Ben Senescu
fec642a039
Add Search Console property links to search performance page (#374)
* Add Search Console property links to performance page

* Simplify GSC property settings link
2026-07-09 22:27:09 -04:00
Ben Senescu
ea95eae368
Sanitize PostHog consent URLs and replay metadata (#376) 2026-07-09 22:23:23 -04:00
Ben Senescu
c11517908d
Handle billing KV failures and preserve Lighthouse start-page matching (#375) 2026-07-09 22:23:03 -04:00
Ben Senescu
dae0067233
feat(gsc): support multiple Google accounts per user (#373) 2026-07-09 21:49:08 -04:00
Ben Senescu
f2c41b1db1 MCP activation + usage PostHog events, North Star dashboard prompt (#369)
* Add MCP activation + usage PostHog events and North Star dashboard prompt

- mcp:authorize_success (server) on OAuth consent completion
- mcp:tool_call (server) on every MCP tool invocation, with tool/success/
  error_code/client_id and mcp_client vs in_app_agent source
- mcp:consent_viewed / mcp:consent_denied on the OAuth consent page
- mcp:setup_url_copy / mcp:setup_command_copy intent events on /ai
- docs/posthog-north-star-dashboard.md: prompt to configure the PostHog
  core-actions dashboard (activation rate, MCP setup funnel, daily MCP users)

* Fix codex review findings: count schema-rejected MCP calls as failures, wrap audit tools

- mcp:tool_call now fires after output validation; schema mismatches report
  success:false with MCP_OUTPUT_VALIDATION (the SDK surfaces them to the
  client as JSON-RPC errors)
- run_site_audit/get_audit_status/get_audit_issues/get_audit_pages were
  registered without instrumentMcpToolHandler, so their usage and failures
  were invisible

* Remove dashboard prompt doc
2026-07-07 22:08:14 -04:00
Ben Senescu
67265a0046 Site audit P0: Issues tab UI, badseo.dev e2e harness, new checks + pages-table polish (#367)
* Site audit P0: issue engine, incremental persistence, block detection

Implements the P0 feature set from docs/site-audit-pm-research.md:

- Issue engine: 24 issue types (shared registry with severity,
  explanation, how-to-fix). Per-page reporters run inside crawl steps;
  cross-page checks (duplicate titles/descriptions/content, broken
  internal links, redirect chains/loops, orphan pages) run at finalize
  as SQL over the persisted crawl.
- New audit_links + audit_issues tables, audit_pages columns (depth,
  content hash, header signals, fetch class, sitemap flag); audit
  tables moved to src/db/audit.schema.ts.
- Incremental persistence: pages/links/issues written to D1 inside
  each crawl-batch step with deterministic row ids + upserts (retry
  idempotent); slim step state; robots.txt checkpointed as step state
  for deterministic replay; merged progress steps keep a 10k-page
  crawl within the Workflows step budget.
- Crawler: manual redirect handling with inline follow of
  normalization-equivalent redirects (slash-canonical sites), response
  header capture (X-Robots-Tag, Link rel=canonical), BFS depth,
  sitemap-last seeding, SSRF check on discovered links, honest
  "we were blocked" classification (403/429/cf-mitigated/challenge).
- UI: Issues tab (default) with severity grouping, per-type
  explanations, drill-down, CSV/JSON/Sheets export, blocked banner.
- MCP: run_site_audit, get_audit_status, get_audit_issues (severity-
  sorted, how_to_fix per issue), get_audit_pages.
- Lighthouse strategies reduced to auto/none (legacy all/manual map on
  read); auto stays 10 URLs x 2 = 20 checks.
- Self-healing: getStatus reconciles audits whose workflow instance
  errored/terminated without reaching mark-failed.

Deploy notes: run db:migrate:prod (additive migration 0022); terminate
running audits before deploying - the workflow step structure changed
and in-flight instances cannot replay under the new code (a finalize
guard fails them loudly instead of completing empty).

* feat(onboarding): hide agent chat step; subscribe after intro steps (#312)

* feat(onboarding): hide agent chat step; subscribe after intro steps

Remove the hosted-only strategy-chat diversion from the onboarding
sequence. After the three intro questions, hosted users now hit the
subscribe paywall directly, then return to the GSC and MCP connect
steps. The chat route and components stay in place but unlinked, to be
revisited later. Preserve the post-payment 'You're in!' interstitial by
carrying checkout=success through validateSearch.

* fix(onboarding): set checkout=success from subscribe route, not speculatively

The previous redirect baked checkout=success into the onboarding return
URL at the point needsSubscription is true — i.e. before the user had
paid. It only worked because the subscribe route gates its redirect on
actual access. Move the marker to the subscribe route's redirect-to-app
path, where checkoutCompleted reflects a real returned-from-Stripe
payment, so the 'You're in!' screen can never show pre-payment.

* website: change link

* fix(rank-tracking): unarchive config when re-adding an archived domain (#313)

* Unify dual-backend DB layer (D1 default + Postgres opt-in) (#238)

* D1 → Postgres data migration (ETL + runbook) (#274)

* Fix Postgres-only rank-tracking & site-audit workflow failures (#317)

* rank-tracking: raise per-project config limit from 20 to 100 (#318)

The cap was only a soft guard against runaway scheduled DataForSEO
workload, not a hard product constraint. Bump it to 100 so projects
tracking many domain/location combos aren't blocked.

Co-authored-by: Claude <noreply@anthropic.com>

* fix(db): add missing indexes and drop redundant ones (#319)

Postgres advisor flagged seq-scans and redundant indexes across both
backends (D1 + Postgres):

- add projects(organization_id) — org-scoped project listings seq-scanned
- add account(account_id, provider_id) — better-auth sign-in lookup
- add verification(expires_at) — expired-token cleanup range scan
- drop saved_keyword_tag_assignments_keyword_idx — covered by unique
  (saved_keyword_id, tag_id) prefix
- drop rank_snapshots_run_idx — covered by unique
  (run_id, tracking_keyword_id, device) prefix

Mirrored in both schema dialects + parity-test required-index guard.

* refactor(keywords): unify keyword-metric fetching behind one helper (#320)

* Fix production errors: onboarding crash hardening + DataForSEO spend/noise cleanup (#282)

* fix(ai-search): use valid Claude model_name and fail fast on unknown ones (#323)

DataForSEO dropped the Claude Sonnet 4.0 family from its llm_responses
catalog, so model_name=claude-sonnet-4-0 was rejected with 'Invalid
Field: model_name' while still billing the failed task. Point Claude at
claude-sonnet-4-5 and validate every model_name against DataForSEO's
accepted catalog before dispatching the paid call.

* fix(mcp): 405 the standalone GET SSE stream to stop /mcp OOM (#325)

The stateless MCP server returns JSON on POST (enableJsonResponse) and
pushes no server-initiated messages, so the optional standalone GET SSE
stream serves no purpose. Left enabled, each GET holds an SSE stream open
indefinitely (25s keepalive, no eventStore) and pins a fresh per-request
McpServer (~5MB of tools + Zod schemas); a few dozen concurrent connected
clients exceed the 128MB isolate limit. This was 100% of the /mcp
exceededMemory OOMs (GET only; POST never OOMed).

Return 405 (spec-compliant 'no standalone stream') before building the
server, so GET allocates nothing. Also removes the bulk of the elevated
GET canceled / responseStreamDisconnected outcomes.

* Re-add free plan as the floor; remove subscribe gate (#321)

* Pin production to Postgres via committed Hyperdrive binding (#329)

* Add Cloudflare Turnstile captcha on email signup (#326)

* Triage production log errors: audit crash, Autumn webhook FK, PostHog capture, auth rate-limit IP, log noise (#327)

* Add badseo.dev: a test site of deliberate SEO mistakes

An open-source Cloudflare Worker that serves ~27 pages, each breaking one
common technical-SEO rule (missing title, redirect loop, orphan page, thin
content, and so on). It doubles as the end-to-end fixture for the OpenSEO
site audit: every page declares the audit issues it should trigger, and
scripts/run-audit.ts drives the real audit engine against a running copy to
check that it does (36/36 checks, 25/25 issue types).

Styled to match the OpenSEO marketing site (web/). Maintained-by-OpenSEO
badge links back to openseo.so.

* badseo.dev: logo in pill, footer/hover polish, SEO-optimized titles

- Use the OpenSEO pine-tree logo (downscaled, base64-embedded, served at
  /openseo-logo.png) in a light chip inside the badge, replacing the ◎ glyph.
- Footer band now fills to the bottom of the page (dropped the mismatched
  body padding strip) with room for the floating badge.
- Index rows: remove the stray full-row underline and the stark white hover
  box; hover is now a soft cream tint with the name underlined.
- Drop the "Maintained by OpenSEO" hero eyebrow; new H1 "A website
  demonstrating common technical SEO problems" and a cleaner subtitle.
- Optimize homepage + catalog <title>/meta around real keywords from OpenSEO
  keyword research (technical seo issues KD25/vol170; technical seo checklist
  KD16/vol390), keeping meta lengths within limits.

* Site audit P0 (1/3): issue engine, incremental persistence, block detection

Server-side foundation of the P0 feature set from docs/site-audit-pm-research.md:

- Issue engine: shared registry of issue types (severity, explanation,
  how-to-fix). Per-page reporters run inside crawl steps; cross-page checks
  (duplicate titles/descriptions/content, broken internal links, redirect
  chains/loops, orphan pages) run at finalize as SQL over the persisted crawl.
- New audit_links + audit_issues tables, audit_pages columns (depth, content
  hash, header signals, fetch class, sitemap flag); audit tables moved to
  src/db/{,pg/}audit.schema.ts; migrations 0029 (D1) / 0006 (PG).
- Incremental persistence: pages/links/issues written inside each crawl-batch
  step with deterministic row ids + upserts (retry idempotent); slim step
  state; robots.txt checkpointed as step state; merged progress steps keep a
  10k-page crawl within the Workflows step budget.
- Crawler: manual redirect handling with inline follow of normalization-
  equivalent redirects, response header capture (X-Robots-Tag, Link
  rel=canonical), BFS depth, sitemap-last seeding, SSRF check on discovered
  links, honest 'we were blocked' classification (403/429/cf-mitigated/
  challenge).
- MCP: run_site_audit, get_audit_status, get_audit_issues, get_audit_pages;
  limitTier resolved via shared AuditService.resolveAuditLimitTier.
- Lighthouse strategies reduced to auto/none (legacy all/manual map on read).
- Self-healing: getStatus reconciles audits whose workflow instance errored/
  terminated without reaching mark-failed.

The Issues UI and the badseo.dev e2e fixture site stack on top of this PR.

Deploy notes: run db:migrate:prod (additive); terminate running audits before
deploying — the workflow step structure changed and in-flight instances cannot
replay under the new code (a finalize guard fails them loudly instead of
completing empty).

* Site audit P0 (2/3): Issues tab UI

- Issues tab (new default) with severity grouping, per-type explanations and
  how-to-fix, drill-down to affected pages, CSV/JSON/Sheets export, and the
  'we were blocked' banner when the crawl was challenged.
- Tabs always render (Issues/Pages, Performance when Lighthouse ran);
  audit route search schema gains the issues tab and defaults to it.

Stacks on claude/audit-p0-server (issue engine + persistence).

* badseo.dev: render the badge logo as a white tree, no chip

The silver source logo was invisible on the dark pill, so it sat in a white
chip. Render it white via a CSS filter instead, so the tree fills the pill
with no backing background.

* badseo.dev: add build (typecheck) step before deploy

- Add 'build'/'typecheck' scripts (tsc --noEmit); 'deploy' now runs the build
  before wrangler deploy.
- Scope the tsconfig typecheck to the Worker source (src/); the e2e harness in
  scripts/ imports the main app and is run with tsx from the repo root.
- Document the deploy flow and first-time custom-domain setup in the README.

* badseo.dev: add trailing-slash redirect-cycle fixture + regression test

Reproduces the 508 "Loop Detected" class of bug from every-app/open-seo#61: a
CMS-style page whose canonical URL ends in a trailing slash, with the non-slash
form 301-redirecting to it. A crawler that strips trailing slashes turns the
canonical /foo/ back into /foo, follows the 301 to /foo/, strips it again, and
loops.

- New fixture at /redirect/trailing-slash: the non-slash form (intercepted in
  index.ts on the raw path) 301s to the slash form, which is served as the
  canonical 200.
- Harness asserts the page is crawled exactly once as a 200 with NO redirect
  loop, plus a dedicated "Trailing-slash cycle -> 200, no loop" guard.

Verified the guard bites: temporarily disabling crawlPage's slash-canonical
inline-follow makes both checks fail (redirect-loop, status 301); with it in
place the harness is 38/38, 25/25 issue types.

* Add webapp-testing skill (installed via /reload-skills)

Vendors the anthropics/skills webapp-testing toolkit: real files under
.agents/skills/webapp-testing, a symlink from .claude/skills/, and skills-lock.json
pinning the source + hash. Matches how the other project skills are tracked.

* Site audit: redesign issues tab as grouped table + calmer page header

- Issues: single bordered table with severity sections (Critical/Warning/Info
  headers carry the counts), dot indicators instead of filled pills, plain
  right-aligned page counts, all rows collapsed by default; expanded rows get
  a severity-colored left rule
- Removed the dead severity-count chips (they looked like filters but were
  inert spans)
- Header: audited hostname is now the H1 with the status badge inline
- Blocked banner: compact tinted panel instead of a full-size alert
- Stats: hairline strip instead of four separate cards; issues stat shows a
  severity breakdown, Lighthouse tile hidden when no tests ran, dropped the
  orange issues-count coloring

* audit: fix trailing-slash redirect cycle at the root (preserve slashes)

Replaces the crawlPage inline-follow workaround with the root-cause fix, so we
don't carry two fixes for the same bug (every-app/open-seo#61).

- normalizeUrl: stop stripping trailing slashes. A trailing slash is the
  canonical form on most CMSes, which 301 the non-slash version to it. Stripping
  rewrote the canonical URL into its own redirect source and looped (508). Now
  /path and /path/ are distinct and the redirect resolves normally.
- crawlPage: remove the isSelfAfterNormalization inline-follow (+ now-unused
  resolveRawUrl). With slashes preserved it's dead code; a trailing-slash
  redirect is recorded as an ordinary hop.
- add canonicalUrlKey (www/http/https-tolerant) and use it for the Lighthouse
  homepage match, which had the same redirect-mismatch vulnerability.
- tests: preserve-trailing-slash + canonicalUrlKey unit tests; badseo harness
  guard is now fix-agnostic (canonical resolves to 200, no loop/error).

Verified: 36 audit unit tests pass, tsc clean, badseo e2e 38/38. Reintroducing
stripping makes the trailing-slash guard fail (redirect-loop), confirming the
regression guard bites.

* Audit: add no-outgoing-links + meta-description-too-short checks, catch empty H1s

Two checks Ahrefs covers that we didn't, plus a fix: <h1></h1> now counts
as missing. badseo.dev gains fixtures for all three (41 checks, 27/27
issue types covered).

* Audit pages table: honest redirect/non-HTML rows, wrapped titles

- 3xx rows show their redirect target (dim →) instead of a red 'missing'
  title, and dash out H1/Words/Images since nothing was analyzed
- red 'missing' only when the engine actually flagged missing-title, so
  200 non-HTML files (security.txt) read as blank, not broken
- URL cells include the host when it differs from the audited site's, so
  apex→www redirect sources no longer render identically to their target
- titles wrap to two lines (line-clamp) in a wider column instead of
  truncating at 220px; PagesTable moved to its own file (lint max-lines)

* Audit pages table: canonical-host display, URL default sort, full title wrap

- host prefix now compares against the site's predominant 2xx host, not
  the typed start URL — auditing apex 12port.com no longer prefixes every
  www row with the host
- default sort by URL so the table opens as a site inventory instead of
  leading with redirects on error-free sites
- titles wrap fully instead of clamping at two lines; long titles are the
  thing being audited, so their tails shouldn't be hidden

* ci: exclude vendored skills from prettier; format test file

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-07 22:08:14 -04:00
Ben Senescu
1c74fded7b Site audit P0 (1/3): issue engine, incremental persistence, block detection (#362)
* Site audit P0 (1/3): issue engine, incremental persistence, block detection

Server-side foundation of the P0 feature set from docs/site-audit-pm-research.md:

- Issue engine: shared registry of issue types (severity, explanation,
  how-to-fix). Per-page reporters run inside crawl steps; cross-page checks
  (duplicate titles/descriptions/content, broken internal links, redirect
  chains/loops, orphan pages) run at finalize as SQL over the persisted crawl.
- New audit_links + audit_issues tables, audit_pages columns (depth, content
  hash, header signals, fetch class, sitemap flag); audit tables moved to
  src/db/{,pg/}audit.schema.ts; migrations 0029 (D1) / 0006 (PG).
- Incremental persistence: pages/links/issues written inside each crawl-batch
  step with deterministic row ids + upserts (retry idempotent); slim step
  state; robots.txt checkpointed as step state; merged progress steps keep a
  10k-page crawl within the Workflows step budget.
- Crawler: manual redirect handling with inline follow of normalization-
  equivalent redirects, response header capture (X-Robots-Tag, Link
  rel=canonical), BFS depth, sitemap-last seeding, SSRF check on discovered
  links, honest 'we were blocked' classification (403/429/cf-mitigated/
  challenge).
- MCP: run_site_audit, get_audit_status, get_audit_issues, get_audit_pages;
  limitTier resolved via shared AuditService.resolveAuditLimitTier.
- Lighthouse strategies reduced to auto/none (legacy all/manual map on read).
- Self-healing: getStatus reconciles audits whose workflow instance errored/
  terminated without reaching mark-failed.

The Issues UI and the badseo.dev e2e fixture site stack on top of this PR.

Deploy notes: run db:migrate:prod (additive); terminate running audits before
deploying — the workflow step structure changed and in-flight instances cannot
replay under the new code (a finalize guard fails them loudly instead of
completing empty).

* Store only internal link edges in audit_links

Both consumers (broken-internal-link and orphan checks) filter on
isInternal; per-page external counts already live on audit_pages.
Dropping external rows cuts stored edges on outbound-heavy sites.
Column stays so P1 external-link checks can re-add rows without a
migration.

* Review fixes: failAudit CAS guard, dedupe hash helpers, cheaper checks

- failAudit only transitions running audits, so the getStatus reconciler
  can't flip a just-completed audit to failed when it races finalize
- collapse the duplicate SHA-256 helper into audit/ids.ts
- finalize integrity guard uses a limit-1 existence probe instead of
  fetching every page row
- get_audit_status MCP tool no longer reads the audit row twice when an
  explicit auditId is given
2026-07-07 22:08:14 -04:00
Ben Senescu
7caaebbbac Shrink eager worker bundle: lazy boundaries at module seams + build-enforced guard (#366) 2026-07-07 21:47:42 -04:00
Ben Senescu
a337f0ac08 perf: bound Autumn retry window + cache customer existence off the hot path (#365) 2026-07-07 21:47:42 -04:00
RDeemer63
ffa6ec5e70
feat(rank-tracking): add city/region-level location targeting (#62) 2026-07-07 21:44:19 -04:00
Ben Senescu
c3caf2009f
Remove backlinks + LLM-mentions access gates 2026-07-05 23:25:33 -04:00
Ben Senescu
3c7e704b4a
Triage active PostHog errors: validator noise, workflow output cap, Autumn retries, deploy-reset noise (#361) 2026-07-05 23:22:44 -04:00
Ben Senescu
828fb17073
perf: cut cross-region latency (Smart Placement, session cookie cache, shared Autumn provider, parallel queries) (#359) 2026-07-05 22:54:33 -04:00
Ben Senescu
c9bcc444b4
Triage unresolved PostHog errors: DataForSEO validation, DO retry, Autumn 429, exception noise (#333) 2026-07-05 22:38:53 -04:00
Ben Senescu
93a373129e
release: v0.0.24 (#358) 2026-07-05 21:05:01 -04:00
Ben Senescu
abc80e1e2b
Fix server-fn deprecation warnings and add CSRF middleware (#357)
* Replace deprecated createServerFn().inputValidator() with .validator()

* Add CSRF middleware for server functions

* Fix prettier formatting after validator rename

* Pass Zod schemas directly to .validator()

Drops the (data: unknown) => schema.parse(data) lambdas in favor of
standard-schema support. This also makes server-fn callers type-checked
against the schema input, which surfaced a dead hideSpam property in
SearchTabStrip (the schema never included it and the server hard-codes
hideSpam for web requests).

* Name the inline project-scoped schema in projects.ts
2026-07-05 20:18:12 -04:00
Ben Senescu
2645750671
Free-plan audit limits: 50 pages, one at a time; remove 'all' lighthouse strategy (#352) 2026-07-05 18:56:19 -04:00
Ben Senescu
053ac4c4cf
Add skeleton loading state for Search Performance page (#355) 2026-07-05 18:53:56 -04:00
Ben Senescu
7bcd8497a0
Upgrade dependencies and add security audit configuration (#349) 2026-07-05 18:52:59 -04:00
Ben Senescu
a3e7bffd5b
Add OpenRouter API key setup gating for SAM AI features (#356) 2026-07-05 18:39:26 -04:00
Ben Senescu
36bd1cd7de
fix: scope SAM sessions to owning user (#354) 2026-07-05 18:36:58 -04:00
Ben Senescu
b22dc13b51
Remove direct POSTGRES_DATABASE_URL fallback; Postgres via Hyperdrive only (#351)
The Worker now connects to Postgres exclusively through the HYPERDRIVE
binding. Local dev uses the binding's localConnectionString (committed in
wrangler.jsonc, pointing at the throwaway Docker Postgres from
docs/LOCAL_POSTGRES.md) instead of a POSTGRES_DATABASE_URL Worker var.

Closes the Codex finding about unpooled direct connections from deployed
Workers: the not-recommended direct-connection config is no longer possible.
POSTGRES_DATABASE_URL remains as a Node-side env var for drizzle-kit and
the D1->Postgres migration script only.
2026-07-05 18:11:11 -04:00
Ben Senescu
86407c97e1
Fix hosted Turnstile captcha fail-open config (#350) 2026-07-05 18:06:19 -04:00
Ben Senescu
e39de22f96
Fix: don't mask DataForSEO 40501 validation errors as empty results (#346)
* Fix: don't mask DataForSEO 40501 validation errors as empty results

isNoResultsTask classified any task with status_code 40501 as a
successful empty result, but 40501 is not unique to "No Search
Results" — DataForSEO also returns it for validation rejections
like "Invalid Field: 'target'.".

After the Business Listings and Q&A endpoints opted into
treatNoResultsAsEmpty, invalid-field 40501 responses (reachable via
the MCP search_local_businesses categories input, which accepts
arbitrary strings) were silently masked as empty successes — billed
and tracked, but hiding the real provider validation error.

Match on the status message ("no search results") instead of the
ambiguous code so charged validation failures flow through to
DataforseoChargedTaskError with the diagnostic message, while genuine
no-results responses still return as empty successes. Fixes all call
sites uniformly (Business, SERP live, SERP task_get polling).

* Format isNoResultsTask per prettier
2026-07-05 17:54:14 -04:00
Ben Senescu
9b6c5640cd
Sanitize Copy keywords clipboard against spreadsheet formula injection (#348)
The Striking Distance 'Copy keywords' action wrote raw selected GSC query
strings to the clipboard, bypassing the formula-injection sanitizer used
by the CSV and Sheets export paths. GSC query strings are untrusted and can
begin with =, +, -, @, tab, CR, or LF; pasting them into Sheets/Excel could
execute them as formulas.

Route the copy path through normalizeExportValue (same OWASP-recommended
guard as CSV/Sheets exports) so dangerous leading characters are prefixed
with a single quote.
2026-07-05 17:50:01 -04:00
Ben Senescu
7c64cd3c5c
Rank tracking: enforce config cap on unarchive, raise cap to 500 (#347)
The archived-config reactivation path in createConfig returned before the
MAX_CONFIGS_PER_PROJECT check, so re-adding previously archived domains
could push a project past the active-config cap (Codex security finding
ba873417). The cap check now runs before both the reactivation and the
new-row insert. Also raises MAX_CONFIGS_PER_PROJECT from 100 to 500.
2026-07-05 17:39:53 -04:00
Ben Senescu
c993ef4931
Bound concurrent D1 upserts in saved-keyword metric refresh (#338)
refreshSavedKeywordMetrics fanned out one Promise per keyword across an
entire location/language group, so a project with thousands of saved
keywords could trigger thousands of simultaneous D1 upserts, pressuring
Worker memory/CPU and D1 concurrency (potential OOM / availability DoS).

Restore bounded DB-write batching: upsert in chunks of 100 so concurrent
writes stay capped regardless of group size.
2026-07-05 17:35:58 -04:00
Ben Senescu
1a74904b67
Cut worker isolate baseline memory: stub just-bash, lazy-load cheerio (#340)
Production OOM triage: every 'Worker exceeded memory limit' burst hits
unrelated cheap routes right after a deploy — the 128MB limit is per
isolate, and the main worker's eagerly-evaluated module graph is what
crowds it, not any single request.

- Alias just-bash to a throwing stub (worker never uses it): removes
  just-bash + turndown + @mixmark-io/domino from the bundle. The chain
  was pulled in eagerly by @cloudflare/think via the SamChatAgent
  re-export in src/server.ts; SAM only uses its own MCP tools.
- Disable Think's workspace bash tool on SamChatAgent so the stub is
  unreachable at runtime.
- Dynamic-import page-analyzer (cheerio) in the site-audit crawl step
  so it evaluates only when an audit runs, not in every isolate.
- Drop the turndown CJS alias workaround (#339): turndown is no longer
  in the graph at all.

Main eager server chunk: 14,434 kB -> 11,712 kB (-19%); cheerio's
503 kB now a lazy chunk.
2026-07-03 22:23:35 -04:00
Ben Senescu
93cf32081a
Fix turndown ESM build compatibility with Cloudflare Workers (#339) 2026-07-03 21:53:55 -04:00
Ben Senescu
41d37792e7
feat: SAM — in-app SEO agent with full MCP toolset (#322) 2026-07-03 18:48:18 -04:00
Ben Senescu
7ed574c408
Sidebar: promote AI & MCP to nav, consolidate footer into account menu (#337) 2026-07-03 13:23:18 -04:00
Ben Senescu
615ecce033
Redesign app shell: Sidebar with cutout content panel (#336) 2026-07-02 19:06:13 -04:00
Ben Senescu
4f17fe4942 Triage production log errors: audit crash, Autumn webhook FK, PostHog capture, auth rate-limit IP, log noise (#327) 2026-07-02 18:15:46 -04:00
Ben Senescu
fffdbc9329 Add Cloudflare Turnstile captcha on email signup (#326) 2026-07-02 18:15:46 -04:00
Ben Senescu
97f026562e Pin production to Postgres via committed Hyperdrive binding (#329) 2026-07-02 18:15:46 -04:00
Ben Senescu
b5bda159ea Re-add free plan as the floor; remove subscribe gate (#321) 2026-07-02 18:15:46 -04:00
Ben Senescu
819e33be07 fix(mcp): 405 the standalone GET SSE stream to stop /mcp OOM (#325)
The stateless MCP server returns JSON on POST (enableJsonResponse) and
pushes no server-initiated messages, so the optional standalone GET SSE
stream serves no purpose. Left enabled, each GET holds an SSE stream open
indefinitely (25s keepalive, no eventStore) and pins a fresh per-request
McpServer (~5MB of tools + Zod schemas); a few dozen concurrent connected
clients exceed the 128MB isolate limit. This was 100% of the /mcp
exceededMemory OOMs (GET only; POST never OOMed).

Return 405 (spec-compliant 'no standalone stream') before building the
server, so GET allocates nothing. Also removes the bulk of the elevated
GET canceled / responseStreamDisconnected outcomes.
2026-07-02 18:15:46 -04:00
Ben Senescu
a49ec51e95 fix(ai-search): use valid Claude model_name and fail fast on unknown ones (#323)
DataForSEO dropped the Claude Sonnet 4.0 family from its llm_responses
catalog, so model_name=claude-sonnet-4-0 was rejected with 'Invalid
Field: model_name' while still billing the failed task. Point Claude at
claude-sonnet-4-5 and validate every model_name against DataForSEO's
accepted catalog before dispatching the paid call.
2026-07-02 18:15:46 -04:00
Ben Senescu
4e48f8344f Fix production errors: onboarding crash hardening + DataForSEO spend/noise cleanup (#282) 2026-07-02 18:15:46 -04:00
Ben Senescu
868924e1fe refactor(keywords): unify keyword-metric fetching behind one helper (#320) 2026-07-02 18:15:46 -04:00
Ben Senescu
116719a019 fix(db): add missing indexes and drop redundant ones (#319)
Postgres advisor flagged seq-scans and redundant indexes across both
backends (D1 + Postgres):

- add projects(organization_id) — org-scoped project listings seq-scanned
- add account(account_id, provider_id) — better-auth sign-in lookup
- add verification(expires_at) — expired-token cleanup range scan
- drop saved_keyword_tag_assignments_keyword_idx — covered by unique
  (saved_keyword_id, tag_id) prefix
- drop rank_snapshots_run_idx — covered by unique
  (run_id, tracking_keyword_id, device) prefix

Mirrored in both schema dialects + parity-test required-index guard.
2026-07-02 18:15:46 -04:00
Ben Senescu
e5e5029e6b rank-tracking: raise per-project config limit from 20 to 100 (#318)
The cap was only a soft guard against runaway scheduled DataForSEO
workload, not a hard product constraint. Bump it to 100 so projects
tracking many domain/location combos aren't blocked.

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-02 18:15:46 -04:00