* Add onboarding agent v1 product spec
* Add onboarding agent implementation plan (Project Think)
* Update onboarding plan: chat + seed function (drop Think/Workflows)
* feat(onboarding): data + metering foundation, Project Context store, MCP tool
* feat(onboarding): site read + DataForSEO signal + OpenRouter strategy seed
* feat(onboarding): strategy + streaming chat UI with update_project_context tool
* fix(onboarding): address review — bound free runs, cap chat, share auth+error helpers, harden scrape
* fix(onboarding): use canonical keyword-locations list, not a separate country list
* Improve onboarding strategy chat
* feat(onboarding): refine upgrade rail UI + fact-checked copy
- Rebuild upgrade sidebar: drop the nested card so the rail itself is the
container (header / plan / features / CTA / progress footer with dividers)
- Remove the 'Free preview' badge + headline pitch; header now reads
'Previewing OpenSEO' with the site domain beneath
- Tighten copy against the fact sheet: fix monthly-vs-top-up credit wording,
drop 'live' rank tracking, add money-back + open-source trust signals,
unify CTAs to 'Upgrade to continue', cut cross-panel feature redundancy
- Replace off-strategy suggested question; add progress bar counter
- FORCE_FREE_PREVIEW flag to always show the preview/limit UI while testing
* feat(onboarding): add 'What do you recommend' strategy chip; revert suggested questions
- Add a highlighted suggestion chip that prompts Sam for the strategy, shown
only when the user hasn't already used the welcome 'Show my strategy' CTA
- Track strategyRequested so the chip isn't re-offered after use
- Restore the original four suggested questions
* feat(onboarding): add OpenSEO Discord CTA + fact-sheet entry
- Discord link in the upgrade sidebar
- Fact-sheet community entry + system-prompt guidance so Sam can point
users to the Discord for community/second-opinion help
* chore(merge-ready): round 1 fixes
- scrape.ts SSRF: validate the initial domain via audit/url-policy
(normalizeAndValidateStartUrl) and re-validate each redirect hop with
redirect:"manual" (one hop, blocked/private/metadata hosts + DoH rebinding).
Replace the content-length-only guard with a bounded streaming read so
chunked/CDN responses can't buffer past MAX_RESPONSE_BYTES. Remove the
unguarded normalizeDomainToUrl helper. Add scrape.test.ts.
- http-errors.ts: map PAYMENT_REQUIRED AppError to HTTP 402 (was 500), so the
onboarding chat paywall backstop surfaces correctly.
- OnboardingStrategyChat: replace the hardcoded FORCE_FREE_PREVIEW=true debug
flag (which forced paid users into the free-preview/paywall UI) with a
safe-by-default ?preview=1 URL override.
- onboardingStrategy.ts: delete the dead generateOnboardingStrategy export
(knip) and its now-unused imports; the chat tool path uses runOnboardingSeed.
- chat.ts: rename inner runOnboardingSeed result to fix no-shadow.
- Extract presentational chat sub-components into OnboardingStrategyChatParts
to satisfy max-lines; reformat Markdown.tsx for prettier.
* chore(merge-ready): round 2 fixes
- chat.ts: validate message role in schema + count total messages (not just user-role) so the free-question gate can't be bypassed with mislabelled roles
- OnboardingStrategyChat.tsx: surface useChat error state with a paywall-aware notice; branch 'Ask about OpenSEO' message text on isPaid
- OnboardingStrategyChatParts.tsx: guard free-preview welcome copy behind !isPaid (paid variant for subscribers)
- onboardingStrategy.ts: reset onboardingRunStatus/onboardingRunAt when the domain changes so a corrected domain can get a fresh free seed
* chore(merge-ready): round 3 fixes
- onboarding chat: count only user-role messages for free-question paywall to match client gate (was counting all messages, firing ~3 turns early)
- ProjectContextStore: drop unused return value/type from saveProjectContextVersion, inline latest-version query into getCurrentProjectContextMarkdown, remove dead toVersion helper and ProjectContextVersion type
- onboarding chat UI: replace 'Why is OpenSEO better than Claude?' suggested chip with 'How does OpenSEO work with Claude?' (Claude is an MCP client, not a competitor)
* feat(onboarding): route post-upgrade to GSC step; drop isPaid from preview chat
- Checkout successUrl now returns to /onboarding?step=3 (GSC connect) instead
of the strategy chat, with a 'You're in!' success banner introducing the
remaining GSC + MCP setup steps. Fixes the post-Stripe 'stuck on paywall'
race since the user leaves the chat entirely.
- The strategy chat is now purely the pre-upgrade free preview: removed the
managed-access query, the isPaid branching, and the ?preview override. The
7-question cap always applies (kept as a conversion funnel).
* feat(onboarding): show post-upgrade success as its own step screen
Instead of a banner stacked above the GSC step, render a 'You're in!' screen
using the standard step layout (logo, title, card, Continue) in place of the
GSC step when ?checkout=success is present. Continue drops the param to reveal
the actual GSC step.
* refactor(nav): remove Project settings from account dropdown
Project settings is now reachable only via the project switcher's 'Manage
projects' → /projects → per-project settings. Drops the dead
projectSettingsLinkOptions helper and the now-unused AccountMenu projectId prop.
* refactor(onboarding): remove project-context persistence + MCP tool
Defers the Project Context store to a later PR to simplify this one.
- Delete ProjectContextStore, the get_project_context MCP tool (+ registration),
the project_context_versions table (schema + migration 0024 + snapshot), and
the update_project_context chat tool.
- generate_initial_strategy now returns the synthesized strategy to the chat
without persisting it; claimRun still gates paid spend (one free run).
- Chat system prompt no longer injects saved context; it just grounds Sam with
the project's domain.
- getOnboardingStrategyState returns only { projectId, domain }.
- Move the agent fact sheet out of docs/ (human docs) to
src/server/features/onboarding/openseo-fact-sheet.md.
* refactor(routing): move /strategy to /onboarding/chat
Rename the onboarding strategy chat route from /strategy to /onboarding/chat.
_authenticated.onboarding.tsx becomes the index route; the chat is a sibling,
so TanStack auto-creates the shared /onboarding parent (Outlet).
* chore(onboarding): clean up leftovers from the persistence removal
- Drop the chat's onFinish=invalidateStrategyState refetch: now that the
strategy state is just { projectId, domain } (no persisted markdown), the
chat can't mutate it, so the post-turn refetch was dead work.
- Fix a stale 'Project Context' doc comment in synthesis.ts.
- Note in specs 0005/0006 that strategy persistence + the MCP read tool were
deferred, so the docs don't contradict the shipped code.
* feat(onboarding): meter LLM (OpenRouter) spend via Autumn track_tokens
Mirror the DataForSEO metering pattern for LLM cost: a best-effort
trackLlmUsage helper emits a PostHog usage event and records token usage on
Autumn's token-tracking endpoint (REST; not in the autumn-js SDK yet), priced
from the model slug. Wired into the chat stream (onFinish) and the strategy
synthesis call. getOnboardingModel now also returns the resolved model slug.
* feat(onboarding): step-styled site form + account menu on chat page
- Restyle the website/country form to match the onboarding step layout (logo,
title, helper) and explain why we ask (read the site + pick the search market).
- Extract OnboardingAccountMenu to a shared component and render it on the
onboarding chat page so signed-in users can reach account actions there too.
* fix(auth): keep verify-email on 'check your inbox' after email sign-up
The post-sign-up redirect always passes ?email=; key the waiting state off it
so a just-signed-up user sees check-your-inbox + resend instead of the sign-in
CTA the verification gate would immediately block, even while the session is
still resolving.
* fix(onboarding): meter total LLM usage across all stream steps
Review caught that streamText onFinish 'usage' is only the last step; with
stopWhen=4 + the strategy tool, multi-step runs under-metered. Use 'totalUsage'
and await the metering so it fires before the stream closes. Also drop the
in-flux cache/reasoning token fields (negligible here).
* feat(onboarding): adopt the 'chat with tools' architecture from agent-onboarding-2
Replace the deterministic seed + synthesis pipeline (and the
onboarding_run_status/run_at columns) with two on-demand tools Sam calls —
read_website and get_seo_metrics — and have Sam write the strategy itself
in-stream, so a mid-stream refresh re-runs cache-backed tools instead of
dead-ending on a 'complete' status. Rename OnboardingStrategy* -> OnboardingChat*.
Preserved from this branch: LLM metering (now via the chat onFinish totalUsage,
covering the in-stream strategy), the account menu on the chat page, the
verify-email fix, and the step-styled site form. Drop columns via migration 0025.
* docs(onboarding): correct spend-bound + stale synthesis comments
Clarify that get_seo_metrics spend is bounded by the question cap + one project
per un-upgraded account (not solely caching, which doesn't cover no-data sites),
and drop 'synthesis' from comments now that Sam writes the strategy in-stream.
* refactor(onboarding): metered LLM via Autumn AI-SDK adapter; drop skipBalanceAssert
Now that every org gets an onboarding_plan with usage credits, onboarding spend
draws down the normal balance — no bypass needed.
- LLM metering: use Autumn's official @useautumn/gateway adapter (withLlmMetering
wraps the model; correct token-pool pricing for cached/reasoning tokens),
replacing the hand-rolled onFinish/track_tokens REST plumbing. Point it at the
existing 'llm_usage' feature (backed by usage_credits + topup_credits) rather
than a to-be-created 'ai_credits' feature.
- DataForSEO: remove skipBalanceAssert end-to-end (chat metering object, the
meter() plumbing in dataforseo/client.ts, and the DomainService override type);
onboarding now asserts balance like every other caller. Kept the email-verified
+ Labs-location gate on get_seo_metrics as the anti-farming bound.
Co-authored with a parallel agent's LLM-metering refactor.
* docs(onboarding): fix stale metering comment + diverged-architecture specs
- DomainService MeteringOverrides comment no longer claims a balance-gate bypass
(skipBalanceAssert + the onboarding seed are gone).
- specs 0005/0006: correct the update notes — the seed/synthesis pipeline,
claimRun, and skipBalanceAssert were replaced by the chat-with-tools design;
flag the bodies as the superseded plan.
- Document the pinned Autumn track_tokens API version.
* feat(onboarding): gate the chat turn on credit balance (LLM included)
Now that every org gets onboarding_plan trial credits and LLM tokens draw from
the same usage/topup balance, assert that balance before streaming — not just
track it. Extract the DataForSEO balance check into subscription.ts
(getUsageCreditsRemaining / assertUsageCreditsAvailable) and reuse it; the chat
throws a friendly PAYMENT_REQUIRED when credits are gone (client shows the
upgrade copy).
* feat(onboarding): make the strategy chat hosted-only
The chat needs the managed LLM + trial credits, so self-hosted has no business
there. Gate the step-2 navigation on hosted mode and add a beforeLoad redirect
on /onboarding/chat so self-hosted lands back in the wizard.
* feat(onboarding): site-form + welcome copy; drop open-source badge
- Site form: 'Tell us about your website.' title, short input labels, no extra
helper descriptions.
- Welcome message: lead with the upgrade ask + a Discord/email escape hatch.
- Remove the 'Open source — self-host for free anytime' badge from the rail.
* fix(onboarding): show typing indicator during the submitted wait
showTyping gated on the last message lacking assistant text, but right after
send the last message is the user's own (which has text), so nothing showed
until the assistant message appeared. Show it whenever busy and the last
message isn't assistant-text-yet.
* refactor(onboarding): drop redundant email-verified gate on get_seo_metrics
The route guard already requires a verified email to reach the chat in hosted
mode, and the trial-credit balance bounds spend — so the in-tool emailVerified
check was redundant for real users and blocked local/bypass testing. Keep the
Labs-location check (functional).
* feat(billing): meter onboarding LLM spend into the shared credit pool
Both DataForSEO and onboarding-LLM now draw from the same usage_credits/
topup_credits pool via one helper, instead of LLM needing a separate Autumn
ai_credit_system.
- Extract trackUsageCreditSpend (markup -> credits -> monthly/topup split ->
autumn.track + usage:credits_consume) into subscription.ts; DataForSEO's
trackDataforseoCost now delegates to it (behavior unchanged, tests pass).
- Enable OpenRouter usage accounting; the chat onFinish sums the real per-step
cost OpenRouter reports and deducts it through the same helper.
- Drop the @useautumn/gateway adapter, llm-metering.ts, track_tokens, and the
AUTUMN_LLM_USAGE_FEATURE_ID constant — no ai_credit_system feature needed.
* feat(onboarding): persist the strategy chat in a Durable Object (AIChatAgent)
Move the onboarding chat from a stateless streamText route to an Agents SDK
AIChatAgent Durable Object, so the conversation persists (DO SQLite) and
survives reloads — one instance per project.
- OnboardingChatAgent.onChatMessage ports the system prompt, read_website +
get_seo_metrics tools, the credit-balance/free-question gate, and the
OpenRouter cost metering. Billing gates surface as a normal assistant message
(staticAssistantResponse) rather than an HTTP 402.
- The Worker authorizes every /agents/* connection (resolve session + verify the
caller's org owns the projectId) before it reaches the DO; the DO derives org
/domain from the project it is named after. Auth stays on the proven path.
- Client swaps useChat -> useAgent + useAgentChat (WebSocket), keyed by projectId.
- Adds the DO binding + new_sqlite_classes migration; pins @cloudflare/ai-chat
0.6.1 to match agents 0.12.3.
* chore(onboarding): bump agents+ai-chat to latest; fix review findings
- Bump agents 0.12.3 -> 0.15.0 and @cloudflare/ai-chat -> 0.8.4 (the supported
pairing; verified MCP, the DO, and the build still compile).
- Thread the per-turn abortSignal into streamText so a user aborting mid-stream
cancels the billable LLM call (was leaking sub-cent cost on abort).
- Ensure the org's Autumn customer exists in the Worker authorize step before
the DO checks the credit balance, avoiding a false 'out of credits' gate on a
brand-new org's first message.
* chore: remove stray reservation-booker-seo-report.html
* chore: drop stale @useautumn/gateway minimumReleaseAge exclusion
The package was removed when LLM metering moved to the shared credit pool.
* perf(onboarding): fetch get_seo_metrics signals in parallel; clarify question-cap
- get_seo_metrics now fetches the domain overview and ranked keywords
concurrently instead of in series (faster tool turn). Trade-off: it always
issues the metered ranked-keywords call now, including for no-ranking sites.
- Correct the FREE_ONBOARDING_QUESTION_LIMIT comment: the server re-check counts
client-supplied history, so the cap is a conversion nudge, not a security
boundary — the credit balance is the real spend bound.
8.5 KiB
Onboarding agent
Status
Proposed (June 2026) — v1 product spec, pending technical design.
Update (June 2026): what shipped is a chat where Sam analyzes the site on demand (via
read_website+get_seo_metricstools) and writes the strategy in-stream, rather than the staged synthesize → persist pipeline described below. Persisting the strategy (the "Project Context" store + R2 versioning) and theget_project_contextMCP tool are deferred to a later PR — the strategy is shown in the chat but not yet saved. Sections below describe the original plan.
Goal
Turn signup into activation. When a new user onboards, an "agent" analyzes their actual website live, then proposes a tailored SEO strategy. The strategy is free and stands on its own; acting on it (rank tracking, content briefs, the ongoing coach) is the paid surface. The strategy becomes the project's durable "Project Context" — readable in the app and over MCP.
This is the top activation priority because it gives every new user a concrete, personalized "here's what to do" moment before they're asked to pay.
What "agent" means here
A guided pipeline narrated live, not an LLM agent loop. The steps are mostly deterministic (scrape, fetch keyword data); a single LLM call at the end synthesizes the strategy narrative. The "agent" feeling comes from streaming the steps in real time ("Reading your homepage… you look like a Notion alternative… 4 keywords ranking, 30 worth targeting…") and streaming the final write-up token-by-token. No Anthropic SDK / agent infra is required for v1.
The experience
Stage 0 — Profile (extend existing onboarding forms). Collect: domain, experience level, and primary goal. Website and default country are captured on the same step, since the country is a property of the site/project being analyzed — keeping them together makes the relationship obvious and avoids a stray standalone country field. That step carries a short note that they can add more projects with different websites (and their own countries) later, so users don't feel they must cram every site into this first one. Keep the whole stage to ~4 fields — the onboarding UX audit flagged form friction. Country and language become the project's default location (see Project defaults) and are reused for every DataForSEO call going forward.
Stage 1 — Discover (live). Sitemap-first using existing robots.txt + sitemap discovery; fall back to a shallow crawl. Produces a map of the site's shape from URLs/titles. Discover many URLs (cheap); scrape few.
Stage 2 — Read (live). Scrape ~3–5 key pages (home + top product/nav pages) to markdown. This is the one net-new capability: page → markdown via Cloudflare Browser Rendering. Honest "we couldn't read your site" flagging with a manual "tell us what you do" fallback so the run never dead-ends.
Stage 3 — Signal (live). See Stage 3 data below.
Stage 4 — Synthesize (streaming LLM). One LLM call over {profile + scraped markdown + keyword data} produces: a positioning statement, 3–5 themes/clusters, a starter keyword table (volume/difficulty), and a prioritized "do this next" list. Must produce a credible strategy even when the site has zero existing rankings — the cold-start case is the default for the indie-founder ICP, not an edge case.
Stage 5 — Persist + gate. Save as the Project Context artifact (markdown, MCP-readable). Present the full strategy free; the paywall lands on executing it.
Stage 3 data (v1 — kept minimal)
Stage 3 is archetype-conditional: detect the site type cheaply, then frame the goal accordingly. For v1 the archetype branches the narrative, not a large matrix of API calls. Two calls baseline:
| Call | Endpoint | Purpose | When |
|---|---|---|---|
domain_rank_overview |
/v3/dataforseo_labs/google/domain_rank_overview/live |
traffic + # ranking keywords; the archetype detector | always |
keyword_ideas |
/v3/dataforseo_labs/google/keyword_ideas/live |
starter keyword list, seeded from scraped themes | always |
ranked_keywords |
/v3/dataforseo_labs/google/ranked_keywords/live |
what they already rank for | only if overview shows real rankings (usually skipped for new sites) |
Archetypes (narrative only in v1): new/pre-traffic, established content site, local business, SaaS/product (core ICP).
Clarifying question: for beginners, skip it and use the archetype's default goal (fewer decisions = less drop-off). For non-beginners, ask one targeted question after detection to confirm intent (e.g. "grow existing rankings, or expand into new topics?") — agency without a survey.
Deferred to v2: serp_competitors, local-business tools
(get_local_serp_results, search_local_businesses,
get_google_business_questions), and Google Ads volume for non-Labs countries.
Cost per free onboarding
The whole run is free (paywall is after), so cost-per-signup must be bounded:
- DataForSEO: ~$0.04–0.08 (2 Labs live calls; metered from the real
costfield in the response envelope, so we can hard-cap). - Browser Rendering scrape: ~negligible.
- LLM synthesis: ~$0.05–0.15 (one call over a few pages of markdown).
- Total ≈ $0.10–0.25 per onboarding.
Guardrails: one run per user, results cached hard, and email verification gates Stage 3 (the DataForSEO spend) to kill drive-by abuse.
Project Context artifact (MCP-readable)
The strategy is stored as a markdown document — a living "Project Context" that is the shared source of truth across the in-app agent, the web UI, and any MCP client.
- Storage: a
project_contextrow in D1 (projectId,markdown,updatedAt,version). Small enough for D1; R2 only if it grows. - MCP: a
get_project_contexttool (and ideally an MCP resource so it auto-loads into the agent's context). The user's own Claude/Codex reads the same strategy the app shows. - The onboarding agent authors the context; the paid execution actions (rank tracking, briefs, coach) read from and append to it. That append-back is the natural shape of the gated surface.
Project defaults (country/language)
Projects don't currently store a default location — only rank_tracking_configs
does (defaults 2840/US, "en"). Add location_code + language_code to
projects, set from the Stage 0 country, and reuse them everywhere project work
needs a location. A small curated country → location_code map is enough for
v1.
Paywall
Free: the full strategy, positioning, themes, capped starter keyword list
(~15–20), and the "do next" list. Gated (execution): rank-tracking the proposed
keywords, full keyword expansion, content briefs per theme, and the ongoing
seo-coach. Every gated action is a named thing tied to a strategy item the
user already believes in — stronger pull than "refine further." GSC stays behind
the gate; we do not prompt for it during onboarding (connecting then hitting
a paywall is exactly the bait-and-switch the UX audit warned against).
Unify the seo-coach skill
Update the seo-coach skill (and onboarding-checklist) so the coach and the
onboarding agent are one continuous experience: the coach reads Project Context
via MCP, picks up where onboarding left off ("here's your strategy — let's work
the backlog"), and can answer both SEO questions and OpenSEO product questions.
The onboarding agent sets the backlog; the coach executes it.
Open questions
- Country →
location_codesource: hand-curate a short list for v1, or reuse an existing DataForSEO locations dataset? - Re-run policy: confirmed one strategy per project, regenerate in place; new domain = new project (a soft upgrade nudge).
- Synthesis model choice (cost vs quality) — decide at technical design.
Out of scope for v1
Conversational/iterative agent loop, competitor and local-business analysis, GSC-enriched strategy, multi-language strategies, automatic content generation. These are incremental improvements once the core activation loop is proven.