# Frontend Guide: URL → Products Flow

For a merchant with neither a CSV nor a product API — they give us their website URL and we read the products off it. Same backend the CSV flow already uses (`/catalog/*`), one new connect endpoint.

## 1. Connect the website

```
POST /sources/website
Content-Type: application/json

{
  "tenant_id": "org_xxx",
  "url": "https://merchant-site.com"
}
```

Only `tenant_id` and `url` are required. No currency field needed — it's inferred automatically from the page if not stated.

**Response:**
```json
{
  "kind": "crawl",
  "external_ref": "merchant-site.com",
  "status": "active",
  "products": "pending_first_crawl",
  "needs": [],
  "job_id": "job_bd08789f384049aa97680293dd0a4830",
  "job_status": "queued"
}
```

This call does two things at once: saves the connection **and** starts the build automatically. Take `job_id` straight to step 2 — there's no separate "press build" step needed in the normal case.

`job_status` tells you what happened:
| Value | Meaning |
|---|---|
| `queued` | A fresh build started — poll `job_id` normally. |
| `already_running_without_this_source` | Another build was already running when this connected. It does **not** include this new source (it read the source list before this one existed). `job_id` here is that *other* build. Once it finishes, call `POST /catalog/build` once more to pick up the new source. |
| `not_started` | Build couldn't be scheduled, but the source **is** still connected. Fall back to `POST /catalog/build` manually. |

## 2. Poll the job

```
GET /catalog/jobs/{job_id}
```

```json
{
  "job_id": "job_bd08789f384049aa97680293dd0a4830",
  "status": "running",
  "percent": 62,
  "step": "loading eligible products",
  "result": null,
  "error": null
}
```

Same polling contract as the CSV flow already uses. `status` is one of `queued | running | done | failed | lost`. Poll every 2-3 seconds; stop on any terminal state.

Realistic timing for a single-site crawl: **~2-5 minutes total** depending on site size (crawling itself is the slow part — it's real page loads, not an API call). The bar now moves incrementally through the crawl instead of sitting frozen — if it's stuck at exactly 0% for several minutes, something is actually wrong; message will show `"lost"` if the job died mid-run.

## 3. List products / show one product

Identical to the CSV flow — the same `GET /catalog/products` and `GET /catalog/products/{product_key}` endpoints, no branching needed on the frontend for source type.

```
GET /catalog/products?tenant_id=org_xxx&limit=50&offset=0
```

**One thing that matters for this source type specifically:** `product_key` for a crawled product is the full page URL (e.g. `crawl:merchant-site.com:https://merchant-site.com/products/shirt`), not a short ID. **Always `encodeURIComponent()` it** before putting it in a path:

```js
fetch(`/catalog/products/${encodeURIComponent(productKey)}`)
```

Without encoding, the request 404s — the key's embedded `/` and `:` characters need to survive as one path segment.

## 3a. Deleting a product

```
DELETE /catalog/products/{product_key}
```

Same `encodeURIComponent()` rule as above:

```js
fetch(`/catalog/products/${encodeURIComponent(productKey)}`, { method: 'DELETE' })
```

**Response:**
```json
{ "product_key": "crawl:merchant-site.com:https://merchant-site.com/products/shirt", "status": "deleted" }
```
`404 {"detail": "Product not found"}` if the key doesn't exist (including calling delete twice — the second call 404s, it's not a silent no-op success).

**What actually gets removed** — exactly this one product, plus everything that references it:
- The product row itself.
- Every pairing it appears in, on either side — as the anchor ("similar to this") *and* as a neighbor (showing up as a recommendation under some other product). Deleting product A also removes its pairing rows with B, C, D if it was paired with them — those other products aren't deleted, just that one relationship.
- Its embedding (used for similarity matching).

**What does NOT happen:**
- Other products are untouched.
- The connected source stays connected — this doesn't disconnect the merchant's site or CSV.
- If the merchant's site still has this product, the next build/re-crawl will find it and re-add it fresh (new embedding, re-paired against whatever exists at that time). Deleting a product isn't a permanent exclusion — there's no "don't re-import this" flag.

## 3b. Wiping a tenant's whole catalogue

```
DELETE /catalog/products?tenant_id=org_xxx
```

No `product_key` in the path this time — this wipes **every** product, pairing, and embedding for the tenant, **and disconnects every source they had connected** (Shopify, CSV, website — all of them), in one call. Use this for a full "reset my catalogue" action, not for removing a single item (use 3a for that).

**Response:**
```json
{ "products": 27, "pairings": 105, "embeddings": 27, "sources_disconnected": 2 }
```
Counts reflect what was actually deleted/disconnected — calling it again on an already-empty, already-disconnected tenant returns all zeros, not an error.

**What happens to connected sources:** they are disconnected, not just left in place — a source stays connected after a plain product wipe would otherwise get silently re-crawled on the next `POST /catalog/build`, repopulating exactly what was just deleted. After this call, `GET /sources` returns none, and a build does nothing until the merchant reconnects (`POST /sources/website`, etc.) — that reconnect is what makes the next build add anything again.

**Edge case, both delete endpoints:** a `tenant_id` that has never been bootstrapped (no schema exists for it at all) returns `500`, not `404` — there's no "does this tenant exist" pre-check. This should never happen for a real, logged-in merchant; it's only reachable by sending a malformed/made-up tenant_id.

## 4. Field completeness — what to expect

Price, brand, and category are populated from the site's own structured data where available, with an AI fallback for pages that have none. Real-world coverage from testing: **100% for price/brand/category** on a well-structured storefront. Two things are *not* filled in yet for crawled products:

- `options` (size/color variants) — always empty for this source type currently.
- `handle` — always null.

Neither blocks display or pairing; just don't build UI that assumes they're populated for crawl-sourced products the way they might be for Shopify-sourced ones.

## 5. Re-crawling

A website source is **not** re-crawled on every build press — crawling is dozens of requests to the merchant's own server, unlike a single API call, so it's throttled to once per 24h by default. If you need to force a re-crawl (e.g. a "refresh" button in a merchant settings screen), call:

```
POST /catalog/build
tenant_id=org_xxx
force=true
```

(form-encoded, not JSON — matches the existing build endpoint's contract).

The throttle only skips a re-crawl when the source's existing products are still there — if the catalogue was emptied (e.g. via 3b) but the source was crawled recently, the next build re-crawls anyway rather than reporting "done" with zero products.

## Quick reference: what's actually new vs. the CSV flow

| | CSV | URL |
|---|---|---|
| Connect call | `POST /catalog/import/csv` (multipart file) | `POST /sources/website` (JSON) |
| Triggers a build automatically | No — import is synchronous, returns immediately | Yes — `job_id` in the same response |
| Everything after connecting | Identical — same jobs, products, pairings endpoints | Identical |
