# Product Recommendations — Frontend Integration

Everything the UI needs, in the order it happens. Every endpoint here has a
matching request in the Postman collection
(`docs/galaxiq-catalogue.postman_collection.json`), named the same way, with a
real saved response you can read before writing any code.

Set two collection variables and the whole thing runs top to bottom:

| Variable | Example |
| --- | --- |
| `base_url` | `http://localhost:8001` |
| `tenant_id` | `org_8c32bf3e-6a18-4739-9b1c-94c0cf11125f` |

Requests save ids (`job_id`, `anchor_key`, `pair_anchor`) into collection
variables as they run, so the steps chain without editing anything.

**Before you ship:** none of these endpoints authenticate. They take a
`tenant_id` in the request and trust it. This service must sit behind your
gateway — it is not safe to call directly from a browser.

---

## The shape of the whole thing

```
STEP 1  Merchant adds products    CSV upload, product API, or their website
STEP 2  Merchant presses ONE button   → returns a job_id
STEP 3  You show a progress bar   poll that job_id
STEP 4  You show the catalogue    categories, products, what goes with what
STEP 5  Merchant approves matches the system was unsure about
STEP 6  Shoppers get recommendations
```

Steps 2 and 3 are one user action. Steps 4 and 5 are the merchant's admin
screens. Step 6 is the storefront.

---

## STEP 1 — Adding products

Three ways in. Use any, or all. **None of them fetches anything or runs the
AI** — that is step 2.

### A. CSV upload

`GET /catalog/sample.csv` → the template to hand the merchant.
`POST /catalog/import/csv` → multipart, with `tenant_id` and `file`.

Returns immediately. There is no job to poll.

```json
{ "dry_run": false, "products": 14, "imported": 14, "replaced": 0, "errors": [] }
```

- Bad rows come back in `errors` with the **spreadsheet row number**. The good
  rows still import — show the merchant which rows to fix rather than failing
  the upload.
- Re-uploading a file with the **same name** replaces that batch. A different
  name adds more products.
- `dry_run=true` validates without saving. Good for a preview screen.
- **Currency** comes from the file's `Currency` column, one per row, so a
  single upload can hold two. Optional, and never guessed — without it, prices
  cannot be compared across sources and nothing else changes.

### B. Connect the merchant's product API

`POST /sources/http` — JSON body. **Only `tenant_id` and `base_url` are
required.**

```json
{
  "tenant_id": "org_...",
  "base_url": "https://api.example.com/v1/products"
}
```

```json
{
  "kind": "http_api",
  "external_ref": "api.example.com",
  "status": "active",
  "mapping": "pending_first_sync",
  "needs": ["currency"]
}
```

Which JSON field is the title, the price, the image — all of it is worked out
automatically on the first fetch. **The merchant never has to describe their
own API.**

`needs` is a prompt for your UI, not an error. The source works without it.
Here it is asking for a currency, because currency is never inferred from
another connected source: a wrong guess silently mis-prices a whole catalogue.

The Postman collection has one request per supported payload:

| Postman request | Adds |
| --- | --- |
| Connect a product API — the minimum | nothing |
| With a currency | `"currency": "USD"` |
| With an API key (bearer) | `"auth": {"style": "bearer"}`, `"api_key"` |
| With an API key (custom header) | `"auth": {"style": "header", "header": "X-API-Key"}` |
| When the products are nested, or paginate differently | `"records_path"`, `"pagination"` |

Notes that matter:

- `base_url` must be **https**, and must not resolve to a private or internal
  address. Either returns **400**.
- Missing `tenant_id` or `base_url` returns **422**.
- Re-posting the **same host updates** that connection rather than adding a
  second one — the source is keyed on the hostname. Use this for "edit
  connection".
- The API key is encrypted at rest and is **never returned** by any endpoint.

### C. Point us at the merchant's website

`POST /sources/website` — for a merchant with neither a spreadsheet nor an API.
They give you their site and we read the products off it.

```json
{
  "tenant_id": "org_...",
  "url": "https://shop.example.com",
  "currency": "USD"
}
```

```json
{
  "kind": "crawl",
  "external_ref": "shop.example.com",
  "status": "active",
  "products": "pending_first_crawl",
  "needs": []
}
```

**Only `tenant_id` and `url` are required.** Same rules as the product API:
https only, publicly resolvable, 400 otherwise; currency optional and reported
in `needs`; re-posting the same host edits that connection.

Prices come from the page's own schema.org markup — the structured data a
storefront already publishes for Google. That is exact, so nothing is guessed.
**A page without markup still becomes a product**, carrying `price` in
`missing_fields` instead of a made-up number. It appears in the catalogue and
still gets similarity matches; only the rules that need a number (upsell, the
price-ratio guard) skip it.

**This source is not re-crawled on every build.** Reading a whole website is
dozens of requests to the merchant's own server, unlike one API call, so it is
re-read at most once every 24 hours. `force=true` on the build forces it, and
`"min_interval_hours": 0` on the source opts out of the throttle entirely.

Optional: `max_pages` (how far the crawl may go) and `min_interval_hours`.

If you show a "last updated" line for this source, read `last_synced_at` from
`GET /sources` — a refresh that skipped the crawl will not have moved it, and
that is the honest thing to show.

---

`GET /sources?tenant_id=...` lists what is connected, when it last synced, and
anything still in `needs`. Credentials are never included. This is also where
you can see what was inferred on the first sync.

---

## STEP 2 — The button

`POST /catalog/build` — form field `tenant_id`.

```json
202 → { "job_id": "job_bd08789f...", "status": "queued" }
```

This is the **only** endpoint the button needs. It fetches, enriches and pairs
as one job.

**This is also the refresh button.** There is no separate refresh endpoint —
the same call re-fetches, re-reads whatever changed, and rebuilds the graph.

### Handle the 409

If the merchant double-clicks, or a run is already going:

```json
409 → { "error": "already_running", "job_id": "job_b71c..." }
```

**Do not show an error.** Take that `job_id` and poll it — the work they asked
for is already happening. This is the single most likely thing to look like a
bug in the UI.

### Two flags you probably will not use

- `wait=true` runs it synchronously and returns the report instead of a
  `job_id`. For scripts and debugging — never the UI, it holds the connection
  open for the whole build.
- `force=true` ignores every cache: it re-reads all products through the model
  and re-crawls a website source even if it was read minutes ago. Full
  first-run cost. Bury it in an admin screen or leave it out — with one
  exception, a "re-read my website now" action for a merchant who just edited
  their site and does not want to wait out the 24-hour window.

---

## STEP 3 — The progress bar

`GET /catalog/jobs/{job_id}` — poll every 2–3 seconds.

```json
{
  "job_id": "job_bd08789f...",
  "kind": "build",
  "status": "running",
  "percent": 85,
  "step": "scoring candidate pairs — 3668 of 8687",
  "result": null,
  "error": null
}
```

| Field | Use |
| --- | --- |
| `percent` | 0–100, only ever goes up. Drive the bar directly off it. |
| `step` | A sentence to print under the bar. Already human-readable. |
| `status` | `queued` \| `running` \| `done` \| `failed` \| `lost` |

Stop polling on `done`, `failed` or `lost`:

- **`done`** → `result` holds the summary. Reload the catalogue.
- **`failed`** → show an error, offer retry. `error` holds a short code.
- **`lost`** → the server restarted mid-job. Treat like `failed`, offer retry.

**404** means the job id is unknown or older than 24 hours.

A finished build's `result`:

```json
{
  "products": 212,
  "pairs": 1247,
  "servable": 1246,
  "queued": 1,
  "by_type": { "similar": 354, "complement": 258, "upsell": 635 }
}
```

`GET /catalog/jobs?tenant_id=...&limit=20` gives recent runs, newest first —
enough for a "last updated" line.

### How long it takes, and why

The build has three stages sharing one bar:

| Stage | Bar | What it does |
| --- | --- | --- |
| sync | 0–10% | fetch every connected source |
| enrich | 10–60% | the model reads each product for colour, material, gender, whether it is an accessory **for** something |
| pair | 60–100% | embed each product, find neighbours, score the pairs |

First run on a real catalogue: a few minutes. **Every run after that: seconds
to a minute**, because each product carries a hash over only the fields that
affect matching. A price change re-fetches and skips everything else. A renamed
product re-enriches that one product.

**A rebuild never touches the merchant's approvals or rejections.** Say so in
your UI copy — a merchant who has spent an hour approving pairs will not press
refresh if they think it wipes their work.

---

## STEP 4 — Showing the catalogue

| Endpoint | Returns |
| --- | --- |
| `GET /catalog/categories?tenant_id=` | `[{ "category": "beauty", "products": 5 }]` |
| `GET /catalog/products?tenant_id=&limit=20&offset=0` | `{ "products": [...], "total": n }` |
| `GET /catalog/products/{product_key}?tenant_id=` | one product, full detail |
| `GET /catalog/products/{product_key}/pairings?tenant_id=` | `{ "similar": [], "complement": [], "upsell": [...] }` |

`product_key` looks like `http_api:dummyjson.com:1` — it encodes the source, so
**URL-encode it** before putting it in a path.

The product detail carries `missing_fields`. A product missing `product_url`
is held back from shoppers entirely — a recommendation nobody can click is
worse than one not shown. Surface that in the catalogue screen so the merchant
can see *why* something is not being recommended, rather than concluding the
system is broken.

---

## STEP 5 — Merchant approval

Two GET endpoints do all the reading — one for **products**, one for **pairs** —
and they take **the same filters**. So the screen keeps one filter bar and
points it at either.

```
GET /catalog/pairings/anchors   ?tenant_id=&status=&category=&score=&limit=
GET /catalog/pairings           ?tenant_id=&status=&category=&score=&limit=&anchor_key=
```

### The filters

| Param | Values |
| --- | --- |
| `status` | `pending` (default) · `approved` · `rejected` · `auto` · `all` |
| `score` | `high` (0.80+) · `medium` (0.60–0.79) · `low` (under 0.60) |
| `category` | exactly as `/catalog/categories` spells it |

They all combine. Any other value returns **422** rather than an empty list —
a typo must never look like "you have no work to do".

### The product list (anchors)

A merchant thinks in products, not pairs. "This shirt has 4 to review" is
workable; a flat list of 232 pairs from 108 products is not.

```json
{
  "anchors": [
    {
      "anchor_key": "http_api:dummyjson.com:85",
      "anchor": { "name": "Man Plaid Shirt", "category": "mens-shirts", "...": "..." },
      "matching": 2,
      "pending": 2, "approved": 0, "rejected": 0, "auto": 5,
      "top_score": 0.6165,
      "avg_score": 0.5817
    }
  ],
  "total": 108,
  "avg_score": 0.579
}
```

- `matching` — how many pairs match the **current filter**. This is your badge.
- `pending` / `approved` / `rejected` / `auto` — the **full breakdown, ignoring
  the filter**. One call gives you "4 to review, 2 approved" per product
  without a second round trip.
- `total` — the count before `limit`, for your pager.

Sorted best-candidate-first, so the merchant meets the most promising decisions
while they still have patience for them.

An empty result is `{"anchors": [], "total": 0, "avg_score": 0.0}` — never
null, never a 404.

### The pairs for one product

Add `anchor_key`. Drop it to browse the whole catalogue.

Each pair carries `state` (`pending` \| `approved` \| `rejected` \| `auto`) and
`decision` — what a person actually recorded, or `null`.

`dropped` counts pairs hidden because a product went out of stock or lost its
URL since the run. Non-zero explains a short list; it is not an error.

### Recording a decision

`POST /catalog/pairings/decide`

```json
{
  "tenant_id": "org_...",
  "decisions": [
    {
      "anchor_key": "...",
      "neighbor_key": "...",
      "pair_type": "similar",
      "decision": "approved",
      "decided_by": "someone@example.com"
    }
  ]
}
```

- `decision` must be exactly **`"approved"` or `"rejected"`**. Anything else is
  a 422.
- Send as many as you like in one call — collect the swipes and post them
  together.
- **One decision covers ONE pair.** Approving a dress with earrings says
  nothing about that dress's other matches.
- Takes effect immediately: a rejected pair stops being served on the very next
  shopper request.
- Decisions survive rebuilds.

---

## STEP 6 — What a shopper sees

`POST /recommendations/event`

```json
{
  "tenant_id": "org_...",
  "visitor_id": "visitor-123",
  "event": "dwell",
  "page": { "url": "https://shop.example.com/products/121", "dwell_seconds": 60 }
}
```

```json
{
  "status": "success",
  "recommendation_id": "rec_1e34183335",
  "recommend": true,
  "message": "Still deciding? You might also like these.",
  "products": [
    {
      "product_id": "http_api:dummyjson.com:106",
      "name": "Apple Watch Series 4 Gold",
      "url": "https://shop.example.com/products/106",
      "image_url": "...",
      "pair_type": "complement",
      "pair_score": 0.7041
    }
  ]
}
```

**Always check `recommend` before rendering.** A 200 does not mean there is
something to show:

```json
{ "status": "success", "recommend": false, "reason": "page_unknown" }
```

`page_unknown` means the URL did not match a product in the catalogue. Render
nothing — no empty state, no skeleton.

---

## Two things that will confuse you if nobody says them

### 1. `score` and `confidence` are different numbers

- **`score`** — how good the match is. You **sort and filter** by this. The
  80% / 60–79% / under-60% bands read this.
- **`confidence`** — how sure the system is that it is right.

**Confidence, not score, decides whether a merchant has to look at it.** A pair
goes live without review when confidence is 0.70 or above, or when the
merchant's own data declared it. Everything else waits in STEP 5. A merchant's
recorded decision beats both numbers, in either direction.

### 2. `status=pending` + `score=high` always returns empty

Not a bug, and not worth debugging. `pending` means low **confidence**, and for
similar-type pairs the score is derived from that same confidence — so a
pending pair cannot arithmetically reach 0.80. If your filter bar allows both
at once, expect an empty state and write copy for it.
