# Online Ranking — Design

**Date:** 2026-08-11
**Status:** Draft, awaiting review
**Scope:** Phase 4 of seven. Depends on Phases 0–3, all built and live-verified.

---

## 1. Purpose

Serve the pairing graph to shoppers.

**The number that justifies this phase: the live widget can recommend for 7 of
218 products.** `related_keys` was only ever populated by the crawler, so when
the catalogue grew to 218 products through Shopify and the HTTP API, the
recommendation feature silently stayed at its original seven. The pairing graph
covers 201.

Every phase so far has been invisible to shoppers. This is the one where the
work reaches them.

**Success criteria:**

1. A visitor on any of 201 products gets a recommendation, not 7.
2. Nothing unapproved, rejected, out of stock or flagged is ever served.
3. The path stays a cheap lookup — no model call, no embedding, tens of
   milliseconds.
4. A tenant who has never run pairing is no worse off than today.

## 2. What changes, precisely

One function: `decide()` in `app/services/pairing/recommendations.py`.

```
today:  find_by_url -> related_products -> related_keys (7 products covered)
after:  find_by_url -> neighbours_for   -> product_neighbors (201 covered)
```

`to_card` is untouched, so the response shape the widget already consumes does
not change. `match_products` on the `search` branch is untouched. The feature
flag, the event vocabulary, the `recommend: false` contract and the
never-raise guarantee are all untouched.

**`catalog/products.py` is not modified.** The neighbour lookup belongs in
`pairing/queries.py`, which already owns every read of
`strategist_product_neighbors` and already imports `is_servable` — putting a
pairing query inside the product store would invert the dependency for no gain.
`related_products`, `compute_relatedness` and `save_related` stay exactly as
they are, because they are the fallback in §6.

The two attribution fields in §9 are attached to the card in
`recommendations.py` after `to_card` returns, so that function stays untouched
too.

## 3. What is served

**Similar and complement, blended, on every event.** The same answer at every
moment rather than a different one per event.

This is a deliberate choice of simplicity over merchandising nuance. The
alternative — alternatives when hesitating, accessories after add-to-cart — is
better targeted but needs a mapping that can be wrong per event and is harder to
reason about. One blend is one thing to verify.

**The known cost, recorded rather than hidden:** on `add_to_cart` and
`checkout_start` a shopper who has already chosen is shown alternatives to the
thing they just chose. If measurement in Phase 5 shows that suppresses
conversion, the per-event mapping is the fix and this section is where to start.

**Upsell is not served.** It is a good signal — arithmetic, high confidence — but
it only makes sense at the moment a shopper is comparing, which is exactly the
per-event targeting this phase is not doing. Serving a pricier version of the
item at every moment, including after purchase intent, would read as pushy.

**Bundles are never served.** Phase 3's live review found the two surviving
bundles are lexical coincidences: "Dress Pea → Yellow Peeler", where the
embedding sees "Pea" in "Peeler". The complement signal that feeds bundles is
not reliable enough on short titles outside electronics. Bundles need behavioural
data or a real merchant catalogue; until then they stay in the admin tool where a
human sees them, and out of the widget where a shopper would.

## 4. The rule that must not leak

**Nothing a merchant has not approved may reach a shopper.**

```
rejected                        -> never served
below the approval threshold,
  with no decision recorded     -> never served
approved                        -> served
merchant-declared               -> served
at or above the threshold       -> served
```

This is `is_servable` from Phase 2, unchanged and reused rather than
reimplemented. The approval queue means nothing if the serving path ignores it.

Servability is evaluated **at serve time, not at pairing time**: a merchant who
rejects a pair expects it to stop appearing immediately, not after the next
pairing run.

## 5. Filters applied at serve time

Beyond servability, three filters, all cheap and all in the same query:

| Filter | Why |
|---|---|
| the current product | never suggest the page the visitor is already on — this rule exists today and is kept |
| `in_stock` false | a recommendation you cannot buy is worse than none |
| `missing_fields` non-empty | a product with no URL cannot be linked to |

Stock and flags are checked **now**, not at pairing time, because both change
between pairing runs. Phase 2 filtered them when building the graph; that is not
enough, because a product can sell out an hour later.

**Price affinity needs no separate rule.** `similar` is already capped at a 3×
price ratio by the pairing rules, and complements are meant to be cheaper than
their anchor. The matching document's price-affinity rule is already satisfied
upstream, and adding it again here would suppress valid accessories.

## 6. Fallback

A tenant who has never run `POST /catalog/pair` has an empty graph. They must not
get a worse experience than today.

```
neighbours from product_neighbors  ->  if none, related_keys  ->  if none, nothing
```

`related_products` stays exactly where it is for this reason. This also makes the
change safe to deploy before every tenant has been paired, and makes rollback a
matter of one constant rather than a revert.

## 7. Ordering

Servable neighbours, sorted by `score` descending, capped at three — the cap the
widget already applies.

**Type is not a tiebreaker and complements are not boosted.** With a uniform
blend, letting type influence order would smuggle the per-event targeting back in
through the side door, without the clarity of having chosen it.

Where scores tie, order by `product_key` so the same page produces the same
answer twice. An unstable order looks like a bug to anyone testing it and makes
Phase 5's measurement noisier.

## 8. Cost

**One query.** The neighbour lookup joins `strategist_product_neighbors` to
`strategist_products` and applies the stock and flag filters in SQL, returning
whole rows ready for `to_card`. Decisions are a second small query.

This path runs on every page view of every visitor on an unauthenticated
endpoint. The existing module docstring is emphatic that nothing here may call a
model or spend `ai_tokens`, and that constraint is carried forward unchanged.
Nothing in this design embeds, scores or infers at request time — the graph was
built once, offline, in Phase 2.

**Measured, not assumed.** An earlier draft of this section claimed the path
answers in tens of milliseconds. That was asserted without measuring and it is
false. On the live tenant:

```
opening one database connection      ~197 ms
the path BEFORE this phase           ~622 ms
the widget end to end after it      ~1018 ms
```

The dominant term is not the query — it is that **every helper opens its own
connection and there is no pooling**, so the cost scales with how many functions
a request touches. Serving the graph adds one connection to a path that already
had two.

Two things follow. Within this phase, the neighbour lookup and its decision
lookup share one connection and the decisions read is scoped to a single anchor
rather than the whole tenant, which took that step from 605 ms to 358 ms.
Beyond it, **connection pooling is the real fix and it is a platform change**
affecting every endpoint, not something to smuggle in here. It is recorded in
§12 as the largest single performance item this project has found.

## 9. Attribution for Phase 5

Each returned card carries the `pair_type` it came from and its `score`.

Without it, Phase 5 cannot answer the only question that matters — which pair
types earn their place — and would need this work done again. It costs one field.

The field is additive to the card, so the widget continues to work unchanged.

## 10. Error handling

| Condition | Behaviour |
|---|---|
| The graph tables do not exist for a tenant | fall back to `related_keys`, as §6 |
| Every neighbour is unservable or out of stock | `recommend: false, reason: "no_match"` — the existing contract |
| The decisions table is unreadable | treat as no decisions recorded, which is the conservative direction: unapproved pairs stay unserved |
| Any unexpected failure | `recommend: false`, logged — `decide()` never raises, and that is unchanged |

The conservative direction matters. If servability cannot be determined, the
answer is to show less, not more.

## 11. Testing

**The one that must not be got wrong:** a rejected pair is never served. Also a
below-threshold pair with no decision, an approved-but-out-of-stock product, and
a flagged product. These are the four ways an unwanted product could reach a
shopper and each gets an explicit test.

Behavioural, with a real database and a seeded graph:

- a product with graph neighbours returns them, ordered by score
- a product with none falls back to `related_keys`
- a product with neither returns `recommend: false`
- the current product never appears in its own recommendations
- the response never exceeds three cards
- every card carries `pair_type` and `score`
- ordering is stable across two identical calls
- the `search` branch is unchanged — it still uses `match_products`
- **no model client is called anywhere under `decide()`**, asserted by
  monkeypatching the embedding and chat clients to raise

**Live verification:** call `POST /recommendations/event` against the live
tenant for a smartphone URL and read what a shopper would actually be shown.
Then reject a pair through the admin tool and confirm it disappears from the
response without re-running the pairing job.

## 12. Known limitations

**No per-event targeting.** See §3. The cost is recorded there.

**No cooldown.** A visitor moving through several pages is suggested something on
each one. This is true today and is not made worse; it belongs with measurement.

**No personalisation.** Recommendations depend on the page, not the visitor.
Behavioural data would change that and does not exist yet.

**Bundles and upsell are computed but unserved.** They remain visible in the
admin tool. Nothing is deleted; the decision is only about what reaches shoppers.

**The threshold is still one number for all tenants.** Carried from Phase 2.

**No connection pooling, and it dominates every request.** ~197 ms to open a
connection, and each helper opens its own. This is pre-existing and affects every
endpoint in the service, not just this one; serving the graph made a slow path
one connection slower. Pooling — or an in-process TTL cache on
`get_tool_settings`, the same pattern already used for `get_platform_config` — is
the largest single performance improvement available in this codebase. It needs
its own decision and its own testing, and is deliberately not done here.

## 13. Out of scope

Measurement (Phase 5); per-variant stock and accessory-target extraction
(Phase 6); the production merchant dashboard; personalisation; cooldowns.
