# Pairing — Design

**Date:** 2026-08-10
**Status:** Draft, awaiting review
**Scope:** Phase 2 of seven. Depends on Phase 0 (ingestion) and Phase 1 (attributes), both built and live-verified.

---

## 1. Purpose

Turn a catalogue of independent products into a graph of related ones, so the
merchant can browse it, correct it, and approve what they are unsure of before it
is ever shown to a shopper.

Phase 1 ended with 218 of 218 products carrying attributes, `is_accessory`
populated correctly by category, and `price_tier` spread across the tenant's own
distribution. Those three are exactly the inputs pairing needs, and none of them
existed a phase ago.

**Success criteria:**

1. Every eligible product has neighbours of each applicable type, with a score
   and a confidence, stored in a table rather than recomputed per request.
2. Low-confidence pairs are not served until a merchant approves them.
3. A merchant can browse categories, open a product, and see its pairings with
   the reason each one was chosen.
4. Re-running the job is cheap and does not discard merchant decisions.

## 2. Two facts from the live catalogue that shape this

**Merchant-declared complements: 0 of 218.** The Shopify Search & Discovery
`complementary_products` metafield is read by the connector and is simply empty —
this merchant has never set one. The matching specification treats merchant
declarations as the highest-confidence source, and that remains true, but it
cannot be the foundation. **The job must generate the entire graph itself and
treat merchant declarations as a rare, welcome override.**

**Existing `related_keys`: 7 of 218.** `compute_relatedness` in `products.py` ran
only over the crawled products and never over the synced catalogue. It is a
similarity-only, top-5, score-less list with no notion of complements.

That old path stays exactly where it is. `products.py` is not modified, nothing
starts reading `product_neighbors` yet, and the live chatbot's behaviour does not
change in this phase. Phase 4 makes the switch deliberately, once both paths can
be compared on the same catalogue. Building the replacement and cutting over to
it in one step would mean a bad pairing run degrades live recommendations with no
fallback.

## 3. Four pairing types

Each answers a different shopper question. All four are computed by the same job
and stored in the same table, distinguished by `pair_type`.

| Type | Shopper question | Direction |
|---|---|---|
| `similar` | "show me another one like this" | symmetric |
| `complement` | "what goes with this" | directional |
| `upsell` | "what is the better version" | directional |
| `bundle` | "give me the whole set" | anchor → set |

### 3.1 `similar` — the substitute

Same leaf category, comparable price tier, high content similarity. The shopper
did not like this one; offer the nearest alternative.

Score combines embedding cosine over title-plus-description with Jaccard overlap
of the attribute set — the term that contributed nothing before Phase 1 and now
has 1,311 attribute values to work with.

**Deliberately excluded:** a product from a different leaf category, however
similar its text. A phone case is not a substitute for a phone.

### 3.2 `complement` — the goes-with

This is the type `is_accessory` was extracted for. The rule is asymmetric and
that asymmetry is the point:

```
anchor is_accessory FALSE, candidate is_accessory TRUE  -> strong complement
both FALSE, different categories, compatible attributes -> weak complement
both TRUE                                                -> not a complement
```

A phone suggests a case; a case does not suggest another case. Where the merchant
has declared a complement, it is stored with confidence 1.0 and never scored —
their own judgement outranks any inference.

**Attribute compatibility** is what makes this work for clothing. A white shirt
and navy trousers are compatible; a white shirt and orange trousers are not. This
is a rule over extracted `color`, `style` and `gender`, not a model call — the
attributes were already extracted once and the pairing job must not re-pay for
them.

**The honest limit:** colour was extracted for only 52 of 218 products and
material for 28, because this catalogue is electronics and groceries whose text
does not state them. Compatibility scoring therefore contributes little here and
must be re-measured against real apparel data before it is trusted.

### 3.3 `upsell` — the trade-up

Same category, materially higher price, and not worse on rating. This is
**arithmetic, not a judgement** — there is no model call in this path and no
reason for one. A product is an upsell of another when it costs more and is at
least as well reviewed, and those are both columns.

Because it is arithmetic it is high-confidence by construction, and under the
rule in §5 it does not queue for approval. That is not special treatment: it
scores high, and high-scoring pairs never queue.

### 3.4 `bundle` — the set

An anchor plus two or more complements from distinct categories. This is the
outfit case: a shirt, trousers and a belt, not three shirts.

Constrained so the result is sellable rather than merely valid: distinct
categories, every member in stock, and the members' combined price within a band
of the anchor's own price, so a £59 shirt does not anchor a £900 set.

Bundles are stored as a set keyed on the anchor, not as pairwise rows.

## 4. Storage

Every per-tenant table in this system is named `strategist_*` and lives in the
tenant's own schema, so these are created as `strategist_product_neighbors` and
`strategist_pairing_decisions`. They are referred to by their short names below.

### 4.1 `product_neighbors`

One row per directed pair per type. Directed because `complement` and `upsell`
are not symmetric, and storing them symmetrically would suggest a phone whenever
a shopper looked at a case.

```sql
CREATE TABLE product_neighbors (
    anchor_key    TEXT NOT NULL,
    neighbor_key  TEXT NOT NULL,
    pair_type     TEXT NOT NULL,   -- similar | complement | upsell | bundle
    score         REAL NOT NULL,   -- 0..1, the strength of the relationship
    confidence    REAL NOT NULL,   -- 0..1, how much we trust the derivation
    source        TEXT NOT NULL,   -- merchant | embedding | attribute | arithmetic
    reasons       JSONB NOT NULL DEFAULT '[]',
    computed_at   TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
    PRIMARY KEY (anchor_key, neighbor_key, pair_type)
);
```

**`score` and `confidence` are separate on purpose.** A pair can be a strong
relationship derived by a weak method, or a weak relationship we are certain
about. Collapsing them into one number would make the approval queue in §5
incoherent — it gates on how much we trust the derivation, not on how good the
pair is.

**`reasons` is what the merchant is shown.** "Same category, 0.82 text
similarity, 3 shared attributes" is reviewable; a bare 0.79 is not. A merchant
asked to approve something must be told why it was proposed.

### 4.2 `pairing_decisions`

```sql
CREATE TABLE pairing_decisions (
    anchor_key    TEXT NOT NULL,
    neighbor_key  TEXT NOT NULL,
    pair_type     TEXT NOT NULL,
    decision      TEXT NOT NULL,   -- approved | rejected
    decided_by    TEXT,
    decided_at    TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
    PRIMARY KEY (anchor_key, neighbor_key, pair_type)
);
```

**Separate from `product_neighbors` for one reason: re-running the job must never
destroy a human decision.** The job rewrites its own table freely. It never
touches this one. A merchant who rejects a pair and re-syncs their catalogue does
not see that pair return.

A rejected pair stays rejected even if a later run scores it highly. A merchant
saying no is information the scorer does not have.

## 5. The approval rule

```
confidence >= APPROVAL_THRESHOLD  -> servable immediately
confidence <  APPROVAL_THRESHOLD  -> queued, not served until approved
merchant-declared                 -> servable immediately, never queued
explicitly rejected               -> never served, never re-queued
```

Only low-confidence pairs queue. Every pair queueing would mean 218 products ×
several neighbours each — over a thousand decisions before the first
recommendation works, which in practice means the queue is never cleared and the
feature stays dark. A queue the merchant can actually finish is worth more than a
complete one they abandon.

The threshold is a constant, tuned once against the live catalogue during
implementation and recorded with the number that was measured. Picking it before
seeing the score distribution would be a guess.

## 6. The job

`POST /catalog/pair` runs it. Like enrichment, it is separate from sync because
it is expensive and must be re-runnable without re-fetching anything.

```
1. load eligible products      (in stock, not missing_fields, has attributes)
2. embed titles+descriptions   (batched, cached on content_hash)
3. block by category           (only compare products that could plausibly pair)
4. score each type
5. write product_neighbors     (replacing this tenant's rows for that type)
6. leave pairing_decisions untouched
7. report counts per type, queued vs servable
```

**Blocking by category in step 3 is what keeps this from being O(n²).** At 218
products the full comparison is 47,000 pairs and would run fine; at 20,000 it is
400 million and would not. Comparing only within and between plausibly-related
categories keeps the job linear in catalogue size for the shape of catalogue this
system actually has.

**Embeddings are cached on `content_hash`**, the same mechanism Phase 1 proved:
its second run made zero model calls. Re-pairing after a small catalogue change
re-embeds only what changed.

**Eligibility is explicit.** A product that is out of stock, or flagged in
`missing_fields`, is not paired at all — recommending something unbuyable is
worse than recommending nothing. Phase 0 established that flag and this is the
first thing that reads it.

## 7. Endpoints

All read-only except the job and the decision.

```
POST   /catalog/pair                       run the job          -> counts per type
GET    /catalog/categories                 browse               -> leaves with counts
GET    /catalog/products                   browse, filterable   -> paged cards
GET    /catalog/products/{key}             one product          -> full record
GET    /catalog/products/{key}/pairings    its pairings         -> grouped by type
GET    /catalog/pairings/pending           the approval queue   -> low-confidence only
POST   /catalog/pairings/decide            approve or reject    -> writes a decision
```

This is the merchant-facing flow you described: browse categories → open a
product → see its pairings → approve the ones the system is unsure of.

`GET /catalog/products/{key}/pairings` returns every type grouped, each pair
carrying its score, confidence, reasons, and whether it is currently servable.
The merchant sees the same evidence the ranker will.

`POST /catalog/pairings/decide` accepts a batch. A merchant working through a
queue decides many pairs in one sitting, and one request per click would make the
screen feel broken.

**Not in this phase:** the shopper-facing recommendation endpoint. Nothing here
changes what the chatbot serves — see §2.

## 8. Error handling

| Condition | Behaviour |
|---|---|
| Embedding call fails for a batch | that batch's products get attribute-only scores, and the degradation is counted in the report |
| A tenant has fewer than 2 eligible products | 200 with a zero report, not an error |
| A pair references a product deleted since the run | filtered at read time, not at write time — the catalogue changes between runs |
| A decision arrives for a pair that no longer exists | recorded anyway; a re-run may recreate the pair and the decision must already be waiting |
| The job is run twice concurrently | the second run's writes win; rows are replaced per type, not appended |

## 9. Testing

**Behavioural, with embeddings stubbed.** The scoring rules are the thing worth
testing and they are deterministic:

- a phone suggests a case; a case does not suggest a phone (asymmetry)
- two accessories are never complements
- an upsell is always more expensive and never worse-rated
- a bundle never contains two products from the same category
- an out-of-stock product is never paired
- a rejected pair does not return after a re-run — the single most important test
  in this phase, because getting it wrong means ignoring a merchant
- a merchant-declared complement is servable without approval
- re-running replaces rows rather than duplicating them

**Assertions are on relationships, not on scores.** Asserting a specific cosine
would make the suite depend on the embedding model, which is how a suite stops
being trusted. Phase 1 established this and it holds here.

**Live run:** pair the real 218-product catalogue and report pairs per type, the
servable/queued split, and a hand-check of the pairings for a smartphone and a
mobile accessory — the two categories where the answer is obvious enough to be
checkable by eye.

## 10. Known limitations

**No behavioural signal.** Every pair here is derived from content. "Frequently
bought together" is the strongest pairing signal in commerce and needs order
history this system does not yet receive. Recorded because it will eventually
outrank most of §3.

**Colour compatibility is under-fed.** 52 of 218 products have a colour. The rule
is built and correct; this catalogue cannot exercise it. See §3.2.

**No per-size stock.** A garment in stock overall but sold out in the shopper's
size is a dead recommendation. `in_stock` is per-product. Carried forward from
Phase 1 and still unsolved; it needs per-variant stock and webhooks.

**No merchant admin UI.** These are endpoints only, and a merchant cannot clear
the approval queue without a client calling them directly. The admin UI is four
screens, not the five the matching document describes: connect and sources,
catalogue browser, product detail with pairings, and the approval queue. The
performance screen is dropped — it has no data until measurement exists, and
building it now would be a mock.

Of those four, only the last two are required for this phase to be usable by a
human. The other two are conveniences over endpoints that already work.

**The threshold is one number for all tenants.** Correct thresholds probably
differ by catalogue shape. One tenant is not enough evidence to make it
per-tenant.

## 11. Out of scope

Online ranking and the shopper-facing recommendation endpoint (Phase 4);
measurement counters (Phase 5); the admin UI (unowned); order-history ingestion.
