# Pairing Implementation Plan


**Goal:** Build the product pairing graph — four relationship types, scored and stored, with an approval queue for the ones we are unsure of and endpoints for a merchant to browse and correct it.

**Architecture:** Rules first, orchestration last. `pairing_rules.py` holds the four scoring functions as pure code with no database and no network, because the rules are the part worth testing and they are deterministic. `pairing_embeddings.py` owns the one expensive thing and caches it on `content_hash`. `pairing.py` is the job. `pairing_decisions.py` owns the human's answers and is the only module that writes them. Two routers expose browse and approval separately.

**Tech Stack:** FastAPI, psycopg2, Azure OpenAI embeddings via the existing client, pytest.

**Spec:** `docs/specs/2026-08-10-pairing-design.md`

**Depends on:** Phase 0 (ingestion) and Phase 1 (attributes), both built and live-verified. The live catalogue holds 218 products, all carrying attributes, with `is_accessory` and `price_tier` populated.

## Global Constraints

- **Never commit, never push, never stage.** The user commits their own work. Every task ends with `git status --short` and a report. Do not run `git commit`, `git push`, or `git add`.
- **No assistant attribution** in any file, message, or commit.
- **Comments explain WHY, not WHAT.** A reviewer who knows Python must find zero noise comments.
- **Do not modify `app/services/products.py`.** Nothing in this phase changes what the chatbot serves. `related_keys` and `compute_relatedness` stay exactly as they are; Phase 4 makes the switch. Importing from `products.py` is fine — modifying it is not.
- **Nothing reads `product_neighbors` for serving in this phase.** The endpoints built here are merchant-facing only.
- **A re-run must never destroy a human decision.** The job rewrites `strategist_product_neighbors` freely and never writes to `strategist_pairing_decisions`. This is the single most important invariant in the phase.
- **Scores are cached, never recomputed per request.** Embeddings are cached on `content_hash`, the mechanism Phase 1 proved when its second run made zero model calls.
- **Tests assert relationships, not scores.** Asserting a specific cosine makes the suite depend on the embedding model, which is how a suite stops being trusted.
- **Never log a full product payload or a raw embedding vector.**
- Per-tenant tables are named `strategist_*` and live in the tenant's own schema, created in `bootstrap_tenant`.
- Match conventions: module-level `logger = logging.getLogger(__name__)`, snake_case, double quotes, 4-space indent, type hints on public signatures, imports stdlib → third-party → `app.*`, `psycopg2` with `sql.SQL`/`sql.Identifier` for interpolated identifiers.
- Tests: `.venv/bin/python -m pytest`. Baselines after Phase 1: `tests/unit/` **546 passing**, `tests/integration/` **70 passing**. Two pre-existing failures under `scripts/` are unrelated and out of scope.
- **Run test suites in the FOREGROUND, one at a time.** Do not start a background run.

---

## File Structure

| File | Responsibility |
|---|---|
| `app/services/database.py` (modify) | Three new per-tenant tables + `migrate_pairing_tables()` |
| `app/services/pairing_rules.py` (new) | The four scoring rules. Pure — no DB, no network. |
| `app/services/pairing_embeddings.py` (new) | Embed and cache on `content_hash`; cosine. |
| `app/services/pairing_candidates.py` (new) | Eligibility and category blocking — who may be compared with whom. |
| `app/services/pairing.py` (new) | The job: load, embed, block, score, write, report. |
| `app/services/pairing_decisions.py` (new) | Read and write human decisions; decide what is servable. |
| `app/services/pairing_queries.py` (new) | Read models for the browse endpoints. |
| `app/api/catalog.py` (new) | The seven routes. |
| `app/main.py` (modify) | Mount the catalog router. |

---

### Task 1: Three tables

**Files:**
- Modify: `app/services/database.py` — `bootstrap_tenant`, plus a new `migrate_pairing_tables`
- Test: `tests/integration/test_pairing_tables.py`

**Interfaces:**
- Consumes: `get_db_connection`, `bootstrap_tenant`.
- Produces: `migrate_pairing_tables(tenant_id: str) -> dict` returning `{"created": [table names]}`.

Three tables, not two. The spec names `product_neighbors` and `pairing_decisions`; the
embedding cache needs somewhere to live too, and putting a 3072-float vector in a column
on `strategist_products` would bloat every catalogue read that never wants it.

- [ ] **Step 1: Write the failing test**

Create `tests/integration/test_pairing_tables.py`:

```python
import pytest

from app.services.database import (
    bootstrap_tenant, get_db_connection, migrate_pairing_tables,
)

TABLES = {"strategist_product_neighbors", "strategist_pairing_decisions",
          "strategist_product_embeddings"}


def _tables(schema):
    conn = get_db_connection()
    try:
        with conn.cursor() as cur:
            cur.execute("SELECT table_name FROM information_schema.tables "
                        "WHERE table_schema = %s", (schema,))
            return {r[0] for r in cur.fetchall()}
    finally:
        conn.close()


def _exec(schema, statement, params=None):
    conn = get_db_connection()
    try:
        with conn.cursor() as cur:
            cur.execute(statement.replace("{S}", f'"{schema}"'), params or ())
        conn.commit()
    finally:
        conn.close()


def test_bootstrap_creates_all_three(temp_tenant):
    bootstrap_tenant(temp_tenant)
    assert TABLES <= _tables(temp_tenant)


def test_migration_is_idempotent(temp_tenant):
    bootstrap_tenant(temp_tenant)
    assert migrate_pairing_tables(temp_tenant)["created"] == []


def test_migration_creates_tables_on_a_tenant_that_predates_them(temp_tenant):
    conn = get_db_connection()
    try:
        with conn.cursor() as cur:
            cur.execute(f'CREATE SCHEMA IF NOT EXISTS "{temp_tenant}"')
        conn.commit()
    finally:
        conn.close()

    assert set(migrate_pairing_tables(temp_tenant)["created"]) == TABLES
    assert TABLES <= _tables(temp_tenant)


def test_a_pair_is_unique_per_type(temp_tenant):
    # The same two products can be both similar and an upsell; they cannot be
    # similar twice. Without pair_type in the key, one type would overwrite
    # another and the graph would silently lose edges.
    bootstrap_tenant(temp_tenant)
    ins = ("INSERT INTO {S}.strategist_product_neighbors "
           "(anchor_key, neighbor_key, pair_type, score, confidence, source) "
           "VALUES ('a','b',%s,0.5,0.5,'embedding')")
    _exec(temp_tenant, ins, ("similar",))
    _exec(temp_tenant, ins, ("upsell",))

    with pytest.raises(Exception):
        _exec(temp_tenant, ins, ("similar",))


def test_a_decision_is_unique_per_pair_and_type(temp_tenant):
    bootstrap_tenant(temp_tenant)
    ins = ("INSERT INTO {S}.strategist_pairing_decisions "
           "(anchor_key, neighbor_key, pair_type, decision) "
           "VALUES ('a','b','similar',%s)")
    _exec(temp_tenant, ins, ("approved",))

    with pytest.raises(Exception):
        _exec(temp_tenant, ins, ("rejected",))


def test_an_embedding_row_is_keyed_by_product(temp_tenant):
    bootstrap_tenant(temp_tenant)
    ins = ("INSERT INTO {S}.strategist_product_embeddings "
           "(product_key, content_hash, vector) VALUES ('a','h',%s)")
    _exec(temp_tenant, ins, ([0.1, 0.2],))

    with pytest.raises(Exception):
        _exec(temp_tenant, ins, ([0.3, 0.4],))
```

- [ ] **Step 2: Run test to verify it fails**

Run: `.venv/bin/python -m pytest tests/integration/test_pairing_tables.py -v`
Expected: FAIL — `migrate_pairing_tables` does not exist.

- [ ] **Step 3: Add the tables to `bootstrap_tenant`**

In `app/services/database.py`, alongside the other `CREATE TABLE IF NOT EXISTS` calls in
`bootstrap_tenant`. Note the brace-doubling convention — these statements go through
`.format()`, so a literal `{}` must be written `{{}}`:

```python
            cur.execute(sql.SQL("""
                CREATE TABLE IF NOT EXISTS {}.strategist_product_neighbors (
                    anchor_key   TEXT NOT NULL,
                    neighbor_key TEXT NOT NULL,
                    pair_type    TEXT NOT NULL,
                    score        REAL NOT NULL,
                    -- Separate from score on purpose: a pair can be a strong
                    -- relationship derived by a weak method. The approval queue
                    -- gates on how much the derivation is trusted, not on how
                    -- good the pair is, so collapsing these would make it
                    -- incoherent.
                    confidence   REAL NOT NULL,
                    source       TEXT NOT NULL,
                    -- What the merchant is shown when asked to approve. A bare
                    -- 0.79 is not reviewable.
                    reasons      JSONB NOT NULL DEFAULT '[]',
                    computed_at  TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
                    PRIMARY KEY (anchor_key, neighbor_key, pair_type)
                )
            """).format(sql.Identifier(tenant_id)))
            cur.execute(sql.SQL(
                "CREATE INDEX IF NOT EXISTS strategist_neighbors_anchor_idx "
                "ON {}.strategist_product_neighbors (anchor_key, pair_type)"
            ).format(sql.Identifier(tenant_id)))
            cur.execute(sql.SQL("""
                CREATE TABLE IF NOT EXISTS {}.strategist_pairing_decisions (
                    anchor_key   TEXT NOT NULL,
                    neighbor_key TEXT NOT NULL,
                    pair_type    TEXT NOT NULL,
                    decision     TEXT NOT NULL,
                    decided_by   TEXT,
                    decided_at   TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
                    PRIMARY KEY (anchor_key, neighbor_key, pair_type)
                )
            """).format(sql.Identifier(tenant_id)))
            cur.execute(sql.SQL("""
                CREATE TABLE IF NOT EXISTS {}.strategist_product_embeddings (
                    product_key  TEXT PRIMARY KEY,
                    content_hash TEXT,
                    vector       REAL[] NOT NULL,
                    computed_at  TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP
                )
            """).format(sql.Identifier(tenant_id)))
```

- [ ] **Step 4: Add the migration**

`bootstrap_tenant` uses `CREATE TABLE IF NOT EXISTS`, so it will not add tables to a tenant
whose schema already exists — the same reason `migrate_products_table` exists. Add
`migrate_pairing_tables` next to it, following its structure: check
`information_schema.tables` first so the returned list reports what was actually created
rather than always claiming all three, then create, commit, and log. Reuse the exact DDL
from Step 3 rather than a paraphrase of it — extract the three statements into a
module-level list of `(table_name, ddl)` pairs that both functions execute, so the two can
never drift apart.

- [ ] **Step 5: Run the tests**

Run: `.venv/bin/python -m pytest tests/integration/test_pairing_tables.py -v`
Expected: PASS, 6 tests

- [ ] **Step 6: Migrate the live tenant**

```bash
PYTHONPATH=. .venv/bin/python -c "
from app.services.database import migrate_pairing_tables
print(migrate_pairing_tables('org_8c32bf3e-6a18-4739-9b1c-94c0cf11125f'))"
```

Run it twice; the second must report `{"created": []}`.

- [ ] **Step 7: Report changed files**

```bash
git status --short
```

Do not commit and do not stage.

---

### Task 2: Eligibility and blocking

**Files:**
- Create: `app/services/pairing_candidates.py`
- Test: `tests/unit/test_pairing_candidates.py`
- Test: `tests/integration/test_pairing_candidates_db.py`

**Interfaces:**
- Consumes: `get_db_connection`.
- Produces:
  - `is_eligible(product: dict) -> bool`
  - `load_eligible(tenant_id: str) -> list[dict]`
  - `blocks(products: list[dict]) -> dict[str, list[dict]]` — leaf category → its products
  - `candidate_pairs(products: list[dict]) -> list[tuple[dict, dict]]` — ordered pairs worth scoring

**Why blocking exists:** 218 products compare in 47,000 pairs and would run fine. 20,000
products compare in 400 million and would not. Only products in the same block, or in
blocks that could plausibly relate, are ever scored.

- [ ] **Step 1: Write the failing test**

Create `tests/unit/test_pairing_candidates.py`:

```python
from app.services.pairing_candidates import (
    blocks, candidate_pairs, is_eligible,
)


def product(key, category="shirts", **kw):
    base = {"product_key": key, "name": key, "category": category,
            "in_stock": True, "missing_fields": [], "attributes": [{"key": "color"}],
            "price_cents": 1000, "is_accessory": False}
    base.update(kw)
    return base


def test_a_normal_product_is_eligible():
    assert is_eligible(product("a")) is True


def test_an_out_of_stock_product_is_not_paired():
    # Recommending something unbuyable is worse than recommending nothing.
    assert is_eligible(product("a", in_stock=False)) is False


def test_a_flagged_product_is_not_paired():
    assert is_eligible(product("a", missing_fields=["product_url"])) is False


def test_a_product_with_no_attributes_is_not_paired():
    # Phase 1 gave every extractable product attributes. One with none has no
    # usable text either, so every score it takes part in would be noise.
    assert is_eligible(product("a", attributes=[])) is False


def test_products_group_by_leaf_category():
    grouped = blocks([product("a", "shirts"), product("b", "shirts"),
                      product("c", "belts")])
    assert set(grouped) == {"shirts", "belts"}
    assert len(grouped["shirts"]) == 2


def test_an_uncategorised_product_gets_its_own_block():
    grouped = blocks([product("a", None)])
    assert len(grouped) == 1


def test_pairs_are_ordered_both_ways():
    # complement and upsell are directional, so (a,b) and (b,a) are different
    # questions and both must be offered to the scorer.
    pairs = candidate_pairs([product("a"), product("b")])
    assert ("a", "b") in [(x["product_key"], y["product_key"]) for x, y in pairs]
    assert ("b", "a") in [(x["product_key"], y["product_key"]) for x, y in pairs]


def test_a_product_is_never_paired_with_itself():
    pairs = candidate_pairs([product("a"), product("b")])
    assert all(x["product_key"] != y["product_key"] for x, y in pairs)


def test_pair_count_is_bounded_by_blocking():
    # 40 products spread over 4 categories must cost far less than 40*39.
    many = [product(f"p{i}", f"cat{i % 4}") for i in range(40)]
    assert len(candidate_pairs(many)) < 40 * 39


def test_products_in_unrelated_blocks_are_still_paired_across():
    # Complements live across categories -- blocking must not make them
    # impossible, only cheaper.
    pairs = candidate_pairs([product("a", "phones"),
                             product("b", "cases", is_accessory=True)])
    assert pairs


def test_two_non_accessories_with_colours_are_paired_across_categories():
    # The outfit case: a shirt and trousers complement each other and neither
    # is an accessory. Without this the weak-complement rule is unreachable and
    # clothing never pairs at all.
    shirt = product("shirt", "shirts",
                    attributes=[{"key": "color", "value": "white"}])
    trousers = product("trousers", "trousers",
                       attributes=[{"key": "color", "value": "navy"}])
    keys = [(a["product_key"], b["product_key"])
            for a, b in candidate_pairs([shirt, trousers])]
    assert ("shirt", "trousers") in keys


def test_colourless_products_are_not_paired_across_categories():
    # Nothing to score them on, so the pair would be pure cost.
    a = product("phone", "phones", attributes=[{"key": "material", "value": "glass"}])
    b = product("bread", "groceries", attributes=[{"key": "material", "value": "flour"}])
    assert candidate_pairs([a, b]) == []
```

- [ ] **Step 2: Run test to verify it fails**

Run: `.venv/bin/python -m pytest tests/unit/test_pairing_candidates.py -v`
Expected: FAIL with `ModuleNotFoundError`

- [ ] **Step 3: Write the implementation**

Create `app/services/pairing_candidates.py`. `is_eligible` and `blocks` are direct. The
design decision is in `candidate_pairs`:

```python
"""Decides which products may be compared with which.

218 products compare in 47,000 pairs and would run fine; 20,000 compare in 400
million and would not. Blocking keeps the job's cost tied to catalog shape rather
than catalog size squared, without making cross-category complements impossible.
"""
import logging

from psycopg2 import sql
from psycopg2.extras import RealDictCursor

from app.services.database import get_db_connection

logger = logging.getLogger(__name__)

UNCATEGORISED = "__uncategorised__"
MAX_CROSS_CATEGORY_PARTNERS = 40


def is_eligible(product: dict) -> bool:
    return bool(
        product.get("in_stock")
        and not product.get("missing_fields")
        and product.get("attributes")
    )


def blocks(products: list) -> dict:
    grouped = {}
    for product in products:
        grouped.setdefault(product.get("category") or UNCATEGORISED, []).append(product)
    return grouped


def candidate_pairs(products: list) -> list:
    """Every ordered pair worth scoring.

    Within a block: everything, both directions -- similar and upsell only ever
    apply here. Across blocks: only where one side is an accessory, because a
    complement is the only cross-category type and it needs that asymmetry
    anyway. Comparing every phone against every grocery item buys nothing.
    """
    grouped = blocks(products)
    pairs = []

    for members in grouped.values():
        for a in members:
            for b in members:
                if a["product_key"] != b["product_key"]:
                    pairs.append((a, b))

    accessories = [p for p in products if p.get("is_accessory")]
    non_accessories = [p for p in products if not p.get("is_accessory")]

    for anchor in non_accessories:
        partners = [a for a in accessories
                    if (a.get("category") or UNCATEGORISED)
                    != (anchor.get("category") or UNCATEGORISED)]
        # Cross-category pairs between two non-accessories are the outfit case:
        # a shirt and trousers complement each other and neither is an
        # accessory. Restricted to products that carry a colour, because that is
        # the only signal the weak-complement rule can score on -- without this
        # the rule would be unreachable and clothing would never pair.
        if _has_colour(anchor):
            partners += [p for p in non_accessories
                         if p["product_key"] != anchor["product_key"]
                         and (p.get("category") or UNCATEGORISED)
                         != (anchor.get("category") or UNCATEGORISED)
                         and _has_colour(p)]

        if len(partners) > MAX_CROSS_CATEGORY_PARTNERS:
            logger.info("Capping %s to %d cross-category partners of %d",
                        anchor["product_key"], MAX_CROSS_CATEGORY_PARTNERS,
                        len(partners))
            partners = partners[:MAX_CROSS_CATEGORY_PARTNERS]
        pairs.extend((anchor, p) for p in partners)

    return pairs


def _has_colour(product: dict) -> bool:
    return any(a.get("key") == "color" and a.get("value")
               for a in (product.get("attributes") or []))


def load_eligible(tenant_id: str) -> list:
    conn = get_db_connection()
    try:
        with conn.cursor(cursor_factory=RealDictCursor) as cur:
            cur.execute(sql.SQL(
                "SELECT product_key, name, description, category, brand, "
                "attributes, price_cents, price_reference_cents, price_tier, "
                "is_accessory, rating, in_stock, missing_fields, content_hash, "
                "tenant_relations "
                "FROM {}.strategist_products"
            ).format(sql.Identifier(tenant_id)))
            rows = [dict(r) for r in cur.fetchall()]
    finally:
        conn.close()

    eligible = [r for r in rows if is_eligible(r)]
    if len(eligible) < len(rows):
        logger.info("%s: pairing %d of %d products; the rest are out of stock, "
                    "flagged, or have no attributes",
                    tenant_id, len(eligible), len(rows))
    return eligible
```

- [ ] **Step 4: Run test to verify it passes**

Run: `.venv/bin/python -m pytest tests/unit/test_pairing_candidates.py -v`
Expected: PASS, 12 tests

- [ ] **Step 5: Write the database test**

Create `tests/integration/test_pairing_candidates_db.py`. Insert four products into a
`temp_tenant` — one normal, one out of stock, one with `missing_fields = '{product_url}'`,
one with empty `attributes` — and assert `load_eligible` returns exactly the first. Use
`bootstrap_tenant(temp_tenant)` and follow the insert helper style from
`tests/integration/test_enrichment_db.py`, which already exists; read it first.

- [ ] **Step 6: Run the database test**

Run: `.venv/bin/python -m pytest tests/integration/test_pairing_candidates_db.py -v`
Expected: PASS

- [ ] **Step 7: Report changed files**

```bash
git status --short
```

Do not commit and do not stage.

---

### Task 3: The four rules

**Files:**
- Create: `app/services/pairing_rules.py`
- Test: `tests/unit/test_pairing_rules.py`

**Interfaces:**
- Consumes: nothing. **Pure — no database, no network, no imports from `app.services` except constants.**
- Produces:
  - `PAIR_TYPES = ("similar", "complement", "upsell", "bundle")`
  - `attribute_jaccard(a: dict, b: dict) -> float`
  - `score_similar(a, b, cosine) -> dict | None`
  - `score_complement(a, b, cosine) -> dict | None`
  - `score_upsell(a, b) -> dict | None`
  - `build_bundles(anchor, complements) -> list[dict]`
  - each scorer returns `{"score", "confidence", "source", "reasons"}` or `None` when the pair is not of that type

This is the heart of the phase and the only part that is genuinely worth testing, because
it is deterministic. Everything else is plumbing around it.

- [ ] **Step 1: Write the failing test**

Create `tests/unit/test_pairing_rules.py`:

```python
import pytest

from app.services.pairing_rules import (
    PAIR_TYPES, attribute_jaccard, build_bundles, score_complement,
    score_similar, score_upsell,
)


def p(key, category="phones", accessory=False, price=50000, rating=4.0,
      tier="mid", attrs=None):
    return {"product_key": key, "name": key, "category": category,
            "is_accessory": accessory, "price_cents": price, "rating": rating,
            "price_tier": tier,
            "attributes": attrs if attrs is not None else [
                {"key": "color", "value": "black"},
                {"key": "material", "value": "glass"}]}


def test_the_four_types():
    assert PAIR_TYPES == ("similar", "complement", "upsell", "bundle")


# --- attribute overlap -------------------------------------------------------

def test_identical_attributes_overlap_fully():
    assert attribute_jaccard(p("a"), p("b")) == 1.0


def test_disjoint_attributes_do_not_overlap():
    other = p("b", attrs=[{"key": "color", "value": "red"},
                          {"key": "material", "value": "cloth"}])
    assert attribute_jaccard(p("a"), other) == 0.0


def test_overlap_is_symmetric():
    a, b = p("a"), p("b", attrs=[{"key": "color", "value": "black"}])
    assert attribute_jaccard(a, b) == attribute_jaccard(b, a)


def test_no_attributes_means_no_overlap_not_a_crash():
    assert attribute_jaccard(p("a", attrs=[]), p("b", attrs=[])) == 0.0


# --- similar -----------------------------------------------------------------

def test_same_category_close_price_is_similar():
    assert score_similar(p("a"), p("b"), cosine=0.9) is not None


def test_a_different_category_is_never_similar():
    # A phone case is not a substitute for a phone, however alike the text.
    assert score_similar(p("a", "phones"), p("b", "cases"), cosine=0.99) is None


def test_a_wildly_different_price_is_not_similar():
    assert score_similar(p("a", price=1000), p("b", price=500000),
                         cosine=0.95) is None


def test_low_text_similarity_is_not_similar():
    assert score_similar(p("a"), p("b"), cosine=0.1) is None


def test_similar_carries_reasons_a_merchant_can_read():
    result = score_similar(p("a"), p("b"), cosine=0.9)
    assert result["reasons"]
    assert all(isinstance(r, str) for r in result["reasons"])


# --- complement --------------------------------------------------------------

def test_a_phone_suggests_a_case():
    anchor, case = p("phone", "phones"), p("case", "cases", accessory=True)
    assert score_complement(anchor, case, cosine=0.4) is not None


def test_a_case_does_not_suggest_a_phone():
    # The asymmetry is the entire point of is_accessory.
    anchor, case = p("phone", "phones"), p("case", "cases", accessory=True)
    assert score_complement(case, anchor, cosine=0.4) is None


def test_two_accessories_are_not_complements():
    a = p("case", "cases", accessory=True)
    b = p("cable", "cables", accessory=True)
    assert score_complement(a, b, cosine=0.5) is None


def test_the_same_category_is_not_a_complement():
    a, b = p("phone1", "phones"), p("phone2", "phones", accessory=True)
    assert score_complement(a, b, cosine=0.5) is None


def test_an_accessory_complement_scores_above_a_weak_one():
    strong = score_complement(p("phone", "phones"),
                              p("case", "cases", accessory=True), cosine=0.4)
    weak = score_complement(p("shirt", "shirts"),
                            p("trousers", "trousers"), cosine=0.4)
    assert weak is None or strong["score"] > weak["score"]


def test_clashing_colours_score_below_matching_ones():
    # The clothing case: a white shirt goes with navy trousers, not orange.
    shirt = p("shirt", "shirts", attrs=[{"key": "color", "value": "white"}])
    navy = p("navy", "trousers", accessory=False,
             attrs=[{"key": "color", "value": "navy"}])
    orange = p("orange", "trousers", accessory=False,
               attrs=[{"key": "color", "value": "orange"}])
    good = score_complement(shirt, navy, cosine=0.4)
    bad = score_complement(shirt, orange, cosine=0.4)
    if good and bad:
        assert good["score"] > bad["score"]


# --- upsell ------------------------------------------------------------------

def test_a_pricier_equally_rated_product_is_an_upsell():
    assert score_upsell(p("a", price=50000, rating=4.0),
                        p("b", price=90000, rating=4.2)) is not None


def test_a_cheaper_product_is_never_an_upsell():
    assert score_upsell(p("a", price=90000), p("b", price=50000)) is None


def test_a_barely_pricier_product_is_not_an_upsell():
    # A 2% difference is noise, not a trade-up.
    assert score_upsell(p("a", price=50000), p("b", price=51000)) is None


def test_a_worse_rated_product_is_never_an_upsell():
    assert score_upsell(p("a", price=50000, rating=4.5),
                        p("b", price=90000, rating=2.0)) is None


def test_an_upsell_across_categories_is_not_an_upsell():
    assert score_upsell(p("a", "phones", price=50000),
                        p("b", "laptops", price=90000)) is None


def test_upsell_is_arithmetic_and_therefore_high_confidence():
    result = score_upsell(p("a", price=50000), p("b", price=90000))
    assert result["source"] == "arithmetic"
    assert result["confidence"] >= 0.9


def test_a_missing_price_cannot_be_an_upsell():
    assert score_upsell(p("a", price=None), p("b", price=90000)) is None


def test_a_missing_rating_does_not_block_an_upsell():
    # 88% of the live catalog has a rating; the rest must still be pairable.
    assert score_upsell(p("a", price=50000, rating=None),
                        p("b", price=90000, rating=None)) is not None


# --- bundle ------------------------------------------------------------------

def test_a_bundle_needs_at_least_two_complements():
    assert build_bundles(p("shirt", "shirts"),
                         [p("belt", "belts", accessory=True)]) == []


def test_a_bundle_never_repeats_a_category():
    anchor = p("shirt", "shirts", price=50000)
    members = [p("belt1", "belts", accessory=True, price=5000),
               p("belt2", "belts", accessory=True, price=6000),
               p("socks", "socks", accessory=True, price=2000)]
    for bundle in build_bundles(anchor, members):
        categories = [m["category"] for m in bundle["members"]]
        assert len(categories) == len(set(categories))


def test_a_bundle_does_not_cost_wildly_more_than_its_anchor():
    # A GBP 59 shirt must not anchor a GBP 900 set.
    anchor = p("shirt", "shirts", price=5900)
    members = [p("watch", "watches", accessory=True, price=90000),
               p("bag", "bags", accessory=True, price=80000)]
    assert build_bundles(anchor, members) == []


def test_a_bundle_is_keyed_on_its_anchor():
    anchor = p("shirt", "shirts", price=50000)
    members = [p("belt", "belts", accessory=True, price=5000),
               p("socks", "socks", accessory=True, price=2000)]
    bundles = build_bundles(anchor, members)
    assert bundles
    assert all(b["anchor_key"] == "shirt" for b in bundles)
```

- [ ] **Step 2: Run test to verify it fails**

Run: `.venv/bin/python -m pytest tests/unit/test_pairing_rules.py -v`
Expected: FAIL with `ModuleNotFoundError`

- [ ] **Step 3: Write the implementation**

Create `app/services/pairing_rules.py`. Keep it pure — this module must import nothing from
`app.services` and touch no connection.

```python
"""The four pairing rules.

Deterministic and dependency-free on purpose: these rules are the part of pairing
worth testing, and a rule that needs a database or a model to exercise stops
being tested honestly.

score and confidence stay separate throughout. A pair can be a strong
relationship derived by a weak method; the approval queue gates on the
derivation, not on the relationship.
"""
import logging

logger = logging.getLogger(__name__)

PAIR_TYPES = ("similar", "complement", "upsell", "bundle")

SIMILAR_MIN_COSINE = 0.55
SIMILAR_MAX_PRICE_RATIO = 3.0
UPSELL_MIN_PRICE_RATIO = 1.15
UPSELL_MAX_RATING_DROP = 0.3
BUNDLE_MIN_MEMBERS = 2
BUNDLE_MAX_MEMBERS = 3
BUNDLE_MAX_PRICE_RATIO = 1.0

# Compatibility is a rule over already-extracted attributes, never a model call:
# the attributes were paid for once in Phase 1 and must not be paid for again.
NEUTRAL_COLORS = {"black", "white", "grey", "gray", "navy", "beige", "cream"}
CLASHING = {frozenset({"orange", "red"}), frozenset({"orange", "pink"}),
            frozenset({"red", "pink"}), frozenset({"brown", "black"})}


def _pairs(product: dict) -> set:
    return {(a.get("key"), str(a.get("value")).lower())
            for a in (product.get("attributes") or [])
            if a.get("key") and a.get("value") is not None}


def attribute_jaccard(a: dict, b: dict) -> float:
    left, right = _pairs(a), _pairs(b)
    if not left or not right:
        return 0.0
    return len(left & right) / len(left | right)


def _same_category(a: dict, b: dict) -> bool:
    return a.get("category") is not None and a.get("category") == b.get("category")


def _price_ratio(a: dict, b: dict):
    low, high = a.get("price_cents"), b.get("price_cents")
    if not low or not high:
        return None
    return max(low, high) / min(low, high)


def _attr(product: dict, key: str):
    for a in product.get("attributes") or []:
        if a.get("key") == key and a.get("value") is not None:
            return str(a["value"]).lower()
    return None


def _colour_compatibility(a: dict, b: dict) -> float:
    """1.0 compatible, 0.5 unknown, 0.0 clashing.

    0.5 rather than 0.0 for unknown: only 52 of 218 live products carry a colour,
    and treating absent data as a clash would suppress most valid pairs.
    """
    left, right = _attr(a, "color"), _attr(b, "color")
    if not left or not right:
        return 0.5
    if left == right or left in NEUTRAL_COLORS or right in NEUTRAL_COLORS:
        return 1.0
    return 0.0 if frozenset({left, right}) in CLASHING else 0.7


def score_similar(a: dict, b: dict, cosine: float):
    if not _same_category(a, b) or cosine < SIMILAR_MIN_COSINE:
        return None
    ratio = _price_ratio(a, b)
    if ratio is not None and ratio > SIMILAR_MAX_PRICE_RATIO:
        return None

    overlap = attribute_jaccard(a, b)
    score = 0.7 * cosine + 0.3 * overlap
    reasons = [f"same category ({a.get('category')})",
               f"{cosine:.2f} text similarity",
               f"{overlap:.2f} attribute overlap"]
    return {"score": round(score, 4), "confidence": round(cosine, 4),
            "source": "embedding", "reasons": reasons}


def score_complement(a: dict, b: dict, cosine: float):
    """Directional: a suggests b. Never the reverse unless scored separately."""
    if _same_category(a, b):
        return None
    if a.get("is_accessory"):
        return None

    compatibility = _colour_compatibility(a, b)
    if b.get("is_accessory"):
        score = 0.6 + 0.2 * compatibility + 0.2 * cosine
        confidence = 0.75
        reasons = [f"{b.get('category')} is an accessory",
                   f"different category from {a.get('category')}"]
    else:
        if compatibility < 0.7:
            return None
        score = 0.3 + 0.3 * compatibility + 0.2 * cosine
        confidence = 0.4
        reasons = [f"different categories ({a.get('category')} / {b.get('category')})",
                   "compatible attributes"]

    if compatibility < 1.0 and _attr(a, "color") and _attr(b, "color"):
        reasons.append(f"colour {_attr(a, 'color')} with {_attr(b, 'color')}")
    return {"score": round(min(score, 1.0), 4), "confidence": confidence,
            "source": "attribute", "reasons": reasons}


def score_upsell(a: dict, b: dict):
    """Directional: b is the trade-up from a. Arithmetic, never a judgement."""
    if not _same_category(a, b):
        return None
    low, high = a.get("price_cents"), b.get("price_cents")
    if not low or not high or high / low < UPSELL_MIN_PRICE_RATIO:
        return None

    a_rating, b_rating = a.get("rating"), b.get("rating")
    if a_rating is not None and b_rating is not None:
        if b_rating < a_rating - UPSELL_MAX_RATING_DROP:
            return None

    lift = min((high / low - 1) / 2, 1.0)
    reasons = [f"{high / low:.1f}x the price", f"same category ({a.get('category')})"]
    if b_rating is not None:
        reasons.append(f"rated {b_rating}")
    # Arithmetic over two columns leaves nothing to be uncertain about, so this
    # never reaches the approval queue. That is the threshold rule applying
    # normally, not an exemption.
    return {"score": round(0.5 + 0.5 * lift, 4), "confidence": 0.95,
            "source": "arithmetic", "reasons": reasons}


def build_bundles(anchor: dict, complements: list) -> list:
    """An anchor plus complements from distinct categories -- the outfit case."""
    anchor_price = anchor.get("price_cents")
    chosen, seen = [], {anchor.get("category")}
    for member in sorted(complements, key=lambda m: m.get("price_cents") or 0):
        category = member.get("category")
        if category in seen:
            continue
        seen.add(category)
        chosen.append(member)
        if len(chosen) == BUNDLE_MAX_MEMBERS:
            break

    if len(chosen) < BUNDLE_MIN_MEMBERS:
        return []

    total = sum(m.get("price_cents") or 0 for m in chosen)
    if anchor_price and total > anchor_price * BUNDLE_MAX_PRICE_RATIO:
        return []

    return [{"anchor_key": anchor["product_key"], "members": chosen,
             "score": round(min(0.5 + 0.1 * len(chosen), 1.0), 4),
             "confidence": 0.5, "source": "attribute",
             "reasons": [f"{len(chosen)} complements from distinct categories",
                         f"set costs {total} against anchor {anchor_price}"]}]
```

- [ ] **Step 4: Run test to verify it passes**

Run: `.venv/bin/python -m pytest tests/unit/test_pairing_rules.py -v`
Expected: PASS, 26 tests

If a threshold constant makes a test fail, **change the constant, not the test** — the tests
encode the relationships the spec requires, and a rule that cannot satisfy them is the
thing that is wrong. Record any constant you changed in your report.

- [ ] **Step 5: Report changed files**

```bash
git status --short
```

Do not commit and do not stage.

---

### Task 4: Embeddings, cached

**Files:**
- Create: `app/services/pairing_embeddings.py`
- Test: `tests/unit/test_pairing_embeddings.py`
- Test: `tests/integration/test_pairing_embeddings_db.py`

**Interfaces:**
- Consumes: `app.services.products._embed_batch` and `_cosine` (import only — do not modify that file), `get_db_connection`.
- Produces:
  - `embedding_text(product: dict) -> str`
  - `load_cached(tenant_id: str) -> dict[str, tuple[str, list[float]]]`
  - `store(tenant_id: str, product_key: str, content_hash, vector: list) -> None`
  - `vectors_for(tenant_id: str, products: list[dict]) -> dict[str, list[float]]` — cached where possible, embedding only the rest
  - `cosine(a, b) -> float`

**The caching contract is the same one Phase 1 proved:** a product is re-embedded only when
its `content_hash` changed. Phase 1's second run made zero model calls and this must too.

The `NO_CONTENT_HASH` trap from Phase 1 applies here verbatim: crawled rows carry a NULL
`content_hash`, and comparing NULL to NULL yields NULL rather than "unchanged", so a naive
check re-embeds them on every run forever. Compare with `COALESCE` on both sides, or reuse
`app.services.enrichment.NO_CONTENT_HASH` and store that sentinel.

- [ ] **Step 1: Write the failing test**

Create `tests/unit/test_pairing_embeddings.py`:

```python
from app.services.pairing_embeddings import cosine, embedding_text


def test_text_combines_title_and_description():
    text = embedding_text({"name": "Oxford Shirt",
                           "description": "A white cotton shirt.",
                           "category": "shirts", "attributes": []})
    assert "Oxford Shirt" in text
    assert "white cotton" in text


def test_text_includes_category_and_attributes():
    # The attributes Phase 1 extracted are the whole reason similarity improved;
    # leaving them out of the embedded text would waste them.
    text = embedding_text({"name": "Shirt", "description": None,
                           "category": "shirts",
                           "attributes": [{"key": "color", "value": "white"}]})
    assert "shirts" in text
    assert "white" in text


def test_text_is_stable_for_the_same_product():
    product = {"name": "Shirt", "description": "d", "category": "c",
               "attributes": [{"key": "b", "value": "2"},
                              {"key": "a", "value": "1"}]}
    assert embedding_text(product) == embedding_text(dict(product))


def test_text_does_not_depend_on_attribute_order():
    # Attribute order is not stable across runs; an order-sensitive text would
    # re-embed the whole catalog for nothing.
    one = embedding_text({"name": "S", "description": "", "category": "c",
                          "attributes": [{"key": "a", "value": "1"},
                                         {"key": "b", "value": "2"}]})
    two = embedding_text({"name": "S", "description": "", "category": "c",
                          "attributes": [{"key": "b", "value": "2"},
                                         {"key": "a", "value": "1"}]})
    assert one == two


def test_cosine_of_identical_vectors_is_one():
    assert cosine([1.0, 0.0], [1.0, 0.0]) == 1.0


def test_cosine_of_a_zero_vector_is_zero_not_a_crash():
    assert cosine([0.0, 0.0], [1.0, 0.0]) == 0.0
```

- [ ] **Step 2: Run test to verify it fails**

Run: `.venv/bin/python -m pytest tests/unit/test_pairing_embeddings.py -v`
Expected: FAIL with `ModuleNotFoundError`

- [ ] **Step 3: Write the implementation**

Create `app/services/pairing_embeddings.py`:

```python
"""Embeds products once and caches the vectors on content_hash.

Phase 1 proved this mechanism: its second run over 218 products made zero model
calls. Re-pairing after a small catalog change must re-embed only what changed.
"""
import logging

from psycopg2 import sql

from app.services.database import get_db_connection
from app.services.enrichment import NO_CONTENT_HASH
from app.services.products import _cosine, _embed_batch

logger = logging.getLogger(__name__)


def cosine(a: list, b: list) -> float:
    return _cosine(a, b)


def embedding_text(product: dict) -> str:
    """Sorted, so attribute order -- which is not stable across runs -- cannot
    change the text and force a needless re-embed."""
    attributes = sorted(
        f"{a.get('key')} {a.get('value')}"
        for a in (product.get("attributes") or [])
        if a.get("key") and a.get("value") is not None)
    parts = [product.get("name") or "", product.get("description") or "",
             product.get("category") or "", " ".join(attributes)]
    return ". ".join(p.strip() for p in parts if p and p.strip())


def load_cached(tenant_id: str) -> dict:
    conn = get_db_connection()
    try:
        with conn.cursor() as cur:
            cur.execute(sql.SQL(
                "SELECT product_key, content_hash, vector "
                "FROM {}.strategist_product_embeddings"
            ).format(sql.Identifier(tenant_id)))
            return {key: (hash_, list(vector))
                    for key, hash_, vector in cur.fetchall()}
    finally:
        conn.close()


def store(tenant_id: str, product_key: str, content_hash, vector: list) -> None:
    conn = get_db_connection()
    try:
        with conn.cursor() as cur:
            cur.execute(sql.SQL(
                "INSERT INTO {}.strategist_product_embeddings "
                "(product_key, content_hash, vector) VALUES (%s, %s, %s) "
                "ON CONFLICT (product_key) DO UPDATE SET "
                "content_hash = EXCLUDED.content_hash, "
                "vector = EXCLUDED.vector, computed_at = CURRENT_TIMESTAMP"
            ).format(sql.Identifier(tenant_id)),
                (product_key, content_hash or NO_CONTENT_HASH, vector))
        conn.commit()
    except Exception:
        conn.rollback()
        logger.error("Could not cache an embedding for %s", product_key,
                     exc_info=True)
    finally:
        conn.close()


def vectors_for(tenant_id: str, products: list) -> dict:
    cached = load_cached(tenant_id)
    vectors, stale = {}, []

    for product in products:
        key = product["product_key"]
        current = product.get("content_hash") or NO_CONTENT_HASH
        entry = cached.get(key)
        # COALESCE-style comparison on both sides: crawled rows carry no
        # content_hash, and a raw NULL comparison would re-embed them every run.
        if entry and entry[0] == current and entry[1]:
            vectors[key] = entry[1]
        else:
            stale.append(product)

    if stale:
        computed = _embed_batch([embedding_text(p) for p in stale],
                                tenant_id=tenant_id)
        if len(computed) != len(stale):
            logger.error("Embedding failed for %s; %d products will be scored "
                         "on attributes alone", tenant_id, len(stale))
            return vectors
        for product, vector in zip(stale, computed):
            vectors[product["product_key"]] = vector
            store(tenant_id, product["product_key"],
                  product.get("content_hash"), vector)

    logger.info("%s: %d embeddings reused, %d computed",
                tenant_id, len(products) - len(stale), len(stale))
    return vectors
```

- [ ] **Step 4: Run test to verify it passes**

Run: `.venv/bin/python -m pytest tests/unit/test_pairing_embeddings.py -v`
Expected: PASS, 6 tests

- [ ] **Step 5: Write the caching test**

Create `tests/integration/test_pairing_embeddings_db.py`. Monkeypatch
`pairing_embeddings._embed_batch` with a counter so no real API call happens, then assert:

```python
def test_a_second_run_embeds_nothing(temp_tenant, monkeypatch):
    # The whole point of the cache. Phase 1's second enrich made zero model
    # calls; this must too.
    import app.services.pairing_embeddings as mod
    calls = {"n": 0}

    def fake(texts, tenant_id=None, user_id=None):
        calls["n"] += 1
        return [[1.0, 0.0] for _ in texts]

    monkeypatch.setattr(mod, "_embed_batch", fake)
    bootstrap_tenant(temp_tenant)
    products = [{"product_key": "a", "name": "A", "description": "d",
                 "category": "c", "attributes": [], "content_hash": "h1"}]

    mod.vectors_for(temp_tenant, products)
    mod.vectors_for(temp_tenant, products)
    assert calls["n"] == 1


def test_a_crawled_product_with_no_content_hash_embeds_only_once(temp_tenant,
                                                                 monkeypatch):
    # The Phase 1 trap: NULL content_hash compared to NULL is NULL, not
    # "unchanged", so a naive check re-embeds and re-bills forever.
    ...same shape, with content_hash=None, asserting calls["n"] == 1


def test_a_changed_content_hash_re_embeds(temp_tenant, monkeypatch):
    ...same shape, second call with content_hash="h2", asserting calls["n"] == 2
```

Write all three in full, following the first one's structure.

- [ ] **Step 6: Run the caching test**

Run: `.venv/bin/python -m pytest tests/integration/test_pairing_embeddings_db.py -v`
Expected: PASS, 3 tests

- [ ] **Step 7: Report changed files**

```bash
git status --short
```

Do not commit and do not stage.

---

### Task 5: Decisions and servability

**Files:**
- Create: `app/services/pairing_decisions.py`
- Test: `tests/integration/test_pairing_decisions.py`

**Interfaces:**
- Consumes: `get_db_connection`.
- Produces:
  - `APPROVAL_THRESHOLD: float` = 0.7
  - `record(tenant_id: str, decisions: list[dict]) -> int`
  - `load(tenant_id: str) -> dict[tuple, str]` keyed `(anchor_key, neighbor_key, pair_type)`
  - `is_servable(pair: dict, decision: str | None) -> bool`

**This module is the only writer of `strategist_pairing_decisions`.** The job never touches
it. That separation is what makes a re-run safe.

`APPROVAL_THRESHOLD` starts at 0.7 and is **re-tuned in Task 8 against the live score
distribution**, with the measured number recorded. Do not treat 0.7 as final.

- [ ] **Step 1: Write the failing test**

Create `tests/integration/test_pairing_decisions.py`:

```python
import pytest

from app.services.database import bootstrap_tenant
from app.services.pairing_decisions import (
    APPROVAL_THRESHOLD, is_servable, load, record,
)


def pair(confidence=0.9, source="embedding"):
    return {"anchor_key": "a", "neighbor_key": "b", "pair_type": "similar",
            "confidence": confidence, "source": source}


def test_a_confident_pair_is_servable_without_approval():
    assert is_servable(pair(confidence=APPROVAL_THRESHOLD), None) is True


def test_an_unsure_pair_is_not_servable_until_approved():
    assert is_servable(pair(confidence=0.2), None) is False
    assert is_servable(pair(confidence=0.2), "approved") is True


def test_a_rejected_pair_is_never_servable_however_confident():
    # A merchant saying no is information the scorer does not have.
    assert is_servable(pair(confidence=0.99), "rejected") is False


def test_a_merchant_declared_pair_never_queues():
    assert is_servable(pair(confidence=0.0, source="merchant"), None) is True


def test_decisions_round_trip(temp_tenant):
    bootstrap_tenant(temp_tenant)
    assert record(temp_tenant, [
        {"anchor_key": "a", "neighbor_key": "b", "pair_type": "similar",
         "decision": "approved", "decided_by": "someone@example.com"}]) == 1
    assert load(temp_tenant)[("a", "b", "similar")] == "approved"


def test_a_decision_can_be_changed(temp_tenant):
    bootstrap_tenant(temp_tenant)
    base = {"anchor_key": "a", "neighbor_key": "b", "pair_type": "similar"}
    record(temp_tenant, [dict(base, decision="approved")])
    record(temp_tenant, [dict(base, decision="rejected")])
    assert load(temp_tenant)[("a", "b", "similar")] == "rejected"


def test_a_batch_is_recorded_in_one_call(temp_tenant):
    # A merchant clearing a queue decides many pairs in one sitting.
    bootstrap_tenant(temp_tenant)
    assert record(temp_tenant, [
        {"anchor_key": "a", "neighbor_key": f"b{i}", "pair_type": "similar",
         "decision": "approved"} for i in range(5)]) == 5


def test_an_invalid_decision_is_refused(temp_tenant):
    bootstrap_tenant(temp_tenant)
    with pytest.raises(ValueError):
        record(temp_tenant, [{"anchor_key": "a", "neighbor_key": "b",
                              "pair_type": "similar", "decision": "maybe"}])


def test_a_decision_survives_for_a_pair_that_does_not_exist_yet(temp_tenant):
    # A re-run may recreate a pair, and the decision must already be waiting.
    bootstrap_tenant(temp_tenant)
    record(temp_tenant, [{"anchor_key": "ghost", "neighbor_key": "x",
                          "pair_type": "similar", "decision": "rejected"}])
    assert load(temp_tenant)[("ghost", "x", "similar")] == "rejected"
```

- [ ] **Step 2: Run test to verify it fails**

Run: `.venv/bin/python -m pytest tests/integration/test_pairing_decisions.py -v`
Expected: FAIL with `ModuleNotFoundError`

- [ ] **Step 3: Write the implementation**

Create `app/services/pairing_decisions.py`. `record` validates `decision` against
`{"approved", "rejected"}` and raises `ValueError` otherwise, then upserts with
`ON CONFLICT (anchor_key, neighbor_key, pair_type) DO UPDATE`. `load` returns the whole
tenant's decisions as a dict — a pairing run needs all of them and one query beats one per
pair. `is_servable` implements the §5 rule exactly:

```python
def is_servable(pair: dict, decision) -> bool:
    if decision == "rejected":
        return False
    if decision == "approved":
        return True
    # A merchant's own declaration outranks any inference, so it never queues.
    if pair.get("source") == "merchant":
        return True
    return (pair.get("confidence") or 0.0) >= APPROVAL_THRESHOLD
```

- [ ] **Step 4: Run the tests**

Run: `.venv/bin/python -m pytest tests/integration/test_pairing_decisions.py -v`
Expected: PASS, 9 tests

- [ ] **Step 5: Report changed files**

```bash
git status --short
```

Do not commit and do not stage.

---

### Task 6: The job

**Files:**
- Create: `app/services/pairing.py`
- Test: `tests/unit/test_pairing.py`
- Test: `tests/integration/test_pairing_job.py`

**Interfaces:**
- Consumes: everything from Tasks 2–5, plus `migrate_pairing_tables`.
- Produces:
  - `MAX_NEIGHBORS_PER_TYPE: int` = 8
  - `score_pair(a, b, vectors) -> list[dict]` — every type that applies to this ordered pair
  - `merchant_pairs(products: list[dict]) -> list[dict]` — from `tenant_relations`
  - `pair_tenant(tenant_id: str) -> dict` — the report

- [ ] **Step 1: Write the failing test**

Create `tests/unit/test_pairing.py`:

```python
from app.services.pairing import MAX_NEIGHBORS_PER_TYPE, merchant_pairs, score_pair


def p(key, category="phones", accessory=False, price=50000, relations=None):
    return {"product_key": key, "name": key, "category": category,
            "is_accessory": accessory, "price_cents": price, "rating": 4.0,
            "attributes": [{"key": "color", "value": "black"}],
            "tenant_relations": relations or {}}


VECTORS = {"a": [1.0, 0.0], "b": [1.0, 0.0], "case": [0.7, 0.7]}


def test_one_pair_can_carry_several_types():
    # The same two products can be both similar and an upsell.
    results = score_pair(p("a", price=50000), p("b", price=90000), VECTORS)
    assert {r["pair_type"] for r in results} >= {"similar", "upsell"}


def test_an_unrelated_pair_yields_nothing():
    assert score_pair(p("a", "phones"), p("b", "groceries"), VECTORS) == []


def test_every_result_carries_the_full_shape():
    for result in score_pair(p("a", price=50000), p("b", price=90000), VECTORS):
        assert set(result) >= {"anchor_key", "neighbor_key", "pair_type",
                               "score", "confidence", "source", "reasons"}


def test_a_missing_vector_does_not_crash_the_pair():
    # An embedding failure degrades to attribute-only scoring, not to an abort.
    assert isinstance(score_pair(p("x"), p("y"), {}), list)


def test_merchant_relations_become_high_confidence_complements():
    products = [p("a", relations={"complementary": ["b"]}), p("b")]
    pairs = merchant_pairs(products)
    assert len(pairs) == 1
    assert pairs[0]["source"] == "merchant"
    assert pairs[0]["confidence"] == 1.0
    assert pairs[0]["pair_type"] == "complement"


def test_a_merchant_relation_to_a_missing_product_is_dropped():
    # The referenced product may have been deleted or filtered as ineligible.
    assert merchant_pairs([p("a", relations={"complementary": ["ghost"]})]) == []


def test_neighbors_are_capped_per_type():
    assert MAX_NEIGHBORS_PER_TYPE == 8
```

- [ ] **Step 2: Run test to verify it fails**

Run: `.venv/bin/python -m pytest tests/unit/test_pairing.py -v`
Expected: FAIL with `ModuleNotFoundError`

- [ ] **Step 3: Write the implementation**

Create `app/services/pairing.py`:

```python
"""Builds a tenant's pairing graph.

The job owns strategist_product_neighbors and rewrites it freely. It never writes
to strategist_pairing_decisions -- a merchant who rejects a pair and then
re-syncs must not see that pair return.
"""
import logging

from psycopg2 import sql
from psycopg2.extras import Json, execute_values

from app.services.database import get_db_connection, migrate_pairing_tables
from app.services.pairing_candidates import candidate_pairs, load_eligible
from app.services.pairing_decisions import is_servable, load as load_decisions
from app.services.pairing_embeddings import cosine, vectors_for
from app.services.pairing_rules import (
    PAIR_TYPES, build_bundles, score_complement, score_similar, score_upsell,
)

logger = logging.getLogger(__name__)

MAX_NEIGHBORS_PER_TYPE = 8


def score_pair(a: dict, b: dict, vectors: dict) -> list:
    left, right = vectors.get(a["product_key"]), vectors.get(b["product_key"])
    # An embedding failure degrades a pair to attribute-only scoring rather than
    # aborting the run: a partial graph beats no graph.
    similarity = cosine(left, right) if left and right else 0.0

    results = []
    for pair_type, scored in (
        ("similar", score_similar(a, b, similarity)),
        ("complement", score_complement(a, b, similarity)),
        ("upsell", score_upsell(a, b)),
    ):
        if scored:
            results.append({"anchor_key": a["product_key"],
                            "neighbor_key": b["product_key"],
                            "pair_type": pair_type, **scored})
    return results


def merchant_pairs(products: list) -> list:
    """A merchant's own declarations outrank every inference in this module."""
    known = {p["product_key"] for p in products}
    pairs = []
    for product in products:
        declared = (product.get("tenant_relations") or {}).get("complementary") or []
        for neighbor in declared:
            if neighbor not in known:
                continue
            pairs.append({"anchor_key": product["product_key"],
                          "neighbor_key": neighbor, "pair_type": "complement",
                          "score": 1.0, "confidence": 1.0, "source": "merchant",
                          "reasons": ["declared by the merchant"]})
    return pairs


def _top_per_anchor_and_type(pairs: list) -> list:
    grouped = {}
    for pair in pairs:
        grouped.setdefault((pair["anchor_key"], pair["pair_type"]), []).append(pair)

    kept = []
    for members in grouped.values():
        members.sort(key=lambda p: p["score"], reverse=True)
        kept.extend(members[:MAX_NEIGHBORS_PER_TYPE])
    return kept


def replace_neighbors(tenant_id: str, pairs: list) -> None:
    """Replaces this tenant's whole graph in one transaction.

    Delete-then-insert rather than upsert: a pair that no longer scores must
    disappear, and an upsert would leave last run's edges behind forever.
    """
    conn = get_db_connection()
    try:
        with conn.cursor() as cur:
            cur.execute(sql.SQL(
                "DELETE FROM {}.strategist_product_neighbors"
            ).format(sql.Identifier(tenant_id)))
            if pairs:
                execute_values(cur, sql.SQL(
                    "INSERT INTO {}.strategist_product_neighbors "
                    "(anchor_key, neighbor_key, pair_type, score, confidence, "
                    "source, reasons) VALUES %s"
                ).format(sql.Identifier(tenant_id)).as_string(cur),
                    [(p["anchor_key"], p["neighbor_key"], p["pair_type"],
                      p["score"], p["confidence"], p["source"],
                      Json(p["reasons"])) for p in pairs])
        conn.commit()
    except Exception:
        conn.rollback()
        logger.error("Could not write the pairing graph for %s", tenant_id,
                     exc_info=True)
        raise
    finally:
        conn.close()


def pair_tenant(tenant_id: str) -> dict:
    migrate_pairing_tables(tenant_id)
    products = load_eligible(tenant_id)
    if len(products) < 2:
        return {"products": len(products), "pairs": 0, "servable": 0,
                "queued": 0, "by_type": {}}

    vectors = vectors_for(tenant_id, products)

    scored = []
    for a, b in candidate_pairs(products):
        scored.extend(score_pair(a, b, vectors))

    by_key = {p["product_key"]: p for p in products}
    for anchor in products:
        complements = [by_key[p["neighbor_key"]] for p in scored
                       if p["anchor_key"] == anchor["product_key"]
                       and p["pair_type"] == "complement"
                       and p["neighbor_key"] in by_key]
        for bundle in build_bundles(anchor, complements):
            for member in bundle["members"]:
                scored.append({"anchor_key": bundle["anchor_key"],
                               "neighbor_key": member["product_key"],
                               "pair_type": "bundle", "score": bundle["score"],
                               "confidence": bundle["confidence"],
                               "source": bundle["source"],
                               "reasons": bundle["reasons"]})

    # Merchant declarations are added last and deduplicated in favour of
    # themselves: their confidence of 1.0 must never be lowered by an inferred
    # duplicate of the same edge.
    declared = merchant_pairs(products)
    declared_keys = {(p["anchor_key"], p["neighbor_key"], p["pair_type"])
                     for p in declared}
    scored = [p for p in scored
              if (p["anchor_key"], p["neighbor_key"], p["pair_type"])
              not in declared_keys] + declared

    kept = _top_per_anchor_and_type(scored)
    replace_neighbors(tenant_id, kept)

    decisions = load_decisions(tenant_id)
    servable = sum(1 for p in kept if is_servable(
        p, decisions.get((p["anchor_key"], p["neighbor_key"], p["pair_type"]))))

    by_type = {t: sum(1 for p in kept if p["pair_type"] == t) for t in PAIR_TYPES}
    report = {"products": len(products), "pairs": len(kept),
              "servable": servable, "queued": len(kept) - servable,
              "by_type": by_type}
    logger.info("Paired %s: %s", tenant_id, report)
    return report
```

- [ ] **Step 4: Run the unit test**

Run: `.venv/bin/python -m pytest tests/unit/test_pairing.py -v`
Expected: PASS, 7 tests

- [ ] **Step 5: Write the invariant test**

Create `tests/integration/test_pairing_job.py`. Insert a small catalogue into a
`temp_tenant` — a phone, a second pricier phone, and a phone case marked
`is_accessory = true` — monkeypatch `pairing.vectors_for` to return fixed vectors so no API
call happens, then assert:

```python
def test_a_rejected_pair_does_not_return_after_a_re_run(temp_tenant, monkeypatch):
    # THE most important test in this phase. Re-running must never overrule a
    # human. If this passes for the wrong reason the whole approval flow is a lie.
    ...run pair_tenant, record a rejection, run pair_tenant again,
    ...assert the decision is still 'rejected' and the pair is not servable


def test_a_re_run_does_not_duplicate_rows(temp_tenant, monkeypatch):
    ...run pair_tenant twice, assert the row count is identical


def test_a_pair_that_stops_scoring_disappears(temp_tenant, monkeypatch):
    ...run, mark a product out of stock, re-run, assert its rows are gone


def test_a_phone_gets_a_case_and_the_case_does_not_get_a_phone(temp_tenant,
                                                               monkeypatch):
    ...assert a complement row exists anchor=phone neighbor=case,
    ...and none exists anchor=case neighbor=phone


def test_an_out_of_stock_product_is_never_paired(temp_tenant, monkeypatch):
    ...assert it appears as neither anchor nor neighbor
```

Write all five in full.

- [ ] **Step 6: Run the invariant test**

Run: `.venv/bin/python -m pytest tests/integration/test_pairing_job.py -v`
Expected: PASS, 5 tests

- [ ] **Step 7: Report changed files**

```bash
git status --short
```

Do not commit and do not stage.

---

### Task 7: The seven endpoints

**Files:**
- Create: `app/services/pairing_queries.py`
- Create: `app/api/catalog.py`
- Modify: `app/main.py`
- Test: `tests/integration/test_catalog_routes.py`

**Interfaces:**
- Consumes: `pair_tenant`, `pairing_decisions.record`/`load`/`is_servable`, `get_db_connection`.
- Produces:
  - `list_categories(tenant_id) -> list[dict]`
  - `list_products(tenant_id, category=None, limit=50, offset=0) -> dict`
  - `get_product(tenant_id, product_key) -> dict | None`
  - `pairings_for(tenant_id, product_key) -> dict` — grouped by type, each pair carrying `servable`
  - `pending_pairings(tenant_id, limit=100) -> list[dict]`
  - the router at `app/api/catalog.py`

```
POST   /catalog/pair
GET    /catalog/categories
GET    /catalog/products
GET    /catalog/products/{key}
GET    /catalog/products/{key}/pairings
GET    /catalog/pairings/pending
POST   /catalog/pairings/decide
```

`POST /catalog/pair` and `POST /catalog/pairings/decide` take `Form(...)` fields, matching
`/catalog/sync` and `/catalog/enrich` in `app/api/endpoints.py` — read those two routes
first and mirror them, including pushing blocking work off the event loop and returning
`502` with only the exception class name on failure. The `GET` routes take `tenant_id` as a
query parameter.

`decide` accepts a JSON body of many decisions, because a merchant clearing a queue decides
many pairs in one sitting and one request per click makes the screen feel broken.

**Route-ordering trap:** `/catalog/pairings/pending` must be registered before any
`/catalog/products/{key}` style path that could shadow it, and `{key}` values contain a
colon (`shopify:123`), so confirm a real product key round-trips through the path.

- [ ] **Step 1: Write the failing test**

Create `tests/integration/test_catalog_routes.py`:

```python
import pytest
from fastapi.testclient import TestClient

from app.main import app

client = TestClient(app)
TENANT = "org_8c32bf3e-6a18-4739-9b1c-94c0cf11125f"


@pytest.mark.parametrize("path", [
    "/catalog/pair", "/catalog/categories", "/catalog/products",
    "/catalog/pairings/pending", "/catalog/pairings/decide",
])
def test_routes_are_registered(path):
    assert path in set(app.openapi()["paths"])


def test_pair_requires_a_tenant():
    assert client.post("/catalog/pair", data={}).status_code == 422


def test_pair_returns_the_report(monkeypatch):
    from app.api import catalog
    monkeypatch.setattr(catalog, "pair_tenant", lambda t: {
        "products": 218, "pairs": 900, "servable": 700, "queued": 200,
        "by_type": {"similar": 400}})
    resp = client.post("/catalog/pair", data={"tenant_id": TENANT})
    assert resp.status_code == 200
    assert resp.json()["queued"] == 200


def test_a_job_failure_returns_502_with_no_detail(monkeypatch):
    from app.api import catalog

    def boom(tenant_id):
        raise RuntimeError("embedding service unavailable")

    monkeypatch.setattr(catalog, "pair_tenant", boom)
    resp = client.post("/catalog/pair", data={"tenant_id": TENANT})
    assert resp.status_code == 502
    assert "unavailable" not in resp.text
    assert "RuntimeError" in resp.text


def test_categories_are_listed(monkeypatch):
    from app.api import catalog
    monkeypatch.setattr(catalog, "list_categories", lambda t: [
        {"category": "smartphones", "products": 16}])
    resp = client.get("/catalog/categories", params={"tenant_id": TENANT})
    assert resp.json()[0]["products"] == 16


def test_products_are_filterable_by_category(monkeypatch):
    from app.api import catalog
    seen = {}
    monkeypatch.setattr(catalog, "list_products",
                        lambda t, category=None, limit=50, offset=0:
                        seen.update(category=category) or {"products": [],
                                                           "total": 0})
    client.get("/catalog/products",
               params={"tenant_id": TENANT, "category": "smartphones"})
    assert seen["category"] == "smartphones"


def test_a_missing_product_is_404(monkeypatch):
    from app.api import catalog
    monkeypatch.setattr(catalog, "get_product", lambda t, k: None)
    resp = client.get("/catalog/products/nope", params={"tenant_id": TENANT})
    assert resp.status_code == 404


def test_a_product_key_containing_a_colon_survives_the_path(monkeypatch):
    # Every key in this system looks like "shopify:123".
    from app.api import catalog
    seen = {}
    monkeypatch.setattr(catalog, "get_product",
                        lambda t, k: seen.update(key=k) or {"product_key": k})
    client.get("/catalog/products/shopify:123", params={"tenant_id": TENANT})
    assert seen["key"] == "shopify:123"


def test_pairings_are_grouped_by_type(monkeypatch):
    from app.api import catalog
    monkeypatch.setattr(catalog, "pairings_for", lambda t, k: {
        "similar": [{"neighbor_key": "b", "score": 0.9, "servable": True,
                     "reasons": ["same category"]}],
        "complement": [], "upsell": [], "bundle": []})
    resp = client.get("/catalog/products/a/pairings", params={"tenant_id": TENANT})
    assert set(resp.json()) == {"similar", "complement", "upsell", "bundle"}


def test_the_pending_queue_is_reachable(monkeypatch):
    # Registered before /catalog/products/{key} could shadow it.
    from app.api import catalog
    monkeypatch.setattr(catalog, "pending_pairings", lambda t, limit=100: [])
    assert client.get("/catalog/pairings/pending",
                      params={"tenant_id": TENANT}).status_code == 200


def test_decisions_are_accepted_as_a_batch(monkeypatch):
    from app.api import catalog
    seen = {}
    monkeypatch.setattr(catalog, "record",
                        lambda t, decisions: seen.update(n=len(decisions))
                        or len(decisions))
    resp = client.post("/catalog/pairings/decide", json={
        "tenant_id": TENANT,
        "decisions": [{"anchor_key": "a", "neighbor_key": f"b{i}",
                       "pair_type": "similar", "decision": "approved"}
                      for i in range(3)]})
    assert resp.status_code == 200
    assert seen["n"] == 3


def test_an_invalid_decision_is_rejected_with_422(monkeypatch):
    from app.api import catalog

    def boom(tenant_id, decisions):
        raise ValueError("maybe is not a decision")

    monkeypatch.setattr(catalog, "record", boom)
    resp = client.post("/catalog/pairings/decide", json={
        "tenant_id": TENANT,
        "decisions": [{"anchor_key": "a", "neighbor_key": "b",
                       "pair_type": "similar", "decision": "maybe"}]})
    assert resp.status_code == 422
```

- [ ] **Step 2: Run test to verify it fails**

Run: `.venv/bin/python -m pytest tests/integration/test_catalog_routes.py -v`
Expected: FAIL — the routes 404.

- [ ] **Step 3: Write the read models**

Create `app/services/pairing_queries.py`. `list_categories` groups
`strategist_products` by `category` with counts. `list_products` pages with `LIMIT`/`OFFSET`
and returns `{"products": [...], "total": n}`. `get_product` returns one row or `None`.

`pairings_for` is the one with real logic: read every row for the anchor, join each against
the tenant's decisions, and attach `servable` using `is_servable` — the merchant must see
the same evidence the ranker will. Return a dict keyed by every entry in `PAIR_TYPES`, with
empty lists for types that have no rows, so the client never has to guess which keys exist.

`pending_pairings` returns only pairs where `is_servable(pair, decision)` is false and no
decision has been recorded — the queue is things needing a human, not things already
refused.

- [ ] **Step 4: Write the router**

Create `app/api/catalog.py` with `router = APIRouter(prefix="/catalog", tags=["catalog"])`
and the seven routes. Import `pair_tenant`, `record`, and the query functions as
module-level names so the tests can monkeypatch them.

`decide` takes a Pydantic body model with `tenant_id: str` and
`decisions: list[DecisionIn]`, and turns a `ValueError` from `record` into a `422` — an
invalid decision value is a bad request, not a server fault.

Mount it in `app/main.py` beside the existing routers:

```python
from app.api.catalog import router as catalog_router
...
app.include_router(catalog_router)
```

- [ ] **Step 5: Run the tests**

Run: `.venv/bin/python -m pytest tests/integration/test_catalog_routes.py -v`
Expected: PASS, 16 tests

- [ ] **Step 6: Run both suites**

Run `.venv/bin/python -m pytest tests/unit/ -q`, then
`.venv/bin/python -m pytest tests/integration/ -q`, sequentially in the foreground.
Expected: no regressions against the 546 / 70 baselines.

- [ ] **Step 7: Report changed files**

```bash
git status --short
```

Do not commit and do not stage.

---

### Task 8: Live pairing and threshold tuning

**Files:** none — verification and one constant.

**Prerequisites:** Postgres reachable, the live tenant holding 218 enriched products.

- [ ] **Step 1: Run the job**

```bash
PYTHONPATH=. .venv/bin/python -c "
import json
from app.services.pairing import pair_tenant
print(json.dumps(pair_tenant('org_8c32bf3e-6a18-4739-9b1c-94c0cf11125f'), indent=2))"
```

Report products, pairs, servable, queued, and the per-type breakdown.

- [ ] **Step 2: Look at the confidence distribution before touching the threshold**

```bash
PYTHONPATH=. .venv/bin/python -c "
from app.services.database import get_db_connection
T='org_8c32bf3e-6a18-4739-9b1c-94c0cf11125f'
c=get_db_connection(); cur=c.cursor()
cur.execute('SELECT pair_type, count(*), round(min(confidence)::numeric,2), '
            'round(avg(confidence)::numeric,2), round(max(confidence)::numeric,2) '
            'FROM \"'+T+'\".strategist_product_neighbors GROUP BY 1 ORDER BY 1')
for r in cur.fetchall(): print(' ', r)
cur.execute('SELECT width_bucket(confidence,0,1,10) b, count(*) '
            'FROM \"'+T+'\".strategist_product_neighbors GROUP BY 1 ORDER BY 1')
print('confidence histogram:', cur.fetchall()); c.close()"
```

- [ ] **Step 3: Set the threshold from what you measured**

Pick `APPROVAL_THRESHOLD` so the queue is a size a merchant would actually work through —
tens, not thousands — while pairs a human would obviously wave through do not queue at all.
Update the constant in `app/services/pairing_decisions.py` and **record in your report both
the number you chose and the distribution that justified it.** A threshold with no measured
distribution behind it is a guess; that is the whole reason this step exists rather than
the number being fixed in Task 5.

Re-run Step 1 afterwards and report the new servable/queued split.

- [ ] **Step 4: Hand-check the two obvious categories**

```bash
PYTHONPATH=. .venv/bin/python -c "
from app.services.database import get_db_connection
T='org_8c32bf3e-6a18-4739-9b1c-94c0cf11125f'
c=get_db_connection(); cur=c.cursor()
cur.execute('''SELECT p.name, n.pair_type, q.name, round(n.score::numeric,2)
               FROM \"'''+T+'''\".strategist_product_neighbors n
               JOIN \"'''+T+'''\".strategist_products p ON p.product_key = n.anchor_key
               JOIN \"'''+T+'''\".strategist_products q ON q.product_key = n.neighbor_key
               WHERE p.category = 'smartphones'
               ORDER BY n.pair_type, n.score DESC LIMIT 25''')
for r in cur.fetchall(): print(' ', r)
c.close()"
```

Expected, and worth stating plainly if it is not what you see: a smartphone's `similar`
neighbours are other smartphones, its `complement` neighbours are mobile accessories, and
its `upsell` neighbours are pricier smartphones. This catalogue makes those three obvious
enough to check by eye, which is the point of choosing it.

- [ ] **Step 5: Confirm the graph is symmetric where it should be and not where it should not**

```bash
PYTHONPATH=. .venv/bin/python -c "
from app.services.database import get_db_connection
T='org_8c32bf3e-6a18-4739-9b1c-94c0cf11125f'
c=get_db_connection(); cur=c.cursor()
cur.execute('''SELECT count(*) FROM \"'''+T+'''\".strategist_product_neighbors a
               JOIN \"'''+T+'''\".strategist_product_neighbors b
                 ON a.anchor_key = b.neighbor_key AND a.neighbor_key = b.anchor_key
                AND a.pair_type = b.pair_type
               WHERE a.pair_type = 'complement' ''')
print('bidirectional complements (must be 0):', cur.fetchone()[0]); c.close()"
```

A non-zero result means the accessory asymmetry is not holding and a phone case would
recommend a phone.

- [ ] **Step 6: Confirm a re-run is cheap and destroys nothing**

Re-run Step 1. Embedding calls must be zero — the cache carries them — and the pair count
must be identical rather than doubled. Then record a rejection through
`POST /catalog/pairings/decide`, re-run the job, and confirm through
`GET /catalog/pairings/pending` and the decisions table that the rejection survived. This
is the invariant the whole phase rests on and it deserves to be checked against the live
database, not only in tests.

- [ ] **Step 7: Report**

Summarise: changed files, test counts, the pairing report, the confidence distribution and
the threshold you chose from it, the smartphone hand-check, the bidirectional-complement
count, and confirmation that a re-run neither doubled the graph nor lost a decision. State
plainly anything that looked wrong. Remind the user to commit.

---

## Follow-on work (not in this plan)

- **Phase 3:** the merchant admin UI — four screens (connect and sources, catalogue browser, product detail with pairings, approval queue), of which only the last two are required for Phase 2 to be usable by a human. The performance screen the matching document describes is dropped: no data exists for it until Phase 5, so it could only be a mock. These endpoints are what the UI calls.
- **Phase 4:** online ranking. `decide()` switches from `related_keys` to `product_neighbors`, and `products.py` finally changes.
- **Phase 5:** measurement — impressions, clicks, and which pair types earn their place.
- **Phase 6:** the known gaps, chiefly per-variant stock and order-history ingestion, which would give pairing its first behavioural signal.
