# Attribute Extraction — Design

**Date:** 2026-08-10
**Status:** Draft, awaiting review
**Scope:** Phase 1 of seven. Depends on Phase 0 (catalogue ingestion), which is built.

---

## 1. Purpose

Give every product the structured attributes the recommendation matching specification
reads from, by extracting them with one language-model call per batch of products.

**The number that justifies this phase: 7 of 218 products currently have any attributes.**

`attribute_jaccard` is one of the four terms in the content-similarity score. Today it
contributes nothing for 97% of the catalogue. `is_accessory` — the strongest complement
signal available — does not exist at all. The matching specification calls this step
"the prerequisite for everything above… no longer optional", and the live catalogue agrees.

**Success criteria:**

1. Every ingested product has `color`, `material`, `style`, `gender`, `size_system`,
   `use_case`, `is_accessory`, `price_tier` and `key_features` where the source material
   supports them.
2. No merchant-supplied attribute is ever overwritten by an inferred one.
3. Re-running costs nothing for products whose content has not changed.
4. A product the model cannot read is flagged, not silently skipped.

## 2. Why this matters most for clothing

A snowboard is discriminated by title and price alone. A garment is not.

```
Oxford Shirt   -> color white,  material cotton,  style formal, gender male
Navy Chinos    -> color navy,   material cotton,  style smart-casual
Orange Chinos  -> color orange, material cotton,  style casual
Leather Belt   -> color brown,  material leather, is_accessory TRUE
```

Without these, all four are a title and a price, and no outfit can be assembled from them.
Two consequences follow directly, both of which Phase 2 depends on:

- **Complements become recognisable.** `is_accessory` is what separates a belt from another
  shirt. Without it, the online ranker's price-affinity rule crushes a £29 belt shown
  against a £59 shirt, because complements are only exempt from that rule if something
  knows they are complements.
- **Colour compatibility becomes possible.** A white shirt goes with navy trousers, not
  orange ones. Phase 2 cannot rank on compatibility that was never extracted.

## 3. What is extracted

One call per batch. Output constrained to a fixed schema:

| Field | Type | Notes |
|---|---|---|
| `color` | string | base palette term, not a marketing name |
| `material` | string | |
| `style` | string | |
| `gender` | enum | `male`, `female`, `unisex`, `kids`, or null |
| `size_system` | enum | `alpha`, `numeric`, `uk`, `eu`, `us`, `volume`, or null |
| `use_case` | string | what it is for |
| `is_accessory` | boolean | **column, not just an attribute** |
| `price_tier` | enum | `budget`, `mid`, `premium` — **column** |
| `key_features` | string[] | at most five short phrases |

Input per product: title, description, category path, brand, and any attributes already
present. Input quality is now adequate — **description is 93% filled** across the live
catalogue, against 11% when only Shopify was connected.

## 4. Three rules that shape the design

### 4.1 Inferred data never overwrites merchant data

If `color` came from a Shopify variant option, the model does not get to change it.
Extraction **fills gaps only**. Everything it produces is stored with `source: "inferred"`
and the model's confidence, so the matching layer can weight a merchant's own value above a
guess — the same precedence table the normaliser already uses.

Where an inferred value contradicts a structured one, the structured one wins and the
disagreement is counted in the run report. A high disagreement rate means the extraction
prompt is wrong, and that is worth knowing.

### 4.2 `price_tier` is relative to the tenant, not absolute

"Premium" has no global meaning. It means expensive *for this catalogue*.

The job therefore computes the tenant's price percentiles first, then classifies:
below the 33rd percentile is `budget`, above the 67th is `premium`, the rest `mid`. The
live catalogue spans 0.79 to 36,999.99 in minor units — any fixed threshold would put
nearly everything in one bucket.

Percentiles are computed over `price_reference_cents` where available, falling back to
`price_cents`, so a two-currency tenant is not banded on unconverted numbers.

### 4.3 `is_accessory` and `price_tier` get columns

Both are filtered on rather than merely read — `is_accessory` by Phase 2's complement
materialisation, `price_tier` by the online ranker. A JSONB lookup for a hot filter is the
wrong shape.

```sql
is_accessory  BOOLEAN,
price_tier    TEXT      -- budget | mid | premium
```

The remaining fields join the existing `attributes` JSONB alongside merchant-supplied ones.

## 5. Batching, caching and cost

**Twenty products per call**, matching the matching specification's guidance. Each call
carries only what the model needs: the batch's products and the tenant's category list.

**Cached on `content_hash`.** A product is re-extracted only when its content actually
changed. This is why removing description from `content_hash` in Phase 0 mattered — without
that correction, a merchant fixing a typo would re-extract the whole catalogue.

Cost on the live catalogue: 218 products, 11 calls, once. A 2,000-product store is 100
calls. A 200,000-product store runs at the category level instead, which the matching
specification already anticipates; that path is out of scope here and recorded in §10.

**The extraction result is stored, not recomputed.** As everywhere else in this system, a
model authors a durable artifact and the artifact is what downstream reads.

## 6. Which products are extracted

**All of them, on first run.** The matching specification proposes gating on
`quality_score < 0.5`, but Phase 0 measured a median of 0.35 on real data — nearly
everything falls below the gate, so it would not be acting as a cost control. Gating is
therefore deliberately not implemented; it belongs in a later phase once a real score
distribution across several merchants is visible.

Subsequent runs extract only products whose `content_hash` changed, which is the effective
cost control.

`POST /catalog/enrich` accepts an optional `force` flag to re-extract regardless of hash,
for use after a prompt change. Without it, a prompt improvement would never reach existing
products.

## 7. Endpoint

```
POST /catalog/enrich     { tenant_id, force? }
```

Separate from `/catalog/sync` because it costs money and must be re-runnable without
re-fetching a catalogue. Runs off the event loop, following the existing precedent in
`app/api/endpoints.py`.

Returns a report:

```json
{
  "products": 218,
  "extracted": 211,
  "skipped_unchanged": 0,
  "failed": 7,
  "conflicts": 3,
  "attribute_coverage": 0.97
}
```

`failed` counts products the model could not read — typically no description and a
title too short to infer from. Those are flagged, not silently skipped: `missing_fields`
gains `attributes`, using the mechanism Phase 0 already established.

## 8. Error handling

| Condition | Behaviour |
|---|---|
| Model call fails for a batch | retry once, then mark that batch's products failed and continue — one bad batch must not abort a catalogue |
| Model returns non-JSON or an unknown field | discard the unknown parts, keep what validates |
| Model returns a value outside a closed set (`gender`, `price_tier`, `size_system`) | discard that field, count it |
| Product has no description and a title under three words | not sent to the model at all; flagged `attributes` in `missing_fields` |
| Tenant has no products | 200 with a zero report, not an error |

Nothing here logs a full product payload or an API key.

## 9. Testing

**Golden fixtures from the live catalogue** — a smartphone, a mobile accessory, a grocery
item and a snowboard, chosen because they exercise different extraction paths.

**Assertions are on shape, not values.** Tests assert the output is constrained — only
known keys, `is_accessory` a real boolean, `price_tier` from the closed set, `gender` from
the closed set, `key_features` at most five entries. Asserting exact values would make the
suite depend on model output and fail on any prompt or model change, which is how a suite
stops being trusted.

**Behavioural tests with the model stubbed:**

- an inferred value never overwrites a merchant-supplied one, and the conflict is counted
- `price_tier` boundaries are computed from the tenant's own distribution, verified by
  constructing a catalogue whose 33rd and 67th percentiles are known
- a batch failure marks only that batch and the run continues
- an unchanged `content_hash` skips extraction, and `force` overrides that
- a product with no usable text is flagged rather than sent to the model

**Live run:** extract the real 218-product catalogue and report coverage before and after.
The meaningful number is attribute coverage moving from 3% to something near complete, and
`is_accessory` becoming populated for the accessory categories — `mobile-accessories`,
`sports-accessories`, `kitchen-accessories` — which is precisely what Phase 2 needs.

## 10. Known limitations

**No category-level extraction.** Very large catalogues are meant to run this at the
category level rather than per product. Not implemented; the per-product path is correct
for catalogues in the thousands and this system has no larger tenant yet.

**No quality gating.** Deliberate — see §6.

**Extraction quality is bounded by input.** A product with no description and a
three-word title yields little, and no prompt fixes that. Such products are flagged so the
gap is visible to the merchant rather than mistaken for a ranking problem.

**Per-size stock is not addressed.** For apparel, a garment in stock overall but sold out
in the shopper's size is a dead recommendation. `in_stock` is per-product. This is a real
gap for clothing tenants, recorded here because extraction is where sizes first become
visible, but solving it needs per-variant stock and webhooks.

## 11. Out of scope

Pairing of any kind; the `product_neighbors` table; complement inference; the approval
queue; online ranking; measurement. Phases 2 through 5.
