# Crawl Product-URL Prioritization Design

**Goal:** Fix the website-crawl product source (`app/services/catalog/crawl.py`)
so it reliably reaches individual product-detail pages on real, complex
e-commerce sites, instead of exhausting its page budget on collection/nav
pages and finding zero products.

**Architecture:** Replace `extract_text_from_url`'s recursive, per-depth
breadth-first crawl with an iterative worklist that prioritizes product-shaped
URLs. A free regex heuristic handles the common case (Shopify/WooCommerce/
Magento-style paths); a one-time LLM classification call is the fallback for
sites whose URL structure doesn't match any known pattern.

**Tech Stack:** Playwright (`app/services/content/browser.py`), the existing
JSON-LD/OG/LLM product extraction pipeline (unchanged by this spec), the
existing LLM call plumbing used by `extract_with_llm`.

## Context

Confirmed by live-testing against `https://cordori.com.au/` (a real Shopify
storefront) before this spec was written:

- `extract_text_from_url` (`app/services/content/website.py:62`) is called by
  both `crawl.py:148` and `endpoints.py:182` with `max_depth=20` — depth is
  effectively unbounded, so the only real limit is `max_pages`
  (`CRAWL_MAX_PAGES`, currently defaulting to 20).
- The traversal is strict breadth-first: it crawls all depth-N links via
  `asyncio.gather` before starting depth-N+1, with no ordering within a
  level. On cordori.com.au, the homepage's own links are almost entirely
  `/collections/*` and `/pages/*` (nav/category pages) — with the default
  budget, all 15-20 of these consumed the entire page budget before any
  `/products/<handle>` link (one hop further) was ever visited.
- Reproduced directly: with `max_pages=15`, the crawl visited 14 pages (all
  collection/nav pages) and found **0 products** — not because extraction is
  broken, but because it never visited a product page. Raising `max_pages` to
  60 alone made the crawler visit 59 pages and find 35 real products with
  correct titles/descriptions/URLs/images, confirming the extraction tiers
  (JSON-LD → OG → LLM, per `app/services/content/product_extraction.py`) work
  correctly once a product page is actually reached.
- `browser.py::get_rendered_content` (line 25) currently returns links as a
  bare list of href strings (`page.eval_on_selector_all("a[href]", "elements
  => elements.map(e => e.href)")`) — no anchor text, which the LLM fallback
  classifier needs.
- This traversal function is shared by `endpoints.py`'s general site-scrape
  path (branding/colors/text) as well as the product crawl — the
  prioritization change benefits both, since product pages carry richer
  content than nav chrome.

## Decisions (settled via clarifying questions)

1. **Both raise the budget and prioritize.** `CRAWL_MAX_PAGES` default moves
   from 20 to 60, and the frontier is reordered so product-shaped URLs are
   visited first regardless of budget size — the ordering fix matters even
   at the higher budget for sites with more nav surface than that.
2. **Prioritization is hybrid: regex heuristic first, one-time LLM fallback
   second.** A free URL-pattern match (`/products/`, `/product/`, `/p/`,
   `/item/`) covers the common Shopify/WooCommerce/Magento case at zero extra
   cost. If, after a threshold number of pages, the heuristic has matched
   nothing, a single LLM call classifies the currently-known unvisited
   candidate links by anchor text + URL, and its verdict re-prioritizes the
   frontier. This generalizes to sites with non-standard URL conventions
   without paying an LLM call on every crawl.
3. **The LLM fallback fires at most once per crawl.** A flag prevents
   re-triggering; if it fails or returns malformed output, the crawl falls
   back to plain insertion-order traversal (today's behavior) rather than
   blocking.

## Data Model / Interfaces

```python
# browser.py — get_rendered_content return type changes:
# before: Tuple[str, List[str], Dict[str, str]]        (content, hrefs, meta)
# after:  Tuple[str, List[Dict[str, str]], Dict[str, str]]
#         each link dict: {"href": str, "text": str}   (text = trimmed anchor text, may be "")

# website.py — new helpers
def _looks_like_product_url(url: str) -> bool:
    """True if the URL path matches a known product-detail pattern."""

async def _classify_product_links(
    candidates: list[tuple[str, str]],  # (url, anchor_text)
    tenant_id: str,
) -> set[str]:
    """One LLM call. Returns the subset of candidate URLs judged to be
    individual product-detail pages. Returns empty set on any failure."""
```

`extract_text_from_url`'s public signature (`url, max_depth, max_pages,
visited`) does not change — only its internal traversal algorithm.

## Traversal Algorithm

Replace the recursive `crawl(url, depth)` + per-depth `asyncio.gather` with
an iterative worklist:

1. `frontier: list[tuple[url, depth]]` seeded with `[(start_url, 0)]`.
   `product_priority: set[str]` starts empty. `fallback_fired: bool = False`.
2. Loop while `frontier` is non-empty and `len(visited) < max_pages`:
   - Sort `frontier` so entries in `product_priority` (or matching
     `_looks_like_product_url`) come first, preserving relative order
     otherwise (stable sort).
   - Pop a batch (existing per-level concurrency stays — batch = all
     currently-sorted frontier entries at the front up to remaining budget)
     and crawl them concurrently via existing `get_rendered_content`.
   - For each crawled page, extract new links (respecting `is_subpath`,
     extension filter, `depth < max_depth`, not already visited/queued) and
     append to `frontier`.
   - Fallback trigger check (only if not yet fired): if `len(visited) >= 3`
     (new constant `CRAWL_PRODUCT_FALLBACK_AFTER_PAGES`) and
     `product_priority` is still empty and zero product-pattern URLs have
     been found among all visited+queued links, call
     `_classify_product_links` once with the current frontier's
     `(url, anchor_text)` pairs, union its result into `product_priority`,
     set `fallback_fired = True`.
3. Existing per-page extraction logic (metadata, colors, images, JSON-LD,
   text cleanup) is unchanged — only the order pages are visited in changes.

## New Prompt

`app/core/prompts.py` gains `PRODUCT_LINK_CLASSIFICATION_PROMPT`: given a
numbered list of `(url, anchor_text)` pairs, return the indices that look
like individual product-detail pages (as opposed to category/listing/nav/
info pages). Mirrors the existing refuse-listing-pages framing already used
in `PRODUCT_EXTRACTION_PROMPT` so the two prompts don't disagree about what
counts as a product page.

The classification call uses `response_format={"type": "json_object"}` (the
same JSON-mode convention already used by `quotas.py`, `analytics.py`,
`fieldmap_proposer.py`, and `enrichment/extractor.py`) rather than relying on
prompt wording alone to produce parseable JSON — more reliable than
`extract_with_llm`'s current approach, which this spec does not change.

## Config

- `app/core/config.py`: `CRAWL_MAX_PAGES` default `20` → `60`.
- New: `CRAWL_PRODUCT_FALLBACK_AFTER_PAGES: int = 3`.

## Error Handling

- LLM classification failure/malformed output → log at WARNING, return empty
  set, crawl continues with insertion-order fallback for the rest of the run.
  This is a prioritization hint, not a correctness dependency — worst case
  behavior matches today's.
- No change to existing per-page error handling (timeout-but-proceed, failed
  navigation logging) in `browser.py`/`website.py`.

## Testing

- Unit tests (synthetic fixtures, no real network/browser — matches existing
  style in `tests/unit/test_crawl_source.py`):
  - `_looks_like_product_url`: positive/negative cases across the pattern
    list.
  - Frontier ordering: given a mixed list of discovered links, product-
    pattern URLs are visited before non-matching ones regardless of
    discovery order.
  - Fallback trigger: with a mocked LLM, verify `_classify_product_links` is
    called exactly once after the threshold, its verdict changes crawl
    order, and a second batch of pages does not trigger it again.
  - Fallback failure path: mocked LLM raises/returns garbage → crawl
    continues, no exception propagates, order falls back to insertion order.
- No existing test asserts on `get_rendered_content`'s link return shape
  directly (confirmed via `tests/unit/test_crawl_source.py` and
  `test_crawl_throttle.py`, which only mock `fetch_products`/higher-level
  calls) — the `List[str]` → `List[Dict[str,str]]` change should not break
  any current test, but this must be re-verified once the change lands.

### Live validation (not automated, run manually before calling this done)

Cordori alone is one data point on one platform (Shopify) with one nav
style. The fix must generalize, not just fix Cordori, so before the
implementation is accepted it gets run live (the same
`fetch_products(config, ...)` throwaway-script approach used to diagnose
this spec) against a small set of real, complex storefronts spanning
different platforms and URL conventions. Each site's actual platform/URL
pattern gets confirmed by inspection at run time (do not assume from the
list below — verify against the live HTML/URLs), and the run records:
pages visited, products found, and whether the regex heuristic or the LLM
fallback did the prioritizing.

Candidate sites — platform/URL pattern noted is a best guess and MUST be
confirmed by inspection at run time, not assumed. Run at least 6-8 of these
(spanning every category below) before implementation is considered
validated; swap out any that turn out unreachable, paywalled, bot-blocked,
or too simple to be a useful test:

*Shopify, `/products/<handle>`, heavy nav/mega-menu:*
- `https://cordori.com.au/` — the original case, regression-guards the fix.
- `https://gymshark.com/` — large catalog, deep collection nesting.
- `https://kith.com/` — notoriously complex mega-menu, good stress test for
  how many nav pages sit between the homepage and a product.
- `https://www.allbirds.com/`
- `https://www.bombas.com/`
- `https://www.fashionnova.com/` — very large catalog, good volume test.

*WooCommerce, `/product/<slug>` singular (different pattern than Shopify's
plural `/products/`) — confirms the regex heuristic isn't Shopify-only:*
- `https://www.athleticbrewing.com/`
- Any other WooCommerce site found during testing (check page source for
  `wp-content`/`woocommerce` markers to confirm platform before using it).

*Magento / BigCommerce / other, often numeric or non-slug product URLs —
these are the important ones for exercising the LLM fallback path, since
their URLs won't match the regex heuristic at all:*
- `https://www.skullcandy.com/` — reported BigCommerce.
- `https://www.hollisterco.com/` — reported Magento-derived, worth checking
  its live URL pattern.
- Any site discovered with `/catalog/product/view/id/...`-style or
  query-param product URLs is a strong candidate — actively look for one if
  none of the above turn out to qualify, since this case is the whole
  reason the LLM fallback exists.

The last category matters most for the "any complex site" requirement —
don't skip it even if the Shopify/WooCommerce sites all pass easily.

Acceptance bar: every site in the validation set finds at least one real
product (title + product_url populated) within the raised `max_pages`
budget, and the log confirms product pages were prioritized over pure
nav/collection pages rather than found only by exhausting the full budget.

## Out of Scope (tracked separately, not part of this spec)

The user has flagged three further gaps found during the same live test,
each to get its own spec once this one ships:
- Price missing on ~74% of extracted products (9/35 had a price).
- Variant options (size/color) never extracted (0/35).
- The `handle` field is always null — needs investigation into whether it's
  a real bug or an unused field before deciding whether to fix it.
