# Crawl Product-URL Prioritization Implementation Plan

> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.

**Goal:** Make the website-crawl product source reliably reach individual product-detail pages on real, complex e-commerce sites (not just Cordori), instead of exhausting its page budget on collection/nav pages and finding zero products.

**Architecture:** Replace `extract_text_from_url`'s recursive, per-depth breadth-first crawl with an iterative worklist that prioritizes product-shaped URLs. A free regex heuristic handles the common case; a one-time LLM classification call is the fallback for sites whose URL structure doesn't match any known pattern.

**Tech Stack:** Playwright (`app/services/content/browser.py`), the existing product-extraction pipeline (unchanged), `app/core/llm_client.py`'s `client` (OpenAI-compatible), pytest with `unittest.mock`.

**Spec:** `docs/superpowers/specs/2026-08-23-crawl-product-url-prioritization-design.md`

## Global Constraints

- `CRAWL_MAX_PAGES` default moves from 20 to 60 (`app/core/config.py`).
- New setting `CRAWL_PRODUCT_FALLBACK_AFTER_PAGES: int = 3`.
- `get_rendered_content`'s link return type changes from `List[str]` to `List[Dict[str, str]]`, each entry `{"href": str, "text": str}`. Every caller of this function must be updated in the same task that changes it — do not leave a caller consuming the old shape.
- The LLM fallback fires **at most once per crawl** and must never raise out of `extract_text_from_url` — any failure degrades to plain insertion-order traversal (today's behavior).
- No changes to the JSON-LD/OG/LLM product-extraction tiers themselves (`product_extraction.py`) — this plan only changes which pages get visited and in what order.
- Follow existing code style: no narrative comments, only "why" comments where genuinely non-obvious (see `product_extraction.py` for the house style).

---

### Task 1: Anchor text on every discovered link

**Files:**
- Modify: `app/services/content/browser.py:25-61` (`get_rendered_content`)
- Modify: `app/services/content/website.py:83-188` (`crawl` inner function — the two places it reads `links`)
- Modify: `app/api/endpoints.py:182` (caller — only needs to keep working, no behavior change there)
- Test: `tests/unit/test_browser_links.py` (new)

**Interfaces:**
- Produces: `get_rendered_content(url, wait_until="load", timeout=30000) -> Tuple[str, List[Dict[str, str]], Dict[str, str]]`. Each link dict is `{"href": str, "text": str}`, `text` is the trimmed inner text of the `<a>` tag (empty string if none).

- [ ] **Step 1: Write the failing test**

Create `tests/unit/test_browser_links.py`:

```python
"""get_rendered_content must hand back anchor text alongside each href --
the crawl-prioritization fallback classifies links by their text, not just
their URL shape."""
import pytest
from unittest.mock import AsyncMock, MagicMock, patch

from app.services.content.browser import BrowserService


@pytest.mark.asyncio
async def test_links_carry_their_anchor_text():
    service = BrowserService()
    fake_page = AsyncMock()
    fake_page.goto = AsyncMock()
    fake_page.content = AsyncMock(return_value="<html></html>")
    fake_page.eval_on_selector_all = AsyncMock(return_value=[
        {"href": "https://shop.example.com/products/oxford-shirt", "text": " Oxford Shirt "},
        {"href": "https://shop.example.com/collections/all", "text": "Shop All"},
    ])
    fake_page.title = AsyncMock(return_value="Home")
    fake_page.close = AsyncMock()

    fake_context = AsyncMock()
    fake_context.new_page = AsyncMock(return_value=fake_page)
    fake_context.close = AsyncMock()

    service._browser = MagicMock()
    service._browser.new_context = AsyncMock(return_value=fake_context)

    content, links, meta = await service.get_rendered_content("https://shop.example.com/")

    assert links == [
        {"href": "https://shop.example.com/products/oxford-shirt", "text": "Oxford Shirt"},
        {"href": "https://shop.example.com/collections/all", "text": "Shop All"},
    ]
```

- [ ] **Step 2: Run test to verify it fails**

Run: `pytest tests/unit/test_browser_links.py -v`
Expected: FAIL — `eval_on_selector_all` mock returns dicts already, but the real implementation still calls `.map(e => e.href)` and returns bare strings, so `links` will be a list of strings, not the mock's dicts (the mock is asserting the *target* shape, the source doesn't produce it yet).

- [ ] **Step 3: Write minimal implementation**

In `app/services/content/browser.py`, replace the link-extraction line inside `get_rendered_content`:

```python
            # Extract internal links directly using Playwright, with anchor
            # text so link-classification callers don't need a second pass.
            links = await page.eval_on_selector_all(
                "a[href]",
                "elements => elements.map(e => ({href: e.href, text: (e.textContent || '').trim()}))",
            )
```

Update the function's return type hint:

```python
    async def get_rendered_content(self, url: str, wait_until: str = "load", timeout: int = 30000) -> Tuple[str, List[Dict[str, str]], Dict[str, str]]:
```

- [ ] **Step 4: Run test to verify it passes**

Run: `pytest tests/unit/test_browser_links.py -v`
Expected: PASS

- [ ] **Step 5: Update the two consumers in `website.py` to the new link shape**

In `app/services/content/website.py`, the `crawl` inner function currently does:

```python
            content, links, pw_meta = await browser.get_rendered_content(current_url)
```

`links` is now `List[Dict[str, str]]`. Update the link-following loop (originally around line 176-185):

```python
            if depth < max_depth and len(visited) < max_pages:
                tasks = []
                for link in links:
                    next_url = link["href"]
                    if len(visited) >= max_pages:
                        break

                    parsed_next = urlparse(next_url)
                    if is_subpath(url, next_url) and next_url not in visited:
                        if not any(next_url.lower().endswith(ext) for ext in ['.pdf', '.jpg', '.png', '.zip', '.gif']):
                            tasks.append(crawl(next_url, depth + 1))

                if tasks:
                    await asyncio.gather(*tasks)
```

This is a placeholder edit for Task 1 only — Task 4 replaces this whole traversal block with the iterative worklist. For now, just keep `extract_text_from_url`'s existing recursive structure working against the new link shape so the test suite stays green between tasks.

- [ ] **Step 6: Run the full existing test suite to confirm nothing else broke**

Run: `pytest tests/unit/ -k "website or crawl or browser" -v`
Expected: PASS (no test currently asserts on `get_rendered_content`'s link shape directly outside the one just added, per the spec's Testing section — confirm this stays true).

- [ ] **Step 7: Commit**

```bash
git add app/services/content/browser.py app/services/content/website.py tests/unit/test_browser_links.py
git commit -m "feat: carry anchor text alongside crawled links"
```

---

### Task 2: Product-URL regex heuristic

**Files:**
- Modify: `app/services/content/website.py` (add `_looks_like_product_url`)
- Test: `tests/unit/test_product_url_heuristic.py` (new)

**Interfaces:**
- Consumes: nothing new.
- Produces: `_looks_like_product_url(url: str) -> bool`, used by Task 4's frontier sort.

- [ ] **Step 1: Write the failing test**

Create `tests/unit/test_product_url_heuristic.py`:

```python
"""Regex-only classification of product-detail URLs -- the free path that
covers Shopify/WooCommerce/Magento-style conventions before the LLM
fallback (see PRODUCT_LINK_CLASSIFICATION_PROMPT) ever runs."""
from app.services.content.website import _looks_like_product_url


def test_shopify_style_plural_products_path():
    assert _looks_like_product_url("https://shop.example.com/products/oxford-shirt")


def test_woocommerce_style_singular_product_path():
    assert _looks_like_product_url("https://shop.example.com/product/oxford-shirt")


def test_short_p_path():
    assert _looks_like_product_url("https://shop.example.com/p/12345")


def test_item_path():
    assert _looks_like_product_url("https://shop.example.com/item/oxford-shirt")


def test_collection_pages_do_not_match():
    assert not _looks_like_product_url("https://shop.example.com/collections/shirts")
    assert not _looks_like_product_url("https://shop.example.com/collections/all")


def test_static_pages_do_not_match():
    assert not _looks_like_product_url("https://shop.example.com/pages/about-us")
    assert not _looks_like_product_url("https://shop.example.com/")


def test_case_and_query_string_are_ignored():
    assert _looks_like_product_url("https://shop.example.com/PRODUCTS/oxford-shirt?variant=1")
```

- [ ] **Step 2: Run test to verify it fails**

Run: `pytest tests/unit/test_product_url_heuristic.py -v`
Expected: FAIL with `ImportError: cannot import name '_looks_like_product_url'`

- [ ] **Step 3: Write minimal implementation**

Add near the top of `app/services/content/website.py`, after the existing `is_subpath` function:

```python
# Known e-commerce product-detail URL conventions. Shopify uses the plural
# /products/, WooCommerce and most others use the singular /product/ -- both
# must match, they are genuinely different platforms' conventions, not a typo.
_PRODUCT_URL_PATTERN = re.compile(r"/(products?|item|p)/[^/?#]+", re.IGNORECASE)


def _looks_like_product_url(url: str) -> bool:
    """Free heuristic for the common case. Returns False for anything that
    doesn't match a known pattern -- callers fall back to the LLM classifier
    for sites this misses, they don't treat False as "not a product"."""
    return bool(_PRODUCT_URL_PATTERN.search(urlparse(url).path))
```

- [ ] **Step 4: Run test to verify it passes**

Run: `pytest tests/unit/test_product_url_heuristic.py -v`
Expected: PASS

- [ ] **Step 5: Commit**

```bash
git add app/services/content/website.py tests/unit/test_product_url_heuristic.py
git commit -m "feat: add regex heuristic for product-detail URLs"
```

---

### Task 3: LLM fallback link classifier

**Files:**
- Modify: `app/core/prompts.py` (add `PRODUCT_LINK_CLASSIFICATION_PROMPT`)
- Modify: `app/services/content/website.py` (add `_classify_product_links`, import `openai_client`, `settings`, `insert_llm_usage`, `PRODUCT_LINK_CLASSIFICATION_PROMPT`)
- Test: `tests/unit/test_product_link_classification.py` (new)

**Interfaces:**
- Consumes: `_PRODUCT_URL_PATTERN`/`_looks_like_product_url` not required here (independent path).
- Produces: `async def _classify_product_links(candidates: List[Tuple[str, str]], tenant_id: str) -> Set[str]`, used by Task 4's fallback trigger. `candidates` is a list of `(url, anchor_text)`. Returns the subset of URLs judged to be individual product-detail pages. Returns `set()` on any failure — never raises.

- [ ] **Step 1: Add the prompt**

In `app/core/prompts.py`, add after `PRODUCT_EXTRACTION_PROMPT`:

```python
PRODUCT_LINK_CLASSIFICATION_PROMPT = """You are looking at a numbered list of links found on an e-commerce website. Identify which ones lead to an individual product's own detail page -- not a category/collection page, not a homepage, not an "our products" listing, not account/cart/info pages.

Return a raw JSON object, no markdown fences, in exactly this shape:

{{"product_indices": [0, 3, 7]}}

Rules:
- A product detail page is dedicated to exactly one item a visitor can buy or act on.
- A page that lists or links to several items (a category, collection, "shop all", search results) is NOT a product page -- exclude it.
- Use both the URL and the link text to judge; a URL with no obvious product slug can still be a product page if the link text names a specific item.
- If nothing in the list looks like a product page, return {{"product_indices": []}}. Do not guess.

LINKS:
{links}
"""
```

- [ ] **Step 2: Write the failing test**

Create `tests/unit/test_product_link_classification.py`:

```python
"""One LLM call classifies discovered links as product-detail pages or not
-- the fallback for sites whose URLs don't match the regex heuristic. Must
never raise: a prioritization hint failing must not break the crawl."""
import json
from unittest.mock import MagicMock, patch

import pytest

from app.services.content import website


def _llm_returning(payload):
    response = MagicMock()
    response.choices = [MagicMock()]
    response.choices[0].message.content = json.dumps(payload)
    response.usage = MagicMock(prompt_tokens=10, completion_tokens=5, total_tokens=15)
    return response


CANDIDATES = [
    ("https://shop.example.com/sku/12345", "Oxford Shirt"),
    ("https://shop.example.com/browse/shirts", "Shop Shirts"),
]


@pytest.mark.asyncio
async def test_classifies_by_index_into_urls():
    with patch.object(website, "openai_client") as client:
        client.chat.completions.create.return_value = _llm_returning({"product_indices": [0]})
        result = await website._classify_product_links(CANDIDATES, "org_x")

    assert result == {"https://shop.example.com/sku/12345"}


@pytest.mark.asyncio
async def test_empty_verdict_is_an_empty_set():
    with patch.object(website, "openai_client") as client:
        client.chat.completions.create.return_value = _llm_returning({"product_indices": []})
        result = await website._classify_product_links(CANDIDATES, "org_x")

    assert result == set()


@pytest.mark.asyncio
async def test_llm_failure_returns_empty_set_not_an_exception():
    with patch.object(website, "openai_client") as client:
        client.chat.completions.create.side_effect = RuntimeError("timeout")
        result = await website._classify_product_links(CANDIDATES, "org_x")

    assert result == set()


@pytest.mark.asyncio
async def test_malformed_json_returns_empty_set_not_an_exception():
    with patch.object(website, "openai_client") as client:
        response = MagicMock()
        response.choices = [MagicMock()]
        response.choices[0].message.content = "not json"
        client.chat.completions.create.return_value = response
        result = await website._classify_product_links(CANDIDATES, "org_x")

    assert result == set()


@pytest.mark.asyncio
async def test_out_of_range_index_is_ignored_not_a_crash():
    with patch.object(website, "openai_client") as client:
        client.chat.completions.create.return_value = _llm_returning({"product_indices": [0, 99]})
        result = await website._classify_product_links(CANDIDATES, "org_x")

    assert result == {"https://shop.example.com/sku/12345"}
```

- [ ] **Step 3: Run test to verify it fails**

Run: `pytest tests/unit/test_product_link_classification.py -v`
Expected: FAIL with `AttributeError: module 'app.services.content.website' has no attribute '_classify_product_links'` (and no `openai_client` attribute yet either)

- [ ] **Step 4: Write minimal implementation**

Add these imports near the top of `app/services/content/website.py` (alongside the existing imports):

```python
from app.core.config import settings
from app.core.llm_client import client as openai_client
from app.core.prompts import PRODUCT_LINK_CLASSIFICATION_PROMPT
from app.services.infra.database import insert_llm_usage
```

Add the function after `_looks_like_product_url`:

```python
async def _classify_product_links(candidates: list, tenant_id: str) -> set:
    """One LLM call, run in a thread since the OpenAI client here is sync
    and this must not block the event loop mid-crawl. Never raises -- a
    failed classification just means the crawl falls back to insertion
    order, which is today's behavior."""
    if not candidates:
        return set()

    numbered = "\n".join(
        f"{i}. {url} -- \"{text}\"" for i, (url, text) in enumerate(candidates)
    )
    prompt = PRODUCT_LINK_CLASSIFICATION_PROMPT.format(links=numbered)

    def _call():
        return openai_client.chat.completions.create(
            model=settings.LLM_MODEL,
            messages=[{"role": "user", "content": prompt}],
            max_completion_tokens=500,
            timeout=30.0,
            response_format={"type": "json_object"},
        )

    try:
        response = await asyncio.to_thread(_call)
        payload = json.loads(response.choices[0].message.content)
    except Exception as ex:
        logger.warning(f"Product-link classification failed: {ex}")
        return set()

    try:
        usage = getattr(response, "usage", None)
        if usage:
            insert_llm_usage(tenant_id, "Product Link Classification", settings.LLM_MODEL,
                             usage.prompt_tokens, usage.completion_tokens, usage.total_tokens)
    except Exception as ex:
        logger.warning(f"Could not log product-link classification usage: {ex}")

    indices = payload.get("product_indices") or []
    return {candidates[i][0] for i in indices if isinstance(i, int) and 0 <= i < len(candidates)}
```

- [ ] **Step 5: Run test to verify it passes**

Run: `pytest tests/unit/test_product_link_classification.py -v`
Expected: PASS

- [ ] **Step 6: Commit**

```bash
git add app/core/prompts.py app/services/content/website.py tests/unit/test_product_link_classification.py
git commit -m "feat: add LLM fallback classifier for non-standard product URLs"
```

---

### Task 4: Iterative, priority-ordered crawl frontier

**Files:**
- Modify: `app/services/content/website.py:62-199` (`extract_text_from_url`, replaces the recursive traversal)
- Modify: `app/core/config.py:84` (`CRAWL_MAX_PAGES` default, add `CRAWL_PRODUCT_FALLBACK_AFTER_PAGES`)
- Test: `tests/unit/test_crawl_prioritization.py` (new)

**Interfaces:**
- Consumes: `_looks_like_product_url` (Task 2), `_classify_product_links` (Task 3), `get_rendered_content` returning `List[Dict[str, str]]` links (Task 1).
- Produces: `extract_text_from_url(url, max_depth=1, max_pages=20, visited=None) -> Dict[str, Any]` — same public signature and same return shape as before (`text`, `pages`, `metadata`, `colors`, `images`, `urls_visited`); only the internal visiting order changes.

- [ ] **Step 1: Update config**

In `app/core/config.py`, change:

```python
    # How much of a tenant's site to crawl. Twenty pages covers a SaaS or
    # donation site; a shop with a large catalogue needs far more, and raising
    # this costs an embedding call per chunk.
    CRAWL_MAX_PAGES: int = 60

    # Pages visited with zero product-pattern URL matches before the crawl
    # tries the one-time LLM link-classification fallback (see
    # website.py::_classify_product_links). Low enough to still leave most
    # of the page budget for the products it finds.
    CRAWL_PRODUCT_FALLBACK_AFTER_PAGES: int = 3
```

- [ ] **Step 2: Write the failing tests**

Create `tests/unit/test_crawl_prioritization.py`:

```python
"""extract_text_from_url must visit product-shaped URLs before plain nav
links, and must fall back to one LLM classification call -- and only one
-- when the regex heuristic finds nothing after a threshold of pages."""
import json
from unittest.mock import AsyncMock, MagicMock, patch

import pytest

from app.services.content import website


def _rendered(url, links, title="Page"):
    """(content, links, meta) tuple shaped like get_rendered_content's
    return value. links is a list of (href, text) tuples."""
    return (
        f"<html><title>{title}</title></html>",
        [{"href": href, "text": text} for href, text in links],
        {"title": title},
    )


@pytest.mark.asyncio
async def test_product_shaped_links_are_visited_before_nav_links():
    # Homepage links to three nav/collection pages and one product page,
    # in that order. Even though the product link is discovered last, it
    # must be visited before the nav pages that were discovered first.
    home = "https://shop.example.com/"
    order = []

    async def fake_render(url, *a, **kw):
        order.append(url)
        if url == home:
            return _rendered(home, [
                ("https://shop.example.com/collections/all", "Shop All"),
                ("https://shop.example.com/collections/shirts", "Shirts"),
                ("https://shop.example.com/pages/about", "About"),
                ("https://shop.example.com/products/oxford-shirt", "Oxford Shirt"),
            ])
        return _rendered(url, [])

    fake_browser = AsyncMock()
    fake_browser.get_rendered_content = fake_render

    with patch.object(website, "get_browser", AsyncMock(return_value=fake_browser)):
        await website.extract_text_from_url(home, max_depth=2, max_pages=3)

    # First visit is always the seed URL. Of the remaining budget (2 more
    # pages), the product page must be one of them despite being discovered
    # last -- the nav pages that would fill a plain-BFS budget must not
    # crowd it out.
    assert order[0] == home
    assert "https://shop.example.com/products/oxford-shirt" in order[1:3]


@pytest.mark.asyncio
async def test_llm_fallback_fires_once_after_threshold_with_no_matches():
    home = "https://shop.example.com/"
    visited_urls = [home] + [f"https://shop.example.com/browse/{i}" for i in range(10)]

    async def fake_render(url, *a, **kw):
        idx = visited_urls.index(url) if url in visited_urls else len(visited_urls)
        # Every page links to the next unvisited /browse/N page -- none of
        # these match the regex heuristic, forcing the fallback to fire.
        next_links = []
        if idx + 1 < len(visited_urls):
            next_links = [(visited_urls[idx + 1], "Next")]
        return _rendered(url, next_links)

    fake_browser = AsyncMock()
    fake_browser.get_rendered_content = fake_render

    classify_calls = []

    async def fake_classify(candidates, tenant_id):
        classify_calls.append(candidates)
        return set()

    with patch.object(website, "get_browser", AsyncMock(return_value=fake_browser)), \
         patch.object(website, "_classify_product_links", fake_classify), \
         patch.object(website.settings, "CRAWL_PRODUCT_FALLBACK_AFTER_PAGES", 3):
        await website.extract_text_from_url(home, max_depth=10, max_pages=10, tenant_id="org_x")

    assert len(classify_calls) == 1


@pytest.mark.asyncio
async def test_llm_fallback_verdict_reprioritizes_the_frontier():
    home = "https://shop.example.com/"

    async def fake_render(url, *a, **kw):
        if url == home:
            return _rendered(home, [
                ("https://shop.example.com/nav/1", "Nav 1"),
                ("https://shop.example.com/nav/2", "Nav 2"),
                ("https://shop.example.com/nav/3", "Nav 3"),
                ("https://shop.example.com/sku/12345", "Oxford Shirt"),
            ])
        return _rendered(url, [])

    fake_browser = AsyncMock()
    fake_browser.get_rendered_content = fake_render

    async def fake_classify(candidates, tenant_id):
        return {"https://shop.example.com/sku/12345"}

    order = []
    real_render = fake_render

    async def tracking_render(url, *a, **kw):
        order.append(url)
        return await real_render(url, *a, **kw)

    fake_browser.get_rendered_content = tracking_render

    with patch.object(website, "get_browser", AsyncMock(return_value=fake_browser)), \
         patch.object(website, "_classify_product_links", fake_classify), \
         patch.object(website.settings, "CRAWL_PRODUCT_FALLBACK_AFTER_PAGES", 1):
        await website.extract_text_from_url(home, max_depth=2, max_pages=3, tenant_id="org_x")

    assert order[0] == home
    assert "https://shop.example.com/sku/12345" in order[1:3]


@pytest.mark.asyncio
async def test_llm_fallback_failure_falls_back_to_insertion_order_without_raising():
    home = "https://shop.example.com/"

    async def fake_render(url, *a, **kw):
        if url == home:
            return _rendered(home, [(f"https://shop.example.com/browse/{i}", "x") for i in range(5)])
        return _rendered(url, [])

    fake_browser = AsyncMock()
    fake_browser.get_rendered_content = fake_render

    async def raising_classify(candidates, tenant_id):
        raise RuntimeError("boom")

    with patch.object(website, "get_browser", AsyncMock(return_value=fake_browser)), \
         patch.object(website, "_classify_product_links", raising_classify), \
         patch.object(website.settings, "CRAWL_PRODUCT_FALLBACK_AFTER_PAGES", 1):
        result = await website.extract_text_from_url(home, max_depth=2, max_pages=4, tenant_id="org_x")

    assert len(result["urls_visited"]) == 4
```

- [ ] **Step 3: Run test to verify it fails**

Run: `pytest tests/unit/test_crawl_prioritization.py -v`
Expected: FAIL — current implementation is pure BFS with no prioritization and never calls `_classify_product_links`, and `extract_text_from_url` does not yet accept a `tenant_id` kwarg.

- [ ] **Step 4: Write the implementation**

Replace the entire body of `extract_text_from_url` in `app/services/content/website.py` (the `async def crawl(...)` inner function and the final `await crawl(url, 0)` call at the bottom) with:

```python
async def extract_text_from_url(url: str, max_depth: int = 1, max_pages: int = 20,
                                  visited: Set[str] = None, tenant_id: str = None) -> Dict[str, Any]:
    """
    Crawls the website starting from URL up to max_depth and max_pages using a headless browser.
    Returns a dictionary containing aggregated text, metadata, and color palette.

    Visits product-shaped URLs before plain nav/collection links -- a strict
    depth-by-depth BFS can exhaust the whole page budget on a site's nav
    surface before ever reaching a product page (see the crawl-prioritization
    spec). A regex heuristic (_looks_like_product_url) handles the common
    case for free; if it finds nothing after CRAWL_PRODUCT_FALLBACK_AFTER_PAGES
    pages, one LLM call (_classify_product_links) re-ranks the frontier for
    sites with non-standard URL conventions.
    """
    if visited is None:
        visited = set()

    base_domain = urlparse(url).netloc
    results = {
        "text": "",
        "pages": [],
        "metadata": [],
        "colors": set(),
        "images": {},
        "urls_visited": []
    }

    browser = await get_browser()

    # Each frontier entry: (candidate_url, depth). product_priority holds
    # URLs known (by heuristic or LLM verdict) to be product pages -- sorted
    # to the front of every batch regardless of discovery order.
    frontier = [(url, 0)]
    product_priority = set()
    fallback_fired = False
    any_product_pattern_seen = False

    async def crawl_one(current_url, depth):
        visited.add(current_url)
        logger.info(f"Crawling URL (Headless): {current_url} (depth {depth}, count {len(visited)})")

        new_links = []
        try:
            content, links, pw_meta = await browser.get_rendered_content(current_url)

            if not content:
                logger.warning(f"No content rendered for {current_url}")
                return new_links

            soup = BeautifulSoup(content, 'html.parser')
            results["urls_visited"].append(current_url)

            page_meta = {
                "url": current_url,
                "title": pw_meta.get("title") or (soup.title.string if soup.title else ""),
                "description": ""
            }
            desc_tag = soup.find("meta", attrs={"name": "description"})
            if desc_tag:
                page_meta["description"] = desc_tag.get("content", "")
            results["metadata"].append(page_meta)

            styles = soup.find_all("style")
            for style in styles:
                hex_colors = re.findall(r'#[0-9a-fA-F]{3,6}', style.string or "")
                results["colors"].update(hex_colors)
                rgb_colors = re.findall(r'rgb\(.*?\)', style.string or "")
                results["colors"].update(rgb_colors)

            page_jsonld = extract_jsonld_blocks(content)

            for element in soup(["script", "style", "nav", "footer", "header", "aside"]):
                element.decompose()

            page_images = {}
            img_tags = soup.find_all("img")
            junk_keywords = ["facebook", "twitter", "linkedin", "instagram", "tiktok", "social", "icon", "arrow", "chevron"]

            for img in img_tags:
                img_url = img.get("src")
                if img_url:
                    full_img_url = urljoin(current_url, img_url)
                    url_lower = full_img_url.lower()
                    if any(key in url_lower for key in junk_keywords):
                        continue

                    label = img.get("alt") or img.get("title")
                    if not label:
                        filename = full_img_url.split("/")[-1].split("?")[0]
                        label = filename.rsplit(".", 1)[0].replace("-", " ").replace("_", " ").capitalize()

                    label = label.strip()[:100]

                    if full_img_url not in results["images"]:
                        results["images"][full_img_url] = label
                    page_images[full_img_url] = label

            raw_text = soup.get_text()
            lines = (line.strip() for line in raw_text.splitlines())
            chunks = (phrase.strip() for line in lines for phrase in line.split("  "))
            cleaned_text = '\n'.join(chunk for chunk in chunks if chunk)

            results["pages"].append({
                "url": current_url,
                "text": cleaned_text,
                "jsonld": page_jsonld,
                "html": content,
                "title": pw_meta.get("title") or (soup.title.string if soup.title else ""),
                "images": [{"url": u, "label": l} for u, l in list(page_images.items())[:10]]
            })

            results["text"] += f"\n--- Content from {current_url} ---\n{cleaned_text}\n"

            if depth < max_depth:
                for link in links:
                    next_url = link["href"]
                    next_text = link.get("text", "")
                    if is_subpath(url, next_url) and next_url not in visited:
                        if not any(next_url.lower().endswith(ext) for ext in ['.pdf', '.jpg', '.png', '.zip', '.gif']):
                            new_links.append((next_url, next_text, depth + 1))

        except Exception as e:
            logger.error(f"Failed to crawl {current_url} with Playwright: {e}", exc_info=True)

        return new_links

    # (url, text) pairs discovered but not yet queued -- used only to build
    # the LLM fallback's candidate list; the frontier itself only needs urls.
    pending_text_by_url = {}

    while frontier and len(visited) < max_pages:
        frontier.sort(key=lambda entry: 0 if entry[0] in product_priority or _looks_like_product_url(entry[0]) else 1)

        batch = []
        while frontier and len(visited) + len(batch) < max_pages:
            candidate_url, depth = frontier.pop(0)
            if candidate_url in visited or candidate_url in [b[0] for b in batch]:
                continue
            batch.append((candidate_url, depth))

        if not batch:
            break

        results_batch = await asyncio.gather(*[crawl_one(u, d) for u, d in batch])

        for new_links in results_batch:
            for next_url, next_text, next_depth in new_links:
                if next_url not in visited and next_url not in [f[0] for f in frontier]:
                    frontier.append((next_url, next_depth))
                    pending_text_by_url[next_url] = next_text
                    if _looks_like_product_url(next_url):
                        any_product_pattern_seen = True

        if (not fallback_fired and not any_product_pattern_seen
                and len(visited) >= settings.CRAWL_PRODUCT_FALLBACK_AFTER_PAGES
                and frontier):
            fallback_fired = True
            candidates = [(f_url, pending_text_by_url.get(f_url, "")) for f_url, _ in frontier]
            try:
                verdict = await _classify_product_links(candidates, tenant_id)
            except Exception as ex:
                logger.warning(f"Product-link classification fallback failed: {ex}")
                verdict = set()
            product_priority |= verdict

    results["colors"] = sorted(list(results["colors"]))
    results["images"] = [{"url": u, "label": l} for u, l in list(results["images"].items())[:20]]

    return results
```

- [ ] **Step 5: Run test to verify it passes**

Run: `pytest tests/unit/test_crawl_prioritization.py -v`
Expected: PASS

- [ ] **Step 6: Update both callers to pass `tenant_id`**

In `app/services/catalog/crawl.py:148`, change:

```python
        result = await extract_text_from_url(url, max_depth=20, max_pages=max_pages)
```
to:
```python
        result = await extract_text_from_url(url, max_depth=20, max_pages=max_pages, tenant_id=tenant_id)
```

Check the enclosing function's signature at `app/services/catalog/crawl.py` around line 130-148 for the exact parameter name holding the tenant/org id (`fetch_products(config, source_ref, tenant_id)` per the spec's Context section — confirm the actual local variable name in that function before editing, it must match what's already in scope there).

In `app/api/endpoints.py:182`, change:

```python
        scrape_result = await extract_text_from_url(url, max_depth=20, max_pages=settings.CRAWL_MAX_PAGES)
```
to:
```python
        scrape_result = await extract_text_from_url(url, max_depth=20, max_pages=settings.CRAWL_MAX_PAGES, tenant_id=tenant_id)
```

Check the enclosing endpoint function for the actual local variable holding the tenant id before editing (grep the function body above line 182 for how it's already named/obtained there — it is a FastAPI endpoint and almost certainly already has one in scope for auth).

- [ ] **Step 7: Run the full unit suite**

Run: `pytest tests/unit/ -v`
Expected: PASS, all tests including the ones from Tasks 1-3.

- [ ] **Step 8: Commit**

```bash
git add app/services/content/website.py app/core/config.py app/services/catalog/crawl.py app/api/endpoints.py tests/unit/test_crawl_prioritization.py
git commit -m "feat: prioritize product-shaped URLs in the crawl frontier"
```

---

### Task 5: Live validation against real, diverse sites

**Files:**
- Create (throwaway, not committed): a local script equivalent to the `/tmp/crawl_test.py` pattern already used during this investigation, calling `app.services.catalog.crawl.fetch_products(config, source_ref, tenant_id)` directly.
- No production files change in this task — it is a manual verification pass per the spec's "Live validation" section.

**Interfaces:**
- Consumes: `fetch_products` from `app/services/catalog/crawl.py` (unchanged signature).

- [ ] **Step 1: Write the throwaway validation script**

```python
import asyncio, json
from app.services.catalog.crawl import fetch_products

SITES = [
    "https://cordori.com.au/",
    "https://gymshark.com/",
    "https://kith.com/",
    "https://www.allbirds.com/",
    "https://www.bombas.com/",
    "https://www.fashionnova.com/",
    "https://www.athleticbrewing.com/",
    "https://www.skullcandy.com/",
    "https://www.hollisterco.com/",
]

async def main():
    for site in SITES:
        config = {"url": site, "max_pages": 60, "currency": "USD"}
        try:
            products = await fetch_products(config, "validation-run", "org_validation")
            print(f"{site}: {len(products)} products found")
            if products:
                print(f"  sample: {products[0].get('title') or products[0].get('name')}")
        except Exception as ex:
            print(f"{site}: EXCEPTION {type(ex).__name__}: {ex}")

asyncio.run(main())
```

- [ ] **Step 2: Run it, per-site, and record results**

Run: `PYTHONPATH=. python3 <script path>`

For each site: confirm the actual platform/URL pattern by opening one product link in the output (or curling the homepage for `wp-content`/`woocommerce`/Shopify markers) — do not assume from the spec's guesses. Record, per site: pages visited, products found, whether it was the regex heuristic or the LLM fallback that found the winning path (check logs for the "Product-link classification fallback" warning/absence and for how early a `/products/`-style URL was hit vs. a non-standard one).

- [ ] **Step 3: Confirm the acceptance bar from the spec**

Per the spec: every site in the validation set finds at least one real product (title + product_url populated) within the raised `max_pages` budget. If any site returns 0 products, treat it as a bug in Task 2/3/4 (not a reason to drop the site) and go back to those tasks — do not weaken the acceptance bar to make a stubborn site pass.

At least one site in the final validated set must have gone through the LLM fallback (not matched by the regex heuristic) to prove that path works against real markup, not just the mocked Task 3 tests. If none of the sites above end up doing so, actively find one with numeric/non-slug product URLs and add it to the set (per the spec's note that this category matters most).

- [ ] **Step 4: Report results**

No commit for this task (it's a manual verification pass, not a code change) — report the per-site table (site, pages visited, products found, heuristic vs. fallback) back in the conversation as the completion signal for this task.
