{
  "info": {
    "_postman_id": "c3f2a6a1-9e7d-4d5e-8e6a-2f4a3b1d5c90",
    "name": "GalaxiQ — Website Crawl → Pairing → Product List",
    "description": "Just the URL-crawl path, end to end: connect a merchant's website (no CSV, no API), run the crawl, poll it, then read back the products and their pairings.\n\nSTEP 1  Connect the website -- this now ALSO starts the build automatically (crawl -> enrich -> pair for the whole tenant), so most of the time STEP 2 is nothing to press\nSTEP 2  (optional) Manually trigger a build -- only needed if auto_build was disabled, or a build was already running when you connected (so this source wasn't included in it), or you just want to force a fresh one later\nSTEP 3  Poll the job until done\nSTEP 4  List products found by the crawl, and one product's pairings\n\nRequests save ids into collection variables, so the steps run in order without editing anything. Set base_url, tenant_id and site_url and go.\n\nCrawl defaults: CRAWL_MAX_PAGES=60, CRAWL_BATCH_SIZE=10 (traversal batches through the frontier in steps of 10 so the LLM fallback -- for sites whose product URLs don't match /products/, /product/, /p/, /item/ -- gets a chance to fire and actually be acted on mid-crawl, not just at the very end of the budget). Extraction also now understands Shopify's ProductGroup JSON-LD schema (brand/category/price straight from markup) and has an LLM fallback for brand/price on pages with no structured data at all.",
    "schema": "https://schema.getpostman.com/json/collection/v2.1.0/collection.json"
  },
  "variable": [
    { "key": "base_url", "value": "http://localhost:8001" },
    { "key": "tenant_id", "value": "org_8c32bf3e-6a18-4739-9b1c-94c0cf11125f" },
    { "key": "site_url", "value": "https://cordori.com.au" },
    { "key": "job_id", "value": "" },
    { "key": "product_key", "value": "" }
  ],
  "item": [
    {
      "name": "STEP 1 — Connect the website",
      "item": [
        {
          "name": "Connect a website as a product source",
          "request": {
            "method": "POST",
            "header": [{ "key": "Content-Type", "value": "application/json" }],
            "url": "{{base_url}}/sources/website",
            "description": "Only tenant_id and url are required. This saves the connection AND starts a build for the whole tenant right away (crawl this new source -> enrich everything -> pair everything) -- take the job_id in the response straight to STEP 3, no separate STEP 2 press needed in the normal case.\n\nurl must be https and publicly resolvable, or this returns 400 (SSRF guard).\n\njob_status in the response tells you what happened to the auto-triggered build:\n  queued                              a fresh build started; job_id is it\n  already_running_without_this_source another build was already running for\n                                       this tenant when you connected -- it\n                                       does NOT include this new source (it\n                                       already read the source list before\n                                       this one existed). job_id is that\n                                       OTHER build; once it finishes, use\n                                       STEP 2 to trigger a fresh one that\n                                       will include this source.\n  not_started                         the build couldn't be scheduled (e.g.\n                                       job store unavailable) -- the source IS\n                                       still connected, just use STEP 2\n                                       manually.\n\nEverything else is optional and left out of the example body below:\n  currency            never required to send. If crawled pages don't state\n                      it via structured markup, an LLM inference pass fills\n                      it in from page context (locale, symbol, shipping/tax\n                      text) instead of asking the caller for it.\n  max_pages           per-source override of the crawl page budget\n                      (server default is CRAWL_MAX_PAGES=60)\n  min_interval_hours  optional rate-limit: skip re-crawling this site if it\n                      was crawled within this many hours (default 24, 0 =\n                      always re-crawl). Add it back to the body only if you\n                      want to override the default while testing.\n  auto_build          set to false to only save the connection, matching the\n                      old two-step behaviour -- useful if you're about to\n                      connect several sources in a row and want one build at\n                      the end instead of one per source.",
            "body": {
              "mode": "raw",
              "raw": "{\n  \"tenant_id\": \"{{tenant_id}}\",\n  \"url\": \"{{site_url}}\"\n}",
              "options": { "raw": { "language": "json" } }
            }
          },
          "response": [
            {
              "name": "200 OK",
              "originalRequest": {
                "method": "POST",
                "header": [{ "key": "Content-Type", "value": "application/json" }],
                "url": "{{base_url}}/sources/website",
                "body": {
                  "mode": "raw",
                  "raw": "{\n  \"tenant_id\": \"{{tenant_id}}\",\n  \"url\": \"{{site_url}}\"\n}",
                  "options": { "raw": { "language": "json" } }
                }
              },
              "status": "OK",
              "code": 200,
              "_postman_previewlanguage": "json",
              "header": [{ "key": "Content-Type", "value": "application/json" }],
              "body": "{\n  \"kind\": \"crawl\",\n  \"external_ref\": \"cordori.com.au\",\n  \"status\": \"active\",\n  \"products\": \"pending_first_crawl\",\n  \"needs\": [],\n  \"job_id\": \"job_bd08789f384049aa97680293dd0a4830\",\n  \"job_status\": \"queued\"\n}"
            }
          ],
          "event": [
            {
              "listen": "test",
              "script": {
                "type": "text/javascript",
                "exec": [
                  "const body = pm.response.json();",
                  "if (body.job_id) pm.collectionVariables.set('job_id', body.job_id);",
                  "console.log('job_status:', body.job_status);"
                ]
              }
            }
          ]
        },
        {
          "name": "Confirm the source is connected",
          "request": {
            "method": "GET",
            "header": [],
            "url": "{{base_url}}/sources?tenant_id={{tenant_id}}",
            "description": "Look for kind=crawl, external_ref matching your site's hostname, status=active. products stays \"pending_first_crawl\" here until the auto-triggered (or manually triggered) build actually finishes -- this endpoint doesn't reflect job progress, only the connection itself."
          },
          "response": []
        }
      ]
    },
    {
      "name": "STEP 2 — (optional) Manually trigger a build",
      "item": [
        {
          "name": "Run crawl -> enrich -> pair",
          "request": {
            "method": "POST",
            "header": [],
            "url": "{{base_url}}/catalog/build",
            "description": "THE BUTTON — same call STEP 1 now fires automatically. Only press this manually if STEP 1's job_status was already_running_without_this_source or not_started, or you disabled auto_build, or you just want to force a fresh run later.\n\nOne call runs all three stages against every connected source for the tenant:\n\n  sync    0-10%   crawls each source (this is where the product-URL\n                  prioritization fix applies — the crawler visits\n                  product-shaped pages first instead of exhausting its\n                  budget on nav/collection pages). A source crawled within\n                  its min_interval_hours window is skipped here, not\n                  re-crawled, unless force=true.\n  enrich  10-60%  the model reads each product page and extracts\n                  attributes\n  pair    60-100% embeds each product and builds the pairing graph\n\nReturns immediately with a job_id — take that to STEP 3.\n\nPress it twice and the second call returns 409 with the job_id already running — poll that one instead.\n\nOptional form fields:\n  force=true   ignore the cache, re-read every product from scratch\n  wait=true    run synchronously and return the report instead of a\n               job_id — fine for this manual flow, but note it holds\n               the connection open for the whole crawl (can be minutes\n               on a large site)",
            "body": {
              "mode": "formdata",
              "formdata": [
                { "key": "tenant_id", "value": "{{tenant_id}}", "type": "text" },
                { "key": "force", "value": "true", "type": "text", "description": "forces a re-crawl even within min_interval_hours" }
              ]
            }
          },
          "response": [
            {
              "name": "202 Accepted",
              "originalRequest": {
                "method": "POST",
                "header": [],
                "url": "{{base_url}}/catalog/build",
                "body": {
                  "mode": "formdata",
                  "formdata": [
                    { "key": "tenant_id", "value": "{{tenant_id}}", "type": "text" },
                    { "key": "force", "value": "true", "type": "text" }
                  ]
                }
              },
              "status": "Accepted",
              "code": 202,
              "_postman_previewlanguage": "json",
              "header": [{ "key": "Content-Type", "value": "application/json" }],
              "body": "{\n  \"job_id\": \"job_bd08789f384049aa97680293dd0a4830\",\n  \"status\": \"queued\"\n}"
            }
          ],
          "event": [
            {
              "listen": "test",
              "script": {
                "type": "text/javascript",
                "exec": [
                  "const body = pm.response.json();",
                  "if (body.job_id) pm.collectionVariables.set('job_id', body.job_id);",
                  "pm.test('started', () => pm.response.to.have.status(202));"
                ]
              }
            }
          ]
        },
        {
          "name": "(alternative) Trigger pairing only",
          "request": {
            "method": "POST",
            "header": [],
            "url": "{{base_url}}/catalog/pair",
            "description": "Runs only the pair stage against products already in the catalogue — use this if products were already crawled/enriched and you just want to rebuild the pairing graph without re-crawling. Not needed in the normal flow; /catalog/build already includes this stage.",
            "body": {
              "mode": "formdata",
              "formdata": [
                { "key": "tenant_id", "value": "{{tenant_id}}", "type": "text" }
              ]
            }
          },
          "response": [
            {
              "name": "202 Accepted",
              "originalRequest": {
                "method": "POST",
                "header": [],
                "url": "{{base_url}}/catalog/pair",
                "body": {
                  "mode": "formdata",
                  "formdata": [{ "key": "tenant_id", "value": "{{tenant_id}}", "type": "text" }]
                }
              },
              "status": "Accepted",
              "code": 202,
              "_postman_previewlanguage": "json",
              "header": [{ "key": "Content-Type", "value": "application/json" }],
              "body": "{\n  \"job_id\": \"job_efd17575540b4b0ead65e9e8a3bb83c8\",\n  \"status\": \"queued\"\n}"
            }
          ]
        }
      ]
    },
    {
      "name": "STEP 3 — Poll the job",
      "item": [
        {
          "name": "Check progress",
          "request": {
            "method": "GET",
            "header": [],
            "url": "{{base_url}}/catalog/jobs/{{job_id}}",
            "description": "Send every 2-3 seconds while running. status: queued | running | done | failed | lost. Stop polling once it's done/failed/lost.\n\nOn done, result holds the summary — for a build job this includes the sync stage's crawl results (pages visited, products found)."
          },
          "response": [
            {
              "name": "200 OK — done",
              "originalRequest": {
                "method": "GET",
                "header": [],
                "url": "{{base_url}}/catalog/jobs/{{job_id}}"
              },
              "status": "OK",
              "code": 200,
              "_postman_previewlanguage": "json",
              "header": [{ "key": "Content-Type", "value": "application/json" }],
              "body": "{\n  \"job_id\": \"job_bd08789f384049aa97680293dd0a4830\",\n  \"kind\": \"build\",\n  \"tenant_id\": \"org_8c32bf3e-\\u2026\",\n  \"status\": \"done\",\n  \"percent\": 100,\n  \"step\": \"writing pairing graph\",\n  \"started_at\": \"2026-08-23T15:04:02Z\",\n  \"updated_at\": \"2026-08-23T15:07:41Z\",\n  \"result\": {\n    \"stages\": {\n      \"sync\": { \"products\": 38 },\n      \"enrich\": { \"products\": 38, \"extracted\": 38 },\n      \"pair\": { \"products\": 38, \"pairs\": 64, \"servable\": 60, \"queued\": 4 }\n    }\n  },\n  \"error\": null\n}"
            }
          ],
          "event": [
            {
              "listen": "test",
              "script": {
                "type": "text/javascript",
                "exec": [
                  "const s = pm.response.json();",
                  "console.log(s.percent + '% - ' + (s.step || s.status));",
                  "pm.test('valid status', () => pm.expect(",
                  "  ['queued','running','done','failed','lost']).to.include(s.status));"
                ]
              }
            }
          ]
        }
      ]
    },
    {
      "name": "STEP 4 — Products found by the crawl",
      "item": [
        {
          "name": "List products (this crawl source)",
          "request": {
            "method": "GET",
            "header": [],
            "url": "{{base_url}}/catalog/products?tenant_id={{tenant_id}}&limit=50&offset=0",
            "description": "Paged. `total` is the full count. Products from a website source have source_kind=\"crawl\" internally (visible on the single-product endpoint below), product_key is crawl:<hostname>:<product_url>."
          },
          "response": [
            {
              "name": "200 OK",
              "originalRequest": {
                "method": "GET",
                "header": [],
                "url": "{{base_url}}/catalog/products?tenant_id={{tenant_id}}&limit=50&offset=0"
              },
              "status": "OK",
              "code": 200,
              "_postman_previewlanguage": "json",
              "header": [{ "key": "Content-Type", "value": "application/json" }],
              "body": "{\n  \"products\": [\n    {\n      \"product_key\": \"crawl:cordori.com.au:https://cordori.com.au/products/nimbus-wool-bomber\",\n      \"name\": \"Nimbus Wool Bomber\",\n      \"category\": \"Bomber Jackets\",\n      \"brand\": \"Cordori\",\n      \"price_cents\": 64900,\n      \"currency\": \"AUD\",\n      \"in_stock\": true,\n      \"image_url\": \"https://cordori.com.au/cdn/shop/files/2_4fe4640a-526c-461a-aa8e-89ca74df1dd3.png\"\n    }\n  ],\n  \"total\": 38\n}"
            }
          ],
          "event": [
            {
              "listen": "test",
              "script": {
                "type": "text/javascript",
                "exec": [
                  "const first = (pm.response.json().products || [])[0];",
                  "if (first) pm.collectionVariables.set('product_key', first.product_key);"
                ]
              }
            }
          ]
        },
        {
          "name": "One product, in full",
          "request": {
            "method": "GET",
            "header": [],
            "url": "{{base_url}}/catalog/products/{{product_key}}?tenant_id={{tenant_id}}",
            "description": "product_key contains colons and slashes (crawl:cordori.com.au:https://...) — URL-encode it (encodeURIComponent) before putting it in the path, or this 404s.\n\nmissing_fields shows which fields couldn't be extracted for this product. price/brand/category now come from structured markup when the page publishes it (including Shopify's ProductGroup schema, used by many current themes for products with variants), with an LLM fallback for pages with no structured data at all -- a field only lands in missing_fields if genuinely neither source had it."
          },
          "response": []
        },
        {
          "name": "What this product goes with",
          "request": {
            "method": "GET",
            "header": [],
            "url": "{{base_url}}/catalog/products/{{product_key}}/pairings?tenant_id={{tenant_id}}",
            "description": "Grouped into similar / complement / upsell. servable=false means it's still waiting for merchant approval (see the general collection's STEP 5 for the approval endpoints — omitted here since this collection is scoped to the crawl path only)."
          },
          "response": []
        }
      ]
    }
  ]
}
