# Background Jobs and Polling — Design

**Date:** 2026-08-11
**Status:** Draft

---

## 1. Purpose

Stop the long catalogue jobs holding an HTTP connection open, and give a client
something to poll.

Measured on the live 218-product catalogue:

```
POST /catalog/sync      seconds
POST /catalog/enrich    ~3½ minutes
POST /catalog/pair      ~90 seconds
```

Every one of those is a synchronous request today. At 218 products it merely
feels slow; at a few thousand the client hits a gateway timeout long before the
work finishes, and there is no way to find out whether it succeeded.

**Success criteria:**

1. The three long endpoints return immediately with a `job_id`.
2. A client can poll one endpoint for status and a percentage.
3. Two jobs of the same kind cannot run on one tenant at once.
4. A job whose process died is reported as dead, not left "running" forever.

## 2. The contract

```http
POST /catalog/enrich          →  202 Accepted
{"job_id": "job_7f3a2c91", "status": "queued"}

GET /catalog/jobs/job_7f3a2c91  →  200
{"job_id":    "job_7f3a2c91",
 "kind":      "enrich",
 "tenant_id": "org_...",
 "status":    "running",
 "percent":   45,
 "step":      "extracting attributes — batch 5 of 11",
 "started_at":"2026-08-11T10:02:11Z",
 "updated_at":"2026-08-11T10:03:48Z",
 "result":    null,
 "error":     null}
```

`status` is one of `queued`, `running`, `done`, `failed`, `lost`.

**On `done`, `result` holds exactly the report the endpoint returns today.**
Nothing downstream has to learn a new shape — a caller that used to read
`{"products": 218, "extracted": 218, ...}` reads the same object from
`result`.

**On `failed`, `error` holds the exception class name only**, matching the
existing rule that no upstream message reaches a caller.

## 3. Where progress comes from

The percentage must be real. A bar that jumps 0 → 100 is worse than none,
because it teaches the user to distrust it.

**Enrich** is already batched, so it is honest arithmetic:
`batch 5 of 11` → 45%.

**Pair** has phases, weighted by their measured share of the 90 seconds:

| Phase | Weight |
|---|---|
| load eligible products | 5% |
| embed (cached where possible) | 40% |
| score candidate pairs | 45% |
| write the graph | 10% |

**Sync** reports per source, and per page within a source.

Each job function takes an optional `progress` callback. **Passing nothing keeps
today's behaviour**, so the functions stay directly callable and every existing
test still exercises them unchanged.

## 4. Execution

In-process: the request hands the work to a thread and returns. No queue, no
worker, no new deployment.

**The honest limits of that**, recorded rather than discovered later:

- A job dies with the process. It cannot be resumed.
- It does not work across replicas — a poller can reach a process that is not
  running the job.

**Both are survivable because of the heartbeat in §6, and neither changes the
API.** Moving to a real queue later swaps the executor and leaves the contract in
place.

## 5. Storage

A Redis hash per job, `job:<id>`, TTL 24 hours. Redis is already a dependency
with a managed async pool, and it already holds the OAuth state nonce.

Job state is deliberately **not** in Postgres: it is ephemeral operational
detail, not tenant data, and a 24-hour TTL is the correct lifetime for something
a client polls for a few minutes.

## 6. The heartbeat

A running job writes `updated_at` on every progress tick.

`GET /catalog/jobs/{id}` reports `lost` when a job claims to be `running` but has
not ticked for `STALE_AFTER_SECONDS` (300). That is how a client learns the
process died, rather than polling `running` forever.

Without this, an in-process executor has no failure story at all — which is the
single thing that would make the design dishonest.

## 7. One job per tenant per kind

A Redis key `job:lock:<tenant>:<kind>` is taken for the job's lifetime.

Starting a second one returns **409** with the running job's id:

```json
{"detail": {"error": "already_running", "job_id": "job_7f3a2c91"}}
```

Not a nicety: two concurrent `pair` runs on one tenant both delete and rewrite
`strategist_product_neighbors`, so one would interleave with the other's writes.

The lock carries the same TTL as the job, so a dead process cannot block a tenant
permanently.

## 8. Endpoints

```
POST /catalog/sync         →  202 {job_id}
POST /catalog/enrich       →  202 {job_id}
POST /catalog/pair         →  202 {job_id}
GET  /catalog/jobs/{id}    →  200 job state, 404 if unknown or expired
GET  /catalog/jobs?tenant_id=…  →  recent jobs for a tenant
```

`POST /catalog/import/csv` stays synchronous. It parses an uploaded file in
memory in under a second, and a client that just uploaded a file expects an
answer about it.

## 9. Backwards compatibility

The three endpoints change from returning a report to returning a `job_id`, and
from `200` to `202`. **That is a breaking change** for anything already calling
them.

`?wait=true` keeps the old behaviour: run synchronously and return the report,
exactly as today. Existing callers add one query parameter rather than adopting
polling. The internal admin page uses it, since its jobs are small.

## 10. Error handling

| Condition | Behaviour |
|---|---|
| Job raises | `status: failed`, `error` is the exception class name, lock released |
| Process dies mid-job | `status: lost` after the heartbeat window |
| Unknown or expired `job_id` | `404` — a client polling past the TTL is told plainly |
| Redis unreachable at start | `503`; the job does not start, because a job whose progress cannot be recorded is worse than no job |
| Redis unreachable mid-job | The work continues; progress ticks are best-effort and logged, never fatal |

That last row matters: losing the *reporting* channel must not abort *work* that
is already spending money on model calls.

## 11. Testing

- a started job returns 202 and a job id, and the work runs
- polling moves through queued → running → done
- `percent` never goes backwards and ends at 100
- a failed job reports `failed` and the exception class, never the message
- a second job for the same tenant and kind gets 409 with the first job's id
- a different kind, or a different tenant, is allowed to run concurrently
- a stale heartbeat reports `lost`
- an unknown job id is 404
- `?wait=true` returns the report inline, with no job id
- the job functions still work when called with no progress callback — which is
  what every existing test does

## 12. Out of scope

A real queue; retries; cancellation; cross-replica execution; progress for the
CSV import; websocket push instead of polling.
