# Job lifecycle — QA note ## Job types → families | `job_type` value | Family | MLC call | Async callback? | |---|---|---|---| | anything else (live / capture / video) | **live** | `POST {mlc}/api/v2/predict_batch` | **Yes** — start and stop both confirmed by callback | | `image_push`, `image_push_type1`, `image_push_type2` | **image_push** | `POST {mlc}/api/v2/push_image` | No — sync accept is the confirmation | | `image_pull` | **image_pull** | `POST {mlc}/api/v2/pull_image` | No — sync accept is the confirmation | ## States | State | Meaning | UI effect | |---|---|---| | `START_PENDING` | Start accepted by the MLC; waiting for its async callback | Both Start/Stop buttons gray | | `RUNNING` | Job confirmed running | Stop enabled (toggle green) | | `STOP_PENDING` | Stop accepted; waiting for the MLC to confirm teardown | Both buttons gray | | `STOPPED` | Cleanly stopped (user stop, abort, or confirmed teardown) | Start re-enabled | | `FAILED` | MLC refused the start (strict mode) or reported FAILED/ERROR in a callback | Start re-enabled (toggle red) | | `INCOMPLETE` | The MLC died mid-flight — health monitor reconciled the job | Start re-enabled | ## Actions (all on `job_app`, JWT required except the callback) - `POST /jobs/start_job` (aliases `/start_live_job`, `/v1/*`, legacy `/live/start_job`) - `POST /jobs/stop_job` (aliases `/stop_live_jobs`, `/v1/*`) — takes `request_id` **or** `device_id` (picks the newest RUNNING/START_PENDING job on that device) - `POST /jobs/abort_job` — operator escape hatch for stuck PENDING rows - `POST /jobs/job_status_callback` (aliases `/v2/job_status_callback`, `/add_benchmark_result`) — **no auth**; the MLC posts `{requestId, state, statusMessage}` (flat or nested under `status`); the uuid `request_id` is the credential - `POST /jobs/mlc/healthcheck` — runs the health sweep on demand (also runs on a timer) - `DELETE /jobs/v2/job/{request_id}` — deletes the row entirely ## Flow per family ### Live / video (async, two-phase) ``` start_job ── MLC predict_batch sync ──┬─ refuses ──────────► FAILED └─ accepts ──► START_PENDING START_PENDING ── callback SUCCESS/RUNNING/COMPLETE ─────────► RUNNING ── callback FAILED/ERROR ─────────────────────► FAILED (capacity released) RUNNING ── stop_job ── MLC accepts stop ────────────────────► STOP_PENDING STOP_PENDING ── any callback ───────────────────────────────► STOPPED any of START_PENDING / STOP_PENDING / RUNNING ── abort_job ─► STOPPED (immediate, no callback wait) ``` Notes for testing: - A job **stays RUNNING** until an explicit stop — later `COMPLETE` callbacks do not stop it (they re-confirm RUNNING). - Terminal rows stay terminal: a late callback must **not** resurrect a STOPPED or FAILED job. - Live jobs consume an MLC capacity slot on start; released on stop dispatch, on a FAILED callback, or by abort. Capacity must never leak or double-decrement (check `mlc.current_capacity`). ### Image push / image pull (sync, single-phase) ``` start_job ── MLC push_image|pull_image (init=start) ──┬─ refuses ─► FAILED └─ accepts ─► RUNNING (immediately, no callback) RUNNING ── stop_job (init=stop) ──────────────────────────────────► STOPPED (immediately) ``` No PENDING phases at all — if you see an image job stuck in START_PENDING, that's a bug. ## Cross-cutting rules to test 1. **One job per camera**: starting while a job on that device is RUNNING / START_PENDING / STOP_PENDING → HTTP 409, no MLC call made. 2. **Dispatch mode** (`settings.mlc_dispatch_mode`): - `strict` — MLC error on start → job FAILED + HTTP 502; MLC error on stop → HTTP 502 and the row **stays RUNNING** (DB mirrors the cluster). - `best_effort` — error recorded on the row, job goes RUNNING anyway. - `off` — no MLC call; job goes straight to RUNNING (bookkeeping only). 3. **Dead MLC reconcile**: kill the MLC → within the health interval (or via `POST /jobs/mlc/healthcheck`) its RUNNING **and** PENDING jobs flip to `INCOMPLETE` with an error message, and `current_capacity` resets to 0. 4. **Abort**: works even when the MLC is unreachable (stop dispatch failure is ignored); row ends STOPPED with `error_message = "aborted by operator"`; capacity slot released. 5. **Callback shape tolerance**: both the flat and the apprunner-nested `{"status": {...}}` payloads must be accepted; unknown `requestId` → 404; unknown states are acknowledged with no transition. 6. Every action publishes to Kafka `dori.jobs.lifecycle` (`job_started` / `job_stopped` / `job_aborted` / `job_status_callback`) — verify events if the test covers integrations.