# Feature list / API gaps Running list of pending features, API gaps, and correctness issues to pick up. Add new items at the bottom; mark items done with `[x]` and the commit/PR. | # | Area | Description | Status | |---|------|-------------|--------| | 1 | mlc_config API | Does not have a `companyId` parameter to filter the configurations based on either the logged-in user's company or a `companyId` parameter. | [x] Done — `list_mlc_config` `?company_id=`, tenant-guarded (`infra.py`) | | 2 | mlc_config API | Returns a single record. Should return all the configurations for the given `mlcId` and/or `companyId`. | [x] Done — returns a list (`infra.py list_mlc_config`) | | 3 | MLC API | No API for listing the MLCs associated with a company. | [x] Done — `GET /infra/mlc?company_id=` returns MLCs linked to a company via the config junction (v4) | | 4 | list_employees API | Returns the users for the `companyId`. `companyId` should be validated against the logged-in user when the user type is `USER_ADMIN`. | [x] Done — `resolve_target_company_id` (`ops.py:278`, `list_employees` route) | | 5 | credentials API | Exception handling is incorrect — returns "Invalid Authorization" in case of invalid input. | [x] Done — 400/422, no auth-error leak (`infra.py set_credentials`) | | 6 | mlc_config API | Exception handling is incorrect — returns "Invalid Authorization" in case of invalid input. | [x] Done — 404/422 (`infra.py create_mlc_config`) | | 7 | Controllers (all) | Check for controller methods raising `ValueError` (audit that they map to proper 4xx responses, not auth errors). | [x] Done — validator ValueErrors→422, JSON→400 | | 8 | MLC update API | Should have the capability to set `serviceBundleId` to `None` (disassociate the MLC and service bundle). | [x] Done — update uses `exclude_unset`; `serviceBundleId:null` clears the column (`infra.py create_mlc`) | | 9 | Camera create | Default setting of `mlcId` in camera creation is to be removed. | [x] Done — no default injected (`ops.py create_device`) | | 10 | service_bundle API | Input validation gaps vs legacy appserver: `POST /config/infra/service_bundle` accepts an empty `name` (creates a nameless bundle row; legacy rejected it). Found by the dash system test negative cases. Audit the other create endpoints for required-field validation the legacy server enforced. | [x] Done — empty→None→422 via `DoriInput._legacy_none_string` | | 10a | Validation audit | More legacy-vs-dash validation gaps from the dash system test negative cases: location create accepts invalid company/timezone/region; company create accepts empty parameters; credentials update accepts an invalid credentials id (upsert semantics); receiver create accepts an invalid email; user create accepts an invalid `userRole`. Each recorded as FAILURE rows in `system_test/dash_sys_test/output_results_mapp_run.csv`. | [x] Mostly done — timezone/email/userRole/provider validated; `region` value still free-form (presence-checked only) | | 10b | event_type API | No write API for the event-type catalog: legacy `POST /configure/events` created per-company event types; dash-app only has read-only `GET /config/ops/event_type` (event types come from service bundles). All 26 add_configure_events system-test cases fail. Decide: port the write API or declare bundle-driven catalogs canonical and retire the cases. | [x] Decided — bundle-driven canonical, write API retired (see #10c) | | 10c | event_type API | Skip execution of write API for the event-type catalog | [x] Done — `infra.py` already makes `/infra/event_type` (and `/ops/event_type`) read-only by design ("Events are a PLATFORM catalog owned by service bundles... no manual create/edit/delete"); events are populated from a bundle archive's `events.json` at `/infra/service_bundle/upload` instead. Formalizing that as the answer to #10b: `add_configure_events` and its 26 `system_test/dash_sys_test` cases test a per-company write path that will not be built — retire those cases rather than porting `POST /configure/events`. | | 11 | Stream status | Implement real stream status for the COD camera card. Today `streamStatus` just echoes the stored `device.status` column, which nothing maintains — UI-created cameras can never show ACTIVE. Port apprunner's semantics: while a job is RUNNING, derive stream health from the latest stream events on the job (`Stream Available` / `Stream Unavailable` / FPS-threshold events, per `_getStreamStatus` in apprunner `dar_model.py`); `UNKNOWN` when no job is running. QA re-confirmed 2026-07-29: "camera state remains Inactive even after a job has been started." | Open | | 12 | MLC health reconcile | Two coupled problems behind "pull jobs stop by themselves after ~2 min" (QA 2026-07-29). (a) The health probe timeout is 2s (`mlc_health.py _HEALTH_TIMEOUT`) with a 60s sweep — a CPU-only MLC busy with inference can miss the 2s window, get marked NOT ACTIVE, and have its in-flight jobs flipped to INCOMPLETE even though it is fine. Make the timeout/`fail-after-N-consecutive-misses` configurable and tolerant. (b) When the sweep marks jobs INCOMPLETE it never dispatches a stop to the MLC — if the MLC was actually alive (or comes back), it keeps running the job, giving the observed dashboard-vs-MLC mismatch. Reconcile should best-effort `dispatch_stop` the orphaned jobs. | [x] Done — configurable timeout + `mlc_health_miss_threshold`; `_reconcile_down_jobs` dispatch_stop (`mlc_health.py`) | | 13 | Stop confirmation timeout | VIDEO JOBS ONLY: the MLC sends start confirmations and image-pull status callbacks fine, but never a stop confirmation for live streams — so a stopped video job sits in STOP_PENDING until manually aborted (QA 2026-07-29; verified on prod: live jobs 354/356 have `api_callback_count=1` — the start callback — and no second callback after stop). Add a stop-confirmation timeout — MINIMUM 1 HOUR (per Ravi 2026-07-29: a legitimate stop confirmation can take a long time; never terminalize earlier): if no callback within the timeout (config default ≥ 3600s) after a successful stop dispatch, terminalize to STOPPED. Manual abort remains the immediate escape hatch. | [x] Done — `mlc_stop_timeout.py` sweep + migration 008 `stop_requested_at`; default 3600s | | 14 | Job status UI refresh | Start confirmation WORKS: prod data shows the MLC's callback flips START_PENDING → RUNNING ~18s after start (`job.job_start` set on jobs 354/356). QA saw "PENDING even though inference is running" because the dashboard doesn't re-poll after start — the Devices page should refresh job status (poll or push) for ~60s after a start/stop action so the PENDING → RUNNING flip is visible without a manual reload. | [x] Done — FE 60s poll while transient (v4); backend jobStatus feed already done | | 14a | Dual MLC health monitors | TWO independent pollers probe `/api/v1/_health` and write the same `mlc.status`/`last_healthcheck`: job_app `mlc_health.py` (60s sweep, 2s timeout, 1 miss → NOT ACTIVE, jobs → INCOMPLETE, capacity reset) and monitor `mlc_check.py` (30s loop with backoff, 5s timeout, 3 misses → jobs → FAILED, Kafka MLC_UP/MLC_DOWN alerts). Consequences: status flapping under load (2s vs 5s timeouts disagree → spurious alerts on every flip), contradictory job outcomes (INCOMPLETE vs FAILED for the same dead MLC), and monitor's reconcile filters on the nonexistent status literal `PENDING` (real states are START_PENDING/STOP_PENDING) so its pending-job cleanup is dead code. Consolidate with feature 12: one health authority (job_app: probe + job/capacity reconcile), monitor alerts off status transitions instead of probing independently. | [x] Done — job_app is sole authority; monitor only emits Kafka on status transitions; dead `PENDING` filter removed (`mlc_check.py`) | | 15 | Bundle picker scoping | Device add/edit form lists ALL service bundles instead of only those associated with the device's company (`service_bundle_company` junction). Filter the dropdown by company; API side may need `GET /config/infra/service_bundle?company_id=`. | [x] Done — FE picker now company-scoped via /config/ops/service_bundle (v4); API filter also available | | 16 | Job-type by media type | Dashboard should only offer job types valid for the device's configured Source Media Type: Video → live only; Image → Pull / Type1 Push / Type2 Push only. Today the picker is unfiltered. | [x] Done — FE DataCard filters job types by Source Media Type (v4); job_app enforces server-side | | 17 | MLC ↔ job ↔ device visibility | Surface which MLCs are running which devices, in both directions (Ravi 2026-07-30). (a) Job views (COD Jobs list / job detail): each running job should show the MLC it is dispatched on AND the camera/device it is running against. (b) ICD MLC list: each MLC should show which company and which camera(s) it is currently serving (MLCs are platform-level, so this is the only place that cross-tenant view can live). Data path exists already — job → device → mlc_config → mlc and job dispatch picks the MLC — so this is join + API shape + UI columns/detail panels; no new state to maintain. | [x] Done — FE JobsCOD Camera+MLC columns; ICD MLC list Cameras column (v4); APIs done earlier | | 18 | MLC list active filter | ICD MLCs page: add a filter control (button/toggle at the top of the list, next to the search box) to show Active / Not Active / All MLCs, driven by `mlc.status`. Default All. (Ravi 2026-07-30) | [x] Done — FE Active/Not Active/All ToggleButtonGroup, verified live 110/18/91 (v4); API ?status= done | | 19 | Consistent field naming | The API accepts and emits a mix of camelCase and snake_case. Request models bridge both via `AliasChoices`; responses emit duplicate mirror keys (snake_case + camelCase); and there are one-off spellings (`queueLowerbound` vs `queueLowerBound`, `mediaServerIP` vs `mediaServerUrl`, `serviceBundleId` vs `service_bundle_id`, `role`/`userRole`/`user_role`). Pick one canonical convention at the API boundary (camelCase for React) and normalize request + response shapes, keeping explicit back-compat aliases only where old appserver/fe-icd clients still depend on them. | Partial / blocked — **request side is already canonical** (camelCase via `DoriInput.alias_generator`, legacy accepted via `AliasChoices`; `pModel.py:8-14`). **Response side cannot drop the mirror keys backend-only**: the React FE (`../icd/smartvision_va`) reads a mix — `mediaServerIP` ×13, `serviceBundleId` ×50 **and** `service_bundle_id` ×8, snake_case `is_active` ×8 / `service_bundle_url` ×2. Removing any mirror breaks specific FE screens. Completing #19 needs a coordinated FE PR: normalize the FE to canonical camelCase, then drop backend mirrors. | | 20 | Rename `d_auth_aws` → `cloud_store_config` | The table now holds multi-provider cloud storage config (aws/azure/gcp: `provider`, plus Azure `account_name`/`container_name`/`endpoint` and GCP `project_id`/`service_account_json`) — it is no longer AWS auth. The `d_auth_aws` name (and the `DAuthAws` model) is misleading. Rename the table to `cloud_store_config` via alembic, rename the model + all references, and keep the `/config/infra/credentials` API and response keys stable (or aliased) so clients don't break. FKs to update: `d_dataset.d_auth_aws_id`, `d_model.d_auth_aws_id`. | [x] Done — migration `010_rename_cloud_store` (table + uq constraint + both FK cols/constraints); model `DCloudStoreConfig` with `DAuthAws` back-compat alias; API/response keys unchanged. Deployed to mapp (5 consumer services), 115 tests green | | 21 | BUG · Drop orphan `d_company_sender` table | Dead table: it exists only in `alembic/sql/create_dori_schema.sql` — no SQLAlchemy model, no code, no API, and nothing FKs into it (it only FKs out to `company`). Drop it via an alembic migration, remove it from the raw schema SQL, and drop the sender row/reference from `docs/company-creation-data-model.html`. Target: next release. | [x] Done — migration `009_drop_company_sender`; removed from schema SQL + data-model docs | | 22 | DB cleanup · drop unused tables + schema | Remove dead/unused data-model tables (and a whole schema) carried over from the canonical dori_db port that the dash app never uses. Requested set: `d_dataset`, `d_dataset_image`, `d_deployment_log`, all "Analytics" tables, `d_training_benchmarks`, `d_use_case`, `d_project`, `pallet_result`, `d_inventory`, `d_inventory_kit`, `d_inventory_location`, `d_job_image`, `d_schedule`, and schema `inference`. Audited 2026-08-11 (prod existence + row counts + code/model/FK refs) — **NOT uniform**: splits into 22a/22b/22c by risk. Drop children before parents. Follow the #21/#20 convention (drop via alembic migration; leave the `001` baseline SQL immutable). | Partial — 22a+22b done (migration 011 dropped 12 orphan tables, v4); 22c still needs decision | | 22a | DB cleanup · orphan tables (safe) | No SQLAlchemy model, no code, present only in the `create_dori_schema.sql` baseline (same profile as #21), empty on prod: `d_dataset_image`, `d_deployment_log`, `d_training_benchmarks`, `d_use_case`, `d_inventory_kit`, `d_inventory_location`, `d_job_image`, `d_schedule`. Safe to drop in one migration. Note `d_dataset_image`→`d_dataset` and `d_inventory_kit`/`d_inventory_location`→`d_inventory` FKs — drop these children with their parents in 22b. | [x] Done — migration 011 dropped all 8 orphan tables (+ dead `d_service_bundle_use_case` junction); empty, no model/code (v4) | | 22b | DB cleanup · model-backed (verify first) | Have a SQLAlchemy model in `dModel.py` but no apparent runtime CRUD — confirm unused before dropping: `d_dataset`, `d_project`, `d_inventory`. Dropping `d_dataset` must first resolve `d_model.dataset_id` (FK from `d_model`, which is **not** in the drop set) and its `d_dataset_image` child; `d_inventory` owns children `d_inventory_kit`/`d_inventory_location`. All empty/near-empty on prod. | [x] Done — migration 011 dropped `d_dataset`/`d_project`/`d_inventory` (+children); confirmed no models/code, empty (v4) | | 22c | DB cleanup · LIVE — needs decision/redesign | **Do NOT drop as-is.** `pallet_result` is warehouse_app's scan-results table (`main.py`, `dashboard.py`, `scan_consumer.py`, `storage.py`, migration `002`) — dropping breaks WMS. Schema `inference` (11 tables incl `inference_event`/`_predicted`/`_alerts` + TimescaleDB hypertables) backs the events/inference pipeline: `dIdoModel.py`, analytics `events.py`/`tables.py`, DIM listener, the `vw_events` views, and the LISTEN/NOTIFY trigger — dropping breaks DIM, analytics, and event views. **"All tables in Analytics"**: no analytics schema exists and analytics_app owns no tables (it reads the inference schema + `vw_events`) — needs the requester to define what "Analytics" means before any action. | Open (needs decision) | | 23 | DB image | Change the Postgres base image to `timescale/timescaledb:latest-pg17` to align with the dori_db docker. (Code review 2026-08-18) | Open — currently `latest-pg16` in all compose files; **pg16→pg17 is a major-version bump** so an existing data dir needs `pg_upgrade`/dump-restore, not just an image swap. | | 24 | Architecture | All logic is implemented in the web layer. Move business logic into a service layer; keep controllers thin. (Code review 2026-08-18) | [x] Done (2026-08-26) — `infra.py`, `ops.py`, `common.py` extracted into 13 `service/*.py` modules (company, credentials, mlc, mlc_config, client, bundle, location, receiver, event_config, alert rule, device, employee, auth, plus the existing users.py extended). Web routers now only parse requests, wire `Depends` auth, and shape responses (the `_*_dict` serializers stay in web.py by design — that's #25, tracked separately). Verified against a **live stack**, not just review: stood up an isolated Postgres + the actual `services/config_app/Dockerfile` image + nginx in Docker, ran migrations + seed, and ran the full `unit_tests/test_config_app.py` suite (84 tests) before and after — 84/84 green both times, including a rebuild of the real Docker image (not just a volume-mounted dev run). `compat.py` intentionally left untouched — it's a separate legacy implementation slated for merge under #33, not #24's scope. | | 25 | Responses | Responses should use pModel classes instead of local dict builders (e.g. `_user_dict`). (Code review 2026-08-18) | Open — ~10 `_*_dict` helpers emit triplicate snake/camel keys; same root as #19, and blocked on the same FE coordination (the mirror keys are load-bearing for specific FE screens — see #19). Deliberately not attempted as a mechanical pass: swapping every handler's return to a pModel risks silently dropping a mirror key an FE screen depends on, with no live stack here to catch it. | | 26 | User create | User creation should also create a corresponding receiver record. (Code review 2026-08-18) | [x] Done — every insert path (`common.py create_user`, `infra.py create_infra_user`, `compat.py compat_create_user`, and the company-create admin-onboarding row) also adds a matching `DReceiver` row. Consolidated 2026-08-25 into `service/users.py insert_user_with_receiver` (#32); that pass also caught `compat.py compat_create_company`'s own separate admin-onboarding insert, which had been missed by the original fix and never got a receiver row either. | | 26a | User create | Receiver row insert should set the created users's id to the receiver record | Open | | 27 | Schema cleanup | Drop the `is_confirmed` flag from `app_user`. (Code review 2026-08-18) | [x] Done — migration `012_drop_is_confirmed`; column removed from `DUser`, `UserCreate`, every `DUser(...)` construction site, and `seed.py`. | | 28 | Company update | Company update + `id`-parameter check. (Code review 2026-08-18) | [x] Done (pre-v4) — `create_or_update_company` upserts on `id` (`service/company_service.py:123-130`), guarded; `test_infra_company_edit` covers it. Flagged in review likely against an older build. | | 29 | MLC/Location update | MLC & Location update + `id`-parameter check. (Code review 2026-08-18) | [x] Done (pre-v4) — `create_or_update_mlc` keys on `req.id` (`service/mlc_service.py:122`); `create_or_update_location` upserts on `id` w/ 404 guard (`service/location_service.py:22`). | | 30 | Identifiers | In create mode the `_id` field should be a server-minted GUID, not an input parameter. (Code review 2026-08-18) | [x] Done — `company_id`/`user_id` (infra.py), `device_id`/`location_id` (ops.py) were already server-minted via `uuid4`, client input dropped. Audit found `compat.py`'s legacy `compat_create_device`/`compat_create_location` were the gap: device honored a client-supplied `deviceId` outright, and location never set `location_id` at all. Both now mint server-side too (2026-08-25). | | 31 | Security bug | Default user created during company creation should be `USER_ADMIN`, not `DORI_ADMIN`. (Code review 2026-08-18) | [x] Done — `create_company`'s admin-onboarding row now mints `USER_ADMIN` (`infra.py`). Audit found the same bug still live in `compat.py compat_create_company`'s separate onboarding-admin insert (a legacy implementation of the same flow the original fix didn't touch) — fixed 2026-08-25. | | 32 | Dedup | `create_user` duplicated across `compat_create_user`, `common/create_user`, `create_infra_user`. (Code review 2026-08-18) | [x] Done — all four insert sites (those three plus `infra.py create_company`'s admin row) now call `service/users.py insert_user_with_receiver`. Upsert-by-id, unique-email prechecks, and role defaults stay in the web layer — those differ per endpoint by design, not by accident. | | 33 | Dedup | `admin/` and `/` APIs are duplicate implementations. (Code review 2026-08-18) | Partial — user-create is deduped via #32. Company/MLC/mlc_config/service_bundle/device/location `admin/*` compat routes still independently reimplement their `/infra` or `/ops` counterparts; the device/location identifier bug that duplication produced is fixed (#30), but merging the implementations themselves is a larger, riskier change (different response envelope — `ok()`/`ok_bare()` vs plain JSON — and unvalidated legacy request bodies vs pydantic models) deferred without a live test environment. Full inventory of the duplication in #33a/#33b (2026-08-25 audit). | | 33a | Dedup | Full pair inventory for #33 — `compat.py` reimplements each of these from a raw `Request` body instead of the shared pydantic model the canonical route uses (18 pairs): login (`common.py login` / `compat.py compat_login`); company list/create/delete (`infra.py list_companies`/`create_company`/`delete_company` / `compat.py compat_list_companies`/`compat_create_company`/`compat_delete_company`); credentials get/set (`infra.py get_credentials`/`set_credentials` / `compat.py compat_get_credentials`/`compat_set_credentials`); company_config get/set (`infra.py get_company_config`/`set_company_config` / `compat.py compat_get_company_config`/`compat_set_company_config`); mlc list/create/delete (`infra.py list_mlc`/`create_mlc`/`delete_mlc` / `compat.py compat_list_mlc`/`compat_create_mlc`/`compat_delete_mlc`); mlc_config list/create/delete (`infra.py list_mlc_config`/`create_mlc_config`/`delete_mlc_config` / `compat.py compat_list_mlc_config`/`compat_create_mlc_config`/`compat_delete_mlc_config`); service_bundle create/delete (`infra.py create_service_bundle`/`delete_service_bundle` / `compat.py compat_create_service_bundle`/`compat_delete_service_bundle`; list is 33b, 3-way); locations list/create/delete (`ops.py list_locations`/`create_location`/`delete_location` / `compat.py compat_list_locations`/`compat_create_location`/`compat_delete_location`); device list/create/delete (`ops.py list_devices`/`create_device`/`delete_device` / `compat.py compat_list_devices`/`compat_create_device`/`compat_delete_device`); user create/delete (`common.py create_user`/`delete_user`, `infra.py create_infra_user`/`delete_infra_user` / `compat.py compat_create_user`/`compat_delete_user`; create shares `insert_user_with_receiver` via #32, the endpoint wrapper logic is still separate). Two of the pairs are more than style duplication — `compat.py compat_create_device` never checks the target location belongs to the caller's company (`ops.py create_device` 404s otherwise), and `compat.py compat_delete_device` performs no company-ownership check at all (`ops.py delete_device` walks device→location→company and 404s) — a live cross-tenant delete via `DELETE /device?id=`. `compat.py compat_delete_company` also lacks `infra.py delete_company`'s guard against orphaning users. | Open | | 33b | Dedup | Triplicate/near-identical logic found alongside #33a (more than 2 copies of the same query): user listing has FOUR independent `select(DUser).where(company_id == …)` implementations serializing different field subsets — `common.py list_users` (own tenant), `infra.py list_infra_users` (cross-company), `compat.py compat_list_users` (aliased to both `/admin/user` and `/list_employees`), `ops.py list_employees` (cross-tenant via `?company_id=`, filters out `DORI_ADMIN`). Service-bundle listing has THREE — `infra.py list_service_bundles` (company-filterable), `ops.py list_company_service_bundles` (tenant-scoped, adds event counts), `compat.py compat_list_service_bundles` (unfiltered). Event-type listing has two structurally-identical bodies over a different scope — `infra.py list_event_types` (platform-wide) vs `ops.py list_event_types` (tenant-scoped via the bundle-company join). | Open | | 34 | Consistency | Inconsistent return-value style (dict creation vs local JSON mapping). (Code review 2026-08-18) | Open — unify with #25. | | 35 | Error handling | Add error handling for SQL errors (insert/delete failures). (Code review 2026-08-18) | [x] Done — `dori_utils/database.py commit_or_4xx` (IntegrityError → 409, other SQLAlchemyError → 400, rollback either way); every bare `await db.commit()` across `common.py`/`infra.py`/`compat.py`/`ops.py` (~47 sites) now goes through it. The two sites with existing bespoke IntegrityError handling (friendlier email-conflict messages) were left as-is. | | 36 | Company create | Seed `company_config` population during company creation. (Code review 2026-08-18) | [x] Done — `create_company` copies company id=1's `company_config` rows to the new company (insert path only, `infra.py`). | | 37 | Architecture | Use the MVC layers. Database interface should be in the Service layer. | [x] Done — see #24; every route's DB interface (queries + mutations) now runs through `service/*.py`. Response-shaping dict builders are the one deliberate exception, tracked separately under #25. | | 38 | Event & Messaging Flow (design "Event and Messaging Flow 20260930") | Implement the documented event pipeline end-to-end: DB insert → `event_insert` NOTIFY → **DIM** → **CIM** `/event_process/` → SMS(Msg91)/MQTT/Kafka/Email, plus WebSocket to the dashboard. System events (DORI_SYSTEM_SERVICE_BUNDLE + `is_trigger`, e.g. MLC_DOWN) → SMS to active `DORI_ADMIN` users; customer events → receiver dispatch scoped by company. JSON config interface (`db_config.json`/`mqtt_config.json`/`kafka_config.json`). CIM configured for MQTT **or** Kafka, not both. | [x] Done — (a) trigger payload now carries `companyId` via device→location→company (migration 017, `create_notify_trigger.std.sql`); (b) system events `is_trigger` backfilled: MLC_DOWN / Stream Unavailable / BARCODE_SCANNER_DOWN on, MLC_UP / Stream Available off (migration 018); (c) JSON config overlay `common/dori_utils/json_config.py` (+ Settings fields) with per-service `config/*.sample.json`; (d) CIM internal `POST /event_process/` routes system vs customer, `_format_phone` (country code, no leading 0/+), real system-event SMS replacing the log-only stub (`event_service.py`); (e) CIM MQTT-or-Kafka exclusivity by `method` (`main.py`); (f) DIM forwards every NOTIFY to CIM (`cim_client.py`, `listener.py`, `CIM_URL`); (g) Monitor gated active-poll mode writing MLC_DOWN events (`mlc_check.py`), `/system_event/` now enqueues into the pipeline; (h) `/cim/event_process/` blocked at nginx (internal-only). SMS/EMAIL test-sink hooks (`SMS_TEST_WEBHOOK`). | | 38a | Event flow · reconcile with #14a | #38(g)'s Monitor active-poll mode (`mlc_check=1` → Monitor probes MLCs and raises MLC_DOWN) reintroduces a second poller alongside job_app's health authority (#14a). Kept **off by default** (`settings.mlc_check=False`) so job_app stays the sole authority unless a deployment explicitly opts in via `db_config.json` `mlc_check:"1"`. | [x] Done by design — gated; default path unchanged (status-transition watcher). Flagged for review if both are ever enabled together (status-flap risk per #14a). | | 39 | Alerts & Notifications (E7 + E4) | Per-company alert config replacing event_config + is_message: `receiver` (channel + address in `config`) + `alert rule` (event+scope → receiver) tables (migrations 020/024; id sequences renamed in 030) with per-company CRUD APIs (`/config/ops/receiver`, `/ops/alert_rule`); event-flag API exposes two lanes — `is_notification` (UI websocket) and `is_alert` (CIM alert lane, migration 029); `is_message`/`is_trigger` removed. Delivery is app-side via `emit_event` (common/dori_utils/events.py), which enqueues a durable `dori.event_outbox` row + `NOTIFY` when `is_notification` OR an active alert rule matches — the DB NOTIFY trigger was dropped (migration 028). DIM's relay drains the outbox → CIM, which dispatches per channel (MQTT/SMS/EMAIL/WEBHOOK) by alert rule. System events are split across two platform bundles (migration 031): `DORI_SYSTEM_SERVICE_BUNDLE` (MLC up/down, attached to the Dori company) and `COMPANY_SYSTEM_SERVICE_BUNDLE` (stream + scanner, attached to a company when its first camera is added); event visibility is strictly by bundle assignment — DORI_SYSTEM is pinned to Dori and is NOT assignable to other companies (its MLC events are visible only in Dori); COMPANY_SYSTEM's stream/scanner events become visible to a company only once that bundle is assigned to it (on its first camera). Edge/MLC events reach the event DB through a Kafka consumer (monitor `InferenceEventConsumer` on `dori.inference.events`) that calls `emit_event`, so an edge event (e.g. BARCODE_SCANNER_DOWN) lands with an alert indication and fires SMS when a alert rule matches. | [x] Done — E4 outbox durability live on mapp (single outbox row per event; CIM idempotency via `event_dispatch`); receivers/alert rules per-company in the React ICD/COD app; standalone Notifications Console removed; system-bundle split + Kafka ingest verified end-to-end (edge BARCODE_SCANNER_DOWN → event DB → outbox isAlert → CIM). Fresh-DB reinstall: unit suite 147 passed / 0 failed. | | 40 | Log out frontend sessions on each deploy | A new build must invalidate existing login sessions so the UI never runs against a stale backend. JWTs now carry `iat`; `decode_token` rejects any token issued before `settings.sessions_valid_since` (401 → the frontend falls back to login). The deploy pipeline (`build-remote-ec2.sh` + `install.sh`) bumps `SESSIONS_VALID_SINCE` to now in the box `.env` and re-applies it stack-wide via `docker compose up -d` (the epoch is on the `&common-env` anchor). | [x] Done — verified on mapp: a pre-deploy token returns 401 after the build, a fresh login returns 200, new tokens carry `iat`, and `.env` shows `SESSIONS_VALID_SINCE` = deploy time. | | 41 | Receivers & Alert Rules (rename + sender + sysadmin) | Clearer names for the alert config: `subscriber`→**`receiver`** (a reusable destination: channel + address in `config`) and `subscription`→**`alert_rule`** (event + optional location/camera scope → receiver), via migration 032 (tables, sequences, `receiver_id` FK) and a from-scratch fresh-DB build — not just a DB rename. APIs are `/config/ops/receiver` + `/config/ops/alert_rule`; React ICD/COD + the `/fe` test console renamed. **Sender**: EMAIL/SMS resolve the "from" as receiver.config.sender → company_config (`SMS_SENDER`/`EMAIL_SENDER`) → global defaults (`sms_sender_default`/`email_sender_default`), so a sender can be per-receiver, per-company, or global. **Visibility by assignment**: a company configures rules only for its assigned SBs' events; Dori is seeded BOTH system bundles (sysadmin → MLC + stream/scanner), DORI_SYSTEM stays Dori-only. Rules scope to company/location/camera. | [x] Done — fresh-DB build head 032; tables `receiver`/`alert_rule`; Dori has both system bundles; live e2e: `/ops/receiver` + `/ops/alert_rule` → injected event → outbox isAlert → CIM dispatched. Unit suite 148 passed / 0 failed. |