Commit Graph
866 Commits
Author SHA1 Message Date
Pouzor 322600f1c9 fix(zigbee): stop canvas imports dying on a proxy read timeout
A Zigbee2MQTT networkmap on a 200+ device mesh takes minutes to build.
Two separate failures fell out of that:

- POST /zigbee/import held the HTTP request open for the whole MQTT
  round-trip, so any reverse proxy in front of the API cut it first
  (Cloudflare returns a 524 at 120 s) and the browser never saw the map.
  It now registers a job, fetches in the background and answers 202; the
  client polls GET /zigbee/import/{job_id} until the payload is ready.
  Job results are transient and live in memory with a 15 min TTL — the
  same single-worker assumption the scheduler already makes. A failed
  fetch replays the status the synchronous route used to raise, so a bad
  broker is still a 502 and a slow mesh still a 504.

- The networkmap wait was hard-coded at 300 s with no way to raise it.
  It now reads ZIGBEE_NETWORKMAP_TIMEOUT, and the shared MQTT round-trip
  used by the Z-Wave import reads MQTT_RESPONSE_TIMEOUT. Both default to
  300 s, fall back to that if misconfigured to a non-positive value, and
  name themselves in the timeout message.

Also corrects the route and doc claims that the wait was 60 s.

The /import tests changed with the contract they cover, not to pass.

Fixes #380

ha-relevant: yes
2026-08-31 11:26:45 +02:00
Pouzor 9c2a1ba0eb docs(canvas): correct a stale comment on where a zone's height is stored
The zone size moved to the nodes.width/height columns; the comment still
described the custom_colors stash it replaced.

ha-relevant: yes
2026-08-27 01:45:23 +02:00
Pouzor 851141951d fix(canvas): store a zone's size in the width/height columns
Every node type persisted its size in `nodes.width` / `nodes.height` except
`groupRect`, which stashed it inside the `custom_colors` JSON next to its
colours. The columns already existed and were simply unused for zones, so
this was an inconsistency rather than a missing-column workaround, and it
put geometry in a blob that otherwise holds style.

The serializer now writes the columns for a zone too, and strips the legacy
`width`/`height` keys out of the blob so the two cannot drift apart and
leave an older canvas reading a stale size.

No data is lost on upgrade:

- `_backfill_zone_size` copies the blob geometry into the columns at
  startup. It only fills a column that is still NULL, so it cannot overwrite
  a size set since; it parses the JSON in Python rather than with
  `json_extract`, so it does not depend on the SQLite build carrying JSON1;
  and an unreadable row is skipped without costing the others their size.
  Re-running it is a no-op.
- the reader still falls back to the blob, covering a payload the backfill
  has not reached — an older server, or an import.

Standalone mode is unaffected: it stores React Flow nodes verbatim, so the
size was always on `node.width` / `node.height` there.

The four serializer tests that pinned the size to the blob now assert the
columns, since that is the behaviour being changed.

ha-relevant: yes
2026-08-27 01:45:23 +02:00
Pouzor f9f88c8fb0 refactor(canvas): drop a dead custom_colors.height write on zone growth
Growing a zone during a subnet import wrote the new height twice: to
`node.height`, the live field, and to `data.custom_colors.height`, which
nothing reads. The blob copy is produced by the serializer at save time
(rebuilt from `node.height`, so the value written here was overwritten
before it reached the API) and consumed at load time off the API payload.
Writing it from the store was invisible, and misleading in a blob that
otherwise holds colours and style.

The test asserting the dead field is replaced by a save/load round-trip
through the real serializer, which is what actually protects the height.

ha-relevant: yes
2026-08-27 01:45:23 +02:00
Pouzor 2bf9f1bca8 fix(canvas): keep a parent ahead of its children after a subnet import
The subnet import only guaranteed the zone preceded the nodes it pulled in.
An arrival that is itself a parent — a Proxmox host with nested VMs, whose
children stay put because they already have a parent — could end up listed
after those children, which React Flow renders detached with a
parent-not-found error. Nothing re-sorts on load, so the broken order
survived a save.

Reorder the whole array instead: every node now follows its parent, which
covers both directions at once. An already-valid list comes back untouched
and a parent cycle terminates rather than recursing.

ha-relevant: yes
2026-08-27 01:45:23 +02:00
Pouzor d9451f35dc feat(canvas): import devices into a zone by subnet
Closes #325.

A zone gains an "Import devices by subnet" action: type a CIDR, and every
free device whose IP falls in that range moves into the zone, laid out on a
grid in its free space.

The CIDR is an argument to a one-shot action, never a property of the zone —
it is not submitted with the form, not persisted, and cleared after each run.
So there is no schema change, no migration, and it works in standalone mode.
The trade-off is that a device scanned later does not join the zone by
itself; the user re-runs the import.

Deliberate rules:
- additive — running it twice with two subnets leaves both sets inside, and
  a device whose IP stops matching is never ejected
- only unparented devices move. A node nested in a group, a container host
  or another zone keeps the parent the user gave it
- canvas furniture (groupRect / group / text) is skipped, and a zone never
  swallows another zone
- one history entry per import, so a single Undo reverses the whole thing

Edit mode gets an Import button; add mode has none, since the zone does not
exist yet and a button would look broken — there the CIDR is applied right
after the zone is created.

IPv4 only: the modal rejects an IPv6 CIDR with a message rather than
silently matching nothing.

ha-relevant: yes
2026-08-27 01:45:23 +02:00
Pouzor 05c0da53e6 fix(scan): reconcile scan runs orphaned by a backend restart
A scan runs on a background thread inside the API process. If that process
dies mid-scan — an OOM kill, docker stop, a crash — the ScanRun row stays
"running" for ever, because nothing is left alive to finish it.

That row is not just cosmetic clutter in Scan History: the trigger endpoints
reject a new scan while one is "running" for the same target, so a single kill
locks that range out permanently.

Nothing can legitimately be "running" the moment we boot, so lifespan() now
marks every such row "error" — the same word run_scan and run_device_scan
write when they fail themselves — with finished_at and an explanatory message.

Reported alongside the OOM itself in #374, which is what produced the orphans.

Fixes #374

ha-relevant: maybe
2026-08-27 00:00:32 +02:00
Pouzor c29680a75f fix(status): stop an endless response body from OOM-killing the backend
Both HTTP paths buffered the whole response body before looking at it. An
endpoint that streams without end and sends no Content-Length — a Freebox
bandwidth-test port, an MJPEG camera, a log tail — grew the backend until the
cgroup limit killed it, every 5-10 minutes at a constant ~1.04 GB RSS.

httpx's timeout does not help: it applies per network operation, not to the
total time spent draining a socket that keeps delivering data.

- status_checker._http_get only needs the status line, so it now uses
  client.stream() and never reads the body at all.
- http_probe._probe_scheme needs at most _MAX_BODY_BYTES to hunt for <title>,
  so it streams and stops there. The cap existed already but was applied to
  resp.text, after the full body had been downloaded.

Regression tests serve an endless, Content-Length-less body through a
MockTransport and assert the read stays bounded. Both fail on the old code.

The existing mocks patched httpx.AsyncClient.get, which neither path calls
now; they are rebuilt on MockTransport — a real client over a fake network —
with every original assertion kept.

Fixes #375

ha-relevant: yes
2026-08-26 23:43:50 +02:00
Pouzor - Rémy JardientandGitHub eb7e74d36b Update README.md 2026-08-25 23:08:03 +02:00
Pouzor 390a712f14 chore: bump version to 3.3.5
ha-relevant: no
v3.3.5
2026-08-21 12:56:17 +02:00
Pouzor 9daea4c4d0 fix(scan): keep an edit, a discovery source and one name for a failed run
Three findings from the PR review.

A finishing deep rescan threw away an edit in progress. The reset effect in
InventoryDeviceModal was keyed on the `device` object, and the parent hands
down a fresh one whenever the row is refreshed — including from the poll's
own onSaved. Same device, new object, so the effect reset the form and left
edit mode, minutes into a scan the user was waiting on. It is keyed on the
id now: a refreshed row is not a different device, and only a different
device is a reason to throw the form away. The fresh row still reaches the
canvas and the grid — gating that behind edit mode would discard the scan
result instead, and a save unions the services server-side anyway.

A rescan tagged every device it touched as "arp"-discovered, so a Proxmox
guest, a rack mount or a hand-added host started answering the network
source filter. It carries its own source through instead.

The two background wrappers wrote status "failed" where the scanners write
"error" for the same condition. Harmonized on "error" — the one Scan
History filters and colours; "failed" showed up unlabelled. The frontend
union keeps 'failed' for rows already recorded.

That last one meant updating an existing assertion in test_scan_run.py: it
encoded the old spelling.

ha-relevant: yes
2026-08-21 12:49:48 +02:00
Pouzor 2e642815f4 fix(scan): stop a scan from repainting hand-picked service icons
merge_services did {**existing, **incoming}, so the fingerprint's guess at
an icon overwrote the one the user chose — and on a port no signature
covers it wrote None, clearing it outright. Since 3.3.0 the inventory row
is the only copy of a device's services, so every "Scan network" repainted
the service on every canvas drawing that device at once.

The scanner now merges with discovered=True: it still adds services and
refreshes what it knows, but leaves an established icon and category alone.
A user edit from the modal or a canvas changes them as before.

Blank incoming values no longer clear established ones either, on both
paths — an absent field is silence, not a reset. Same rule merge_properties
already follows.

ha-relevant: yes
2026-08-21 12:49:48 +02:00
Pouzor c6d6b525ea feat(scan): choose the port range before a deep scan
The Deep scan link on a device now opens a small dialog instead of firing
straight away. It is prefilled with the full 1-65535 range — that is still
the point of the feature — but a user who knows where a service lives can
narrow it and get an answer in seconds instead of minutes.

The dialog takes an nmap-style spec: a port, a range, or a comma list
(80,443,8000-9000). It shows the port count live, refuses to start on
something nmap could not use, and carries three presets (all / 1-1024 /
1-10000).

Backend:
- _parse_port_spec / _port_chunks generalize the slicing that used to be
  full-range only. Ranges are merged before slicing, so an overlapping
  spec is never scanned twice, and small ranges are packed into one nmap
  call instead of one call each. _deep_port_chunks still yields the same
  eight slices as before.
- run_device_scan takes ports=; it wins over full_ports.
- The retry-free flags (--max-retries 0 --min-rate 2000) now key on the
  total port count rather than on full_ports. They pay for themselves over
  thousands of ports on a lossy host; over a handful they only cost
  accuracy.
- RescanDeviceRequest.ports validates the spec — 422 rather than handing
  nmap a bad -p. Blank means the full sweep.

Frontend utils/portSpec.ts mirrors the backend parser so a typo is caught
before the request; the backend validates again because it is the one
calling nmap.

The existing rescan tests now go through the dialog: the click path
changed, so the start is two steps.

ha-relevant: yes
2026-08-21 12:49:48 +02:00
Pouzor fbb660504e fix(scan): slice the deep rescan instead of timing out the host
A deep rescan of a slow host came back with nothing at all: the run took
its full 600s ceiling and the device's services were unchanged, so a
service deleted by hand was never rediscovered.

nmap answers --host-timeout with "Skipping host <ip> due to host timeout"
and discards every port it had already found — the ceiling turned a slow
scan into one that reports nothing. What costs the time is a host that
drops packets: 8188 of 8192 ports filtered, each waiting out its probe.

- No --host-timeout on the deep discovery pass, ever.
- The full range runs as 8 slices of 8192 ports, one nmap call each,
  unioning the open ports. A slice that overruns costs its own ports, not
  all of them, and the loop has somewhere to notice a stop request.
- `scanner_deep_host_timeout` is now a total budget checked between
  slices (default 2700s), not an nmap flag. The first slice always runs.
- A partial sweep is reported rather than passed off as complete: the run
  finishes `done` carrying "Scanned 3/8 port ranges …", and the modal
  toasts a warning instead of success.
- Deep slices use --max-retries 0 --min-rate 2000. Measured against a
  dropping host, 8192 ports took 329s at --max-retries 1 and 164s at 0,
  finding the same ports; capping the RTT changed nothing. The range scan
  keeps nmap's default retries on its curated port list.

ha-relevant: yes
2026-08-21 12:49:48 +02:00
Pouzor 1be96d1045 feat(scan): deep-rescan one device from the inventory detail
Devices added before the scanner knew a service showed an empty Services
section with no way to refresh it (#350). The detail modal now starts a
full-port scan of that single device.

- `process_host` lifted out of `run_scan` so the range scan and the new
  single-device scan share the same match / merge / dedupe rules — a
  rescan unions services, it never replaces what the user added by hand.
- `run_device_scan`: no ping sweep, no mDNS, straight to the phase-2 nmap
  pass on the device IP over all 65535 TCP ports.
- `POST /scan/pending/{id}/rescan` records a normal ScanRun
  (`kind=device`, `ranges=["<ip>/32"]`), so stop, progress and Scan
  History work unchanged. One run per device at a time — a second request
  while the first is scanning is a 409. 404 unknown, 409 no-IP or hidden.
- `GET /scan/runs/{id}` so a caller can poll the run it started.
- Deep scan button in the Services section of the device detail, hidden
  without an IP and for Zigbee. Swaps to a stop control while running,
  folds the fresh services back in on completion (never over an edit in
  progress).

ha-relevant: yes
2026-08-21 12:49:48 +02:00
Pouzor 9bc02d61c2 fix(canvas): give a valueless property label the full node width
The property line caps its label at max-w-15 so a long key cannot crowd
out the value drawn beside it. When the value is empty there is nothing
to protect, but the cap still applied, truncating the label against
empty space.

Drop the cap when the value is blank and let the label truncate against
the node width instead. Same fix in BaseNode and ProxmoxGroupNode.

Closes #361

ha-relevant: yes
2026-08-17 20:26:29 +02:00
Pouzor 21a64e52ab fix(auth): accept the app's own origin in the OIDC CSRF check
A same-origin deployment behind a reverse proxy needs no CORS, so
CORS_ORIGINS is left at its localhost default — but browsers still send
Origin on unsafe methods, and OIDCCSRFMiddleware validated it against
CORS_ORIGINS alone. Every POST/PUT/DELETE came back 403 while reads
worked, so creating a canvas failed with no usable error.

The OIDC callback URL is served by this app, so the origin of
OIDC_REDIRECT_URI is the app's own and is always a legitimate CSRF
origin. origin_is_allowed now accepts it in addition to CORS_ORIGINS,
and settings validation warns at startup when CORS_ORIGINS omits it so
the misconfiguration is visible rather than silent.

Fixes #356

ha-relevant: no
2026-08-17 20:13:04 +02:00
Pouzor ad0b1b427d chore: bump version to 3.3.4
ha-relevant: no
v3.3.4
2026-08-17 19:19:25 +02:00
Pouzor 4c460cec1c fix(scan): stop a re-scan from deleting a device's hand-added services
Since 3.3.0 the inventory row is the only copy of a device's services and
every canvas drawing it reads that list, but the scanner still assigned
`keep.services = fingerprint_ports(...)` — the pre-split behaviour, when a
node held its own copy. One re-scan therefore deleted every service the user
added by hand, on every canvas at once, and brought back the ones they had
deleted. It unions now, like every other writer of that field.

Two more things #347 turned up:

* A node whose view matches nothing the row still holds drew nothing at all —
  the row had been replaced under it, so every key was gone and every key was
  new, and `apply_view` hid the lot. Such a view says nothing about the list
  that replaced it, so it is treated as no view: the row is drawn. An empty
  view is untouched, being a real answer ("this canvas draws none of them").

* The 3.3.3 view seed recovers a node's arrangement from the pre-3.3.0 backup,
  which is 3.2.0-era and cannot know about a property added afterwards. On
  3.3.0-3.3.2 the row was the only place to add one and every canvas drew it,
  so seeding strictly from the backup took it off all of them at once. Those
  are appended visible, keeping the recovered order and hidden flags for
  everything the backup does know. Services keep the strict recovery: holding
  back what a scan fingerprinted is the whole point of the view.

ha-relevant: maybe
2026-08-17 19:05:13 +02:00
Pouzor 1aa88de491 chore: bump version to 3.3.3
ha-relevant: no
v3.3.3
2026-08-17 14:09:01 +02:00
Pouzor d0679ad72b fix(inventory): stop the view seed from re-reading the backup every boot
The seed selected any node with a `device_id` and no view. Deleting a device
leaves that column dangling — SQLite runs with foreign keys off, so the
`ON DELETE SET NULL` never fires on an existing table — and such a node can be
seeded from nothing, so it stayed selected. The query matched it again on every
later start, re-opening the pre-upgrade backup and logging a recovery that
recovered nothing.

Narrowing the query to nodes whose row actually exists restores the invariant
the seed is documented to have: it runs on one boot, and never opens a file
again.

ha-relevant: no
2026-08-17 14:06:09 +02:00
Pouzor 5df70ce128 feat(canvas): give each node its own view of a device's services and properties
Since 3.3.0 the inventory row owns a device's services and properties, and
every node drawing that device rendered the row wholesale. One row shared by
several canvases meant one rendering: a service the scanner fingerprinted
appeared on every canvas at once (users reported Uptime Kuma and Synology DSM
on hosts running neither — both are port-only signatures), and a property added
on one schematic showed up on all the others.

Order and visibility are presentation, so they move to the node. `display_view`
records, per node, which of the row's services and properties it draws and in
what order, keyed by `port|protocol|name` and by lowercased property key so the
view survives an edit to the fact itself. The facts stay on the row: hiding is
per node, deleting is still device-wide.

An item the view does not list is reported hidden rather than dropped, so what
a scan finds next is one toggle away instead of pushed onto every canvas.

The wire shape is unchanged. A client already sends its lists in display order
with their `visible` flags, so the view is read back out of them rather than
asking for a second field, and `visible` is only stamped when something is
hidden.

Upgrades keep what each canvas showed:

* from 3.2.0, the backfill seeds each node's view from its own legacy columns
  before they are dropped;
* from 3.3.0-3.3.2 those columns are gone and the row holds the union of every
  canvas, so the layout is recovered from the backup `_backup_db` took before
  the 3.3.0 migration — the newest one whose `nodes` table still has the
  columns, read read-only and matched by node id;
* with no usable backup the row is the seed, so every canvas keeps showing
  exactly what it shows today and only later additions are held back.

Refs #347

ha-relevant: maybe
2026-08-17 14:06:09 +02:00
Pouzor a58725c2ee chore: bump version to 3.3.2
ha-relevant: no
v3.3.2
2026-08-17 11:37:54 +02:00
Pouzor 466df59b0c fix(inventory): finish the 3.3.0 backfill on databases carrying timestamps
3.3.1 kept `nodes` writable when the backfill failed, but the backfill still
failed — so a canvas upgraded from 3.2.0 shows every device stripped of its ip,
services, hostname, notes and hardware. The facts are intact in the legacy
columns; nothing was ever reading them across.

The backfill reads those columns with raw SQL, so SQLite hands `last_seen` and
`last_scan` back as text. Writing text into a DateTime column raises
`TypeError: SQLite DateTime type only accepts Python datetime and date objects`,
which is neither IntegrityError nor OperationalError — so it went straight past
the per-node savepoint and aborted the whole run, exactly as before.

- Parse the legacy timestamps back into datetimes, the way `_decode_json`
  already handles the JSON columns. An unparseable stamp becomes None rather
  than an error: the status checker refreshes both within a minute.
- Catch every exception per node, not two chosen classes. Picking the exception
  types was the defect; a node that fails for any reason must cost only itself.

The same run would also have destroyed data. A canvas saved while an earlier
migration was stuck minted a blank inventory row from a UI that had no facts to
show and linked the node to it. The backfill only looked at unlinked nodes, so
it skipped that one, counted the canvas migrated, and dropped the columns
holding its only copy of the ip and services.

- Read linked nodes too, and fill only what their row is missing. A row that
  already holds a fact is never overwritten, so an edit made after the migration
  survives, and the pass stays a no-op once everything has moved across.

Verified end to end against a rebuilt 3.2.0 database stuck the way the reports
describe: one boot restores ip, services, hostname, mac, notes, check method,
hardware, properties, status and last_seen, then drops the legacy columns.

Refs #347, #348, #351

ha-relevant: no
2026-08-17 11:36:32 +02:00
Pouzor 7eec0b3f9a chore: bump version to 3.3.1
ha-relevant: no
v3.3.1
2026-08-17 01:58:26 +02:00
Pouzor b76519b0dd fix(db): keep nodes insertable when the 3.3.0 backfill cannot finish
Upgrading 3.2.0 -> 3.3.0 could leave a database where approving a device —
or creating any node — failed with `NOT NULL constraint failed: nodes.status`.

The 3.3.0 migration moves the device facts off `nodes` onto `device_inventory`
and then drops the columns. The drop is skipped when a node is still unlinked,
because those columns are the only remaining copy of its facts. But a 3.2.0
`nodes` declares `status`, `services`, `properties` and `show_hardware` NOT NULL
with no server-side default, and the 3.3.0 model no longer writes them — so a
skipped drop bricks every later INSERT.

It skipped because one node killed the whole backfill: the loop shared a single
session, so an IntegrityError raised by autoflush lost every link made before it.

- Give each node its own savepoint. A node whose merge violates a constraint is
  skipped and logged on its own; the rest still link. `skipped` joins the stats
  reported at boot.
- Relax NOT NULL on the retained legacy columns when the drop is skipped, so the
  database stays writable while later boots retry the backfill. Values are kept;
  only the constraint goes. Guarded on the PRAGMA notnull flag, so it runs once.
- Share one `_rebuild_nodes` between the drop and the relax: it restores
  `PRAGMA foreign_keys = ON` in a finally (the connection returns to the pool)
  and clears a `nodes_new` left by a failed attempt.
- Match IEEE addresses case-insensitively, and never write one another inventory
  row already owns — `device_inventory.ieee_address` is UNIQUE, and a duplicate
  is one way the backfill was raising in the first place.

Tests build the real v3.2.0 schema from the tag: the existing legacy-migration
fixture declares `status` nullable, which is why this reached a release.

Fixes #351

ha-relevant: no
2026-08-17 01:55:51 +02:00
Pouzor a2870a6e29 docs(changelog): match the house style for contributor credit
ha-relevant: no
2026-08-15 02:48:06 +02:00
Pouzor d6308a5573 chore: bump version to 3.3.0
ha-relevant: no
v3.3.0
2026-08-15 02:39:49 +02:00
Pouzor fafbaeffcd feat(sidebar): link the version row to the public changelog
ha-relevant: maybe
2026-08-15 02:32:40 +02:00
Pouzor 824aa5d044 docs(install): document the pre-built GHCR images
The quick starts already pull ready-made images through install.sh, but
nothing said so — issue #310 asked for something that already ships.

Add a "Pre-built Docker images" section to INSTALLATION.md listing the
four GHCR images, their tag scheme and a manual docker-compose.prebuilt.yml
setup, and note that the root docker-compose.yml builds from source.

Closes #310

ha-relevant: no
2026-08-15 02:18:57 +02:00
Pouzor 416b5fd1cb feat(properties): make the property value optional
A property may now carry a label alone: the add and edit forms require
only the key, and every surface that prints one — the badge list, node
and Proxmox group renderers, rack cable annotations — drops the "· value"
separator when the value is empty.

ha-relevant: yes
2026-08-15 02:10:02 +02:00
Pouzor fcc2927b69 feat(canvas): let a zone parent the nodes dropped inside it
Dropping a node onto a groupRect zone now asks to add it to the zone and
sets a real React Flow parentId, so moving the zone moves its contents.

Zone children are deliberately not extent-clamped, unlike group and
container children: a zone has no side panel to release a child from, so
dragging the node back out of the zone is what detaches it.

Deleting a zone now releases its children onto the canvas in absolute
coordinates instead of cascade-deleting them; groups and containers keep
the cascade. Collapse counts and hides both parented and merely
overlapping nodes, and the deserializer restores zone parenting on load.

ha-relevant: yes
2026-08-14 20:09:11 +02:00
Pouzor be46c21dd4 fix(deps): bump undici and ip-address to patch high severity advisories
Resolves the two high Dependabot alerts on the frontend lockfile:

- undici 7.28.0 -> 7.29.0 (transitive via jsdom): cross-user information
  disclosure and parse-time crash via degenerate private cache directives
- ip-address 10.2.0 -> 10.5.0 (transitive via shadcn ->
  @modelcontextprotocol/sdk -> express-rate-limit): leading-zero octets
  decoded as decimal, allowing SSRF and trust-boundary bypass

Also closes the moderate advisories on the same two packages. Lockfile only,
both are dev dependencies.

ha-relevant: no
2026-08-14 18:36:36 +02:00
Pouzor 1c9d5791b7 fix(inventory): resolve the remaining PR review findings
Three unrelated fixes from the review of this branch.

`create_pending` merging into a hidden row left it hidden, so adding a device
by hand whose IP collides with one hidden weeks ago answered 201 while nothing
appeared in the inventory — the add read as a no-op. An explicit add outranks
the earlier hide, exactly like restore. Only `hidden` is lifted; an approved
row is not walked back down the lifecycle.

Standalone stored a node's live reachability in the inventory row's `status`,
which is the pending/approved/hidden lifecycle — the backend keeps the two
apart as `status` vs `status_live`. Reachability now goes to `status_live`, and
`normalizeLifecycle` repairs rows already written the old way: the node still
reads its status on load, and the next save rewrites the row correctly.

Status checks stay device-scoped and keep covering a device no canvas draws.
That is deliberate — the inventory row is what's monitored, so a rack mount or
an inventory-only entry reports state without being drawn anywhere, and
deleting a node no longer silently stops monitoring the host. `hide`, or
clearing the check method, is what ends the checks. Documented and pinned with
tests rather than changed.

ha-relevant: maybe
2026-08-14 18:02:02 +02:00
Pouzor 0c7fd3c127 fix(canvas): stop a canvas save from reverting a device edited elsewhere
A canvas node carries a full copy of its Device Inventory row, hydrated when
the canvas loads. The save routed all of it back, so a save made for nothing
but a moved node rewrote the row from a snapshot that could be hours old —
silently reverting an edit made meanwhile in the inventory modal, on another
canvas, or by the scanner.

Diffing the payload against the row server-side cannot fix this: it can't tell
"I edited this" from "the row moved on since I loaded it". Only the client
holds the baseline.

- canvasStore keeps `factsBaseline` — the device facts as received — set on
  load, refreshed on save, rebased per field by `applyDeviceFacts`.
- `serializeNode` sends `changed_facts`: what this canvas actually edited.
  `link_facts(changed_fields=)` writes nothing outside that list, even where
  the values differ. Absent (older client, YAML import, MCP) keeps the previous
  full-write behaviour.
- `changed_facts()` additionally drops facts already equal to the row, so a
  no-op save writes nothing. Identity matching still uses the full payload.
- Live fields (status, last_seen…) bypass the filter — the status checker owns
  reachability, and a save only ever fills a row never checked.

An inventory edit also lands on the canvases already on screen:
`applyDeviceFacts` pushes the saved row onto every node drawing it, without
marking the canvas unsaved. A fact the canvas has edited but not saved is left
alone — work in progress wins locally and still saves.

ha-relevant: maybe
2026-08-14 18:02:02 +02:00
Pouzor 19758b2ba1 Revert "chore: bump version to 3.3.0"
This reverts commit 779e09ecb8.

The bump does not belong to this PR. 3.3.0 ships several PRs, so the version
files and the release notes are raised once, in a release PR of their own,
covering the whole diff against main.

ha-relevant: no
2026-08-14 18:02:02 +02:00
Pouzor 1f6812ef2a chore: bump version to 3.3.0
The Device Inventory row becomes the source of a device's facts, and the
device columns are dropped from `nodes` in a one-way migration on first start,
so the release notes lead with a Breaking section: what moved, what the
migration does, when it refuses to drop, and that `homelab.db.back-3.3.0` —
written before any migration runs — is the way back.

ha-relevant: no
2026-08-14 18:02:02 +02:00
Pouzor 9db8b2a7a9 test(standalone): cover a canvas stored in the pre-split shape
An install that predates the split holds canvases whose nodes carry the device
facts inline, with no `devices` key at all. Reading one already worked; nothing
pinned that the next save converts it in place rather than writing the old
shape back.

Covers the load-then-save round trip with every fact intact, the second save
not minting a duplicate row, furniture staying out of the conversion, and the
legacy bare-key canvas carried through `ensureSeed`.

ha-relevant: no
2026-08-14 18:02:02 +02:00
Pouzor d40f73b85c fix(mcp): read a canvas node's facts where the API puts them
`_slim_canvas` filtered `n["data"]`, the React Flow shape the frontend store
holds. `GET /api/v1/canvas` has never sent that: a node is reported flat, so
the filter matched nothing and `get_canvas` returned an id and a node type per
node — no label, no address, no services. The three tests covering it fed a
hand-built nested payload, so they passed against a shape the backend does not
produce.

Fold `data` over the node and read both, so either form slims the same way.
`type` leaves `NODE_KEEP` — flat it is the node type, not a device fact, and it
is already reported as `node_type`; `parentId` becomes `parent_id`, the key the
API uses, normalized from either spelling.

The new tests are built on a full `NodeResponse` copied from a real response,
and the backend now pins that wire shape from its side so the two move together.

ha-relevant: no
2026-08-14 18:02:02 +02:00
Pouzor 71e29362aa fix(inventory): stop a partial node update from blanking the other list
`merge_facts_into_device` applied one `replace_lists` flag to properties
and services alike, taking each wholesale from `facts`. But a PATCH only
carries the keys the client sent, so `{"properties": [...]}` replaced
services with `[]` — losing them for every canvas drawing the device and
stopping the scheduler from service-checking it. Symmetric the other way.

Gate the replace per list: a list absent from `facts` falls through to the
merge, which is a no-op against `None`. An explicit `[]` still clears, and
the canvas save path is unchanged since it always sends both keys.

ha-relevant: no
2026-08-14 18:02:02 +02:00
Pouzor 7a5762e1a0 fix(inventory): read a device's curated type on its inventory card
A device documented straight on a canvas carries a `label` and a `type`
and no discovery guess at all — no scan ever saw it. The card read
`suggested_type` alone, so a printer drawn on a canvas fell back to
'generic': unnamed, filed under no type, and drawn with a plain circle.

Two faults compounded:

- The card decided the kind from `suggested_type`, ignoring the curated
  `type`. That also meant retyping a scanned device in the detail modal
  changed nothing here, and approving it still built a node of the type
  the scanner had guessed.
- It carried its own table of eleven icons, unrelated to the ~40
  NodeType values, so anything outside that list drew a circle even with
  the right type in hand.

`deviceType()` now resolves `type` before `suggested_type` for the icon,
the badge, the type filter and both approve paths, and the icons come
from the shared `NODE_TYPE_DEFAULT_ICONS` — one map, the same one the
canvas and the detail modal use. `deviceLabel` prefers the curated
`label` for the same reason, otherwise a canvas-created device shows as
its IP.

ha-relevant: maybe
2026-08-14 18:02:02 +02:00
Pouzor be2c5ce7e2 feat(inventory): group the device sheet by what a reader is after
The services sat in the curation column, beside the properties, while the
addresses that reach them were two columns away. They move under
Identity: what answers on a host belongs with how to reach it, and the
curation column now carries the properties alone.

Every section heading takes an icon — identity, services, monitoring,
activity, notes, properties, make & model — so a column is scannable
without reading each title. `Section` requires it, so a new section
cannot be added without one. The properties editor keeps its own plain
heading: it comes from the shared `PropertyList`, which the rack cable
panel also renders.

The three view columns carry a data-testid, which is what lets the
layout be asserted rather than eyeballed.

ha-relevant: maybe
2026-08-14 18:02:02 +02:00
Pouzor b4b22cc68f feat(inventory): make the device detail modal readable
The detail modal was a narrow column: a 15px icon, eighteen key-value
rows in one grey block, and a form that grew the dialog to `max-w-lg`
while the header and the buttons scrolled away with the content.

It reads as a device sheet now, following NodeModal's shape:

- One wide dialog with three fixed regions — header, scrolling body,
  footer — so the title and the actions stay put.
- The header identifies the row at a glance: a type-coloured icon tile
  (the accent the node wears on the canvas), then a badge per discovery
  source, the type, live reachability with a dot, and how many canvases
  draw the device.
- Three columns instead of one: identity, operations, curation. A
  section stays put when it has nothing to show and says why, rather
  than disappearing and leaving the reader unsure whether the field is
  absent or unsupported.
- Timestamps read relative with the exact value on hover, matching the
  inventory cards. IP / hostname / MAC / IEEE offer a copy button.
- The form uses the same control sizing as the canvas modals, and gains
  inputs for `friendly_name` and `device_subtype` — both were already
  submitted, so they could not be corrected here.

Hardware is a property like any other, so the modal no longer edits it:
nothing draws `cpu_*` / `ram_gb` / `disk_gb` / `show_hardware`, and the
CPU / RAM / Disk keys are property suggestions instead. The columns are
still sent back untouched — Proxmox and the YAML import write them and
the YAML export reads them.

Two assertions in the edit tests covered the removed section: the
curated-fields test drops the hardware line, and the numeric-payload
test stops editing a field that no longer exists while still checking
the values leave as numbers.

ha-relevant: maybe
2026-08-14 18:02:02 +02:00
Pouzor 21465eee65 fix(canvas): import YAML hardware specs as properties
A node draws its `properties`; `cpu_model` / `cpu_count` / `ram_gb` /
`disk_gb` are inventory data that nothing renders — BaseNode reads them
only when `properties` is absent, which the API never allows (it returns
`device.properties or []`, and the model defaults to a list).

The YAML import filled those fields and set `show_hardware`, so an
imported node did show its specs — until the first save, when the
serializer sent `properties: []`, the fallback died and the specs
silently vanished.

It now mints the same four properties the 3.x migration and the Proxmox
import already create (CPU Model, CPU Cores, RAM, Disk), visible, so the
specs survive the save. The structured fields stay written: the YAML
export reads them back.

`show_hardware` is no longer set — it drives nothing.

The two assertions that covered the old behaviour move with it: the
scalar-field test drops `show_hardware`, and "sets show_hardware only
when hardware fields present" becomes three tests over the minted
properties.

ha-relevant: yes
2026-08-14 18:02:02 +02:00
Pouzor 9b77e43ca1 feat(inventory): drop the device columns from nodes
Completes the split: `nodes` now holds only how a device is drawn on one
canvas, and every device fact reaches the API from the inventory row.

- The device columns are removed from `nodes` (SQLite table rebuild, the
  same shape as the existing device_inventory and canvas_state rebuilds).
  The drop is skipped, and logged, while any non-furniture node is still
  unlinked: those columns are the last copy of that node's facts. The
  backfill therefore reads them with raw SQL — by the time it runs, the
  model no longer declares them.
- The status checker iterates devices, not nodes: one check per device
  however many canvases draw it, writing `status_live` / `last_seen` /
  `response_time_ms` on the row. `/ws/status` messages carry `device_id`
  and the node ids they light up. Hidden devices are not probed.
- Readers repointed: scanner (last_scan lands on the row), proxmox (its
  node tier collapses into the inventory tier, keeping only the cluster
  handles), zigbee/zwave (one property refresh serves every canvas), rack
  inventory, liveview, stats, node dedupe.
- `POST /scan/pending` merges into the row that already describes the host
  instead of minting a second one — one device is one row, whichever way it
  was documented.
- Standalone keeps parity: the canvas blob gains `devices`, split on save
  and hydrated on load. A blob written before the split still reads.

The rack inventory had a related bug: a mount that names a node explicitly
printed the mount's device rather than the pinned node's. It now reads the
node's own row.

Tests that built a node with device columns are ported to the link; where a
behaviour genuinely moved (properties refresh once on the row, last_scan is
the device's) the assertion moved with it rather than being dropped.

ha-relevant: yes
2026-08-14 18:02:02 +02:00
Pouzor e53197d2cb feat(inventory): make the inventory row the source of a node's facts
A device drawn on three canvases was three independent copies of the same
facts. Point every node at the Device Inventory row it draws, and let that
row own what the device *is* — the node keeps only how it is drawn.

- nodes.device_id -> device_inventory.id, ON DELETE SET NULL. NULL for
  canvas furniture (group / groupRect / text), which describes nothing
  physical. Deleting a node never deletes the row.
- services/inventory_sync holds the shared rules: matching by ieee > ip >
  mac (per token, so 10.0.0.4 never matches 10.0.0.40), a property union on
  key, a service union on (port, protocol, name), and the backfill that
  links every pre-existing node.
- The backfill is non-destructive by construction: it writes device_id and
  fills the row, and deletes nothing. Nodes are visited oldest-edit-first,
  so where two canvases disagree on a scalar the most recently edited wins,
  while properties and services stay unioned — nothing any canvas recorded
  is lost. A second boot finds nothing to do.
- The wire shape does not change: GET /canvas hydrates the device fields
  from the row, and a save routes them back to it. Editing a node's IP on
  one canvas now shows on every other canvas holding that device.
- approve / bulk-approve set device_id instead of owning a copy, and a new
  canvas node joins (or mints) its row. Rows minted this way are tagged
  with a new `canvas` discovery source and get their own inventory filter.
- DetailPanel offers "Open in inventory"; standalone, which has no
  inventory, is not offered it.

test_racks' "reports what the canvas node knows" seeded a second inventory
row for a host that already had one — a state a node create can no longer
produce. Its seeding order is swapped so the node links to the row; every
assertion is unchanged.

ha-relevant: yes
2026-08-14 18:02:02 +02:00
Pouzor 4393ac92c6 feat(inventory): show and edit a device from the inventory
The Device Inventory was a read-only discovery log: rows arrived from a
scan or an import and went stale. Everything a user curates — properties,
services, notes, hardware, check method — could only be edited on a canvas
node, and never came back.

Give the inventory row the fields a node carries and a way to write them:

- device_inventory gains label, type, notes, cpu_count/cpu_model/ram_gb/
  disk_gb, show_hardware, check_method/check_target, last_seen, last_scan,
  response_time_ms and updated_at. Live reachability goes in a new
  status_live: `status` already means the pending/approved/hidden lifecycle
  and the two must not be conflated.
- PATCH /api/v1/scan/pending/{id} applies only what the client sends, so
  editing one field never clears the rest. Lifecycle and discovery
  bookkeeping stay owned by the approve/hide routes and the importers.
- InventoryDeviceCreate carries the same fields, so a hand-made entry needs
  no create-then-PATCH round trip.
- InventoryDeviceModal becomes show *and* edit, reusing the canvas editors
  (PropertyList, ServiceModal) rather than growing a second implementation.
- InventoryEntry moves to types/index.ts, its home; the modal re-exports it
  so existing call sites are untouched. NODE_TYPE_GROUPS moves to
  utils/nodeTypeGroups so both type pickers share one vocabulary.

ha-relevant: maybe
2026-08-14 18:02:02 +02:00
Pouzor 59079d90ca fix(canvas): seed a group child near its group and refuse nested groups
Follow-ups to the grouping fix. `handleAddNode` still gated its
"position near the parent" branch on `container_mode`, so a node added
straight into a group got a viewport-centred coordinate that `addNode`
then made group-relative — it landed far outside the box.

`isValidParentNode` also accepted a group or a zone as the child of a
group, which the drag-onto-a-group path refuses; both now agree.

ha-relevant: yes
2026-08-13 19:16:47 +02:00
Pouzor 26e18dafdd fix(canvas): keep a node inside its group when editing it
Editing any field of a grouped node dropped it out of the group on the
next save. NodeModal validated the parent against the type rules
(proxmox/vm/lxc/docker_host) or `container_mode`, and a `group` node is
neither — so it cleared `parent_id` on submit, `updateNode` detached the
node and the save persisted `parent_id: null`.

A shared `isValidParentNode` now treats a group as a valid parent for
any child type; the type rules only govern container nesting. Two
siblings of the same bug go with it: `updateNode` and `addNode` only
nested under a `container_mode` parent, so assigning or adding a node to
a group left it visually unparented until the next reload.

ha-relevant: yes
2026-08-13 19:16:47 +02:00
Pouzor c796cb39bd fix(services): split a comma-listed host override and type the wire host
The frontend takes the first host of a comma-separated override
(`splitFirstHost`), the backend took the whole string — so a two-host
override checked a hostname the UI never links to. `_parse_override`
now splits the same way.

The WebSocket service-status entry also declares `host` (nullable, as
the backend sends it), so the overlay key stops relying on an
undeclared field.

ha-relevant: yes
2026-08-13 19:16:47 +02:00