A Zigbee2MQTT networkmap on a 200+ device mesh takes minutes to build.
Two separate failures fell out of that:
- POST /zigbee/import held the HTTP request open for the whole MQTT
round-trip, so any reverse proxy in front of the API cut it first
(Cloudflare returns a 524 at 120 s) and the browser never saw the map.
It now registers a job, fetches in the background and answers 202; the
client polls GET /zigbee/import/{job_id} until the payload is ready.
Job results are transient and live in memory with a 15 min TTL — the
same single-worker assumption the scheduler already makes. A failed
fetch replays the status the synchronous route used to raise, so a bad
broker is still a 502 and a slow mesh still a 504.
- The networkmap wait was hard-coded at 300 s with no way to raise it.
It now reads ZIGBEE_NETWORKMAP_TIMEOUT, and the shared MQTT round-trip
used by the Z-Wave import reads MQTT_RESPONSE_TIMEOUT. Both default to
300 s, fall back to that if misconfigured to a non-positive value, and
name themselves in the timeout message.
Also corrects the route and doc claims that the wait was 60 s.
The /import tests changed with the contract they cover, not to pass.
Fixes#380
ha-relevant: yes
Every node type persisted its size in `nodes.width` / `nodes.height` except
`groupRect`, which stashed it inside the `custom_colors` JSON next to its
colours. The columns already existed and were simply unused for zones, so
this was an inconsistency rather than a missing-column workaround, and it
put geometry in a blob that otherwise holds style.
The serializer now writes the columns for a zone too, and strips the legacy
`width`/`height` keys out of the blob so the two cannot drift apart and
leave an older canvas reading a stale size.
No data is lost on upgrade:
- `_backfill_zone_size` copies the blob geometry into the columns at
startup. It only fills a column that is still NULL, so it cannot overwrite
a size set since; it parses the JSON in Python rather than with
`json_extract`, so it does not depend on the SQLite build carrying JSON1;
and an unreadable row is skipped without costing the others their size.
Re-running it is a no-op.
- the reader still falls back to the blob, covering a payload the backfill
has not reached — an older server, or an import.
Standalone mode is unaffected: it stores React Flow nodes verbatim, so the
size was always on `node.width` / `node.height` there.
The four serializer tests that pinned the size to the blob now assert the
columns, since that is the behaviour being changed.
ha-relevant: yes
Growing a zone during a subnet import wrote the new height twice: to
`node.height`, the live field, and to `data.custom_colors.height`, which
nothing reads. The blob copy is produced by the serializer at save time
(rebuilt from `node.height`, so the value written here was overwritten
before it reached the API) and consumed at load time off the API payload.
Writing it from the store was invisible, and misleading in a blob that
otherwise holds colours and style.
The test asserting the dead field is replaced by a save/load round-trip
through the real serializer, which is what actually protects the height.
ha-relevant: yes
The subnet import only guaranteed the zone preceded the nodes it pulled in.
An arrival that is itself a parent — a Proxmox host with nested VMs, whose
children stay put because they already have a parent — could end up listed
after those children, which React Flow renders detached with a
parent-not-found error. Nothing re-sorts on load, so the broken order
survived a save.
Reorder the whole array instead: every node now follows its parent, which
covers both directions at once. An already-valid list comes back untouched
and a parent cycle terminates rather than recursing.
ha-relevant: yes
Closes#325.
A zone gains an "Import devices by subnet" action: type a CIDR, and every
free device whose IP falls in that range moves into the zone, laid out on a
grid in its free space.
The CIDR is an argument to a one-shot action, never a property of the zone —
it is not submitted with the form, not persisted, and cleared after each run.
So there is no schema change, no migration, and it works in standalone mode.
The trade-off is that a device scanned later does not join the zone by
itself; the user re-runs the import.
Deliberate rules:
- additive — running it twice with two subnets leaves both sets inside, and
a device whose IP stops matching is never ejected
- only unparented devices move. A node nested in a group, a container host
or another zone keeps the parent the user gave it
- canvas furniture (groupRect / group / text) is skipped, and a zone never
swallows another zone
- one history entry per import, so a single Undo reverses the whole thing
Edit mode gets an Import button; add mode has none, since the zone does not
exist yet and a button would look broken — there the CIDR is applied right
after the zone is created.
IPv4 only: the modal rejects an IPv6 CIDR with a message rather than
silently matching nothing.
ha-relevant: yes
A scan runs on a background thread inside the API process. If that process
dies mid-scan — an OOM kill, docker stop, a crash — the ScanRun row stays
"running" for ever, because nothing is left alive to finish it.
That row is not just cosmetic clutter in Scan History: the trigger endpoints
reject a new scan while one is "running" for the same target, so a single kill
locks that range out permanently.
Nothing can legitimately be "running" the moment we boot, so lifespan() now
marks every such row "error" — the same word run_scan and run_device_scan
write when they fail themselves — with finished_at and an explanatory message.
Reported alongside the OOM itself in #374, which is what produced the orphans.
Fixes#374
ha-relevant: maybe
Both HTTP paths buffered the whole response body before looking at it. An
endpoint that streams without end and sends no Content-Length — a Freebox
bandwidth-test port, an MJPEG camera, a log tail — grew the backend until the
cgroup limit killed it, every 5-10 minutes at a constant ~1.04 GB RSS.
httpx's timeout does not help: it applies per network operation, not to the
total time spent draining a socket that keeps delivering data.
- status_checker._http_get only needs the status line, so it now uses
client.stream() and never reads the body at all.
- http_probe._probe_scheme needs at most _MAX_BODY_BYTES to hunt for <title>,
so it streams and stops there. The cap existed already but was applied to
resp.text, after the full body had been downloaded.
Regression tests serve an endless, Content-Length-less body through a
MockTransport and assert the read stays bounded. Both fail on the old code.
The existing mocks patched httpx.AsyncClient.get, which neither path calls
now; they are rebuilt on MockTransport — a real client over a fake network —
with every original assertion kept.
Fixes#375
ha-relevant: yes
Three findings from the PR review.
A finishing deep rescan threw away an edit in progress. The reset effect in
InventoryDeviceModal was keyed on the `device` object, and the parent hands
down a fresh one whenever the row is refreshed — including from the poll's
own onSaved. Same device, new object, so the effect reset the form and left
edit mode, minutes into a scan the user was waiting on. It is keyed on the
id now: a refreshed row is not a different device, and only a different
device is a reason to throw the form away. The fresh row still reaches the
canvas and the grid — gating that behind edit mode would discard the scan
result instead, and a save unions the services server-side anyway.
A rescan tagged every device it touched as "arp"-discovered, so a Proxmox
guest, a rack mount or a hand-added host started answering the network
source filter. It carries its own source through instead.
The two background wrappers wrote status "failed" where the scanners write
"error" for the same condition. Harmonized on "error" — the one Scan
History filters and colours; "failed" showed up unlabelled. The frontend
union keeps 'failed' for rows already recorded.
That last one meant updating an existing assertion in test_scan_run.py: it
encoded the old spelling.
ha-relevant: yes
merge_services did {**existing, **incoming}, so the fingerprint's guess at
an icon overwrote the one the user chose — and on a port no signature
covers it wrote None, clearing it outright. Since 3.3.0 the inventory row
is the only copy of a device's services, so every "Scan network" repainted
the service on every canvas drawing that device at once.
The scanner now merges with discovered=True: it still adds services and
refreshes what it knows, but leaves an established icon and category alone.
A user edit from the modal or a canvas changes them as before.
Blank incoming values no longer clear established ones either, on both
paths — an absent field is silence, not a reset. Same rule merge_properties
already follows.
ha-relevant: yes
The Deep scan link on a device now opens a small dialog instead of firing
straight away. It is prefilled with the full 1-65535 range — that is still
the point of the feature — but a user who knows where a service lives can
narrow it and get an answer in seconds instead of minutes.
The dialog takes an nmap-style spec: a port, a range, or a comma list
(80,443,8000-9000). It shows the port count live, refuses to start on
something nmap could not use, and carries three presets (all / 1-1024 /
1-10000).
Backend:
- _parse_port_spec / _port_chunks generalize the slicing that used to be
full-range only. Ranges are merged before slicing, so an overlapping
spec is never scanned twice, and small ranges are packed into one nmap
call instead of one call each. _deep_port_chunks still yields the same
eight slices as before.
- run_device_scan takes ports=; it wins over full_ports.
- The retry-free flags (--max-retries 0 --min-rate 2000) now key on the
total port count rather than on full_ports. They pay for themselves over
thousands of ports on a lossy host; over a handful they only cost
accuracy.
- RescanDeviceRequest.ports validates the spec — 422 rather than handing
nmap a bad -p. Blank means the full sweep.
Frontend utils/portSpec.ts mirrors the backend parser so a typo is caught
before the request; the backend validates again because it is the one
calling nmap.
The existing rescan tests now go through the dialog: the click path
changed, so the start is two steps.
ha-relevant: yes
A deep rescan of a slow host came back with nothing at all: the run took
its full 600s ceiling and the device's services were unchanged, so a
service deleted by hand was never rediscovered.
nmap answers --host-timeout with "Skipping host <ip> due to host timeout"
and discards every port it had already found — the ceiling turned a slow
scan into one that reports nothing. What costs the time is a host that
drops packets: 8188 of 8192 ports filtered, each waiting out its probe.
- No --host-timeout on the deep discovery pass, ever.
- The full range runs as 8 slices of 8192 ports, one nmap call each,
unioning the open ports. A slice that overruns costs its own ports, not
all of them, and the loop has somewhere to notice a stop request.
- `scanner_deep_host_timeout` is now a total budget checked between
slices (default 2700s), not an nmap flag. The first slice always runs.
- A partial sweep is reported rather than passed off as complete: the run
finishes `done` carrying "Scanned 3/8 port ranges …", and the modal
toasts a warning instead of success.
- Deep slices use --max-retries 0 --min-rate 2000. Measured against a
dropping host, 8192 ports took 329s at --max-retries 1 and 164s at 0,
finding the same ports; capping the RTT changed nothing. The range scan
keeps nmap's default retries on its curated port list.
ha-relevant: yes
Devices added before the scanner knew a service showed an empty Services
section with no way to refresh it (#350). The detail modal now starts a
full-port scan of that single device.
- `process_host` lifted out of `run_scan` so the range scan and the new
single-device scan share the same match / merge / dedupe rules — a
rescan unions services, it never replaces what the user added by hand.
- `run_device_scan`: no ping sweep, no mDNS, straight to the phase-2 nmap
pass on the device IP over all 65535 TCP ports.
- `POST /scan/pending/{id}/rescan` records a normal ScanRun
(`kind=device`, `ranges=["<ip>/32"]`), so stop, progress and Scan
History work unchanged. One run per device at a time — a second request
while the first is scanning is a 409. 404 unknown, 409 no-IP or hidden.
- `GET /scan/runs/{id}` so a caller can poll the run it started.
- Deep scan button in the Services section of the device detail, hidden
without an IP and for Zigbee. Swaps to a stop control while running,
folds the fresh services back in on completion (never over an edit in
progress).
ha-relevant: yes
The property line caps its label at max-w-15 so a long key cannot crowd
out the value drawn beside it. When the value is empty there is nothing
to protect, but the cap still applied, truncating the label against
empty space.
Drop the cap when the value is blank and let the label truncate against
the node width instead. Same fix in BaseNode and ProxmoxGroupNode.
Closes#361
ha-relevant: yes
A same-origin deployment behind a reverse proxy needs no CORS, so
CORS_ORIGINS is left at its localhost default — but browsers still send
Origin on unsafe methods, and OIDCCSRFMiddleware validated it against
CORS_ORIGINS alone. Every POST/PUT/DELETE came back 403 while reads
worked, so creating a canvas failed with no usable error.
The OIDC callback URL is served by this app, so the origin of
OIDC_REDIRECT_URI is the app's own and is always a legitimate CSRF
origin. origin_is_allowed now accepts it in addition to CORS_ORIGINS,
and settings validation warns at startup when CORS_ORIGINS omits it so
the misconfiguration is visible rather than silent.
Fixes#356
ha-relevant: no
Since 3.3.0 the inventory row is the only copy of a device's services and
every canvas drawing it reads that list, but the scanner still assigned
`keep.services = fingerprint_ports(...)` — the pre-split behaviour, when a
node held its own copy. One re-scan therefore deleted every service the user
added by hand, on every canvas at once, and brought back the ones they had
deleted. It unions now, like every other writer of that field.
Two more things #347 turned up:
* A node whose view matches nothing the row still holds drew nothing at all —
the row had been replaced under it, so every key was gone and every key was
new, and `apply_view` hid the lot. Such a view says nothing about the list
that replaced it, so it is treated as no view: the row is drawn. An empty
view is untouched, being a real answer ("this canvas draws none of them").
* The 3.3.3 view seed recovers a node's arrangement from the pre-3.3.0 backup,
which is 3.2.0-era and cannot know about a property added afterwards. On
3.3.0-3.3.2 the row was the only place to add one and every canvas drew it,
so seeding strictly from the backup took it off all of them at once. Those
are appended visible, keeping the recovered order and hidden flags for
everything the backup does know. Services keep the strict recovery: holding
back what a scan fingerprinted is the whole point of the view.
ha-relevant: maybe
The seed selected any node with a `device_id` and no view. Deleting a device
leaves that column dangling — SQLite runs with foreign keys off, so the
`ON DELETE SET NULL` never fires on an existing table — and such a node can be
seeded from nothing, so it stayed selected. The query matched it again on every
later start, re-opening the pre-upgrade backup and logging a recovery that
recovered nothing.
Narrowing the query to nodes whose row actually exists restores the invariant
the seed is documented to have: it runs on one boot, and never opens a file
again.
ha-relevant: no
Since 3.3.0 the inventory row owns a device's services and properties, and
every node drawing that device rendered the row wholesale. One row shared by
several canvases meant one rendering: a service the scanner fingerprinted
appeared on every canvas at once (users reported Uptime Kuma and Synology DSM
on hosts running neither — both are port-only signatures), and a property added
on one schematic showed up on all the others.
Order and visibility are presentation, so they move to the node. `display_view`
records, per node, which of the row's services and properties it draws and in
what order, keyed by `port|protocol|name` and by lowercased property key so the
view survives an edit to the fact itself. The facts stay on the row: hiding is
per node, deleting is still device-wide.
An item the view does not list is reported hidden rather than dropped, so what
a scan finds next is one toggle away instead of pushed onto every canvas.
The wire shape is unchanged. A client already sends its lists in display order
with their `visible` flags, so the view is read back out of them rather than
asking for a second field, and `visible` is only stamped when something is
hidden.
Upgrades keep what each canvas showed:
* from 3.2.0, the backfill seeds each node's view from its own legacy columns
before they are dropped;
* from 3.3.0-3.3.2 those columns are gone and the row holds the union of every
canvas, so the layout is recovered from the backup `_backup_db` took before
the 3.3.0 migration — the newest one whose `nodes` table still has the
columns, read read-only and matched by node id;
* with no usable backup the row is the seed, so every canvas keeps showing
exactly what it shows today and only later additions are held back.
Refs #347
ha-relevant: maybe
3.3.1 kept `nodes` writable when the backfill failed, but the backfill still
failed — so a canvas upgraded from 3.2.0 shows every device stripped of its ip,
services, hostname, notes and hardware. The facts are intact in the legacy
columns; nothing was ever reading them across.
The backfill reads those columns with raw SQL, so SQLite hands `last_seen` and
`last_scan` back as text. Writing text into a DateTime column raises
`TypeError: SQLite DateTime type only accepts Python datetime and date objects`,
which is neither IntegrityError nor OperationalError — so it went straight past
the per-node savepoint and aborted the whole run, exactly as before.
- Parse the legacy timestamps back into datetimes, the way `_decode_json`
already handles the JSON columns. An unparseable stamp becomes None rather
than an error: the status checker refreshes both within a minute.
- Catch every exception per node, not two chosen classes. Picking the exception
types was the defect; a node that fails for any reason must cost only itself.
The same run would also have destroyed data. A canvas saved while an earlier
migration was stuck minted a blank inventory row from a UI that had no facts to
show and linked the node to it. The backfill only looked at unlinked nodes, so
it skipped that one, counted the canvas migrated, and dropped the columns
holding its only copy of the ip and services.
- Read linked nodes too, and fill only what their row is missing. A row that
already holds a fact is never overwritten, so an edit made after the migration
survives, and the pass stays a no-op once everything has moved across.
Verified end to end against a rebuilt 3.2.0 database stuck the way the reports
describe: one boot restores ip, services, hostname, mac, notes, check method,
hardware, properties, status and last_seen, then drops the legacy columns.
Refs #347, #348, #351
ha-relevant: no
Upgrading 3.2.0 -> 3.3.0 could leave a database where approving a device —
or creating any node — failed with `NOT NULL constraint failed: nodes.status`.
The 3.3.0 migration moves the device facts off `nodes` onto `device_inventory`
and then drops the columns. The drop is skipped when a node is still unlinked,
because those columns are the only remaining copy of its facts. But a 3.2.0
`nodes` declares `status`, `services`, `properties` and `show_hardware` NOT NULL
with no server-side default, and the 3.3.0 model no longer writes them — so a
skipped drop bricks every later INSERT.
It skipped because one node killed the whole backfill: the loop shared a single
session, so an IntegrityError raised by autoflush lost every link made before it.
- Give each node its own savepoint. A node whose merge violates a constraint is
skipped and logged on its own; the rest still link. `skipped` joins the stats
reported at boot.
- Relax NOT NULL on the retained legacy columns when the drop is skipped, so the
database stays writable while later boots retry the backfill. Values are kept;
only the constraint goes. Guarded on the PRAGMA notnull flag, so it runs once.
- Share one `_rebuild_nodes` between the drop and the relax: it restores
`PRAGMA foreign_keys = ON` in a finally (the connection returns to the pool)
and clears a `nodes_new` left by a failed attempt.
- Match IEEE addresses case-insensitively, and never write one another inventory
row already owns — `device_inventory.ieee_address` is UNIQUE, and a duplicate
is one way the backfill was raising in the first place.
Tests build the real v3.2.0 schema from the tag: the existing legacy-migration
fixture declares `status` nullable, which is why this reached a release.
Fixes#351
ha-relevant: no
The quick starts already pull ready-made images through install.sh, but
nothing said so — issue #310 asked for something that already ships.
Add a "Pre-built Docker images" section to INSTALLATION.md listing the
four GHCR images, their tag scheme and a manual docker-compose.prebuilt.yml
setup, and note that the root docker-compose.yml builds from source.
Closes#310
ha-relevant: no
A property may now carry a label alone: the add and edit forms require
only the key, and every surface that prints one — the badge list, node
and Proxmox group renderers, rack cable annotations — drops the "· value"
separator when the value is empty.
ha-relevant: yes
Dropping a node onto a groupRect zone now asks to add it to the zone and
sets a real React Flow parentId, so moving the zone moves its contents.
Zone children are deliberately not extent-clamped, unlike group and
container children: a zone has no side panel to release a child from, so
dragging the node back out of the zone is what detaches it.
Deleting a zone now releases its children onto the canvas in absolute
coordinates instead of cascade-deleting them; groups and containers keep
the cascade. Collapse counts and hides both parented and merely
overlapping nodes, and the deserializer restores zone parenting on load.
ha-relevant: yes
Resolves the two high Dependabot alerts on the frontend lockfile:
- undici 7.28.0 -> 7.29.0 (transitive via jsdom): cross-user information
disclosure and parse-time crash via degenerate private cache directives
- ip-address 10.2.0 -> 10.5.0 (transitive via shadcn ->
@modelcontextprotocol/sdk -> express-rate-limit): leading-zero octets
decoded as decimal, allowing SSRF and trust-boundary bypass
Also closes the moderate advisories on the same two packages. Lockfile only,
both are dev dependencies.
ha-relevant: no
Three unrelated fixes from the review of this branch.
`create_pending` merging into a hidden row left it hidden, so adding a device
by hand whose IP collides with one hidden weeks ago answered 201 while nothing
appeared in the inventory — the add read as a no-op. An explicit add outranks
the earlier hide, exactly like restore. Only `hidden` is lifted; an approved
row is not walked back down the lifecycle.
Standalone stored a node's live reachability in the inventory row's `status`,
which is the pending/approved/hidden lifecycle — the backend keeps the two
apart as `status` vs `status_live`. Reachability now goes to `status_live`, and
`normalizeLifecycle` repairs rows already written the old way: the node still
reads its status on load, and the next save rewrites the row correctly.
Status checks stay device-scoped and keep covering a device no canvas draws.
That is deliberate — the inventory row is what's monitored, so a rack mount or
an inventory-only entry reports state without being drawn anywhere, and
deleting a node no longer silently stops monitoring the host. `hide`, or
clearing the check method, is what ends the checks. Documented and pinned with
tests rather than changed.
ha-relevant: maybe
A canvas node carries a full copy of its Device Inventory row, hydrated when
the canvas loads. The save routed all of it back, so a save made for nothing
but a moved node rewrote the row from a snapshot that could be hours old —
silently reverting an edit made meanwhile in the inventory modal, on another
canvas, or by the scanner.
Diffing the payload against the row server-side cannot fix this: it can't tell
"I edited this" from "the row moved on since I loaded it". Only the client
holds the baseline.
- canvasStore keeps `factsBaseline` — the device facts as received — set on
load, refreshed on save, rebased per field by `applyDeviceFacts`.
- `serializeNode` sends `changed_facts`: what this canvas actually edited.
`link_facts(changed_fields=)` writes nothing outside that list, even where
the values differ. Absent (older client, YAML import, MCP) keeps the previous
full-write behaviour.
- `changed_facts()` additionally drops facts already equal to the row, so a
no-op save writes nothing. Identity matching still uses the full payload.
- Live fields (status, last_seen…) bypass the filter — the status checker owns
reachability, and a save only ever fills a row never checked.
An inventory edit also lands on the canvases already on screen:
`applyDeviceFacts` pushes the saved row onto every node drawing it, without
marking the canvas unsaved. A fact the canvas has edited but not saved is left
alone — work in progress wins locally and still saves.
ha-relevant: maybe
This reverts commit 779e09ecb8.
The bump does not belong to this PR. 3.3.0 ships several PRs, so the version
files and the release notes are raised once, in a release PR of their own,
covering the whole diff against main.
ha-relevant: no
The Device Inventory row becomes the source of a device's facts, and the
device columns are dropped from `nodes` in a one-way migration on first start,
so the release notes lead with a Breaking section: what moved, what the
migration does, when it refuses to drop, and that `homelab.db.back-3.3.0` —
written before any migration runs — is the way back.
ha-relevant: no
An install that predates the split holds canvases whose nodes carry the device
facts inline, with no `devices` key at all. Reading one already worked; nothing
pinned that the next save converts it in place rather than writing the old
shape back.
Covers the load-then-save round trip with every fact intact, the second save
not minting a duplicate row, furniture staying out of the conversion, and the
legacy bare-key canvas carried through `ensureSeed`.
ha-relevant: no
`_slim_canvas` filtered `n["data"]`, the React Flow shape the frontend store
holds. `GET /api/v1/canvas` has never sent that: a node is reported flat, so
the filter matched nothing and `get_canvas` returned an id and a node type per
node — no label, no address, no services. The three tests covering it fed a
hand-built nested payload, so they passed against a shape the backend does not
produce.
Fold `data` over the node and read both, so either form slims the same way.
`type` leaves `NODE_KEEP` — flat it is the node type, not a device fact, and it
is already reported as `node_type`; `parentId` becomes `parent_id`, the key the
API uses, normalized from either spelling.
The new tests are built on a full `NodeResponse` copied from a real response,
and the backend now pins that wire shape from its side so the two move together.
ha-relevant: no
`merge_facts_into_device` applied one `replace_lists` flag to properties
and services alike, taking each wholesale from `facts`. But a PATCH only
carries the keys the client sent, so `{"properties": [...]}` replaced
services with `[]` — losing them for every canvas drawing the device and
stopping the scheduler from service-checking it. Symmetric the other way.
Gate the replace per list: a list absent from `facts` falls through to the
merge, which is a no-op against `None`. An explicit `[]` still clears, and
the canvas save path is unchanged since it always sends both keys.
ha-relevant: no
A device documented straight on a canvas carries a `label` and a `type`
and no discovery guess at all — no scan ever saw it. The card read
`suggested_type` alone, so a printer drawn on a canvas fell back to
'generic': unnamed, filed under no type, and drawn with a plain circle.
Two faults compounded:
- The card decided the kind from `suggested_type`, ignoring the curated
`type`. That also meant retyping a scanned device in the detail modal
changed nothing here, and approving it still built a node of the type
the scanner had guessed.
- It carried its own table of eleven icons, unrelated to the ~40
NodeType values, so anything outside that list drew a circle even with
the right type in hand.
`deviceType()` now resolves `type` before `suggested_type` for the icon,
the badge, the type filter and both approve paths, and the icons come
from the shared `NODE_TYPE_DEFAULT_ICONS` — one map, the same one the
canvas and the detail modal use. `deviceLabel` prefers the curated
`label` for the same reason, otherwise a canvas-created device shows as
its IP.
ha-relevant: maybe
The services sat in the curation column, beside the properties, while the
addresses that reach them were two columns away. They move under
Identity: what answers on a host belongs with how to reach it, and the
curation column now carries the properties alone.
Every section heading takes an icon — identity, services, monitoring,
activity, notes, properties, make & model — so a column is scannable
without reading each title. `Section` requires it, so a new section
cannot be added without one. The properties editor keeps its own plain
heading: it comes from the shared `PropertyList`, which the rack cable
panel also renders.
The three view columns carry a data-testid, which is what lets the
layout be asserted rather than eyeballed.
ha-relevant: maybe
The detail modal was a narrow column: a 15px icon, eighteen key-value
rows in one grey block, and a form that grew the dialog to `max-w-lg`
while the header and the buttons scrolled away with the content.
It reads as a device sheet now, following NodeModal's shape:
- One wide dialog with three fixed regions — header, scrolling body,
footer — so the title and the actions stay put.
- The header identifies the row at a glance: a type-coloured icon tile
(the accent the node wears on the canvas), then a badge per discovery
source, the type, live reachability with a dot, and how many canvases
draw the device.
- Three columns instead of one: identity, operations, curation. A
section stays put when it has nothing to show and says why, rather
than disappearing and leaving the reader unsure whether the field is
absent or unsupported.
- Timestamps read relative with the exact value on hover, matching the
inventory cards. IP / hostname / MAC / IEEE offer a copy button.
- The form uses the same control sizing as the canvas modals, and gains
inputs for `friendly_name` and `device_subtype` — both were already
submitted, so they could not be corrected here.
Hardware is a property like any other, so the modal no longer edits it:
nothing draws `cpu_*` / `ram_gb` / `disk_gb` / `show_hardware`, and the
CPU / RAM / Disk keys are property suggestions instead. The columns are
still sent back untouched — Proxmox and the YAML import write them and
the YAML export reads them.
Two assertions in the edit tests covered the removed section: the
curated-fields test drops the hardware line, and the numeric-payload
test stops editing a field that no longer exists while still checking
the values leave as numbers.
ha-relevant: maybe
A node draws its `properties`; `cpu_model` / `cpu_count` / `ram_gb` /
`disk_gb` are inventory data that nothing renders — BaseNode reads them
only when `properties` is absent, which the API never allows (it returns
`device.properties or []`, and the model defaults to a list).
The YAML import filled those fields and set `show_hardware`, so an
imported node did show its specs — until the first save, when the
serializer sent `properties: []`, the fallback died and the specs
silently vanished.
It now mints the same four properties the 3.x migration and the Proxmox
import already create (CPU Model, CPU Cores, RAM, Disk), visible, so the
specs survive the save. The structured fields stay written: the YAML
export reads them back.
`show_hardware` is no longer set — it drives nothing.
The two assertions that covered the old behaviour move with it: the
scalar-field test drops `show_hardware`, and "sets show_hardware only
when hardware fields present" becomes three tests over the minted
properties.
ha-relevant: yes
Completes the split: `nodes` now holds only how a device is drawn on one
canvas, and every device fact reaches the API from the inventory row.
- The device columns are removed from `nodes` (SQLite table rebuild, the
same shape as the existing device_inventory and canvas_state rebuilds).
The drop is skipped, and logged, while any non-furniture node is still
unlinked: those columns are the last copy of that node's facts. The
backfill therefore reads them with raw SQL — by the time it runs, the
model no longer declares them.
- The status checker iterates devices, not nodes: one check per device
however many canvases draw it, writing `status_live` / `last_seen` /
`response_time_ms` on the row. `/ws/status` messages carry `device_id`
and the node ids they light up. Hidden devices are not probed.
- Readers repointed: scanner (last_scan lands on the row), proxmox (its
node tier collapses into the inventory tier, keeping only the cluster
handles), zigbee/zwave (one property refresh serves every canvas), rack
inventory, liveview, stats, node dedupe.
- `POST /scan/pending` merges into the row that already describes the host
instead of minting a second one — one device is one row, whichever way it
was documented.
- Standalone keeps parity: the canvas blob gains `devices`, split on save
and hydrated on load. A blob written before the split still reads.
The rack inventory had a related bug: a mount that names a node explicitly
printed the mount's device rather than the pinned node's. It now reads the
node's own row.
Tests that built a node with device columns are ported to the link; where a
behaviour genuinely moved (properties refresh once on the row, last_scan is
the device's) the assertion moved with it rather than being dropped.
ha-relevant: yes
A device drawn on three canvases was three independent copies of the same
facts. Point every node at the Device Inventory row it draws, and let that
row own what the device *is* — the node keeps only how it is drawn.
- nodes.device_id -> device_inventory.id, ON DELETE SET NULL. NULL for
canvas furniture (group / groupRect / text), which describes nothing
physical. Deleting a node never deletes the row.
- services/inventory_sync holds the shared rules: matching by ieee > ip >
mac (per token, so 10.0.0.4 never matches 10.0.0.40), a property union on
key, a service union on (port, protocol, name), and the backfill that
links every pre-existing node.
- The backfill is non-destructive by construction: it writes device_id and
fills the row, and deletes nothing. Nodes are visited oldest-edit-first,
so where two canvases disagree on a scalar the most recently edited wins,
while properties and services stay unioned — nothing any canvas recorded
is lost. A second boot finds nothing to do.
- The wire shape does not change: GET /canvas hydrates the device fields
from the row, and a save routes them back to it. Editing a node's IP on
one canvas now shows on every other canvas holding that device.
- approve / bulk-approve set device_id instead of owning a copy, and a new
canvas node joins (or mints) its row. Rows minted this way are tagged
with a new `canvas` discovery source and get their own inventory filter.
- DetailPanel offers "Open in inventory"; standalone, which has no
inventory, is not offered it.
test_racks' "reports what the canvas node knows" seeded a second inventory
row for a host that already had one — a state a node create can no longer
produce. Its seeding order is swapped so the node links to the row; every
assertion is unchanged.
ha-relevant: yes
The Device Inventory was a read-only discovery log: rows arrived from a
scan or an import and went stale. Everything a user curates — properties,
services, notes, hardware, check method — could only be edited on a canvas
node, and never came back.
Give the inventory row the fields a node carries and a way to write them:
- device_inventory gains label, type, notes, cpu_count/cpu_model/ram_gb/
disk_gb, show_hardware, check_method/check_target, last_seen, last_scan,
response_time_ms and updated_at. Live reachability goes in a new
status_live: `status` already means the pending/approved/hidden lifecycle
and the two must not be conflated.
- PATCH /api/v1/scan/pending/{id} applies only what the client sends, so
editing one field never clears the rest. Lifecycle and discovery
bookkeeping stay owned by the approve/hide routes and the importers.
- InventoryDeviceCreate carries the same fields, so a hand-made entry needs
no create-then-PATCH round trip.
- InventoryDeviceModal becomes show *and* edit, reusing the canvas editors
(PropertyList, ServiceModal) rather than growing a second implementation.
- InventoryEntry moves to types/index.ts, its home; the modal re-exports it
so existing call sites are untouched. NODE_TYPE_GROUPS moves to
utils/nodeTypeGroups so both type pickers share one vocabulary.
ha-relevant: maybe
Follow-ups to the grouping fix. `handleAddNode` still gated its
"position near the parent" branch on `container_mode`, so a node added
straight into a group got a viewport-centred coordinate that `addNode`
then made group-relative — it landed far outside the box.
`isValidParentNode` also accepted a group or a zone as the child of a
group, which the drag-onto-a-group path refuses; both now agree.
ha-relevant: yes
Editing any field of a grouped node dropped it out of the group on the
next save. NodeModal validated the parent against the type rules
(proxmox/vm/lxc/docker_host) or `container_mode`, and a `group` node is
neither — so it cleared `parent_id` on submit, `updateNode` detached the
node and the save persisted `parent_id: null`.
A shared `isValidParentNode` now treats a group as a valid parent for
any child type; the type rules only govern container nesting. Two
siblings of the same bug go with it: `updateNode` and `addNode` only
nested under a `container_mode` parent, so assigning or adding a node to
a group left it visually unparented until the next reload.
ha-relevant: yes
The frontend takes the first host of a comma-separated override
(`splitFirstHost`), the backend took the whole string — so a two-host
override checked a hostname the UI never links to. `_parse_override`
now splits the same way.
The WebSocket service-status entry also declares `host` (nullable, as
the backend sends it), so the overlay key stops relying on an
undeclared field.
ha-relevant: yes