A Zigbee2MQTT networkmap on a 200+ device mesh takes minutes to build.
Two separate failures fell out of that:
- POST /zigbee/import held the HTTP request open for the whole MQTT
round-trip, so any reverse proxy in front of the API cut it first
(Cloudflare returns a 524 at 120 s) and the browser never saw the map.
It now registers a job, fetches in the background and answers 202; the
client polls GET /zigbee/import/{job_id} until the payload is ready.
Job results are transient and live in memory with a 15 min TTL — the
same single-worker assumption the scheduler already makes. A failed
fetch replays the status the synchronous route used to raise, so a bad
broker is still a 502 and a slow mesh still a 504.
- The networkmap wait was hard-coded at 300 s with no way to raise it.
It now reads ZIGBEE_NETWORKMAP_TIMEOUT, and the shared MQTT round-trip
used by the Z-Wave import reads MQTT_RESPONSE_TIMEOUT. Both default to
300 s, fall back to that if misconfigured to a non-positive value, and
name themselves in the timeout message.
Also corrects the route and doc claims that the wait was 60 s.
The /import tests changed with the contract they cover, not to pass.
Fixes#380
ha-relevant: yes
Three findings from the PR review.
A finishing deep rescan threw away an edit in progress. The reset effect in
InventoryDeviceModal was keyed on the `device` object, and the parent hands
down a fresh one whenever the row is refreshed — including from the poll's
own onSaved. Same device, new object, so the effect reset the form and left
edit mode, minutes into a scan the user was waiting on. It is keyed on the
id now: a refreshed row is not a different device, and only a different
device is a reason to throw the form away. The fresh row still reaches the
canvas and the grid — gating that behind edit mode would discard the scan
result instead, and a save unions the services server-side anyway.
A rescan tagged every device it touched as "arp"-discovered, so a Proxmox
guest, a rack mount or a hand-added host started answering the network
source filter. It carries its own source through instead.
The two background wrappers wrote status "failed" where the scanners write
"error" for the same condition. Harmonized on "error" — the one Scan
History filters and colours; "failed" showed up unlabelled. The frontend
union keeps 'failed' for rows already recorded.
That last one meant updating an existing assertion in test_scan_run.py: it
encoded the old spelling.
ha-relevant: yes
The Deep scan link on a device now opens a small dialog instead of firing
straight away. It is prefilled with the full 1-65535 range — that is still
the point of the feature — but a user who knows where a service lives can
narrow it and get an answer in seconds instead of minutes.
The dialog takes an nmap-style spec: a port, a range, or a comma list
(80,443,8000-9000). It shows the port count live, refuses to start on
something nmap could not use, and carries three presets (all / 1-1024 /
1-10000).
Backend:
- _parse_port_spec / _port_chunks generalize the slicing that used to be
full-range only. Ranges are merged before slicing, so an overlapping
spec is never scanned twice, and small ranges are packed into one nmap
call instead of one call each. _deep_port_chunks still yields the same
eight slices as before.
- run_device_scan takes ports=; it wins over full_ports.
- The retry-free flags (--max-retries 0 --min-rate 2000) now key on the
total port count rather than on full_ports. They pay for themselves over
thousands of ports on a lossy host; over a handful they only cost
accuracy.
- RescanDeviceRequest.ports validates the spec — 422 rather than handing
nmap a bad -p. Blank means the full sweep.
Frontend utils/portSpec.ts mirrors the backend parser so a typo is caught
before the request; the backend validates again because it is the one
calling nmap.
The existing rescan tests now go through the dialog: the click path
changed, so the start is two steps.
ha-relevant: yes
Devices added before the scanner knew a service showed an empty Services
section with no way to refresh it (#350). The detail modal now starts a
full-port scan of that single device.
- `process_host` lifted out of `run_scan` so the range scan and the new
single-device scan share the same match / merge / dedupe rules — a
rescan unions services, it never replaces what the user added by hand.
- `run_device_scan`: no ping sweep, no mDNS, straight to the phase-2 nmap
pass on the device IP over all 65535 TCP ports.
- `POST /scan/pending/{id}/rescan` records a normal ScanRun
(`kind=device`, `ranges=["<ip>/32"]`), so stop, progress and Scan
History work unchanged. One run per device at a time — a second request
while the first is scanning is a 409. 404 unknown, 409 no-IP or hidden.
- `GET /scan/runs/{id}` so a caller can poll the run it started.
- Deep scan button in the Services section of the device detail, hidden
without an IP and for Zigbee. Swaps to a stop control while running,
folds the fresh services back in on completion (never over an edit in
progress).
ha-relevant: yes
Three unrelated fixes from the review of this branch.
`create_pending` merging into a hidden row left it hidden, so adding a device
by hand whose IP collides with one hidden weeks ago answered 201 while nothing
appeared in the inventory — the add read as a no-op. An explicit add outranks
the earlier hide, exactly like restore. Only `hidden` is lifted; an approved
row is not walked back down the lifecycle.
Standalone stored a node's live reachability in the inventory row's `status`,
which is the pending/approved/hidden lifecycle — the backend keeps the two
apart as `status` vs `status_live`. Reachability now goes to `status_live`, and
`normalizeLifecycle` repairs rows already written the old way: the node still
reads its status on load, and the next save rewrites the row correctly.
Status checks stay device-scoped and keep covering a device no canvas draws.
That is deliberate — the inventory row is what's monitored, so a rack mount or
an inventory-only entry reports state without being drawn anywhere, and
deleting a node no longer silently stops monitoring the host. `hide`, or
clearing the check method, is what ends the checks. Documented and pinned with
tests rather than changed.
ha-relevant: maybe
A canvas node carries a full copy of its Device Inventory row, hydrated when
the canvas loads. The save routed all of it back, so a save made for nothing
but a moved node rewrote the row from a snapshot that could be hours old —
silently reverting an edit made meanwhile in the inventory modal, on another
canvas, or by the scanner.
Diffing the payload against the row server-side cannot fix this: it can't tell
"I edited this" from "the row moved on since I loaded it". Only the client
holds the baseline.
- canvasStore keeps `factsBaseline` — the device facts as received — set on
load, refreshed on save, rebased per field by `applyDeviceFacts`.
- `serializeNode` sends `changed_facts`: what this canvas actually edited.
`link_facts(changed_fields=)` writes nothing outside that list, even where
the values differ. Absent (older client, YAML import, MCP) keeps the previous
full-write behaviour.
- `changed_facts()` additionally drops facts already equal to the row, so a
no-op save writes nothing. Identity matching still uses the full payload.
- Live fields (status, last_seen…) bypass the filter — the status checker owns
reachability, and a save only ever fills a row never checked.
An inventory edit also lands on the canvases already on screen:
`applyDeviceFacts` pushes the saved row onto every node drawing it, without
marking the canvas unsaved. A fact the canvas has edited but not saved is left
alone — work in progress wins locally and still saves.
ha-relevant: maybe
Completes the split: `nodes` now holds only how a device is drawn on one
canvas, and every device fact reaches the API from the inventory row.
- The device columns are removed from `nodes` (SQLite table rebuild, the
same shape as the existing device_inventory and canvas_state rebuilds).
The drop is skipped, and logged, while any non-furniture node is still
unlinked: those columns are the last copy of that node's facts. The
backfill therefore reads them with raw SQL — by the time it runs, the
model no longer declares them.
- The status checker iterates devices, not nodes: one check per device
however many canvases draw it, writing `status_live` / `last_seen` /
`response_time_ms` on the row. `/ws/status` messages carry `device_id`
and the node ids they light up. Hidden devices are not probed.
- Readers repointed: scanner (last_scan lands on the row), proxmox (its
node tier collapses into the inventory tier, keeping only the cluster
handles), zigbee/zwave (one property refresh serves every canvas), rack
inventory, liveview, stats, node dedupe.
- `POST /scan/pending` merges into the row that already describes the host
instead of minting a second one — one device is one row, whichever way it
was documented.
- Standalone keeps parity: the canvas blob gains `devices`, split on save
and hydrated on load. A blob written before the split still reads.
The rack inventory had a related bug: a mount that names a node explicitly
printed the mount's device rather than the pinned node's. It now reads the
node's own row.
Tests that built a node with device columns are ported to the link; where a
behaviour genuinely moved (properties refresh once on the row, last_scan is
the device's) the assertion moved with it rather than being dropped.
ha-relevant: yes
A device drawn on three canvases was three independent copies of the same
facts. Point every node at the Device Inventory row it draws, and let that
row own what the device *is* — the node keeps only how it is drawn.
- nodes.device_id -> device_inventory.id, ON DELETE SET NULL. NULL for
canvas furniture (group / groupRect / text), which describes nothing
physical. Deleting a node never deletes the row.
- services/inventory_sync holds the shared rules: matching by ieee > ip >
mac (per token, so 10.0.0.4 never matches 10.0.0.40), a property union on
key, a service union on (port, protocol, name), and the backfill that
links every pre-existing node.
- The backfill is non-destructive by construction: it writes device_id and
fills the row, and deletes nothing. Nodes are visited oldest-edit-first,
so where two canvases disagree on a scalar the most recently edited wins,
while properties and services stay unioned — nothing any canvas recorded
is lost. A second boot finds nothing to do.
- The wire shape does not change: GET /canvas hydrates the device fields
from the row, and a save routes them back to it. Editing a node's IP on
one canvas now shows on every other canvas holding that device.
- approve / bulk-approve set device_id instead of owning a copy, and a new
canvas node joins (or mints) its row. Rows minted this way are tagged
with a new `canvas` discovery source and get their own inventory filter.
- DetailPanel offers "Open in inventory"; standalone, which has no
inventory, is not offered it.
test_racks' "reports what the canvas node knows" seeded a second inventory
row for a host that already had one — a state a node create can no longer
produce. Its seeding order is swapped so the node links to the row; every
assertion is unchanged.
ha-relevant: yes
The Device Inventory was a read-only discovery log: rows arrived from a
scan or an import and went stale. Everything a user curates — properties,
services, notes, hardware, check method — could only be edited on a canvas
node, and never came back.
Give the inventory row the fields a node carries and a way to write them:
- device_inventory gains label, type, notes, cpu_count/cpu_model/ram_gb/
disk_gb, show_hardware, check_method/check_target, last_seen, last_scan,
response_time_ms and updated_at. Live reachability goes in a new
status_live: `status` already means the pending/approved/hidden lifecycle
and the two must not be conflated.
- PATCH /api/v1/scan/pending/{id} applies only what the client sends, so
editing one field never clears the rest. Lifecycle and discovery
bookkeeping stay owned by the approve/hide routes and the importers.
- InventoryDeviceCreate carries the same fields, so a hand-made entry needs
no create-then-PATCH round trip.
- InventoryDeviceModal becomes show *and* edit, reusing the canvas editors
(PropertyList, ServiceModal) rather than growing a second implementation.
- InventoryEntry moves to types/index.ts, its home; the modal re-exports it
so existing call sites are untouched. NODE_TYPE_GROUPS moves to
utils/nodeTypeGroups so both type pickers share one vocabulary.
ha-relevant: maybe
POST /api/v1/settings persisted the new status check interval and echoed it
back, but never rescheduled the running job: reschedule_status_checks() is
defined in app/core/scheduler.py and has no caller anywhere. The UI and
scan_config.json show the new value while the scheduler keeps firing at the
old one until the backend restarts.
On our install this went unnoticed for three weeks: 3600s was configured but
30s stayed in effect, so ~120 nodes were pinged 120x more often than intended
with no visible symptom.
Also add ge=10 to interval_seconds, mirroring the floor already enforced by
reschedule_status_checks(). Without it a value below 10 is written to
scan_config.json first and only then raises, which surfaces as a 500 and
leaves the invalid value to be reloaded on the next boot.
Both new tests fail without the fix: the first with AttributeError (the route
does not import the function), the second with 200 instead of 422.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The rack's link picker read `/nodes`, so it only ever offered devices
someone had already approved onto a logical canvas — two rows on a homelab
holding 74 inventory entries. A device on no canvas is still the record of
a real box, and is exactly what a rack is built out of.
`DevicePickerModal` replaces `NodePickerModal` and lists the Device
Inventory itself. Picking an entry calls the new `relinkDevice`, which
repoints the mount's `deviceId`, adopts that entry's node, status and —
unless the user renamed the plate — its label. One entry, one mount: a row
another plate stands for is not offered, and the store refuses it anyway.
The placeholder a rack-created plate left behind is dropped through the new
`DELETE /api/v1/scan/pending/{id}`, which refuses a device a rack still
mounts (409): foreign keys are off at runtime, so the mount would be left
naming a row that no longer exists.
`LinkedDevicePanel` becomes "Linked device" and now prints what discovery
found even when nothing on a canvas answers for the device; only the
canvas-side rows go missing, under a "Not on a logical canvas." note.
Also renames `pending_devices` to `device_inventory` (and
`pending_device_links` to `device_inventory_links`), with the Python and
TypeScript names that followed it. "Pending devices" was the scanner's word
for a queue of finds awaiting approval; the rows outlive approval, are
edited by hand and are what a rack mounts. Routes, payload keys and MCP
tool names are a published contract and are unchanged — `/scan/pending/*`
and the `pending_devices` key in `/stats` stay as they are.
The rename migration runs before `create_all`, or an empty new table would
be created beside the populated old one and every scanned device would read
as gone; it repairs that state too, for anyone whose app already started
mid-upgrade. Foreign keys are switched on for the rename so SQLite rewrites
the `REFERENCES` clause in `rack_devices`.
ha-relevant: maybe
The Edit Device modal left the column under the port list empty, while
the logical canvas already held every technical fact about the same box.
A mount that stands for a Device Inventory entry now prints them there:
canvas name, type, hostname, IP, MAC, OS, the status check the node runs,
the canvas it is drawn on, when it was last seen, and the services
discovery fingerprinted on it. Read-only — the logical view owns them.
`/racks/inventory` ships the node half as `node_*` alongside the
inventory row's own mac/hostname/os/services, and the panel prefers the
node value: the node is what the user curates, the inventory row is what
discovery last saw and goes stale after a rename or a DHCP move. Rows
neither side can fill are dropped, so a device on no canvas still shows
what was reported and an accessory shows no panel at all.
The new API fields are optional client-side, so an older backend reads
back as "not on a canvas" rather than an empty node.
ha-relevant: maybe
The number input's min/max are hints the browser does not enforce on a
typed value, and updateRack wrote the patch straight through. Two ways
out of a usable canvas:
- Shrinking below a mounted device left the plate drawn above the
chassis, outside the React Flow node, with nothing in the UI able to
drag it back. freeUnits only counts 1..uHeight, so the modal's "used
of" label under-counted it in silence.
- A height over 100 made RackSave reject every save with a 422, so both
the explicit Save and each autosave tick reported "Save failed" with
no cause.
updateRack now clamps to [MIN_RACK_U, MAX_RACK_U] and relocates the
mounts a shrink pushes past the top rail, one at a time so two never
land on the same slot. It returns false — changing nothing — when one
has nowhere to go, and the modal says so.
Two server-side guards that would have contained it: RackSaveRequest
cross-checks every mount against the rack it names, and RackDeviceSave
checks col_start + col_span against the grid, which each field passing
its own bounds never caught.
Also normalizes the MAC on POST /scan/pending. Every other write path
canonicalizes it and dedup compares by equality, so a hand-typed
AA-BB-CC-11-22-33 never matched the scanned aa:bb:cc:11:22:33 and
approve built a duplicate node.
ha-relevant: maybe
Four more defects from the branch review.
`POST /racks/save` fetched every row by primary key alone and then wrote
`design_id` from the payload, so a save on design B carrying an id that belongs
to design A moved A's rack — devices and cables with it — into B, with no error
and no way to notice. Ids come from the client, so a copied design or a stale
tab is enough. `_owned` now refuses a row owned by another design with a 409.
`commitEdit` dropped `applyFaceplate`'s return. Picking a plate the rack has no
room for, then shrinking the height by hand, let `updateDevice` succeed on its
own and `setPorts` write the new plate's ports onto the old faceplate — the
device ended up wearing a plate it did not have, silently. The failure is now
reported and nothing is committed.
Two smaller ones. The inventory select held a dead id after the list was
refetched, so submitting reported "No free slot in this rack" when the real
cause was an entry that had been racked elsewhere or deleted; it now falls back
to its placeholder and says so. And both create paths patched the form's own
geometry over the slot `findSlot` had just chosen, undoing a relocation the user
was never shown — the mount is handed the whole geometry up front, and only the
overrides are patched afterwards.
ha-relevant: maybe
Four defects found reviewing the last two commits.
Reverting a faceplate choice took every cable on the device down with it.
Browsing to another plate and back left `plateChanged` set with freshly seeded
port ids, while `commitEdit` skipped `applyFaceplate` because the id matched —
so `setPorts` received ids the device never had and dropped every cable
attached to the old ones, with no warning, since from the user's side the plate
had not changed. Coming back to the device's own plate now restores its ports,
ids included.
The inventory select preselects the first unracked entry, but only
`pickInventory` seeded the status from it, so submitting without touching the
select wrote `unknown` over the status the mount had just copied. Status now
seeds from that same entry.
Creating a device POSTs to the Device Inventory before the mount is attempted,
so a rack with no free slot stranded the row — and stranded another on every
retry. The rack is checked for a slot before anything is created.
And `_service_name` accepted a service with no port, where the inventory
requires one (`None not in _COMMON_PORTS` is True): a portless service named a
device in the rack picker but not in the Device Inventory, the exact mismatch
the naming fix set out to remove.
ha-relevant: maybe
The rack picker listed "192.168.1.63 · 192.168.1.63" where the inventory
showed "jellyfin", so a device the user picked out by one name turned up here
under another.
Two causes. The inventory route labelled an entry
`friendly_name or hostname or ip or id`, skipping the app-name step: `_device_label`
now mirrors `deviceLabel` in PendingDevicesModal — friendly name, host, the app
its ports say it runs (ssh/http/https ignored, generic web deprioritized), IP,
IEEE, id. And the option line appended the IP unconditionally, doubling it
whenever the label already was the IP; it now adds the IP only when it says
something new, plus the type as the tiebreaker between look-alikes.
Also fixes the port row: the name field and the type select both carried
`w-full` from the shared input class, so the select took the row and the name
field collapsed into an unlabelled box.
ha-relevant: maybe
The rack canvas loses its right rail: a mount is now edited in its own dialog,
like a network node, and the canvas keeps the full width.
Devices:
- `RackDeviceModal` adds and edits a mount — label, faceplate, U position,
height, column, width, status, colour, ports, unmount. Adding offers three
sources: an existing Device Inventory entry, a new device created here, or a
rack-only accessory. Opened by `+ Device` in the left rail or a double-click
on a plate.
- `RackSettingsModal` takes over rack chrome (name, location, capacity, 19"/10",
numbering, colours, delete), on a double-click of the chassis.
- The left rail's `+ Rack` becomes `+ Device`; `+ Rack` stays in the header.
Inventory:
- Gear created from a rack canvas lands in the Device Inventory with
`discovery_source: "rack"`, so it inherits the existing search, filters, hide
and delete instead of a parallel list in the sidebar. New "Rack devices"
source filter and badge.
- It is never placed on a logical canvas: approve returns 409, bulk-approve
skips it with `match: "rack"`, and the UI drops the Approve button and reports
the skip.
- `InventoryTray` becomes `AccessoryTray` — blanks, shelves and cable managers
only, still a drag source.
Fixes:
- A taller faceplate (or a hand-typed height) no longer silently keeps the old
geometry when it collides: the device relocates to the nearest slot that takes
the new size, and only a rack with no such slot refuses the edit. This is why
a 1U and a 2U plate rendered at the same height and the field looked locked.
- The empty-state buttons sat under `.react-flow__renderer` (z-index 4), so the
pane swallowed their clicks as a canvas drag.
Tests: rack device/settings modals, empty-state stacking, relocate-vs-refuse in
the store, rack source bucketing, inventory approve gating (frontend + backend).
ha-relevant: maybe
The modal-based editors and the rack store are portable; the inventory tagging
rides on the standalone FastAPI scan routes. Discuss before porting.
The rack prototype lived behind /racklab with no backend. It now becomes a
third design_type alongside network and electrical, sharing the app shell:
same header, same left rail, same right panel, same design switcher. App
branches the renderer on design_type.
Backend
- racks / rack_devices / rack_cables, all design_id-scoped.
- rack_devices.device_id -> pending_devices is the primary link. Those rows
survive approval and node deletion, so unracking never removes the device
from the Device Inventory. node_id -> nodes is a second, optional link used
only for live status and for matching network-link imports. Both are
ON DELETE SET NULL and label is denormalized, so a rack keeps rendering
after an inventory purge.
- GET /racks, POST /racks/save (full-state upsert + prune, same contract as
canvas/save), GET /racks/inventory.
- POST /scan/pending adds an inventory entry by hand, for hardware no scan
can discover.
- design_type is now validated and actually branched on; the schema comment
calling it vestigial is gone.
- Design delete unwinds rack rows explicitly (SQLite does not always honour
the cascade); design copy duplicates racks, mounts and patches with fresh
ids and re-points the cables at the copies.
Frontend
- /racklab and RackLab are removed; no dual maintenance.
- Rack types lifted into @/types; rackSerializer maps the snake_case wire
shapes to them, narrowing every enum on the way in.
- The store gains load/save, dirty tracking and an editSeq, so racks get the
same explicit Save and the same opt-in autosave as the logical canvas.
- Header, sidebar and inspector adapt in rack mode: canvas-only actions are
hidden, Add Rack / Patch / cable visibility / type filter / Import links
take their place, the sidebar body becomes the inventory tray, and the
footer counts racks, mounts, cables and free U.
- rackTheme derives the rack palette from the active app theme rather than
declaring one per theme, so a new theme works here for free. Per-rack
chrome stays user-editable; only its default comes from the theme.
- New Canvas gains a Kind picker (hidden when copying, since a copy inherits
the source kind).
Tests: 24 backend, 5 frontend files (store, persistence, networkLinks,
rackTheme, rackSerializer) plus DesignModal coverage for the Kind picker.
ha-relevant: maybe
Frontend rack renderer is portable; the FastAPI routes and SQLAlchemy models
are standalone-only by the persistence/transport split. Discuss before porting.
Add a label substring filter to GET /api/v1/nodes, plus MCP tools
list_node_summaries (id/label/type/status only) and get_node
(by id -> single object, by label -> array) so an LLM client can look
up nodes without paying the token cost of the full node payload.
Cleared canvas re-showed the demo, and backend errors silently fell back to
demo — hiding real outages and forcing users to wipe demo nodes before bulk
edits. Now:
- backend load reports `initialized` (CanvasState row exists = ever saved)
- decideCanvasLoad() picks real | empty | demo; empty is kept empty
- backend down/error shows an error banner + toast, never the demo
- data-new-user flag reserved as the Getting Started walkthrough hook
ha-relevant: yes
Mirror the Proxmox auto-sync pattern for the Zigbee2MQTT and Z-Wave JS
UI mesh imports. Connection config + MQTT credentials live in .env only,
are never persisted to scan_config.json, and are never returned by any
API or shown in the UI (single source of truth).
- config: ZIGBEE_* / ZWAVE_* env settings; only sync_enabled+interval
are persisted, connection/credentials stay env-only
- routes: GET/POST /config, POST /sync-now; auto-sync reuses the exact
manual _background_*_import + _persist_pending_import path (fresh
import when empty, update-in-place when nodes exist, ScanRun trace)
- scheduler: zigbee_sync / zwave_sync jobs with live enable + reschedule
- frontend: reusable MeshAutoSync section in Settings (Zigbee, Z-Wave)
- .env.example: documented both blocks
- tests: scheduler jobs, router config/sync-now/auth, credential-never-
persisted, SettingsModal sections
Manual Zigbee/Z-Wave import behaviour is unchanged.
ha-relevant: no
Add backend API tests for the two issue #265 features that shipped
untested: auto-positioning root nodes into a free grid slot (first at
origin, collision avoidance, child origin default, explicit coords and
explicit zero preserved) and auto-assigning edge handles from absolute
canvas Y (source above/below/equal target, partial handle fill, child
abs-Y resolved through parent).
Drop the max(0, round()) clamp in _find_free_position: it folded
negative-positioned nodes onto cell (0,0), falsely blocking or freeing
the origin slot. Negative nodes now keep their true cells and never
intersect the positive search space.
Node.ip stores several comma-separated addresses once a user adds an
IPv6 (e.g. "fe80::1, 192.168.1.5"). All placement/inventory matching did
exact string equality on that field, so a device scanned as the plain
IPv4 looked absent from the canvas: canvas_count stayed 0, the "In N
canvas" badge and hide-on-canvas button vanished, and bulk-approve
re-placed a duplicate node.
Match per-address instead, and add MAC (a stable identifier immune to IP
edits) as a cumulative match key alongside ieee/ip:
- _canvas_correlation indexes ip per token + by_mac
- bulk_approve skip detection tokenizes ip + tracks mac
- find_duplicate_node narrows with Node.ip.contains then confirms per
token in Python (guards the 10.0.0.4 / 10.0.0.40 substring false match)
ha-relevant: maybe
Single-device approve now guards duplicates per-design the same way bulk
approve does, and asks the user instead of failing or silently merging:
- create_node and approve_device reject a same-design duplicate (ieee, ip
or mac) with 409 + the existing node; a force flag creates it anyway.
- Frontend shows a confirm dialog: go to existing node, add duplicate
anyway, or cancel.
- approve_device no longer rejects a device already on another canvas
(status is global, canvas membership is per-design) — it can be placed
on a new design, matching bulk approve.
- IEEE (Zigbee/Z-Wave) devices now use the same prompt as ip/mac instead
of auto-merging into the existing node.
- bulk approve reports which devices it skipped as duplicates.
Closes#260
ha-relevant: maybe
Add a 'Re-sync now' button to the Proxmox auto-sync settings section that
triggers an immediate inventory import (POST /proxmox/sync-now) using the
server env config — the manual counterpart to scheduled auto-sync.
Also fix a dual-source-of-truth bug: Proxmox connection config (host, port,
token, verify_tls) is now env-only and never persisted to scan_config.json.
Previously save_overrides() dumped host/port/verify alongside scan settings,
so saving an unrelated setting wrote an empty proxmox_host that load_overrides
then clobbered PROXMOX_HOST with on every boot. Only the auto-sync activation
(sync_enabled + sync_interval) stays user-editable and persisted.
ha-relevant: no
Add a 'Copy from existing' option to the New Canvas modal. It lists every
canvas with node/group/text counts; picking one deep-copies its nodes, edges,
parent/child links and canvas state (viewport, custom style, floor plan) into
a fresh design.
Backend: POST /designs/{source_id}/copy remaps node ids, re-points edges and
parent links, and clones canvas state; GET /designs now returns per-design
counts for the picker. Standalone mode clones the localStorage canvas.
Closes#216
ha-relevant: maybe
Match the validated filename against iterdir() entries with == instead of
building a path from the user string, so no tainted value ever reaches a
filesystem sink. Satisfies CodeQL py/path-injection.
ha-relevant: no
Resolve media filenames through a shared _resolve_media_path() barrier that
confirms the resolved path sits directly under the resolved media dir, so
CodeQL can trace the sanitization the regex already guaranteed.
ha-relevant: no
Approving the same zigbee/zwave/proxmox devices onto a second design
placed the nodes but drew no edges. _resolve_pending_links_for_ieee
deleted each pending_device_link after materializing its edge, so the
first approve consumed the whole topology and later approves had nothing
to resolve.
Links are topology, not one-shot: every importer wipes+reinserts its
link set on each import, so they can safely persist across approvals.
- keep the link rows (drop both db.delete(link) calls)
- scope resolution to the target design (Node.ieee_address + design_id)
so a re-approve links that canvas's nodes, not another's
- thread design_id through all three approve call sites
Existing links deleted by the old code do not come back on their own;
a re-import repopulates them.
ha-relevant: maybe
Reconcile the same physical device discovered by both the nmap IP scan and
the Proxmox importer into a single inventory row, keyed on MAC. Previously
each path only deduped by IP, and the importer captured no MAC (and no IP for
stopped guests), so most guests double-listed.
Backend:
- mac_utils.normalize_mac: canonical MAC (lowercase, ':'-separated), the
cross-source join key. Normalized on write and on compare.
- proxmox_service: capture the guest NIC MAC agent-free from the net0 config
(qemu virtio=<MAC>, lxc hwaddr=<MAC>); works for stopped guests. Resolver
now returns (ip, mac).
- proxmox persist: match existing Node/PendingDevice by ieee OR ip OR MAC;
fill mac, keep the vm/lxc type, union sources.
- scanner persist: match PendingDevice by ip OR MAC; fill the IP a Proxmox
import lacked, keep a pve row's type, union the scan source. Stamp query
matches raw + normalized MAC (legacy-safe).
- Multi-source tags: new PendingDevice.discovery_sources JSON column so a
merged device shows under both the IP and Proxmox filters. Idempotent
migration backfills from discovery_source (legacy NULL-scalar rows with an
IP become ["arp"]). _sources_after_merge preserves a scanned row's IP origin
through the merge without tagging a pure Proxmox guest.
- Import now broadcasts a scan update on completion so an open inventory
reloads without a manual refresh.
Frontend:
- pendingSources: sourceBuckets/orderedSources map discovery_sources to filter
buckets; a device with ["arp","proxmox"] matches both filters and renders
both badges. PendingDevicesModal filter + badges use them.
Tests: MAC normalization, config MAC capture, cross-source merge both
directions, legacy-row IP-tag preservation, no-false-IP-tag guard, refresh
broadcast, and the frontend bucket mapping.
ha-relevant: maybe
Cluster edges created via the pending -> approve path rendered on the top
handle instead of left/right, because the edge and its endpoints lost their
handle information on the way to the canvas.
- Approve resolver (scan.py) now returns each edge's type + source/target
handle. Handle IDs are the bare slot-0 side names ('right'/'left'), the
canonical stored form React Flow resolves to the correct side; a '-t' target
id fails to resolve and falls back to the top handle.
- Frontend injectAutoEdges no longer hardcodes iot/bottom/top-t. It injects
each edge with its real type + handles and bumps the referenced nodes'
left/right handle counts (which default to 0) so the cluster endpoints exist.
Logic extracted to a pure, tested util (applyAutoEdges).
- clusterEdges direct-import path uses the bare 'left' target to match.
Tests: new autoEdges unit tests; updated backend handle assertions.
ha-relevant: maybe
Scan History:
- Proxmox is now a first-class scan kind (badge, filter chip, Server icon,
completion toast) instead of being mislabeled as an IP scan.
- A done run carrying a non-fatal advisory renders amber (info) with a
warning toast, distinct from red failures.
Import diagnostics:
- test-connection probes /access/permissions and warns when the API token
has no ACL (VMs/LXC would be invisible) — points at the PVEAuditor grant.
- import surfaces an advisory when hosts import but no guests are visible
(privilege-separated token whose rights are the intersection with the user),
rather than a silent "done".
Node style:
- Proxmox container mode is now opt-in (container_mode === true), matching the
rest of the codebase (App.tsx nesting logic). Imported proxmox nodes leave
the flag unset and render like a manually-created node instead of an empty
group container.
Cluster edges:
- Hosts from one import are chained with 'cluster' edges via left/right handles,
distinct from the vertical host->guest 'virtual' edges. Wired for both the
direct "Add to Canvas" path and the pending -> approve path (host<->host
proxmox_cluster links, resolved to cluster edges on approve; cluster hosts
get left/right handles).
Tests added on both sides.
ha-relevant: maybe
Add a Proxmox VE importer that reads the /api2/json REST API with a read-only
API token and drops hosts (proxmox), VMs (vm) and LXC containers (lxc) onto the
canvas as typed nodes with run state and hardware specs (vCPU/RAM/disk).
- Backend: proxmox_service (httpx) + proxmox routes (test-connection, import,
import-pending, config). Two-tier dedupe — merge onto an existing scanned node
by IP, else synthetic pve-{host}-{vmid} identity. Update-in-place, never
deletes. Host->guest rendered as a 'virtual' edge via the pending-link flow.
- Security: token is env-only (PROXMOX_TOKEN_*), never written to disk by the
app, never returned by any endpoint; errors are credential-sanitized.
- Auto-sync: optional scheduled re-import into pending (APScheduler job).
- PendingDevice.properties carries specs through approve (+ migration).
- Frontend: ProxmoxImportModal, sidebar entry, pending inventory source filter,
Settings auto-sync section, proxmoxApi client.
- Docs: docs/proxmox-import.md, README + FEATURES sections, .env.example keys.
- Tests: backend service/router/scheduler, frontend modal/client/pending.
ha-relevant: maybe
A device already on a canvas (Node exists) but with no pending_devices row
never showed in the discovery inventory — the inventory lists pending_devices,
not nodes. This stranded legacy auto-placed coordinators: on a canvas yet
invisible in the inventory list.
Both mesh imports now, in the already-approved-node branch, ensure an
inventory row exists: create one as status="approved" when missing, or
refresh its metadata (preserving status) when present. New/unplaced devices
still land as status="pending"; hidden rows stay hidden.
Tests updated: approved-node path now backfills an approved inventory row
(zigbee + zwave), plus a regression that a hidden inventory row is not
revived.
ha-relevant: yes
The coordinator was special-cased in both mesh imports to auto-create a
canvas Node, so it never appeared in the pending inventory and users could
not approve/hide/type it like every other device.
Remove the auto-placement: the coordinator now flows through the shared
pending path (upsert into pending_devices with suggested_type
zigbee_coordinator / zwave_coordinator). An already-approved coordinator
Node still gets its properties refreshed on re-import via the shared
approved-node path. The response's coordinator/coordinator_already_existed
fields are retained (now always unset) for backward-compatible shape; the
frontend already ignored them.
Tests updated: coordinator lands in pending (counts include it), no Node
auto-created, pending metadata carried, approved coordinator refresh + no
pending re-list.
ha-relevant: yes
Zigbee and Z-Wave imports looked up canvas nodes by ieee_address with
scalar_one_or_none(), assuming one node per IEEE globally. A device placed
on two designs (one Node per canvas — a supported feature) made re-import
crash with MultipleResultsFound.
- Both mesh imports now refresh properties on every matching node instead
of a single row (loop over .scalars().all()).
- approve_device guards against a true duplicate: same IEEE already on the
SAME design reuses that node instead of inserting a second one.
- New node_dedupe service: loss-free repair keyed on (ieee, design_id).
Collapses only genuine same-canvas duplicates (merges properties/services/
missing fields, re-points edges + parent_id, drops self-loops/parallel
edges). Cross-design placements preserved. Runs at start of both imports
and bulk-approve. No-op on healthy DBs.
- IP correlation path already handled multiple nodes; left unchanged.
Tests: dedupe unit tests (collapse, cross-design preservation, edge/parent
re-point, idempotent), zigbee + zwave multi-canvas regression, approve
no-dupe guard.
ha-relevant: yes
- Render floor plan inside React Flow ViewportPortal so it pans/zooms with
nodes (was screen-fixed, desynced on pan/zoom); zoom-stable resize handles.
- Move floor plan config from the left panel into the canvas (design) edit
modal; attach per-design and fix cross-design bleed on load.
- Store images via a new generic backend media endpoint (POST/GET/DELETE
/api/v1/media) on disk under <data_dir>/uploads, not base64 in the canvas.
- Disable floor plans in standalone mode (no backend to upload/serve); drop
base64 localStorage persistence. See ADR-001 in CLAUDE.md.
- Tests: backend media route, DesignModal floor plan + upload, store floorMap.
ha-relevant: maybe
create_node/create_edge persisted rows with design_id=null when the
client omitted it (the MCP write tools), so they existed in the DB but
never rendered on the canvas until a container restart reconciled them.
Both routes now fall back to the first design, matching bulk-approve.
Also fix MCP resource reads (homelable://canvas, homelable://edges):
the framework passes a pydantic AnyUrl, not a str, which raised
"'AnyUrl' object has no attribute 'startswith'". Coerce to str.
EdgeResponse now exposes design_id for symmetry with NodeResponse.
ha-relevant: no
Stop button had no effect: run_scan only checked the cancel flag between
CIDR ranges and between hosts, never inside the per-range nmap call. For a
single /24 the whole scan is one blocking call, so cancel was ignored for
minutes and the run status stayed 'running'.
- thread run_id into _nmap_scan/_ping_sweep/_nmap_port_scan; check cancel
before each phase and skip queued hosts once cancelled
- flip ScanRun status to 'cancelled' eagerly in the stop endpoint so the UI
reacts immediately instead of waiting for a checkpoint
Fixes#218
ha-relevant: yes
Bulk-approve filtered status=='pending', so a device already approved onto
another canvas (status is global, canvas membership is per-design) — or whose
node was later deleted — was silently skipped. Selecting 64 devices on an
empty canvas produced only the ~28 still pending.
Approve now places a node on the target design for any selected, non-hidden
device that isn't already on that design (deduped by ip/ieee_address, including
within the batch). Returned device_ids/node_ids stay index-aligned so the
client places them all.
ha-relevant: yes
Surface the same lifecycle timestamps on the Device Inventory cards as on the
detail panel. Tiles for devices placed on a canvas show their linked node's
created / last scan / last modified / last seen (correlated by ip or
ieee_address, aggregated across matches: created = oldest, others = newest).
Devices not yet on any canvas fall back to their discovered_at.
Rendered as compact relative times ("2d ago") with the full date on hover, in
a tight two-column footer so the tile keeps its original footprint.
Backend: PendingDeviceResponse gains node_created_at / node_last_scan /
node_last_modified / node_last_seen; the canvas correlation now also pulls node
timestamps in the same single query. Frontend: shared timeFormat util
(absolute + relative), reused by the detail panel.
ha-relevant: maybe
Approve (single + bulk) created the canvas Node under the first design
instead of the design the user is viewing, so approved devices were
invisible on the active canvas and got wiped on the next save — and a
re-approve returned "0 approved" because the rows were already approved.
- bulk-approve now accepts design_id; the UI sends the active design.
- single approve already honoured design_id; the UI now sends it too.
- generalize the wireless branch (status=online, mesh props, no ICMP
check) to Z-Wave as well as Zigbee, using build_zwave_properties.
ha-relevant: yes
Import a Z-Wave JS UI (zwavejs2mqtt) network over the MQTT gateway API,
mirroring the existing Zigbee pipeline:
- New Z-Wave Import modal + sidebar entry (broker, prefix, gateway name)
- coordinator/router/end-device typing with mesh tree from node neighbors
- import to Pending section or straight to canvas
- Pending Devices gains a Z-Wave source filter
- shared mqtt_common helpers extracted from the zigbee service
ha-relevant: yes
Reworks the "Pending Devices" panel into a "Device Inventory": scanned
devices already placed on a canvas are no longer suppressed — they stay
listed and badged with how many canvases they appear on.
- scanner: stop deleting/skipping on-canvas IPs (hidden still suppressed)
- scan API: /pending returns all non-hidden devices; compute canvas_count
by correlating ip/ieee_address against nodes grouped by design
- frontend: rename to "Device Inventory", top-right canvas-count corner,
toggle to show/hide on-canvas devices (default show)
ha-relevant: maybe
Adds scanner_http_ranges / scanner_http_probe_enabled / scanner_http_verify_tls
to Settings (persisted in scan_config.json, Options page defaults). /scan/trigger
accepts an optional body to override these per-scan; /scan/config GET/POST read
and persist the defaults. Port ranges validated at the API boundary.
ha-relevant: yes
The status WebSocket pool removed connections with list.remove(), which
raises ValueError on a double-remove (broadcast already dropped a dead
socket, then disconnect tries again), and only released a slot on
WebSocketDisconnect — any other error leaked the socket into the
broadcast pool. Centralise removal in an idempotent _drop() called from
a finally block and from _broadcast.
ha-relevant: yes