Commit Graph
900 Commits
Author SHA1 Message Date
Pouzor 4e8d841807 feat(deploy): serve Homelable under a configurable base path
Homelable could only own the root of an origin. Behind an existing proxy
at `https://home.example/homelab/` the built `index.html` still asked for
`/assets/*`, which fell through to whatever owned the root — and when
that answered `text/html` for a `<script>` under `nosniff`, the browser
failed the load on an HTTP 200. The only workaround was patching the
checkout and rebuilding on every upgrade.

One build-time knob, `VITE_BASE_PATH`, becomes Vite's `base`. Vite
rewrites the asset URLs it emits; everything the app builds by hand goes
through the new `utils/basePath.ts` — the axios instances, the WebSocket
URL, the live-view route, the local brand icons, and the OIDC login href.
`resolveServerPath` covers what the *backend* hands back, which is always
root-absolute because it cannot know where the SPA is mounted: uploaded
floor-plan URLs already stored in a canvas are resolved at render time,
so plans predating the move keep loading. The OIDC callback used to
redirect to `/`, dropping subpath users at the origin root after login;
it now reads the prefix back out of `OIDC_REDIRECT_URI`.

The default is `/`, and stays a no-op there by construction: every helper
returns the string it returned before, the root build output is unchanged
and both nginx sites are byte-identical to what shipped — the Docker
image copies `docker/nginx.conf` verbatim and the installer keeps its
original heredoc. Only a non-default prefix takes the generated config.

Those generated configs use `root`, never `alias`, since `alias` plus
`try_files` mis-resolves `$uri` — the Docker build lands the bundle in
the matching subdirectory, and the installer symlinks it under
`/var/www/homelable`. Both accept either reverse-proxy style, prefix
forwarded intact or already stripped, with no redirect loop between them,
and `absolute_redirect off` stops the no-slash 301 from eating the port.

Closes #334.

ha-relevant: maybe
2026-09-04 11:23:16 +02:00
Pouzor 3cbda60277 feat(icons): add the Homelable logo to the brand icon picker
The brand catalog is the generated homarr-labs dashboard-icons manifest,
which carries no Homelable mark, so a user cannot label the app itself on
a canvas. Serve a small set of logos from `public/brand/` instead: the
picker prepends `LOCAL_BRAND_ICONS` to the manifest and `brandIconUrl`
routes those slugs at the local asset rather than the jsDelivr CDN.

`dashboardIcons.json` stays untouched — `scripts/fetch-dashboard-icons.mjs`
overwrites it wholesale.

Closes part of #397 (the Docker host polling request is untouched).

ha-relevant: yes
2026-09-04 10:21:32 +02:00
Pouzor 03e5ae939e ci(security): give pip-audit the same registry-outage retry as npm
pip-audit queries PyPI's advisory API and exits non-zero when it cannot
reach it, exactly as `npm audit` does — the same class of outage that
reddened #409 would have failed the Python half too.

Both now go through .github/scripts/audit-with-retry.sh, which retries
only when the output names a transport or availability failure and lets
a real finding fail on the first attempt, unretried. The npm loop added
in the previous commit is folded into it.

ha-relevant: no
2026-09-04 10:01:35 +02:00
Pouzor 37a3aee9cc ci(security): retry the npm audit when the registry endpoint is down
`npm audit` exits 1 both for a real advisory and for a registry that
refuses to answer, so a 503 from the audit endpoint reddened PRs that
changed no dependency at all — as it did on #409.

Retry up to three times, but only when the output names a transport or
availability failure. An advisory at or above high severity still fails
on the first attempt, unretried.

ha-relevant: no
2026-09-04 10:01:35 +02:00
Pouzor 37f893abe6 fix(rack): keep the saved pan/zoom when switching back to a rack canvas
Leaving a rack design for another canvas and coming back put the racks
at the wrong pan/zoom until a hard reload. Two causes, both fixed here.

The rack flow shared the App-level `ReactFlowProvider` with the logical
canvas, so that canvas' transform and pane size carried over into it.
`RackCanvas` now wraps its own provider.

The restore effect was also spent before it could work: a switch back
remounts the component before `App` calls `loadDesign`, so the first
pass applied the pre-load viewport and then locked itself out for that
design. The marker now clears whenever a load starts, and the apply
waits a frame so React Flow has measured its pane.

Closes #408

ha-relevant: no
2026-09-04 10:01:35 +02:00
Pouzor f51dd131c1 refactor(rack): tidy the device modal's last row and its actions
Status and the colour override shared a column each, one under the other,
while the pair fits on one row like the placement fields above them. The
footer sat flush against the last section, so Save read as part of it, and
a section rule carrying the faceplate button sat lower than the plain rule
beside it.

- Status and Colour override share a two-column grid.
- A rule separates the form from Unmount / Cancel / Save.
- SectionHeader is a fixed-height row, so an `aside` button no longer
  pushes its border down.

ha-relevant: yes
2026-09-04 01:40:34 +02:00
Pouzor 4b90c60cc6 refactor(rack): put the plate picker on the faceplate rule
The Faceplate field sat in the Information column, a two-line block whose only
job was to open the catalog — a long way from the plate it changes, and the
last thing left between the label and the placement fields.

It is now a small button on the right end of the Faceplate rule, above the
drawing it affects: the current plate names itself, and `Change` says the name
is a button rather than a read-out. The name truncates first, so a long one
("Switch 24 ports + 2 SFP") never pushes the verb off. The warning about a plate
swap dropping its ports moved with it, between the rule and the plate.

ha-relevant: yes
2026-09-04 01:40:34 +02:00
Pouzor b26e9d3137 refactor(rack): rebuild the device modal around a headed strip
The plate used to sit alone above the form at a fixed 660px, and what the
mount actually stands for was buried under the port list at the bottom of the
right column — the two things the user checks first, at opposite ends of the
dialog.

They now share a strip across the top: the linked device on the left, the plate
on the right. The plate fills the half it is given instead of a constant, which
is what made port placement so fiddly — a full-width switch drew at two thirds
of the space with sockets to match. `PortPositionEditor` measures its container
and derives the scale from that; `fullWidth` stays as an explicit override, and
both fall back to the old default where nothing can be measured.

The form is split by the same section rules the logical canvas' `NodeModal`
draws — Linked device, Faceplate, Information, Placement, Ports — so the two
editors read the same way. `SectionHeader` moves to its own component and takes
an `aside`, which is where the mount status now rides. An accessory has no
ports and no inventory row, so it gets neither heading.

Port rows go from two columns to three; a 24-port switch no longer runs past
the fold.

ha-relevant: yes
2026-09-04 01:40:34 +02:00
Pouzor 4ccad61e73 feat(rack): let a mount say when it shows its ports, and never end a cable on hidden metal
Two things about when a faceplate draws its sockets.

A mount now carries `portVisibility`: `auto` leaves the faceplate in charge —
switches and patch panels permanently, everything else on focus — while
`always` and `hover` override that either way. The selector sits in the port
section of the device modal. It is a canvas display choice, so it lives on the
rack mount and never reaches the inventory row the mount points at; `auto` is
the default, so nothing about existing racks changes.

The fix underneath: with the cabling overlay on `hover`, hovering a switch drew
its runs, but the far end landed on a plate that was not itself focused and so
drew no socket — a cable stopping on blank metal. A port a *drawn* cable ends
on is now always rendered, on both plates, whatever the plate rule or the mode
says. `cableVisibility.ts` owns the question of which cables are on screen and
`CableLayer` draws from the same answer, so the two can no longer disagree.

Fixes #403

ha-relevant: yes
2026-09-04 01:40:34 +02:00
Pouzor 6fa1b1a3d0 fix(scan): never send an HTTP GET to a raw printing port
A printer treats whatever bytes land on TCP/9100 as a print job, so our
probe's "GET / HTTP/1.1" came out of the tray as a page of plaintext.

Two paths reached it, and neither could be configured away — 9100 is
hard-coded in the scanner's port list, and user ranges only ever add to
that list:

- the deep-scan HTTP probe, once per scan;
- the status checker, every 60 s for as long as the node exists, which is
  the one that turns a scan annoyance into a ream of paper.

Both already had a non-HTTP port denylist; 515 and 9100-9107 were simply
missing from it. nmap ships `Exclude T:9100-9107` for this exact reason,
which is why the -sV pass never triggered it and only our own probes did.

Node Exporter also lives on 9100 and loses its status check as a result.
A grey dot is the cheaper mistake. Its fingerprint entry is port-match
only, so discovery and labelling are unaffected.

Fixes #404

ha-relevant: yes
2026-09-03 23:15:36 +02:00
Pouzor caaa408618 fix(rack): stop a clamped size from shrinking the device everywhere
The write-through was asymmetric. A load applies the inventory row's
height and width to a mount only where they still fit that rack, so a
mount can legitimately be drawn smaller than the row says. A save then
copied the mount's size back onto the row unconditionally — so opening a
rack where a 4U device sits too high, changing anything at all and
saving turned that device into a 1U everywhere, silently. The rack store
did the same in the live session.

Size now travels only when the save actually changed it; an unchanged
size is an echo of the clamp and is dropped. The plate, its colour and
its ports keep going through unconditionally: none of them can overflow
a rail.

The other half of the overlay had no collision check, so a model bigger
than the mount could be applied on top of the neighbour above it —
drawing two plates on the same U, a state the canvas itself refuses and
`RackSaveRequest` happily stored. The overlay now obeys the placement
rule too, and the save endpoint rejects overlapping mounts in one rack
(devices in different racks, and half-width plates side by side on one
U, still pass).

ha-relevant: no
2026-09-03 20:58:20 +02:00
Pouzor bffe1927dd refactor(rack): lead with the faceplate, and make the port list a chip grid
Three things the port editor got wrong on a real switch:

The plate sat under the form, so the thing being edited was the thing you
had to scroll to. It leads now, in both the rack device modal and the
inventory one.

The arrow keys only moved a port while its handle held focus. Selecting a
port in the list and reaching for them — the obvious gesture — did
nothing. Placement mode now listens for as long as it is on and moves the
selected port, stepping aside for a text field being typed in.

And the list drew one full-height row per port: a 24-port switch ran a
metre down the modal and buried the panels under it. Each port is now a
28px chip (borderless name, bare type select, remove) in a two-column
grid, so the same switch is a list rather than a page.

ha-relevant: no
2026-09-03 20:58:20 +02:00
Pouzor f933f774fc feat(inventory): draw and edit a device's rack faceplate from its inventory row
The row already owns the front panel; until now only a rack canvas could
show it. The device detail modal draws the plate full size for any device
some rack has modelled, and its edit mode carries the whole model —
faceplate, height, width, colour override and the port list, placement
mode included — writing back through `PATCH /scan/pending/{id}`. A device
therefore no longer has to be racked, or racked again, to have its ports
named and placed.

The port list is now one component (`PortListEditor`) instead of a copy
in each modal, since both edit the same list on the same row.

A device with no modelisation gets no section and no `rack_*` fields in
the PATCH: editing an ordinary host must not silently give it a plate.
Creating a model from the inventory is deliberately left out — a plate is
still born on a rack.

ha-relevant: no
2026-09-03 20:58:20 +02:00
Pouzor b658a68a09 feat(rack): let the inventory own the front panel, and place ports by hand
The Device Inventory row now owns the rack modelisation — faceplate, U
height, column span, colour override and the port list, positions
included. A device therefore wears the same front panel in every rack it
is mounted in; only its placement, pinned status and label stay per
canvas. `POST /racks/save` writes it through, `GET /racks` overlays it on
load, and size is applied only where it still fits the mount's own rack:
geometry is global, a rack is not.

Ports were also unplaceable. Adding one dropped it on the middle of the
plate, so three added ports drew as one socket, and the only preview was
a 160×18px thumbnail that lost its label and stacked its ports. The
modal is wider, the plate is drawn full size below the form, and a
Position mode turns it into the port editor: drag to place, arrow keys
to nudge, per-axis snap to the nearest peer so banks stay in line.

Accessories stand for no inventory row, so they carry no ports and get no
port UI. Standalone has nothing to hang a global panel on, so ports stay
per mount there.

ha-relevant: no
2026-09-03 20:58:20 +02:00
Pouzor 6e2ccf7b7f fix(canvas): make the basic edge animation direction independent of geometry
The basic dash flipped to `animation-direction: reverse` whenever the source
sat below the target on screen, so the same edge marched one way or the other
depending only on where its nodes had been dragged.

The flip was compensating for the keyframes: homelable-basic-dash decreased
stroke-dashoffset while homelable-snake and homelable-flow increase it, so the
basic dash ran against the other two modes. Flip the keyframes to match and
drop the geometry test.

Does not address the underlying handle bug: the invisible oversized target
handle in SideHandles wins the pointerdown, so a drag from A to B is stored as
source=B, target=A. That is why all three modes now march target to source.

Closes #395

ha-relevant: yes
2026-09-03 02:46:58 +02:00
Pouzor fb39663293 refactor(inventory): rename the Approve action to "Put in current Canvas"
"Approve" described a moderation step the inventory no longer really has —
a row is not judged, it is placed on the canvas the user is looking at. The
new label says what the button does and which canvas it targets.

Renames the button in the device detail modal and the bulk button in the
inventory grid, plus the success toasts and the walkthrough step that spoke
of "approved devices". The API paths, the `status` value and every handler
name keep the word: it is a published contract the MCP server also calls.

ha-relevant: no
2026-09-02 21:14:44 +02:00
Pouzor 54176983b6 fix(proxmox): no host link for a guest nested in container mode
A nested guest was still getting the host -> guest 'virtual' edge, which
loops out of the container box and back into it — the screenshot on #399.
Being inside the box already says the host runs it.

Both entry points drop the link only for guests they actually nest: the
import modal path skips the edge in App's handler, and the inventory path
filters the host <-> guest pairs out of the edges bulk-approve returns.
Linked mode, loose guests and cluster edges are untouched.

ha-relevant: maybe
2026-09-02 21:14:44 +02:00
Pouzor deff7a34d0 feat(proxmox): optionally nest guests inside their host on the canvas
Two entry points now offer container mode, where a Proxmox host is drawn as
a box holding its VMs and LXCs instead of sitting beside them.

Import modal: an opt-in "Nest guests inside their host" checkbox, shown only
in "Inventory + canvas" mode. `onAddToCanvas` carries the choice as a
`ProxmoxCanvasMode` ('linked' | 'container').

Device Inventory: approving a Proxmox host opens a new ProxmoxApproveModal
asking whether its guests come along, and nested or linked. The guests come
from a new `GET /scan/pending/{id}/proxmox-children`, which resolves the
host -> guest `device_inventory_links` the Proxmox import records (cluster
links excluded, hidden rows skipped). A host with no guests, or a failed
lookup, falls through to the plain single approve.

Placement lives in the new `utils/proxmoxContainerLayout`: `groupProxmoxGuests`
splits a selection into hosts, guests and loose nodes (one parent per guest),
and `layoutProxmoxContainers` grids the guests inside their host box, hosts
side by side, anything left over below. Positions are absolute — `addNode`
turns them relative when it nests. A host skipped as a same-canvas duplicate
leaves its guests loose rather than orphaning a parent_id.

Also extracts `inventoryNodeData()` so the single, bulk and Proxmox approve
paths build node data from one place.

ha-relevant: maybe
2026-09-02 21:14:44 +02:00
Pouzor e988427bb2 docs: explain why Docker bridge networking yields no MAC addresses
The default compose files put the backend on a bridge network, where ARP
only ever reaches the Docker gateway — so scans report IPs, ports and
services but never a MAC. NET_RAW does not help: it grants raw sockets,
not a place on the LAN. Rescan matching prefers MAC over IP, so this also
makes a DHCP device reappear as a new inventory entry when its lease
changes.

None of this was documented anywhere. Adds an INSTALLATION.md section
covering the network_mode: host fix, the Docker Desktop macOS/Windows
limitation and the macvlan alternative, a pointer from the README scanner
section, and a commented-out network_mode: host in both compose files so
it is a one-line change.

Reported in discussion #368.

ha-relevant: no
2026-09-02 15:58:07 +02:00
Pouzor a1e7eeb309 chore(deps): patch dependabot alerts in the frontend lockfile
Five open advisories, all transitive through devDependencies (shadcn CLI
and @vitejs/plugin-react), none reachable from the shipped bundle:

- browserslist  4.28.2 -> 4.28.8  (high, GHSA prototype write in normalizeStats)
- postcss-selector-parser 7.1.1 -> 7.1.5  (low, AST recursion DoS)
- hono 4.12.33 -> 4.13.5  (memo() SSR leak, Connection header, i18n DoS)

Lockfile only; no package.json range changed.

ha-relevant: no
2026-09-02 14:26:01 +02:00
Pouzor 97472b65ba fix(ci): run the install smoke test under bash, not dash
The runner hands container steps `sh -e {0}` unless told otherwise, so
every step died on `set: Illegal option -o pipefail` before the installer
ran. Pin the job's shell to bash; `[[ … ]]` in the assertions needs it too.

ha-relevant: no
2026-09-02 02:15:10 +02:00
Pouzor c5a49537ed test(ci): smoke-test the bare-metal installer in a debian:12 container
ShellCheck reads the installer but never runs it, so it cannot catch a
`set -euo pipefail` abort — the failure class that broke this script on a
fresh host. This job executes it for real in debian:12, which carries no
Node and gives the job no TTY, so both the nodesource path and the prompt
fallbacks are exercised on every run.

It asserts what the script promises: SECRET_KEY generated, the bcrypt hash
and the JSON values single-quoted (the systemd EnvironmentFile trap),
.env at mode 600, SQLITE_PATH under the install dir, the venv and the Vite
build present, the service user created, the unit's ExecStart and
EnvironmentFile correct. It then boots the backend by the unit's own
ExecStart and waits on /api/v1/health, which proves the generated .env
actually parses, and re-runs the installer to check the idempotency claim
leaves .env untouched.

systemd is out of reach in a container: systemctl is stubbed, so the unit
is written but never started, and the EnvironmentFile parse itself stays
untested. The job is gated on paths, since it costs roughly six minutes.

Also adds gnupg to the installer's apt list — the nodesource setup script
needs it and a minimal Debian does not have it.

Verified by running the same steps locally in debian:12: all assertions
pass, the backend answers /api/v1/health, and the second run keeps .env.

ha-relevant: no
2026-09-02 02:15:10 +02:00
Pouzor cefb96ddb4 fix(install): stop the bare-metal installer aborting on a fresh host
Two ways install-baremetal.sh failed on exactly the hosts it targets, both
caused by `set -euo pipefail` turning an expected non-zero into an exit.

Node detection ran `node --version` in a command substitution. On a host
with no Node — the normal case, since the base apt list does not install
it — that exits 127, pipefail propagates it, and the script died at that
line, before the block that installs Node. `|| true` plus a numeric
fallback, and the version is stripped to digits so `-lt` cannot error on
unexpected output. The same abort applied to the `ip` and `hostname -I`
substitutions, which now tolerate failure too.

The prompts had no TTY guard, so the documented `curl … | sudo bash` form
hit EOF on the first `read` and aborted. Both prompts are now skipped when
stdin is not a terminal, falling back to their defaults with a warning.
Two `[[ … ]] && …` one-liners were the same trap in miniature — a false
test is a non-zero list — and are now if/fi.

Docs follow the behaviour: the piped install is shown with ADMIN_PASSWORD
and SCANNER_RANGES set, and the fallback (password `admin`, guessed range)
is stated rather than implied.

Verified by running the patched blocks under `env -i` with no node on PATH
and stdin closed: both reach the end, exit 0. Shellcheck and `bash -n`
clean.

ha-relevant: no
2026-09-02 02:15:10 +02:00
Pouzor 9998b3784f feat(install): native bare-metal install path, no Docker
Adds scripts/install-baremetal.sh: installs Homelable natively on a
Debian/Ubuntu host — Python venv plus a homelable systemd unit for the
backend on 127.0.0.1:8000, the built frontend served by nginx on :3000.
Modeled on scripts/lxc-mcp-install.sh, with the install steps taken from
the community-scripts/ProxmoxVE recipe so the two stay recognisably the
same install.

The script clones into INSTALL_DIR when empty, creates the service user,
builds the venv and the frontend, generates backend/.env with a random
SECRET_KEY and a bcrypt hash, writes the systemd unit and the nginx site,
then waits on /api/v1/health. Re-running is safe and is the upgrade path:
an existing .env is kept, everything else is rebuilt. Every prompt has an
environment-variable override, so a non-interactive install is one line.

Two details worth calling out:

- JSON values in the generated .env are single-quoted. systemd's
  EnvironmentFile parser strips bare double quotes, which would hand
  pydantic [http://...] instead of ["http://..."] and fail startup.
- The admin password reaches Python through the environment rather than
  argv, which is world-readable in ps.

The unit runs unprivileged, so nmap falls back to a TCP connect scan;
AmbientCapabilities=CAP_NET_RAW is shipped commented out with the
trade-off spelled out in the unit and in the docs.

INSTALLATION.md gains a "Bare metal — no Docker" section covering the
quick start, the upgrade, the option table, the non-root scan trade-off
and the host-nginx blocks for anyone bringing their own reverse proxy.
README links it from the install line.

No test suite: shell installers have none in this repo, and CI's
lint-scripts job shellchecks scripts/. Shellcheck is clean.

Closes #333

ha-relevant: no
2026-09-02 02:15:10 +02:00
Pouzor 39cad9efc7 fix(canvas): keep a nested container ahead of its children
A container dropped into a zone left its own children draggable anywhere
on the canvas: they lost the box they were pinned inside (#366
follow-up).

The store reordered nodes for React Flow with a two-bucket split —
parentless nodes first, everything else after. That is only correct while
nesting is one level deep. A zone can hold a container, so the container
and its own children now sit in the same bucket, and a child whose row
predates its parent's lands ahead of it. React Flow then drops the
child's parent binding entirely: no parent-relative position, no
`extent: 'parent'` clamp, so it renders detached and drags free.

`orderParentsFirst` — a topological sort that already backs the zone drop
— now covers the other five reorder sites: loadCanvas, applyLayout,
updateNode's re-parent path, setProxmoxContainerMode and pasteNodes. The
child's stored state was never wrong; only the array order was.

ha-relevant: yes
2026-09-02 01:24:12 +02:00
Pouzor efcafee774 fix(mcp): keep the zone tools inside one design
/api/v1/nodes has no design filter, so list_zones reported every
design's zones and add_to_zone would happily parent a node from design B
under a zone on design A — which hides it on the canvas it belongs to.

list_zones takes an optional design_id and reports each zone's own, and
add_to_zone skips a node whose design differs from the zone's, alongside
the other skip rules. remove_from_zone needs neither: a node and the
zone it sits in are on the same design by construction.

ha-relevant: no
2026-09-01 08:52:31 +02:00
Pouzor 6c55cbfe58 fix(canvas): detach a selection from a zone in one undo step
Dragging a multi-selection out of a zone called removeFromGroup once per
node, and each call pushed its own history entry — so an add of three
nodes undid in one step but the matching detach took three, with the
intermediate states leaving part of the selection out.

removeNodesFromGroup detaches the batch in a single set, mirroring
addNodesToZone; removeFromGroup delegates to it.

ha-relevant: yes
2026-09-01 08:52:31 +02:00
Pouzor 1ec7ea7e8e feat(mcp): expose zones to the MCP server
An AI client could create nodes but had no way to create the areas that
organize them, or to move a node into one (#365).

Four tools, all on the existing /api/v1/nodes endpoints: create_zone
(a groupRect node, with size and border/text colours), list_zones (each
zone plus the ids it contains), add_to_zone and remove_from_zone, both
taking a list of node ids.

The canvas stores a child's position relative to its parent, so
add_to_zone rebases each node on the zone and remove_from_zone restores
absolute coordinates — otherwise a node jumps by the zone's offset the
next time the canvas loads. add_to_zone skips what it cannot move (the
zone itself, a node already inside it, an unknown id, and any node the
zone descends from, which would build a cycle) and reports them back
rather than failing the batch.

groupRect stays out of NODE_TYPES: create_node remains a device tool,
and the enum sync tests keep passing.

ha-relevant: no
2026-09-01 08:52:31 +02:00
Pouzor 3301f00e57 feat(canvas): add a whole selection to a zone in one drop
Dropping a multi-selection on a zone, a group or a container only moved
the node under the cursor: handleNodeDragStop read `dragNode` and ignored
`dragNodes`. Dragging a parent and its children in was especially
tedious (#365).

The drag handler now uses the whole dragged selection, with `dragNode`
only picking the destination, and the store grows batch counterparts
(addNodesToZone / addNodesToGroup / addNodesToContainer) so the move
lands in a single history entry and one undo takes it all back. The
singular actions delegate to them.

Ineligible children are skipped rather than failing the batch: the
destination itself, a node already parented to it, a node whose own
parent is in the same selection (it rides along with that parent), and
any node the destination descends from. Leaving a zone detaches every
dragged child too, not just one.

The confirmation modal takes labels instead of a label, and counts and
lists them ("Add 3 nodes (Router, Switch, NAS) to the zone DMZ?"),
summarizing the tail past five.

ha-relevant: yes
2026-09-01 08:52:31 +02:00
Pouzor 895d00a499 fix(canvas): stop a self-parented node from freezing the canvas
A node recorded as its own parent (`parent_id = id`) made the child walk in
`translateWaypointsForMovedNodes` recurse forever. The resulting stack overflow
was thrown inside the `onNodesChange` reducer, so the whole state update was
discarded: the node still selected but would never move. It also renders
unparented, sitting wherever its stored coordinates put it rather than inside
the container it appears to belong to.

Guard the walk, then close every path that could write the row:

- `propagate` keeps a `walked` set, matching the cycle guard `orderParentsFirst`
  already carries. Covers longer cycles (a -> b -> a) too, not just self-parent.
- `updateNode` / `addNode` drop a `parent_id` equal to the node's own id. The
  key is removed rather than nulled, so the rest of the edit still lands and an
  existing real parent is left alone.
- `importYaml` resolves parents by label, so a node naming its own label — or a
  duplicate label mapping back to it — imported as its own parent. Warn and skip.
- `node_dedupe` re-pointed every child of a duplicate at the canonical node,
  including the canonical node itself when it had been nested under one of its
  own duplicates. Detach instead, mirroring the self-loop edge deletion below it.
- `NodeSave` and the node PATCH route normalize it away. Dropped rather than
  rejected: a canvas that already carries the bad row must still be able to save,
  and a 422 would cost the user the whole save.

`_repair_self_parent_nodes` clears what is already persisted at startup. The
real parent is not recoverable from the row, so NULL returns the node to the top
level where the user can re-nest it. Idempotent, and never fatal to boot.

A longer cycle is deliberately left alone by the repair — there is no single
right link to cut, and the runtime guard keeps the canvas usable either way.

ha-relevant: yes
2026-09-01 01:53:03 +02:00
Pouzor 28f3ac3878 fix(services): stop appending the scanned port to a custom service host
A service `host` override usually names a reverse proxy, where the scanned
port is an internal detail: `https://proxy.example.com:8083/` is not a URL
that resolves to anything. The override now suppresses `svc.port` in both the
link the canvas opens and the URL the status checker probes — only a port
typed into the override itself (`host:port`) is printed.

The scanned port still gates the non-HTTP/SSH skip and the http/https guess,
so a MySQL or SSH service behind an override stays unlinked and unchecked.

Closes #382

ha-relevant: yes
2026-09-01 00:50:37 +02:00
Pouzor 2413ac3950 feat(imports): link a canvas-direct import to its Device Inventory row
A canvas-direct Zigbee / Z-Wave / Proxmox import placed nodes and left the
inventory alone. The first canvas save then minted a row for each of them
(`link_facts`), but the save wire shape carries no `ieee_address` — so the row
landed IEEE-less and tagged `canvas`, and the next mesh import, matching by
IEEE, created a second row for the same device.

The canvas import now upserts the inventory first, through the same
`_persist_pending_import` the pending path uses, and stamps each returned node
with the id of the row it draws. The canvas already round-trips `device_id`, so
the save links to that row instead of minting a rival. A row a node draws also
leaves the pending queue — the mirror of the existing "approved but no longer
drawn -> pending" revival. A hidden row stays hidden.

The mode selector is relabelled to match what it now does: both modes reach the
Device Inventory, so the choice is "Device inventory only" vs "Inventory +
canvas". The three modals also gain the width and spacing the two-line labels
needed.

Pre-existing split rows (a `canvas` row and an import row for one device) are
left alone: with no shared ieee/ip/mac they could only be matched on hostname,
and welding inventory rows together on a name match is worse than the drift.

ha-relevant: maybe
2026-08-31 13:26:37 +02:00
Pouzor f85130a27a test(backend): run background-task sessions against the test database
Background tasks open their own session through `AsyncSessionLocal`, which the
request-scoped `get_db` override never reached — so any test exercising one ran
against the real configured database, inserting rows and, through the imports'
wipe-and-reinsert of `device_inventory_links`, deleting real ones.

Point the factory at the test engine in `app.db.database` and in every route
module that imported it by name, and restore it when the fixture tears down.

ha-relevant: no
2026-08-31 13:26:37 +02:00
Pouzor c85bb46373 fix(mcp): make resources/read and resources/templates/list work
Two bugs made the Resources tab unusable in any MCP client.

read_resource returned mcp.types.TextContent, but the low-level
Server.read_resource decorator expects str, bytes or an iterable of
ReadResourceContents and reads .content / .mime_type off each item.
TextContent carries .text, so every resources/read failed at
serialisation with "'TextContent' object has no attribute 'content'".
The handler itself runs, so the error only ever reached the client and
never showed up in the server log.

No list_resource_templates handler was registered either, so the SDK
answered resources/templates/list with "Method not found". That hid the
homelable://nodes/{node_id} template read_resource already serves — it
is absent from RESOURCE_LIST because it is a template, not a concrete
URI, so nothing advertised it at all.

ha-relevant: no
2026-08-31 11:46:53 +02:00
Pouzor 0acfe3acc3 fix(zigbee): stop canvas imports dying on a proxy read timeout
A Zigbee2MQTT networkmap on a 200+ device mesh takes minutes to build.
Two separate failures fell out of that:

- POST /zigbee/import held the HTTP request open for the whole MQTT
  round-trip, so any reverse proxy in front of the API cut it first
  (Cloudflare returns a 524 at 120 s) and the browser never saw the map.
  It now registers a job, fetches in the background and answers 202; the
  client polls GET /zigbee/import/{job_id} until the payload is ready.
  Job results are transient and live in memory with a 15 min TTL — the
  same single-worker assumption the scheduler already makes. A failed
  fetch replays the status the synchronous route used to raise, so a bad
  broker is still a 502 and a slow mesh still a 504.

- The networkmap wait was hard-coded at 300 s with no way to raise it.
  It now reads ZIGBEE_NETWORKMAP_TIMEOUT, and the shared MQTT round-trip
  used by the Z-Wave import reads MQTT_RESPONSE_TIMEOUT. Both default to
  300 s, fall back to that if misconfigured to a non-positive value, and
  name themselves in the timeout message.

Also corrects the route and doc claims that the wait was 60 s.

The /import tests changed with the contract they cover, not to pass.

Fixes #380

ha-relevant: yes
2026-08-31 11:46:36 +02:00
Pouzor 9c2a1ba0eb docs(canvas): correct a stale comment on where a zone's height is stored
The zone size moved to the nodes.width/height columns; the comment still
described the custom_colors stash it replaced.

ha-relevant: yes
2026-08-27 01:45:23 +02:00
Pouzor 851141951d fix(canvas): store a zone's size in the width/height columns
Every node type persisted its size in `nodes.width` / `nodes.height` except
`groupRect`, which stashed it inside the `custom_colors` JSON next to its
colours. The columns already existed and were simply unused for zones, so
this was an inconsistency rather than a missing-column workaround, and it
put geometry in a blob that otherwise holds style.

The serializer now writes the columns for a zone too, and strips the legacy
`width`/`height` keys out of the blob so the two cannot drift apart and
leave an older canvas reading a stale size.

No data is lost on upgrade:

- `_backfill_zone_size` copies the blob geometry into the columns at
  startup. It only fills a column that is still NULL, so it cannot overwrite
  a size set since; it parses the JSON in Python rather than with
  `json_extract`, so it does not depend on the SQLite build carrying JSON1;
  and an unreadable row is skipped without costing the others their size.
  Re-running it is a no-op.
- the reader still falls back to the blob, covering a payload the backfill
  has not reached — an older server, or an import.

Standalone mode is unaffected: it stores React Flow nodes verbatim, so the
size was always on `node.width` / `node.height` there.

The four serializer tests that pinned the size to the blob now assert the
columns, since that is the behaviour being changed.

ha-relevant: yes
2026-08-27 01:45:23 +02:00
Pouzor f9f88c8fb0 refactor(canvas): drop a dead custom_colors.height write on zone growth
Growing a zone during a subnet import wrote the new height twice: to
`node.height`, the live field, and to `data.custom_colors.height`, which
nothing reads. The blob copy is produced by the serializer at save time
(rebuilt from `node.height`, so the value written here was overwritten
before it reached the API) and consumed at load time off the API payload.
Writing it from the store was invisible, and misleading in a blob that
otherwise holds colours and style.

The test asserting the dead field is replaced by a save/load round-trip
through the real serializer, which is what actually protects the height.

ha-relevant: yes
2026-08-27 01:45:23 +02:00
Pouzor 2bf9f1bca8 fix(canvas): keep a parent ahead of its children after a subnet import
The subnet import only guaranteed the zone preceded the nodes it pulled in.
An arrival that is itself a parent — a Proxmox host with nested VMs, whose
children stay put because they already have a parent — could end up listed
after those children, which React Flow renders detached with a
parent-not-found error. Nothing re-sorts on load, so the broken order
survived a save.

Reorder the whole array instead: every node now follows its parent, which
covers both directions at once. An already-valid list comes back untouched
and a parent cycle terminates rather than recursing.

ha-relevant: yes
2026-08-27 01:45:23 +02:00
Pouzor d9451f35dc feat(canvas): import devices into a zone by subnet
Closes #325.

A zone gains an "Import devices by subnet" action: type a CIDR, and every
free device whose IP falls in that range moves into the zone, laid out on a
grid in its free space.

The CIDR is an argument to a one-shot action, never a property of the zone —
it is not submitted with the form, not persisted, and cleared after each run.
So there is no schema change, no migration, and it works in standalone mode.
The trade-off is that a device scanned later does not join the zone by
itself; the user re-runs the import.

Deliberate rules:
- additive — running it twice with two subnets leaves both sets inside, and
  a device whose IP stops matching is never ejected
- only unparented devices move. A node nested in a group, a container host
  or another zone keeps the parent the user gave it
- canvas furniture (groupRect / group / text) is skipped, and a zone never
  swallows another zone
- one history entry per import, so a single Undo reverses the whole thing

Edit mode gets an Import button; add mode has none, since the zone does not
exist yet and a button would look broken — there the CIDR is applied right
after the zone is created.

IPv4 only: the modal rejects an IPv6 CIDR with a message rather than
silently matching nothing.

ha-relevant: yes
2026-08-27 01:45:23 +02:00
Pouzor 05c0da53e6 fix(scan): reconcile scan runs orphaned by a backend restart
A scan runs on a background thread inside the API process. If that process
dies mid-scan — an OOM kill, docker stop, a crash — the ScanRun row stays
"running" for ever, because nothing is left alive to finish it.

That row is not just cosmetic clutter in Scan History: the trigger endpoints
reject a new scan while one is "running" for the same target, so a single kill
locks that range out permanently.

Nothing can legitimately be "running" the moment we boot, so lifespan() now
marks every such row "error" — the same word run_scan and run_device_scan
write when they fail themselves — with finished_at and an explanatory message.

Reported alongside the OOM itself in #374, which is what produced the orphans.

Fixes #374

ha-relevant: maybe
2026-08-27 00:00:32 +02:00
Pouzor c29680a75f fix(status): stop an endless response body from OOM-killing the backend
Both HTTP paths buffered the whole response body before looking at it. An
endpoint that streams without end and sends no Content-Length — a Freebox
bandwidth-test port, an MJPEG camera, a log tail — grew the backend until the
cgroup limit killed it, every 5-10 minutes at a constant ~1.04 GB RSS.

httpx's timeout does not help: it applies per network operation, not to the
total time spent draining a socket that keeps delivering data.

- status_checker._http_get only needs the status line, so it now uses
  client.stream() and never reads the body at all.
- http_probe._probe_scheme needs at most _MAX_BODY_BYTES to hunt for <title>,
  so it streams and stops there. The cap existed already but was applied to
  resp.text, after the full body had been downloaded.

Regression tests serve an endless, Content-Length-less body through a
MockTransport and assert the read stays bounded. Both fail on the old code.

The existing mocks patched httpx.AsyncClient.get, which neither path calls
now; they are rebuilt on MockTransport — a real client over a fake network —
with every original assertion kept.

Fixes #375

ha-relevant: yes
2026-08-26 23:43:50 +02:00
Pouzor - Rémy JardientandGitHub eb7e74d36b Update README.md 2026-08-25 23:08:03 +02:00
Pouzor 390a712f14 chore: bump version to 3.3.5
ha-relevant: no
v3.3.5
2026-08-21 12:56:17 +02:00
Pouzor 9daea4c4d0 fix(scan): keep an edit, a discovery source and one name for a failed run
Three findings from the PR review.

A finishing deep rescan threw away an edit in progress. The reset effect in
InventoryDeviceModal was keyed on the `device` object, and the parent hands
down a fresh one whenever the row is refreshed — including from the poll's
own onSaved. Same device, new object, so the effect reset the form and left
edit mode, minutes into a scan the user was waiting on. It is keyed on the
id now: a refreshed row is not a different device, and only a different
device is a reason to throw the form away. The fresh row still reaches the
canvas and the grid — gating that behind edit mode would discard the scan
result instead, and a save unions the services server-side anyway.

A rescan tagged every device it touched as "arp"-discovered, so a Proxmox
guest, a rack mount or a hand-added host started answering the network
source filter. It carries its own source through instead.

The two background wrappers wrote status "failed" where the scanners write
"error" for the same condition. Harmonized on "error" — the one Scan
History filters and colours; "failed" showed up unlabelled. The frontend
union keeps 'failed' for rows already recorded.

That last one meant updating an existing assertion in test_scan_run.py: it
encoded the old spelling.

ha-relevant: yes
2026-08-21 12:49:48 +02:00
Pouzor 2e642815f4 fix(scan): stop a scan from repainting hand-picked service icons
merge_services did {**existing, **incoming}, so the fingerprint's guess at
an icon overwrote the one the user chose — and on a port no signature
covers it wrote None, clearing it outright. Since 3.3.0 the inventory row
is the only copy of a device's services, so every "Scan network" repainted
the service on every canvas drawing that device at once.

The scanner now merges with discovered=True: it still adds services and
refreshes what it knows, but leaves an established icon and category alone.
A user edit from the modal or a canvas changes them as before.

Blank incoming values no longer clear established ones either, on both
paths — an absent field is silence, not a reset. Same rule merge_properties
already follows.

ha-relevant: yes
2026-08-21 12:49:48 +02:00
Pouzor c6d6b525ea feat(scan): choose the port range before a deep scan
The Deep scan link on a device now opens a small dialog instead of firing
straight away. It is prefilled with the full 1-65535 range — that is still
the point of the feature — but a user who knows where a service lives can
narrow it and get an answer in seconds instead of minutes.

The dialog takes an nmap-style spec: a port, a range, or a comma list
(80,443,8000-9000). It shows the port count live, refuses to start on
something nmap could not use, and carries three presets (all / 1-1024 /
1-10000).

Backend:
- _parse_port_spec / _port_chunks generalize the slicing that used to be
  full-range only. Ranges are merged before slicing, so an overlapping
  spec is never scanned twice, and small ranges are packed into one nmap
  call instead of one call each. _deep_port_chunks still yields the same
  eight slices as before.
- run_device_scan takes ports=; it wins over full_ports.
- The retry-free flags (--max-retries 0 --min-rate 2000) now key on the
  total port count rather than on full_ports. They pay for themselves over
  thousands of ports on a lossy host; over a handful they only cost
  accuracy.
- RescanDeviceRequest.ports validates the spec — 422 rather than handing
  nmap a bad -p. Blank means the full sweep.

Frontend utils/portSpec.ts mirrors the backend parser so a typo is caught
before the request; the backend validates again because it is the one
calling nmap.

The existing rescan tests now go through the dialog: the click path
changed, so the start is two steps.

ha-relevant: yes
2026-08-21 12:49:48 +02:00
Pouzor fbb660504e fix(scan): slice the deep rescan instead of timing out the host
A deep rescan of a slow host came back with nothing at all: the run took
its full 600s ceiling and the device's services were unchanged, so a
service deleted by hand was never rediscovered.

nmap answers --host-timeout with "Skipping host <ip> due to host timeout"
and discards every port it had already found — the ceiling turned a slow
scan into one that reports nothing. What costs the time is a host that
drops packets: 8188 of 8192 ports filtered, each waiting out its probe.

- No --host-timeout on the deep discovery pass, ever.
- The full range runs as 8 slices of 8192 ports, one nmap call each,
  unioning the open ports. A slice that overruns costs its own ports, not
  all of them, and the loop has somewhere to notice a stop request.
- `scanner_deep_host_timeout` is now a total budget checked between
  slices (default 2700s), not an nmap flag. The first slice always runs.
- A partial sweep is reported rather than passed off as complete: the run
  finishes `done` carrying "Scanned 3/8 port ranges …", and the modal
  toasts a warning instead of success.
- Deep slices use --max-retries 0 --min-rate 2000. Measured against a
  dropping host, 8192 ports took 329s at --max-retries 1 and 164s at 0,
  finding the same ports; capping the RTT changed nothing. The range scan
  keeps nmap's default retries on its curated port list.

ha-relevant: yes
2026-08-21 12:49:48 +02:00
Pouzor 1be96d1045 feat(scan): deep-rescan one device from the inventory detail
Devices added before the scanner knew a service showed an empty Services
section with no way to refresh it (#350). The detail modal now starts a
full-port scan of that single device.

- `process_host` lifted out of `run_scan` so the range scan and the new
  single-device scan share the same match / merge / dedupe rules — a
  rescan unions services, it never replaces what the user added by hand.
- `run_device_scan`: no ping sweep, no mDNS, straight to the phase-2 nmap
  pass on the device IP over all 65535 TCP ports.
- `POST /scan/pending/{id}/rescan` records a normal ScanRun
  (`kind=device`, `ranges=["<ip>/32"]`), so stop, progress and Scan
  History work unchanged. One run per device at a time — a second request
  while the first is scanning is a 409. 404 unknown, 409 no-IP or hidden.
- `GET /scan/runs/{id}` so a caller can poll the run it started.
- Deep scan button in the Services section of the device detail, hidden
  without an IP and for Zigbee. Swaps to a stop control while running,
  folds the fresh services back in on completion (never over an edit in
  progress).

ha-relevant: yes
2026-08-21 12:49:48 +02:00
Pouzor 9bc02d61c2 fix(canvas): give a valueless property label the full node width
The property line caps its label at max-w-15 so a long key cannot crowd
out the value drawn beside it. When the value is empty there is nothing
to protect, but the cap still applied, truncating the label against
empty space.

Drop the cap when the value is blank and let the label truncate against
the node width instead. Same fix in BaseNode and ProxmoxGroupNode.

Closes #361

ha-relevant: yes
2026-08-17 20:26:29 +02:00