Skip to main content

Installing the noBGP Agent

The noBGP agent is a lightweight background process that connects your machines to noBGP networks, enabling secure encrypted communication with other nodes. Your AI assistant can execute authorized commands, monitor system status, and create secure tunnels for your services.

Getting Started​

Run the install script and follow the prompts — it handles everything: installation, registration, and service setup.

curl -fsSL https://downloads.nobgp.com/agent/install.sh | sh

Or on macOS with Homebrew:

brew install nobgp/tap/nobgp

The install script will:

  1. Download and install the correct package for your platform
  2. Open your browser to sign in with your noBGP account (Google, GitHub, etc.)
  3. Let you pick, on that same page, the organization and network the node joins
  4. Install and start the agent as a system service

No keys or tokens needed — just run the script and follow the prompts.

tip

When using your AI assistant with noBGP integration, it can generate a ready-to-run install command using the register_node tool.

Platform Details​

Installation Scripts​

The install scripts from Getting Started detect your system and install the appropriate agent package. Here are platform-specific details and alternative installation methods.

Which version they install. With nothing set, the scripts ask noBGP for the stable channel version and install that. There is no fallback to whatever the download site happens to hold, so an answer they cannot read — a captive portal, a proxy error page — stops the install and names NOBGP_VERSION rather than installing something unexpected. Set NOBGP_VERSION to pin a version:

curl -fsSL https://downloads.nobgp.com/agent/install.sh | NOBGP_VERSION=0.4.153 sh

From agent 0.4.154 install.sh leaves an agent that is ahead of the channel alone. Re-running the one-line installer on a machine whose agent is newer than the stable-channel version keeps the installed binary — it prints Installed version <x> is newer than channel version <y>. Keeping it. and carries on to registration and service setup — so a machine following a pre-release channel is no longer quietly downgraded by an install command copied from these pages. NOBGP_VERSION still installs exactly what you name, downgrades included. Earlier releases, and the Windows install.ps1 at any release, install the channel version over a newer one; to move a node down deliberately, use sudo nobgp upgrade -f.

Linux​

Using curl:

# Install latest released version
curl -fsSL https://downloads.nobgp.com/agent/install.sh | sh
# Keep the downloaded package file. Only NOBGP_* variables survive the
# script's own elevation, so this one needs sudo in front.
curl -fsSL https://downloads.nobgp.com/agent/install.sh | sudo KEEP_PACKAGE=true sh

Or using wget:

# Using wget instead of curl
wget -qO- https://downloads.nobgp.com/agent/install.sh | sh

The install.sh script supports:

  • Alpine Linux 3.16+
  • Arch Linux
  • Amazon Linux 2023+
  • Debian 11+ / Ubuntu 20.04+
  • OpenWRT
  • Synology DSM (see Synology DSM below)
  • Buildroot and other Linux images with no package manager, from agent 0.4.99 (see Linux with no package manager below)

On RPM hosts — RHEL, Rocky Linux, AlmaLinux, Oracle Linux, CentOS Stream and Amazon Linux — the package links /usr/bin/nobgp at the installed binary from agent 0.4.124, so sudo nobgp … works by name. Those distributions leave /usr/local/bin, where the binary itself lives, out of sudo's secure_path, which is why a bare sudo nobgp there used to answer command not found; see The install script finished but the node never registered. The link is created only where nothing already occupies that path and is removed with the package.

On Debian and Ubuntu, the script does not require apt-get update to succeed (an EOL release with archived repositories fails it permanently), and if apt refuses the install because of pre-existing broken dependencies elsewhere on the host, it falls back to dpkg -i. See Debian and Ubuntu hosts for the same behavior during upgrades.

The script also installs the packages the shared drive needs — the FUSE kernel module and its userspace helper (fuse3, or kmod-fuse and fuse-utils on OpenWrt), falling back to davfs2 where the FUSE package cannot be installed. That fallback is only tried on Debian/Ubuntu, Alpine and RPM hosts, and only if FUSE fails, which it rarely does — Arch and OpenWrt install FUSE alone, and on Synology DSM — and on a Linux image with no package manager — the script installs neither, because neither host has anything to install it with. So most Linux nodes have no davfs2 and therefore no webdav backend to fall back to; install it by hand if you want that third option. This is best-effort either way: the filesystem is optional, so a host that cannot supply those packages still installs and runs, without the shared drive. See Shared drive is empty and never mounts if that is where you end up.

From agent 0.4.58 it also installs an NFSv4 client, so the node can take the nfs backend — which on Linux is what a host with no usable FUSE falls back to, rather than dropping all the way to WebDAV:

PlatformWhat is installedNotes
Debian / Ubuntunfs-common, without recommended packagesAbout 2 MB. Dropping the recommendations leaves out python3
Amazon Linux / RHEL / Oracle / Rocky / AlmaLinuxnfs-utils, without weak dependencies
Arch Linuxnfs-utils
macOSnothingmount_nfs is part of the base system
Alpine, OpenWrt, Synology DSMnothingDeliberate — see below. FUSE already works on all three
Linux with no package managernothingThere is nothing to install it with. The node gets whatever the image was built with

It is installed beside the FUSE helper, not instead of it: the agent probes every backend and picks per host, so both are worth having — and from agent 0.4.65 a Linux node that has both prefers FUSE, with NFS as the fallback. Alpine is left out on cost — its nfs-utils hard-depends on rpcbind and python3, taking the install from 27 MB to 80 MB — and OpenWrt and Synology because FUSE serves them already.

From agent 0.4.70 no package is needed for NFS on Linux at all. With no mount helper on the host the agent makes the mount itself, so the nfs backend is available wherever the kernel has an NFSv4 client — a bare container included, with nothing installed. The packages above are still installed where the script installs them, and a helper that is present is still used, so nothing changes on a host that already has one; what changes is that Alpine, OpenWrt and Synology no longer need a hand-installed client to reach this backend. On those three, check the kernel half if nobgp status still reports nfs as unavailable — that is the only requirement left.

rpcbind is confined to loopback

On Debian, Ubuntu and Arch the NFS client package pulls in rpcbind, and a fresh install of it listens on every interface, TCP and UDP, on port 111 — a daemon NFSv4 does not use at all. When the installer is what brought rpcbind in, it writes /etc/systemd/system/rpcbind.socket.d/10-nobgp-loopback.conf to bind it to 127.0.0.1 and [::1] only, and restarts the socket. Local RPC consumers keep working; the node simply stops offering RPC to the network.

A machine that already had rpcbind is left exactly as its operator configured it. To undo the confinement, delete that file and run sudo systemctl daemon-reload && sudo systemctl restart rpcbind.socket.

FUSE is not the only backend a Linux node can use. The agent probes what the machine can actually do and mounts with the first of FUSE, NFSv4 and WebDAV that can serve — so a host with no fusermount but an NFSv4 client in its kernel mounts over NFS and needs nothing installed at all, which is why the missing FUSE helper is best-effort rather than fatal. The davfs2 fallback is the least of the three and the least available: the script reaches for it only when FUSE cannot be installed, which almost never happens on Debian, Alpine or Amazon Linux, and Synology cannot install it at all — synopkg has no WebDAV package, and DSM registers no NFSv4 client either, so a Synology node has FUSE and nothing behind it. Run nobgp status on the host to see every backend it could use and why each can or cannot serve, and see Which filesystem you get for the full picture.

macOS​

Using Homebrew (recommended):

brew install nobgp/tap/nobgp

Using the install script:

curl -fsSL https://downloads.nobgp.com/agent/install.sh | sh

macOS 13+ (Ventura and later) is supported on both amd64 and arm64 (Apple Silicon).

The shared drive needs nothing installed on macOS — it mounts over NFSv4 using mount_nfs, which is part of the operating system. FUSE and WinFsp have no macOS build at all, and installing macFUSE does not change that.

macOS: the agent's own service cannot read the mount

File operations and commands the agent runs for you against /Volumes/nobgp fail with Operation not permitted, at every identity including root. This is macOS TCC consent rather than a permission problem — running elevated fails identically, and a background LaunchDaemon has no way to ask you for consent.

Your own account reads the mount normally from Finder or Terminal; only the agent's service is blind. Granting the agent Full Disk Access removes the denial: System Settings → Privacy & Security → Full Disk Access → +, then ⌘⇧G (the picker hides /usr/local) and enter /usr/local/bin/nobgp — the daemon's own binary. The running daemon picks it up at once; there is nothing to restart or remount. Without the grant, reach the same files through the router — a network's share by network name, or the node's own area with storage: true — instead of by a path under the mount point.

Whether a given node is in this state is reported rather than assumed: from router 0.4.83, network_directory publishes info.mount_readable per node, so a granted Mac shows true and its paths behave like any other node's. The node is what answers, from agent 0.4.85 — below that the field is absent whatever the grant.

From agent 0.4.93 the Mac itself says so too: sudo nobgp status prints an fs.unreadable line under fs naming the denial and this grant, and prints nothing there once the grant is in place. Before that release the same node reported fs.mounted: true and named no condition, which reads as a healthy drive.

From agent 0.4.144 registering the Mac asks you for the grant, so a new node does not have to discover this later. Where its output goes to a terminal on the Mac, sudo nobgp register has the agent check its own mount and, where it is refused, open System Settings on the Full Disk Access pane itself — on your own screen, with the terminal naming "nobgp" (/usr/local/bin/nobgp) as the row to switch on. It never fails the registration, and it says nothing when the grant is not needed. On a Mac that is already enrolled, running sudo nobgp register again makes the same offer.

From agent 0.4.145 the install script asks as well. The offer reads nothing from you, so it is the command's output being a terminal that arms it, not its input: curl … | NOBGP_KEY=<KEY> sh at your own terminal on the Mac opens the same pane. An enrolment whose output is not a terminal — an MDM, a script logging to a file — arms nothing, and on agent 0.4.144 no registration with a key did, because the offer was gated on standard input and the script hands registration /dev/null there. Grant it by hand on such a node, or run sudo nobgp register at the machine.

Synology DSM​

Run the install script over SSH on your Synology NAS:

curl -fsSL https://downloads.nobgp.com/agent/install.sh | sh

DSM has a busybox-based userland with no native package manager, so the installer detects Synology and drops the static Alpine binary into /usr/local/bin/nobgp with a config directory at /etc/nobgp/. Only amd64 and arm64 NAS models are supported.

DSM 7

On DSM 7 (which ships systemd), nobgp register installs and starts the agent as a service automatically — no extra steps required.

DSM 6 and older

DSM 6 and older have no systemd, so the install script prints a warning and skips service auto-install. To start the agent on boot, create a Triggered Task in DSM Control Panel → Task Scheduler:

  • Trigger: Boot-up
  • User: root
  • Run command: /usr/local/bin/nobgp agent

Both manual nobgp upgrade and router-driven auto-upgrade work on DSM. Since DSM has no package manager, the agent fetches the raw static binary (instead of the .apk) and atomically swaps it in place — see Upgrading.

Linux with no package manager (Buildroot and similar)​

From agent 0.4.99 the install script also serves a Linux image that has no package manager it can drive — an appliance or router firmware built with Buildroot, for example. Synology DSM was the first machine of this shape; this is the same treatment generalised.

curl -fsSL https://downloads.nobgp.com/agent/install.sh | sh

How the script decides. It takes this path only when both halves are true: /etc/os-release reports ID=buildroot, and none of apk, apt-get, dnf, opkg, pacman or yum is on the path. A Buildroot image that ships opkg is an OpenWrt-shaped box and installs like one. rpm and dpkg are deliberately not consulted — some firmware carries both against an empty package database, where installing a .deb fails on dependencies the static binary does not need. A host with no package manager and some other ID is still reported as unsupported rather than guessed at.

What it installs. The raw static Linux binary — every Linux build is statically linked, so it needs no package manager and no matching libc — published beside the Debian package for the same architecture: amd64, arm64, armv6 or armv7. Anything else stops with No static binary published for <arch>.

WhatWhere
Binary/usr/bin/nobgp — not /usr/local/bin, which such an image need not have at all, and need not have on PATH
Configuration/etc/nobgp/

Set INSTALL_DIR to put the binary somewhere else, which is what a read-only rootfs needs. The script tests the directory before it writes and names it if it cannot. Put sudo in front for this one: the script carries only NOBGP_* variables through its own elevation.

curl -fsSL https://downloads.nobgp.com/agent/install.sh | sudo INSTALL_DIR=/opt/bin sh

Pick a directory that survives a reboot — a tmpfs path leaves the node without its binary at the next boot. Whatever you choose is put on PATH for the rest of the run, so the registration step at the end still finds the binary.

No systemd here, but still a service

These images usually run busybox init rather than systemd. From agent 0.4.100 the agent recognises that shape — init scripts, but neither Debian's service command nor a boot directory of symlinks (/etc/rc.d or /etc/rcN.d) — and installs its init script at /etc/init.d/nobgp, driving it by path and adding the S99nobgp boot symlink that busybox's rcS actually reads. So nobgp register, nobgp service start|stop|restart|status and start-at-boot all work, and the installer says so during the run.

It is a capability check rather than a distribution check: a box that has either helper keeps the ordinary SysV path.

From agent 0.4.101 the link does not always go beside the init script. On an appliance whose root filesystem is an overlay — GL.iNet's KVM appliance is the machine this was measured on — rcS expands its for i in /etc/init.d/S??* list once, before the script that mounts the overlay has run, so anything installed later is in no list and never starts. Where the firmware ships a hook script (/etc/init.d/S99custom) that enumerates a directory of its own when it runs, the agent puts the boot link there instead (/etc/kvmd/user/scripts/S99nobgp), creating the directory if the machine has never used it. It reads that hook and requires it to name the directory rather than trusting the filename, so unfamiliar firmware falls back to /etc/init.d/S99nobgp. Machines without such a hook are unaffected.

⚠ Agent 0.4.99 did not manage this correctly, even though the rest of the packageless install landed there. On such a host it failed the service step outright with failed to install service: symlink /etc/init.d/nobgp /etc/rc.d/S50nobgp: no such file or directory — a directory these images do not have — leaving the init script written and nothing to start it, then or at boot. See Buildroot node registers but the service will not install to repair one.

⚠ Nothing is installed for the shared drive, and there is no mechanism here even in principle — that is what having no package manager means. The node gets whatever filesystem support the image was built with, and the installer prints that rather than passing over it in silence. Run nobgp status to see which backends the host can actually serve; see Which backend mounts the drive.

The script loads the tun module itself, since there is no package dependency here to pull it in.

Upgrades work the same way they do on DSM: with no package manager, both nobgp upgrade and router-driven auto-upgrade fetch the raw static binary published beside the Debian package and swap it in place — see Upgrading.

Windows​

Run the following. It is one command for every Windows shell — Command Prompt, PowerShell 5.1 and PowerShell 7 — so there is nothing to pick:

powershell -NoProfile -Command "irm https://downloads.nobgp.com/agent/install.ps1 | iex"

The installer requires Administrator privileges to write to C:\Program Files\nobgp, update the system PATH, and register a Windows Service. If not already elevated, it automatically requests UAC elevation and forwards any NOBGP_* environment variables to the elevated session.

Windows 10/11 and Windows Server 2019+ are supported on amd64 and arm64.

⚠ From agent 0.4.82, auto mounts the shared drive with WinFsp where the driver is installed, and falls back to the WebDAV client built into Windows where it is not. Below 0.4.82, auto always used WebDAV.

Why the change: a WebDAV drive belongs to exactly one logon session, so a service, a scheduled task or an elevated Administrator prompt sees no drive at all on a machine where nobgp status says it is mounted. A WinFsp volume is machine-wide, visible to every account and every service, and is not labelled DavWWWroot.

⚠ It can reach a machine that never chose it. WinFsp is a shared driver — rclone, sshfs-win, Cygwin and MSYS2 all install it — so a node can have it for entirely unrelated reasons and still swap its working WebDAV drive for it at the next upgrade, with nothing but a changed backend in nobgp status to say so. fs: webdav in the profile (Filesystem keys) keeps the drive exactly as it is; fs: off mounts nothing at all, which is a different thing and not a fallback. An explicit choice is honoured or refused, never silently replaced.

It needs WinFsp 1.10 or newer — an older version is reported as unavailable by name and version rather than mounting and failing, and auto then uses WebDAV as before. What it costs is one operation: renaming a file was measured about 1.7× slower than WebDAV, and agent 0.4.82 removes a wasted network round trip from every rename, create and folder creation, so the gap should now be smaller than that measurement suggests. ⚠ Its file locks stay on that machine, so two programs on two different nodes can each believe they hold an exclusive lock on one file — the same is true of the WebDAV drive it replaces, which additionally lets the loser's write overwrite the winner's; here the losing write is refused instead.

⚠ A machine that already has the driver moves at its next upgrade, on its own, with nothing but a changed fs.type in nobgp status to announce it — so it is worth knowing what moves with it. Locks on the WinFsp volume stop at that machine (though the WebDAV drive's lock left it and protected nothing either), renaming files is somewhat slower while walking, listing, creating and deleting are quicker, and a cross-node write that loses is refused rather than blended. See winfsp on Windows for the whole list, and set fs: webdav in the profile (Filesystem keys) to keep the old drive on a machine that wants it.

WinFsp is a shared driver that rclone, sshfs-win, Cygwin and MSYS2 all install, so many machines have it without ever having chosen it for noBGP. It needs 1.10 or newer: from agent 0.4.73 the version is part of the probe, so an older install reports as unavailable in nobgp status — naming the version it found and the one it needs — and the node stays on WebDAV rather than being selected and failing every mount. To put a machine that has no driver onto the volume, install it (winget install WinFsp.WinFsp) and restart the agent.

WinFsp went read-only in agent 0.4.62 and read-write in 0.4.63, though a real mount only accepts writes from agent 0.4.65; through agent 0.4.81 auto declined it and fs: winfsp was how you asked for it.

The installation scripts support the following architectures:

  • amd64 (x86_64)
  • arm64 (aarch64)
  • armv6 and armv7 (Debian/Ubuntu, for Raspberry Pi and similar devices; armv7 also on OpenWRT)

The install scripts guide you through registration (browser sign-in) and service installation. After the script completes, the agent is running and connected. If you installed via Homebrew or a manual package install, run sudo nobgp register to complete setup.

Windows Firewall and traffic from other nodes​

From agent 0.4.124.

By default, the agent adds one inbound rule to Windows Firewall, named noBGP-<profile> (for example noBGP-default), with a description that names noBGP. The rule allows all inbound traffic from your noBGP network, on every protocol and every port. So any node in your network can reach every service that listens on this machine — for example file sharing (SMB, port 445), RPC (port 135) and Remote Desktop (port 3389) — even when no other Windows Firewall rule allows those ports. A block rule that you added still wins over this rule.

The per-profile windows-firewall setting decides this. Set it with nobgp config --windows-firewall <value> or in the profile's configuration file:

ValueWhat the agent does
allow-overlay (default)Adds the allow rule described above. This is what every earlier agent release did.
respect-windowsAdds no allow rule, and removes the rule that an earlier start added. Windows Firewall and your own rules decide what other nodes can reach on this machine.
nobgp config --windows-firewall respect-windows
# C:\ProgramData\nobgp\default.yml
windows-firewall: respect-windows
  • A change takes effect within about 5 seconds. When nobgp config or an edit writes the configuration file, the agent reloads, and it applies the setting when the tunnel starts again. The reload closes the node's connections for a short time.
  • nobgp config refuses a value that is not one of the two.
  • A clean stop of the agent keeps the rule. nobgp uninstall removes it, and from agent 0.4.126 nobgp remove removes the rule of the profile it removes. With respect-windows, the agent removes the rule at its next start.
  • The rule applies to every network interface, also while the agent is stopped.
  • If the agent cannot read the host's firewall rules, allow-overlay leaves the rule exactly as it is until the next start rather than risking a second copy of it (agent 0.4.126). respect-windows still removes it, by name.
From agent 0.4.126 the rule is named after the profile

Releases up to 0.4.125 named it after the tunnel adapter — noBGP-nobgp0 — and that adapter's number can change between starts, so a rule kept across a stop could end up belonging to no profile, or to the wrong one on a machine running several.

You do not have to do anything: at the next agent start the device-named rule belonging to a profile is renamed to noBGP-<profile>, keeping every edit you made to it — whether it is enabled, which firewall profiles it applies to, ports, interface types, and a remote address range you narrowed by hand. This is true when one rule has that name and its remote address range is inside the profile's slice, or its name is the adapter that this start opened. A name with duplicate rules (left by a crash), or a rename that fails, gets a new rule without your edits. On a Windows whose display language is not English, the agent cannot read the remote address range, so it sets the range to the profile's slice. Any other device-named rule that profile owns is removed, so the migration leaves one rule where there was one. Each profile on the machine has its own rule.

With respect-windows, the noBGP tools still work: command, the file tools and published services go through the agent, not through an inbound firewall rule. Through agent 0.4.139 what changed was direct traffic from other nodes to this machine's address, such as Remote Desktop, file sharing or ping: to allow one of those you added your own inbound rule for the noBGP network address range.

From agent 0.4.140 this setting no longer keeps peers off TCP, UDP and ping

A connection from another node now ends at the agent and is opened again from this machine, so Windows Firewall sees it as local traffic rather than traffic from the network and does not filter it. Agents 0.4.140 to 0.4.143 still applied respect-windows to everything that arrived on the tunnel adapter — other IP protocols, IP fragments, and ICMP other than ping — and from agent 0.4.144 the setting filters nothing from another node, because none of that is written into the adapter any more: the agent drops what it does not deliver.

Another node therefore reaches what a program on this machine reaches — including a service bound to 127.0.0.1 — whichever value this setting holds. nobgp show says so on the windows-firewall line of a machine delivering that way. See How another node reaches this machine's own services.

The setting has no effect on Linux and macOS.

Docker and Automation​

For Docker containers, CI/CD pipelines, and other non-interactive environments where a browser login isn't possible, use a registration key to register the agent automatically.

Get your registration key from the noBGP Add Node page or ask your AI assistant to retrieve it.

warning

Treat registration keys as secrets. Don't commit them to git repositories or share them publicly.

Who may read a key, from router 0.4.179

A registration key is the network's reusable credential — whoever holds it can join a machine to the network — so reading, creating, disabling and deleting one is now the network's owner or an Owner or Admin of the organization. Any other member who asks an assistant for install commands gets the keyless form instead: it runs nobgp register and prints a sign-in link, and the person who signs in owns the machine that results. See network_key_list and register_node.

⚠ A machine enrolled with the team's key belongs to the network's owner, not to the person who ran the command — so that person cannot set it to private. Enrol a personal machine with the keyless, signed-in form if you want to own it.

Docker Run​

docker run --detach --restart unless-stopped \
--pull always \
--cap-add NET_ADMIN \
--cap-add SYS_ADMIN \
--device=/dev/net/tun \
--device=/dev/fuse \
--security-opt apparmor=unconfined \
-e NOBGP_KEY=<YOUR_REGISTRATION_KEY> \
-e NOBGP_NAME=docker-node-1 \
nobgp/nobgp

Docker Compose Example​

services:
nobgp:
image: nobgp/nobgp
pull_policy: always
restart: unless-stopped
cap_add:
- NET_ADMIN
- SYS_ADMIN
devices:
- /dev/net/tun
- /dev/fuse
environment:
NOBGP_KEY: "<YOUR_REGISTRATION_KEY>"
NOBGP_NAME: "docker-node-1"
Without NET_ADMIN and /dev/net/tun, the node still starts

Some platforms will not grant either — AWS Fargate, Cloud Run, Azure Container Instances, and Kubernetes pods under a restrictive PodSecurity profile. From agent 0.4.142 a Linux container there comes up without an overlay instead of exiting and restarting: the node registers and stays online, and command, file, the MCP tools and nobgp status all work, but it has no tunnel interface, no overlay DNS and no peer traffic — so nothing reaches it by its overlay name and it reaches no peer by theirs. nobgp status names the reason under network.overlay. Add the capability and the device to get the overlay back, then restart the container.

Headless Registration with Key​

Sign in from your phone instead

From agent 0.4.131, sudo nobgp register prints the sign-in link as a QR code you can scan with a phone — the sign-in completes there, because the command polls noBGP for the result rather than waiting for a browser on the machine. That covers a box with a terminal but no browser; a container or a cron job, which has no terminal at all, still needs a key. From agent 0.4.154 a terminal always gets a code: the tall painted drawing where the window fits it, and a half-block one about 41×21 where it does not, so a default 80×24 window is enough. Agent 0.4.153 drew only the tall form and left the code out of a window too small for it — which was most of them. From agent 0.4.158 the code is black and white on every terminal that shows colour, and a window too short for the tall drawing is asked to grow so it gets that one instead. The terminal also prints a short code to type on the page; from router 0.4.182 the approval page shows no code back and names the machine instead, so read the card before you approve. See nobgp register.

On any system where interactive browser login isn't available, register with a key:

sudo nobgp register --key <YOUR_REGISTRATION_KEY> --name "web-server-1"

Or give the key to the install script for a single-command setup:

curl -fsSL https://downloads.nobgp.com/agent/install.sh | NOBGP_KEY=<YOUR_REGISTRATION_KEY> NOBGP_NAME=<NODE_NAME> sh

The script elevates itself, so there is no sudo in the command.

With NOBGP_KEY set, the script never prompts, so it runs unattended anywhere a shell can run — a cron job, a systemd unit, a cloud-init script, or a Synology DSM scheduled task, none of which have a terminal attached. Without a key, the script needs a terminal so you can complete the browser sign-in; if it can't reach one, it stops after the install and tells you to run sudo nobgp register yourself.

A registration key enrolls any machine into your network until you rotate it. After a bulk install, rotate the key.

Rotate the key after a bulk install​

When you have installed many machines with one registration key, rotate the key. Every machine and every script that saw the key could enroll another machine into your network for as long as the key is valid.

In the web app — or with network_key_create, router 0.4.179+ — add a new registration key and move your scripts to it. Then disable or delete the old key (network_key_update / network_key_delete). A key's own secret never changes, so rotating is always adding one and retiring the other. Nodes that are already registered stay connected, because after enrollment a node authenticates with its own credentials, not with the key. A machine that must enroll again uses the new key: a reinstall, a node that was offline for longer than its credentials last (30 days), or a drop-in configuration file that no agent has read yet.

Rotate the key also when you think that someone else has seen it, for example when it was on a command line, in a shared script, or in a log.

Drop-in Provisioning​

For fleet management or temporary network access, you can provision agents by placing a config file directly in the config directory. The process manager detects new profiles within 5 seconds and automatically starts the agent.

Create a YAML file in the config directory (e.g., /etc/nobgp/support.yml):

registration-key: "<YOUR_REGISTRATION_KEY>"
router: "router.nobgp.com"
node-name: "web-server-1"

The agent will:

  1. Detect the new profile automatically
  2. Register with the specified network
  3. Remove the registration key from the config file
  4. Start running as a new profile

The agent then takes the file over: it rewrites it with its own recorded state, drops keys whose value is the default (router: "router.nobgp.com" above is one of them — it becomes a commented reference line), and does not preserve comments you wrote. Provision the values you need and read the result back with nobgp show rather than expecting the file to stay as you typed it.

This is useful for temporarily attaching a node to a support or monitoring network without affecting existing profiles.

To remove the profile later:

sudo nobgp remove support

Network Requirements​

The agent requires outbound connectivity only — no inbound ports need to be opened.

DestinationProtocolPurpose
noBGP control planeQUIC (UDP 443)Preferred control plane transport (transport: auto); the endpoint is learned from the router automatically
router.nobgp.com:443WebSocket (wss://)Control plane transport and automatic fallback
downloads.nobgp.comHTTPSPackage downloads and updates

If you're behind a corporate firewall or restrictive network, ensure these destinations are allowed for outbound traffic on port 443 (TCP for the HTTPS/WebSocket destinations, UDP for QUIC). If UDP 443 is blocked, the agent automatically falls back to the WebSocket transport over TCP, so connectivity is preserved — but allowing UDP 443 to the noBGP control plane lets the agent use the faster QUIC transport.

The shared drive needs no destination of its own: the agent mounts it over HTTPS against the same router.nobgp.com:443 endpoint listed above, so a host that can reach the control plane can mount it.

Configuration​

The noBGP agent can be configured through multiple methods:

  1. Web Interface: Manage certain configuration options through the noBGP web application
  2. Configuration File: Local settings stored in a YAML file
  3. Command Line Options: Options passed directly to the nobgp command
  4. Interactive Configuration: Guided setup process
  5. Environment Variables: System environment variables for automation

Configuration File​

The noBGP agent uses a YAML configuration file. Default locations by platform:

  • Linux: /etc/nobgp/default.yml
  • macOS: /usr/local/etc/nobgp/default.yml
  • Windows: C:\ProgramData\nobgp\default.yml

The agent owns this file and rewrites it whenever it records something for itself — and, since agent 0.4.36, once at startup, so a node upgraded from an older agent is converted on its first restart rather than waiting for a write that may never come. Since agent 0.4.35 it writes a key only when its value differs from the default — every other key appears as a commented # key: value line under a short explanatory block, so the file reads as the choices someone made rather than a copy of every default. Uncomment a line to pin that value; leave it commented and the node follows the default, including one that changes in a later release. Comments you add by hand do not survive the next rewrite. See Configuration File for the full shape.

So after a stock registration the settings section may be empty, with the file holding only what the agent recorded — the node's overlay slice, its local MCP port and token, the account captured at install:

overlay-cidr: 100.64.0.0/20
mcp-port: 51843
user: alice
note

Agent identity (keys and tokens) is stored in separate .key and .jwt files alongside the config, not in the YAML file itself. The registration key and node name are automatically removed from the config after successful registration.

tip

The noBGP agent monitors the configuration file for changes. When you modify the configuration file while the agent is running, the new settings will be automatically applied with minimal network disruption.

info

When you have a single profile, the agent uses default.yml in the config directory. The agent also supports multiple profiles (see Multi-Profile Support below).

Configuration Options​

Config File KeyCommand Line OptionEnvironment VariableDescriptionDefault
router-r
--router
NOBGP_ROUTERRouter FQDN, optionally host:port. Connects over wss:// (TLS); legacy wss:///ws:// URLs are accepted and normalized to the bare FQDNrouter.nobgp.com
insecure--insecureNOBGP_INSECUREAllow a plaintext (ws://) control channel — for local or self-hosted routers without TLS onlyfalse
transport--transportNOBGP_TRANSPORTControl-channel transport: auto, quic, or wss. auto negotiates the fastest available transport and falls back to WebSocket automaticallyauto
transport-family--transport-familyNOBGP_TRANSPORT_FAMILYWhich IP address family this node's outbound connections lead with: auto, v4, or v6. auto races both families, exactly as every node always has. v4 or v6 dials that family first and on its own, falling back to the other only if it fails — the way to keep a family off the wire of a machine whose uplink misbehaves on it. A preference, never a force: both values fall back, and an unrecognised value reads as auto. See Address family. Agent 0.4.113+auto
quic-router--quic-routerNOBGP_QUIC_ROUTERQUIC control endpoint (host:port) override for self-hosted routers. Leave unset to use the endpoint learned from the router, falling back to the router's own domain(learned from router)
debug-d
--debug
NOBGP_DEBUGEnable debug modefalse
log-level-l
--log-level
NOBGP_LOG_LEVELLogging level (error, warning, info, debug, trace)info
compress-z
--compress
NOBGP_COMPRESSEnable compressiontrue
encrypt-e
--encrypt
NOBGP_ENCRYPTRequire encryptiontrue
idle-ttl--idle-ttlNOBGP_IDLE_TTLRecycle a peer session after this much idle time with no data traffic (0 = defer to peer). Negotiated end-to-end as the shorter of the two peers' values; clamped to a 1m floor1h
interface-i
--interface
NOBGP_INTERFACENetwork interface to useauto
ping-interval--pingNOBGP_PING_INTERVALInterval for connection health checks58s
mount-m
--mount
Filesystem mount point. From agent 0.4.94 the value is tidied once before anything reads it, so a trailing slash (/mnt/nobgp/) is the same setting as /mnt/nobgp — on agents 0.4.82–0.4.93 the raw form failed the nfs mount's health check against the kernel's own mount table on every poll, so a working drive was torn down and remounted about once a minute indefinitely. A bare Windows drive letter (N:) is left exactly as writtenPlatform-dependent
fsWhich shared drive backend this node mounts with: auto, off, nfs, fuse, winfsp (agent 0.4.62+, Windows) or webdav. auto probes the host and takes the best one available, winfsp included from agent 0.4.82 — so a Windows machine with the WinFsp driver mounts that volume rather than its WebDAV drive, and fs: webdav is what puts it back. Through agent 0.4.81 auto declined winfsp and pinning it was how you asked for it; the reasons came off one at a time, the last two being node-local locks and the cost of creating a file. A pinned backend the host cannot serve is refused and reported, never quietly swapped for another; a value that isn't one of the six is a typo and falls back to auto with a warning. off mounts nothing and stops the local proxy too. Agent 0.4.57+auto
fs-cache-ttl--fs-cache-ttlNOBGP_FS_CACHE_TTLHow long the shared drive may serve a cached directory listing before re-asking noBGP, clamped to [0, 10m]. A staleness window rather than a performance dial: it is how long this node may show a directory as it was before another node's newer write, and equally how often every mount re-asks about a directory nothing changed. 0 re-asks every time and still serves unchanged file contents from the local cache. Requires a unit — 30 means thirty nanoseconds and is ignored as a missing unit. Read by the nfs, fuse and winfsp backends from agent 0.4.67 (nfs alone in 0.4.66), including the timeouts the kernel caches with, which are mount options and so pick a changed value up on the next mount. Agent 0.4.66+30s
nfs-portAgent-written, not a setting you edit: the loopback port the nfs backend's server bound, recorded so a restart binds the same one and an existing mount is not orphaned. Per profile, so two profiles on a machine cannot collide. A remembered port that is taken is replaced with a new one and recorded. Agent 0.4.59+(remembered)
last-mount-pointAgent-written, not a setting you edit: where this profile last had the shared drive actually mounted. It is never mounted on — it is read by one path only, the startup cleanup that removes a mount a killed agent left behind on a node that is not going to mount (fs: off, or no mount point), which is the one case that cannot name the point any other way. A stale value is harmless: only a mount whose source belongs to this agent is ever removed. Agent 0.4.78+(remembered)
user--userNOBGP_USERAccount that command sessions, dispatched bus commands and file operations run as when the caller does not ask to be elevated. Installation records the account that ran the install (on Windows, whoever is signed in at the desktop), so this is typically that user. On a headless or container install where none was captured and the agent runs as root, unelevated calls are refused rather than run as root — set this, or have callers pass admin (unelevated never means root). An account that does not exist on the machine is refused, as is one that resolves to uid 0. The environment variable outranks the config file, which is how a container image or task definition ships a node with an account already named. On Windows the account is the node's second identity, governing file operations first and execution as well from agent 0.4.46 (0.4.44 and 0.4.45 refused an unelevated execution call there)(the installing account)
overlay-cidr--overlay-cidrNOBGP_OVERLAY_CIDROverlay TUN CIDR: a /20 within 100.64.0.0/10 (auto-picked if unset)auto-picked
overlay-ipv6--overlay-ipv6NOBGP_OVERLAY_IPV6Whether this node takes an IPv6 overlay slice beside its IPv4 /20, so peer names resolve to both families. false is the off switch for the IPv6 datapath on this node; the IPv4 overlay and everything else are unaffected. A host that cannot take a slice — IPv6 off in the kernel, or a /64 collision — runs on IPv4 anyway and says why in nobgp status. Agent 0.4.132+true
overlay-cidr6NOBGP_OVERLAY_CIDR6Agent-written, not a setting you edit: the unique local /64 inside fd00::/8 this node picked for its IPv6 overlay, recorded so a restart reuses it and peers keep the addresses they had. A value that is not a /64 in fd00::/8, or one already routed on the machine, is replaced by a fresh pick with a warning. Agent 0.4.132+auto-picked
mtu--mtuNOBGP_MTUOverlay TUN MTU in bytes (clamped to [1280, 8000]; lower for a constrained path)8000
runtime-dirNOBGP_RUNTIME_DIRDirectory for ephemeral runtime files (locks, PID). Point it at the config directory on a mostly read-only rootfs where only the nobgp directory is writablePlatform-dependent (e.g. /run/nobgp)
auto-upgradeSet to false to disable automatic router-driven upgradestrue
windows-firewall--windows-firewallWindows only. allow-overlay adds an inbound Windows Firewall rule that allows all traffic from your noBGP network; respect-windows adds no rule and lets Windows Firewall decide. Applied within seconds, when the agent reloads. See Windows Firewall and traffic from other nodesallow-overlay
lan-names--lan-namesNOBGP_LAN_NAMESWhich host names on this machine's own LAN it resolves for the rest of the network — comma-separated patterns with * and ?, where !pattern excludes and wins whatever the order, and an empty value (or off / none) means no name at all. What this machine's own applications look up is never filtered. See Which LAN names this node serves. Agent 0.4.128+* (any name)
mdns-report-services--mdns-report-servicesNOBGP_MDNS_REPORT_SERVICESWhether this node tells the network which services its own host offers — file sharing, screen sharing, remote login, RDP — so other nodes can show them to their local applications. false is the owner's veto: nothing is looked for and an empty list is reported, which clears any earlier one. See mDNS keys. Agent 0.4.128+true
mdns-announce--mdns-announceNOBGP_MDNS_ANNOUNCEmacOS and Linux: register each peer as <name>-nobgp.local with this machine's own mDNS responder, local-only, so applications here — Finder among them — see peers and the services they report. Nothing is announced on the LAN, and Windows has no local-only mechanism and announces nothing. See mDNS keys. Agent 0.4.128+false

Local MCP server options (see Local MCP Server). What starts the agent's MCP server on 127.0.0.1 — so a local Claude Desktop or Claude Code can drive this node and its peers — is a node_grant from an Owner or Admin at the router. There is no on-device key that turns the endpoint on or off; both keys below are agent-written state, not settings you edit:

Config File KeyDescriptionDefault
mcp-portLoopback port to prefer. The agent binds it when free, otherwise takes one from the OS and records it here, so restarts usually keep the same port(remembered)
mcp-tokenBearer token for that server, minted by the agent on first start and recorded here so registrations written by nobgp mcp install survive restarts(remembered)

An mcp-proxy key left in a profile by an older agent is removed automatically, as is a windows-user-identity key left by agent 0.4.43 — on Windows a configured user is now the whole of the second identity, with no separate switch.

Capability options — the node owner's say over what this machine serves remotely (see Monitoring & Events). They cover MCP callers, event-bus sources and, since agent 0.4.33, published terminal services too, so dropping command closes the browser terminal along with the MCP tools. See Node Access Control.

Config File KeyDescriptionDefault
allow-toolsCapability domains served on this node: fs, command, logs (flag: --allow-tools). Dropping one refuses both its event source and its direct tool, with a reason, rather than going quiet. An explicit [] serves none of them. Values that aren't domains are dropped with a log line and, from agent 0.4.153, if none survive the node serves nothing — it fails closed, where an earlier agent restored the permissive default. nobgp config refuses such a value outright. logs is a domain from 0.4.153, serving node_logs; a profile that already named its domains gets it added once. Renamed from event-sources[fs, command, logs]
allow-rootsDirectories the fs domain may reach on this node, e.g. [/etc, /var/log] — drive-qualified on Windows, e.g. [C:\ProgramData\app] (flag: --allow-roots). Anything outside them is refused, with symlinks resolved first, as is the agent's own configuration directory whatever this is set to. The whole-filesystem value is / (or \; both are accepted on either platform). Renamed from event-watch-roots[/]
allow-rootWhether callers may ask for execution and file access as the superuser (flag: --allow-root, env NOBGP_ALLOW_ROOT). false refuses such a request rather than silently downgrading it, and drops the rest to user. Where there is no unprivileged identity to drop to — Windows, or a root-running agent with no user — false refuses remote execution on the node entirely, since the unelevated path is refused there too. ⚠ allow-admin is its old name, still read, and where both keys are set the stricter value wins — see The veto on root was renamed. Agent 0.4.153 for the new nametrue
local-mcpWhether this node runs its local MCP server at all — on or off (flag: --local-mcp). off and the server never starts, whatever the router grants; the owner's off switch, which no remote path can change. A value the agent cannot parse reads as off. Agent 0.4.153+on
peer-loopbackWhether an overlay peer may reach a port bound only on this node's loopback (flag: --peer-loopback). false keeps a loopback-only service off the network — see Peer reach into loopback. Agent 0.4.153+true
api-socket-groupGroup (name or numeric gid) granted access to the agent's local API socket. Unix only(unset)
api-socket-modeOctal mode for that socket, e.g. "0660". Only meaningful with api-socket-group set0600
webdav-proxy-ownersNumeric uids allowed to reach the agent's local shared drive proxy on 127.0.0.1, which carries the node's credentials. ⚠ Leave it unset. The proxy listens on loopback TCP, where no platform will tell the agent which account is calling, so from agent 0.4.97 a non-empty list refuses every caller and takes the drive off the node — on Linux and macOS as well as Windows. Empty observes rather than allows and refuses nothing. See Shared drive proxy keys. Agent 0.4.53+(unset — observe only)

Registration-specific options (used with nobgp register, cleaned up after registration):

OptionEnvironment VariableDescription
--keyNOBGP_KEYRegistration key for the network
--nameNOBGP_NAMENode name (defaults to hostname)

Using the Register Command​

If you installed via Homebrew or a manual package install (not the install script), register separately:

# Register interactively via OAuth browser login (recommended)
sudo nobgp register

# Register with a specific name and network
sudo nobgp register --name "web-server-1" --network "production"

# Register with a key (for automation — see Headless Registration with Key)
sudo nobgp register --key <YOUR_REGISTRATION_KEY> --name "web-server-1"

After successful registration, the agent installs and starts the system service automatically — no prompt. Exit code 3 means the node enrolled but the service could not be started; fix the cause and run sudo nobgp service install.

Using the Config Command​

Update runtime configuration for a registered agent:

# Set the router (bare FQDN, optionally host:port)
sudo nobgp config --router "custom-router.example.com"

# Enable debug mode
sudo nobgp config --debug

The command saves your settings to the configuration file for future runs.

Environment Variables​

Environment variables override configuration settings without modifying the configuration file. This is useful for automation or containerized environments.

Example:

export NOBGP_DEBUG=true
export NOBGP_LOG_LEVEL="debug"
sudo -E nobgp agent

For key-based registration in automated environments (see Docker and Automation):

export NOBGP_KEY="<YOUR_REGISTRATION_KEY>"
export NOBGP_NAME="api-server-1"
sudo -E nobgp register

Command Line Options​

Override configuration settings with command line options when running nobgp:

sudo nobgp agent --debug --log-level trace

Web Application Configuration​

Manage your noBGP network through the web interface at https://app.nobgp.com:

  • Network creation and management
  • Registration key management
  • Node authorization and discovery
  • Network topology visualization
  • Connection monitoring

Changes made through the web application to your network configuration automatically propagate to your nodes, potentially causing them to restart to apply new settings.

Multi-Profile Support​

The noBGP agent supports running multiple profiles simultaneously, allowing you to connect a single machine to multiple noBGP networks.

How it works:

  • Each profile is stored as a separate YAML file in the config directory (e.g., /etc/nobgp/production.yml, /etc/nobgp/staging.yml)
  • Each profile has its own .key and .jwt credential files alongside the config
  • When running as a service, the process manager automatically starts an agent instance for each profile
  • The default profile is named default (file: default.yml)

Managing profiles:

# Register a named profile (interactive)
sudo nobgp register production --name "prod-server-1"

# Or with a key (for automation)
sudo nobgp register production --key <YOUR_REGISTRATION_KEY> --name "prod-server-1"

# List all profiles
nobgp list

# Show details of a specific profile
nobgp show production

# Show default profile
nobgp show

# Check status of all running profiles
nobgp status

# Remove a profile (deletes .yml, .key, and .jwt files)
sudo nobgp remove staging

Running multiple profiles:

When installed as a service, all profiles in the config directory are automatically managed:

# Install service (manages all profiles)
sudo nobgp service install

# Start all profiles
sudo nobgp service start

# Check status
nobgp status

To run a specific profile manually:

# Run the production profile
sudo nobgp agent production

# Run the default profile
sudo nobgp agent
note

The profile name is specified as a positional argument (e.g., nobgp agent production), not as a flag.

Running the noBGP Agent​

Running as a Standalone Process​

Start the noBGP agent as a standalone process:

sudo nobgp agent

This connects to the router and sets up the required networking components.

For verbose output, add the debug flag or set the log level:

sudo nobgp agent --debug

Run the noBGP agent as a system service for automatic startup. This is recommended for production environments. The service automatically starts on system boot and can be managed using standard system commands.

info

Most service commands do not work in Docker containers — use process management or orchestration tools instead. The exception is nobgp service logs: a containerized agent running in the foreground mirrors its output to an on-disk log file, so nobgp service logs works inside the container (as does docker logs).

The nobgp service subcommand names are consistent across platforms (the underlying service manager differs — systemd on Linux, launchd on macOS, Windows SCM on Windows):

# Install the service
sudo nobgp service install

# Check service status
sudo nobgp service status

# Start the service
sudo nobgp service start

# Stop the service
sudo nobgp service stop

# Restart the service
sudo nobgp service restart

# Uninstall the service
sudo nobgp service uninstall

On Windows, run these commands in an Admin PowerShell session (without sudo). The agent is registered as a Windows Service and managed through the Windows Service Manager.

On macOS, the agent is registered as a launchd daemon.

On Linux, the agent is registered as a systemd service.

When running as a service, the noBGP agent uses the settings from your configuration file.

Verifying Connection​

Once the agent is running, verify it's connected:

  1. Check service status:

    sudo nobgp service status
  2. Ask your AI assistant:

    Show me all nodes in my default network

You should see your node listed as "online".

System Requirements​

The noBGP agent requires:

  • Operating System:
    • Linux: Alpine Linux 3.16+, Arch Linux, Amazon Linux 2023+, Debian 11+ / Ubuntu 20.04+, OpenWRT
    • Buildroot and other Linux images with no package manager, from agent 0.4.99 (amd64, arm64, armv6 or armv7)
    • Synology DSM (amd64 or arm64; DSM 7 recommended for service auto-install)
    • macOS 13+ (Ventura and later), amd64 or arm64 (Apple Silicon)
    • Windows 10/11 or Windows Server 2019+, amd64 or arm64
  • Privileges: Root or sudo access (Linux/macOS); Administrator (Windows)
  • Network: Connectivity to the noBGP router
  • Architecture: amd64 (x86_64) or arm64 (aarch64) on all platforms; armv6/armv7 on Debian/Ubuntu (Raspberry Pi)
  • Resources:
    • 512MB RAM minimum
    • 1GB disk space
    • Network interface with IP connectivity

Usage​

When the noBGP agent runs on your system, you can access any other node on the same network by simply using their name. For example, if you have nodes named web-server and api-server in the same network, they can reach each other using those names.

Your AI assistant can also:

  • Run commands on the node
  • Start interactive shell sessions
  • Publish services running on the node
  • Monitor system status
  • Transfer files

Troubleshooting​

If you encounter issues:

  1. Enable debug mode: sudo nobgp agent --debug
  2. Check the logs (all platforms, needs root — an Administrator terminal on Windows): sudo nobgp service logs -f, or platform-specific:
    • Linux (systemd): sudo journalctl -u nobgp.service -f
    • Linux (Alpine/OpenRC): tail -f /var/log/nobgp.err
    • macOS: tail -f /var/log/nobgp.err.log
    • Windows: Get-Content -Path C:\ProgramData\nobgp\nobgp.log -Tail 50 -Wait (the agent mirrors its logs to this file, since the Windows Service Control Manager does not capture service output)
  3. Check registration status: nobgp show
  4. Verify the configuration file exists in the config directory
  5. Check that the router URL is accessible from your machine
  6. Verify that you have root/administrator privileges when running the noBGP agent

Common Issues and Solutions​

IssuePossible CauseSolution
Connection failedNot registeredRun sudo nobgp register
Service won't startMissing permissionsCheck sudo access
Node not visibleFirewall blockingCheck router access
Agent crashesIncompatible OSVerify OS version support
High CPU usageDebug mode enabledDisable debug in production
Windows CLI command refuses to runWindows does not self-elevateRe-run the named command from an Administrator terminal (Root privileges)
macOS service failsSIP or permissionsCheck /var/log/nobgp.err.log
macOS agent works but nothing starts at bootService was never installed by an older buildUpgrade to 0.4.21+, then sudo nobgp service install (see below)
Buildroot-class node registers, but the service install fails on a symlink into /etc/rc.dAgent 0.4.99 did not recognise busybox initInstall 0.4.100+, then sudo nobgp service start (see below)
Buildroot appliance runs fine, but the agent never comes back after a rebootOn an overlay-root appliance, 0.4.100 put the boot link in a directory that machine's boot does not readUpgrade to 0.4.101+, then sudo nobgp service restart (see below)
Node never comes back after a reboot or a Wi-Fi drop, and the agent exits about once a minuteAgent 0.4.106 and older treated a host with no usable network interface as a fatal startup errorUpgrade to 0.4.107+ (see below)
Node flaps online/offlineDuplicate agent installStop/uninstall the redundant service (see below)
Node deauthorizedCredentials rejectedRe-enroll with sudo nobgp register (see below)
Node stays offline and nobgp status names an expired tokenThe registration token expired while the node was offline, and an expired token cannot refresh itselfRe-enroll with sudo nobgp register on agent 0.4.72+ (see below)
Agent version no longer supportedAgent older than 0.3.51Reinstall the current agent (see below)
nobgp register or nobgp login opens a page that says the sign-in was refusedAgent older than 0.4.153, which noBGP's routers cannot tie to the browser pageUpgrade to 0.4.153+, or register with a key (see below)
Certificate error during install or upgradeAntivirus HTTPS scanning or a corporate proxy re-signing the connectionExclude the download and router domains from inspection (see below)
Windows drive keeps disconnecting and reappearingMount health checked from the wrong logon session by an older buildUpgrade to 0.4.31+ (see below)
Shared drive is empty on Linux, node otherwise healthyThe FUSE userspace helper is not installedInstall fuse3 (fuse-utils on OpenWrt) and restart the agent (see below)
Node online, but every command and file call fails with a DNS errorThe host's resolver is broken and some of the agent's connections could not fall back to public DNSUpgrade to 0.4.81+ (see below)
Node is online and healthy but no peer is reachable, on a Linux machine also running Tailscale or another overlay VPNThe other product claims the overlay's 100.64.0.0/10 range in a routing table agents through 0.4.110 did not readUpgrade to 0.4.111+, then free space in 100.64.0.0/10 on that machine (see below)
Node is online and answers command and file, but has no TUN interface and no peer trafficIt could not pick an overlay /20 — nobgp status names the reason under network.overlayFree space in 100.64.0.0/10, or pin overlay-cidr to a free /20, then restart (see below)
Agent exits during startup, over and over, logging no free /20 slice available in 100.64.0.0/10Agent 0.4.111 treated an unpickable overlay slice as a fatal startup errorUpgrade to 0.4.112+, which comes up without an overlay and stays reachable (see below)
In a container, the node answers command and file but has no TUN interface and no peer trafficThe container has no NET_ADMIN capability or no /dev/net/tun — nobgp status names it under network.overlayAdd --cap-add NET_ADMIN --device=/dev/net/tun and restart the container (see below)
Container starts and exits again seconds later, over and over, with no NET_ADMINAgents through 0.4.141 treated a refused tunnel adapter as a fatal startup errorOn Linux, upgrade to 0.4.142+, which comes up without an overlay and stays reachable. On macOS and Windows a refused adapter still fails startup — grant the adapter (see below)
rm -rf on the drive deletes nothing and says "Directory not empty"A davfs2 (webdav) mount's lock token was refused by the deleteUpgrade to 0.4.79+ (see below)
macOS drive never returns after changing fs:The previous backend left an empty directory at /Volumes/nobgp, which the next mount reads as a mount pointsudo rmdir /Volumes/nobgp, then sudo nobgp service restart (see below)

Detailed Troubleshooting​

Agent Won't Start​

Symptoms: Service fails to start or exits immediately

Solutions:

  • Check system requirements are met
  • Verify the agent is registered: nobgp show
  • Look for error messages in logs
  • Try running in standalone mode with --debug flag

Agent Exits Repeatedly While the Host Has No Network​

Symptoms: The service starts and exits again a few seconds later, over and over, at a steady rate of roughly 60 times an hour. The log shows the same startup failure each time — no interface found, no addresses found on interface <name>, or configured interface not found: <name>. It typically begins after a reboot, or after the machine's Wi-Fi or cellular link dropped while the agent happened to be restarting.

Cause: Through agent 0.4.106, a host with no usable network interface failed agent startup. The condition is usually temporary — the link is not up yet at boot — but the exit is not: the service manager restarts the agent, the next start runs the same probe, and the cycle continues for as long as the link is down. Measured on a Buildroot appliance whose Wi-Fi had died: 1,991 identical exits over 31.5 hours.

Solutions:

  • Upgrade to 0.4.107+, where the agent waits for a usable interface instead of exiting. It re-probes on a widening interval (2 seconds, doubling to a 30-second ceiling), keeps its process, its log and its pid file, and attaches the moment a link appears. See Waiting for an interface at startup.
  • Bring the link back — this is still a machine with no network, and the agent cannot reach the router until it has one.
  • If the log says configured interface not found, the node has an explicit interface: set to a name the host does not have. Check it with nobgp show and either correct it or set it back to auto. From 0.4.107 the agent waits for that name indefinitely rather than exiting, so a typo shows up as a node that never connects.
  • While the agent is waiting, nobgp status reports the profile as unavailable: the local API socket is opened later in startup, so the log is the only place the wait is visible.

A Container Gets No Tunnel Adapter​

Symptoms: two shapes, depending on the agent version and the platform.

  • Agent 0.4.141 and earlier. The container starts and exits again seconds later, over and over, and the node never appears online.
  • Agent 0.4.142 and later, on Linux. The node registers and stays online. command, file, the MCP tools and nobgp status all answer, but it has no tunnel interface, no overlay DNS and no peer traffic — nothing reaches it by its overlay name, and it reaches no peer by theirs. network.overlay reads no permission to create a TUN device or this host has no TUN device, and network.dns reads unavailable.
  • On macOS and Windows the first shape is still the one you get, from agent 0.4.143: a refused tunnel adapter fails startup there. The refusal can be temporary on those platforms — an antivirus holding the adapter, or one still being removed after an upgrade — and the agent does not try the adapter again while it runs, so a machine that carried on without one would sit with no overlay until somebody restarted it. Agent 0.4.142 alone came up without an overlay on all three platforms.

Cause: the overlay needs a tunnel adapter, which needs the NET_ADMIN capability and the /dev/net/tun device. Several container platforms withhold one or both by default — AWS Fargate, Cloud Run, Azure Container Instances, and Kubernetes pods under a restrictive PodSecurity profile. On Fargate the device can be opened and only the configuration call is refused; under the PodSecurity restricted profile there is no device at all.

Solutions:

  • Grant both where you can. --cap-add NET_ADMIN --device=/dev/net/tun on docker run, or securityContext.capabilities.add: [NET_ADMIN] plus a /dev/net/tun device on a pod. Restart the container afterwards. See Docker and Automation.
  • Where a Linux platform will not grant them, decide whether a node without an overlay is useful to you. It is a real node: it is listed in the network, it runs command and the file tools, and it can use the shared drive — none of which go over the overlay. What it cannot do is carry traffic to or from a peer.
  • ⚠ Upgrade before you diagnose this on an older agent. Through 0.4.141 the restart loop left nothing to read: the node never reached the point of having a status to report, so the log on the machine was the only evidence, and on a managed container platform that is often the one thing hardest to reach.

Node Shows Offline in Dashboard​

Symptoms: Agent is running but node appears offline

Solutions:

  • Check firewall rules allow websocket connections
  • Verify router URL is accessible: curl https://router.nobgp.com
  • Verify registration status: nobgp show
  • Restart the agent: sudo nobgp service restart

Node Repeatedly Flaps Online/Offline​

Symptoms: The node cycles between online and offline every few seconds to a minute. The agent logs may show a rejection reason such as "re-registered more than 8 times in 10m0s while its existing connection is healthy".

Cause: Two agent processes are running with the same identity on one host. This happens when a machine ends up with more than one live install — a duplicate install (for example, both a package install and a script/manual install left a running service), an orphaned pre-upgrade process that was never stopped, or a legacy nobgp-agent.service unit from a pre-0.3 install still enabled alongside the current nobgp.service. Each process evicts the other's router connection on connect, so the node flaps.

Two safeguards catch this: the agent takes a per-profile single-instance lock, so a second process on the same host fails fast with an error like profile "default" is already running on this host … duplicate install or orphaned instance; and the router rejects a node that re-registers too rapidly while its existing connection is healthy, sending the reason in a close frame that lands in the rejected agent's local logs.

Solutions:

  • Check for more than one running noBGP process/service on the host and stop or uninstall the extra one (sudo nobgp service uninstall for the redundant install):
    • Linux (systemd): systemctl status nobgp.service and look for stray nobgp agent processes (ps aux | grep nobgp)
    • macOS: sudo launchctl list | grep nobgp
    • Windows: check for a duplicate nobgp service in Services and any manual nobgp agent processes
  • If you recently upgraded, restart the service so the old process is fully replaced: sudo nobgp service restart
  • On Linux, a leftover legacy nobgp-agent.service unit is disabled and removed automatically the next time the agent upgrades or the service restarts — no manual action needed.
  • On Windows, only one nobgp service can exist, so if it detects any other nobgp supervisor process started outside it (for example a Task Scheduler job, a Startup-folder shortcut, or a leftover console run) it logs a warning naming that process and the autostart to remove. Delete that autostart source so it does not return on the next reboot.
  • Confirm each host registers under a unique node name. ⚠ Two different machines configured with the same name no longer flap and are no longer refused, from router 0.4.113: the second one enrols under a suffixed name — web-1 — beside the first, and the machine goes on calling itself web locally, so the extra node appears only in noBGP. See A contested name costs the name, not the connection. What this section describes is the other case, two processes sharing one identity on one host, which still flaps and is still what the flap guard refuses.

macOS: Agent Works, but No Service Was Installed​

Symptoms: On macOS, registration completes and the agent runs fine when you start it by hand (sudo nobgp agent), but sudo nobgp service status reports the service is not installed and the node does not come back after a reboot.

Cause: Older agent builds wrote a launchd property list that launchd rejected when the service was installed, so the install step logged the failure and continued — leaving a working agent with no launchd daemon behind it. Fixed in 0.4.21.

Solutions:

  • Upgrade the agent, then install the service explicitly:
    sudo nobgp upgrade
    sudo nobgp service install
    sudo nobgp service start
  • Confirm with sudo nobgp service status (exit code 0 means running) and sudo nobgp service logs.
  • If the agent is too old to upgrade itself, reinstall it with the install script or brew upgrade nobgp.

macOS: The Node Went Offline and Stayed Offline Until Someone Restarted It​

Symptoms: A Mac that is up, awake and reachable the whole time shows its node offline for hours, and sudo nobgp service status reports the service stopped. Nothing in /var/log/nobgp.err.log names a failure — the log simply ends. Starting the service by hand brings the node straight back. Agents 0.4.107 and earlier.

Cause: The agent exited 0 for every shutdown, whatever caused it. The launchd job restarts the agent on an unsuccessful exit only — which is what makes nobgp service stop stick — so a stop nobody asked for was indistinguishable from one somebody did, and launchd left the node down. Measured on a Mac: the agent took a signal it had not been asked for, exited 0, and stayed stopped for 8 hours on a host that was reachable throughout.

Solutions:

  • Upgrade to agent 0.4.108+ (sudo nobgp upgrade). The agent now exits non-zero when nothing asked it to stop, so launchd restarts it about 8 seconds later, and it logs this stop was not asked for by any nobgp command; the exit code will say so before it goes.
  • ⚠ A habit worth changing on macOS from 0.4.108: sudo launchctl stop nobgp, launchctl kill TERM system/nobgp and sudo pkill -TERM nobgp leave the job loaded, so they no longer stop the node — it comes back seconds later. Use sudo nobgp service stop, which unloads the job.
  • On an older agent, sudo nobgp service start brings the node back; there is nothing to configure, and the fix ships with the binary rather than with the launchd job, so it takes effect on the next upgrade with no reinstall.
  • Linux and Windows never had this: systemd's Restart=always and OpenWrt's respawn restart a clean exit already.

Buildroot Node Registers, but the Service Will Not Install​

Symptoms: On a Linux image with no package manager — appliance or router firmware built with Buildroot, running busybox init — the install script gets as far as registration and then fails the service step:

failed to install service: symlink /etc/init.d/nobgp /etc/rc.d/S50nobgp: no such file or directory

/etc/init.d/nobgp exists afterwards, but nothing starts it, then or at the next boot.

Cause: Agent 0.4.99 shipped the packageless install without handling busybox init, so it tried to start the agent at boot through a directory of symlinks (/etc/rc.d) that these images do not have. Fixed in 0.4.100, which drives the init script by path and installs the S99nobgp link busybox's rcS actually reads.

Solutions:

  • Re-run the install script to put 0.4.100 or later on the machine. The node is already registered, so it keeps its identity.
  • Then start it: sudo nobgp service start. That command also repairs the missing boot link on 0.4.100+, so a machine half-installed by 0.4.99 is corrected without reinstalling the service — as does sudo nobgp service restart, which is what the agent's own upgrade runs.
  • Confirm with sudo nobgp service status, and check that the boot link points at /etc/init.d/nobgp — ls -l /etc/init.d/S99nobgp on an ordinary machine, or see the entry below for an appliance with an overlay root.

Symptoms: On a Linux image with no package manager the agent installs, registers and runs. sudo nobgp service status reports it running, starting and stopping it by hand works, and ls -l /etc/init.d/S99nobgp shows a link that is present, executable and pointing at /etc/init.d/nobgp. Then the machine reboots and the node never comes back — with nothing in sudo nobgp service logs and no error anywhere, because the link was never read.

Cause: The appliance's root filesystem is an overlay, and the firmware script that mounts it runs from inside rcS's own loop:

for i in /etc/init.d/S??* ;do ... done

A shell expands that list once, before the first iteration — so the loop is working from the firmware's /etc/init.d as it stood before the writable layer existed. Every script the firmware ships is in the list and starts; anything installed afterwards lives only in the writable layer, is in no list, and runs nothing. Agent 0.4.100 installed a correct link into a directory this machine's boot does not enumerate. Fixed in 0.4.101, which puts the link in the directory the firmware's own /etc/init.d/S99custom hook enumerates after the overlay is mounted.

Solutions:

  • Upgrade the agent to 0.4.101 or later: sudo nobgp upgrade. There is nothing to reinstall and no identity to re-enrol.
  • Then sudo nobgp service restart (or start, if it is down). Both repair the boot link before running anything on 0.4.101+, which is what moves it into place.
  • Confirm the link is in the hook directory: ls -l /etc/kvmd/user/scripts/S99nobgp should point at /etc/init.d/nobgp.
  • Reboot and check the node comes back online.
  • The link 0.4.100 left at /etc/init.d/S99nobgp can be left alone — that directory is precisely what this machine's boot does not read, so it does nothing. sudo nobgp service uninstall removes it along with the live one.

Appliance: The Agent's Log Starts Empty After Every Reboot​

Symptoms: On appliance or router firmware, sudo nobgp service logs and node_logs work perfectly while the machine is up, and hold nothing older than the current boot. Whatever the node was doing when it went down — the panic, the disconnect, the failed mount — is gone, and gone precisely when you go looking for it. ls -l /var/log shows the file the init script writes and it is as old as the boot, not as old as the install.

Cause: the firmware keeps /var/log in RAM — a tmpfs, or a symlink into one. The service manager captures the agent's output faithfully, into a file the next boot erases. Measured on a GL.iNet KVM appliance: a live log 82 KB long and nine hours old, and one reboot erased all of it.

Solutions:

  • Upgrade the agent to 0.4.103 or later (sudo nobgp upgrade). Such a node then also writes the agent's own durable mirror at <config-dir>/nobgp.log — capped at 10 MiB with one rotated generation, so about 20 MiB at most — and reads it first, ahead of the init script's own file. Nothing to configure: the filesystem is measured at startup, and a machine whose /var/log is durable, or cannot be measured, is left exactly as it was.
  • The change is Linux only, and it takes effect on the agent's next start. History begins accumulating from then — it cannot recover what earlier boots already erased.
  • Both readers follow it: sudo nobgp service logs and node_logs consult one description of where this node keeps its log, so node_logs reports the mirror's path in path and there is nothing to switch by hand.
  • The init script's own file is still written and still offered, one place further down — it remains the current boot's log and the only place a message from before the agent's own logger started can appear.
  • ⚠ Prefer agent 0.4.152 on a machine whose service is started by an init script — busybox or sysv init, which is what this firmware class uses. Through 0.4.151 the mirror stopped being written at the first restart the agent made for itself: an auto-upgrade, or a nobgp service restart issued through command. The file kept everything up to that moment and grew no further, so sudo nobgp service logs and node_logs answered with pre-upgrade lines and nothing since — on the machines where this mirror is the only log that survives a reboot. Starting the service from a clean shell, or rebooting, restored it until the next upgrade. Machines whose service manager is systemd or procd were never affected.
  • ⚠ Prefer agent 0.4.105 if the machine also has a ring buffer — the systemd journal, or logread on OpenWRT. Agents 0.4.103 and 0.4.104 read the mirror ahead of every other source, ring buffers included, and the mirror starts empty at the upgrade that begins writing it: the first read after upgrading could reach the bottom of a nearly empty file and answer truncated: false with no cursor, presenting a short page as the whole log at exactly the moment the log matters. It filled in on its own over the following hours, which is what made it easy to miss. 0.4.105 lifts the mirror above the /var/log files only, and a ring buffer holding none of this agent's lines is skipped exactly as before.

Certificate Error During Install or Upgrade​

Symptoms: The install script or nobgp upgrade stops with a certificate validation failure against downloads.nobgp.com or your router domain, on a machine whose browser reaches the same sites fine.

Cause: Something on the machine or the network is re-signing HTTPS connections — consumer antivirus with a "web shield", or a corporate inspection proxy. The replacement certificate is not one the agent trusts.

Solutions:

  • The installers and the agent name the certificate's owner in the error, and inspection products identify themselves there (Norton's certificate literally says it was generated for SSL/TLS scanning). That name tells you which product to configure.
  • Exclude downloads.nobgp.com and your router domain from HTTPS/SSL scanning, or ask IT to add them to the proxy's inspection bypass list.
  • Full per-vendor instructions are in Security software and VPNs.
  • On a machine already running the agent, nobgp status reports the same finding continuously as router.tls_intercepted — see Environment interference.

Windows: Network Drive Keeps Disconnecting​

Symptoms: On Windows, the noBGP drive letter disappears from Explorer and comes back every 30 seconds or so, and nobgp status reports fs.mounted: false even while file operations succeed.

Cause: Windows drive mappings belong to the logon session that created them. Once the agent moved the mapping into the interactive desktop session so Explorer could see it, its own health check — running as the SYSTEM service — could no longer see the mapping, judged the mount dead, and remounted it, on repeat. Fixed in 0.4.31.

Solutions:

  • Upgrade the agent from an Administrator terminal (nobgp upgrade), then nobgp service restart.
  • If the drive is mounted on a different letter than the one you configured, that is deliberate and unrelated: when the configured letter is already taken by something that is not a noBGP mount, the agent leaves it alone and mounts on the first free letter instead. nobgp status reports the letter actually in use under fs.mount.

Windows: nobgp Says "Access is denied" for a Non-Administrator​

Symptoms: On a Windows node, nobgp runs normally from an Administrator terminal but answers an ordinary account with a bare Access is denied — before any command runs, and with no mention of elevation. The agent itself is fine: the service is running, the node is online, and every remote call works. The machine is one that has taken at least one agent upgrade; a freshly installed node behaves.

Cause: the upgrade moved the new binary into C:\Program Files\nobgp from a temporary folder. An NTFS move carries the file's permissions across verbatim and Windows never re-evaluates them against the new parent, so the replaced nobgp.exe kept the temporary folder's permissions and never picked up the install directory's grant to BUILTIN\Users. The tell is that wintun.dll, placed once at install time and never replaced by an upgrade, sits in the same directory and is readable by ordinary users. Fixed in agent 0.4.97, which stages the new binary inside the install directory so it inherits that directory's permissions.

Solutions:

  • Upgrade the agent to 0.4.97 or later from an Administrator terminal (nobgp upgrade). The upgrade repairs an already-broken install rather than preserving it, so the next one puts the permissions right on its own.
  • Or re-run the install script from an Administrator PowerShell, which copies the binary into place the same way.
  • Until either has run, use nobgp from an Administrator terminal.

Windows: Creating a File on a winfsp Drive Says "Access is denied"​

Symptoms: On a Windows node pinned to the winfsp backend, the drive mounts and lists and reads perfectly, but every attempt to create a file on it is refused with Access is denied — from an Administrator prompt, from a scheduled task, from Explorer, from the agent itself. Nothing appears in the agent log, because nothing reached the agent. Agents 0.4.63 and 0.4.64.

Cause: Windows checks access against a security descriptor that WinFsp builds from what the filesystem reports about ownership, before any of the agent's code runs. The share has no per-file owner to report, and what the volume did report granted full access to an account identifier that no token on the machine holds — so the only permissions any real caller matched were read and execute. Reads passed, creates did not, for everyone.

Solutions:

  • Upgrade the agent to 0.4.65 or later from an Administrator terminal (nobgp upgrade), then nobgp service restart. The mount is given an explicit descriptor granting full access to LocalSystem, Administrators and Authenticated Users.
  • Make sure WinFsp is 1.10 (2022) or newer. Older builds reject the option that carries the descriptor, which refuses the mount outright — visible as fs.error in nobgp status through agent 0.4.72. From agent 0.4.73 the probe checks the version first, so such a host reports winfsp: unavailable (WinFsp <version> is installed at … but this backend needs 1.10 or newer …) in the backends list and the backend is not selected at all.
  • To get a working drive back on an older agent, set fs: webdav in the profile and restart the agent: the node returns to its WebDAV mount, which is read-write. ⚠ Set it explicitly — removing the pin is no longer enough. Through agent 0.4.81 an absent fs meant WebDAV on Windows; from 0.4.82 auto selects WinFsp wherever the driver is present, so clearing the pin on a current agent lands you back on the backend you were trying to leave.

Windows: Appending to a File on a winfsp Drive Overwrites It​

Symptoms: On a Windows node pinned to the winfsp backend, Add-Content, >> or any other append to a file the drive has recently written replaces the file's contents instead of adding to them. Every command reports success, and the bytes upload correctly — they are simply written at the wrong offset. Agents 0.4.67 and 0.4.68, the two releases where the volume both wrote and deferred its uploads. webdav, which every unpinned Windows node mounts, is not affected.

Cause: Windows asks for a file's size by name as well as through an open handle, and computes an append's write offset from the answer. The by-name answer came from the cached directory listing, which still held the size from before the write — 0 for a file just created — so the append landed at offset 0.

Solutions:

  • Upgrade the agent to 0.4.69 or later from an Administrator terminal (nobgp upgrade), then nobgp service restart. A by-name size is answered from the open handle while one is open, and from the pending upload after the last handle closes. See Appending to a file the volume just wrote.
  • On an older agent, set fs: webdav in the profile and restart: the node returns to its WebDAV mount. ⚠ Not "remove the pin" — from agent 0.4.82 an absent fs selects WinFsp on Windows wherever the driver is installed.
  • Either way, write a file whole rather than appending to it on this backend, or write it through the file tools or the share's URL, until you are on 0.4.69 or later. The fix was confirmed on real hardware on 2026-08-11, and from agent 0.4.73 nobgp status stopped naming it among the reasons auto declined winfsp — reasons that are gone entirely from 0.4.82, where auto selects it. Two reasons remained from agent 0.4.75 — node-local locks and a slow create — and both are settled: the create cost was measured and WinFsp won it (402 ms/op against WebDAV's 600 on the same machine), and node-local locks stayed a limitation but stopped being a reason, since the WebDAV drive they held the fleet on reports a lock it does not enforce across nodes. auto takes the backend from 0.4.82.

macOS: Duplicate noBGP Volumes After Sleep​

Symptoms: After the Mac wakes from sleep, the Computer view in Finder shows two nobgp volumes, one of which is dead. The agent logs an unmount that timed out or was refused around the same time.

Cause: The mount's health check reaches the router over the network, so a wake — or any brief connection reset — could fail the check on a mount that was perfectly healthy. The agent then tried to tear it down, the unmount was refused because the mount was live, and a second volume was mounted over the top, leaving Finder holding a row for a volume nobody unmounted. Fixed in 0.4.48: a mount must fail three consecutive checks, about a minute apart in total, before it is remounted, and the unmount on shutdown is given long enough to complete.

Solutions:

  • Upgrade the agent (sudo nobgp upgrade), then sudo nobgp service restart.
  • Clear an existing ghost by ejecting it in Finder, or with sudo diskutil unmount force /Volumes/nobgp. A dead row also disappears at the next reboot.
  • nobgp status reports the mount the agent believes in under fs.mount and fs.mounted, which is the one to trust.
Use diskutil, not umount -f, on macOS

If sudo umount -f answers Operation not permitted — as root, with nothing open on the volume — that is macOS objecting to the volume going away rather than the mount being busy, and -f only overrides busy. sudo diskutil unmount force /Volumes/nobgp goes through the machinery that raised the objection and clears it. From agent 0.4.61 the agent escalates the same way on its own.

Shared Drive Times Out After an Agent Restart​

Symptoms: On a node mounting over nfs, every operation on the drive hangs after the agent restarts or upgrades and then answers Operation timed out. nobgp status shows fs.type empty and fs.mounted: false while fs.backends still lists nfs: available — the node saying both that the backend it wants can serve and that it is not serving. On macOS a second nobgp volume may appear per restart. On 0.4.59 the same node reports fs.error: nfs: mount point /Volumes/nobgp: mkdir /Volumes/nobgp: file exists.

Cause: The agent's NFS server took a fresh loopback port on every start, while the kernel remembers the port the existing mount was made against. After a restart the mount pointed at a port nothing was listening on, and that leftover then held the mount point so the replacement mount failed too — a node retrying indefinitely with no way to repair itself. Agents 0.4.57 and 0.4.58 only, and only on the nfs backend.

Fixed in 0.4.59: the port is remembered per profile as nfs-port, so a restart binds the same one and the existing mount keeps working, and a mount left behind at the node's own mount point is cleared before mounting rather than stacked on top of.

Completed in 0.4.60. On 0.4.59 that clearing ran after the mount point was created, and a wedged mount makes its own mount point unanswerable — so creating it failed first and the clearing never ran, which is exactly the case it was written for. A node already stuck this way stayed stuck on 0.4.59 and needed an unmount by hand; from 0.4.60 the clearing runs first and the node repairs itself on its next mount cycle.

Completed on macOS in 0.4.61. On 0.4.60 the clearing ran at the right moment and could still be refused: a macOS unmount answers Operation not permitted — to root, with nothing holding the volume — when something has objected to the volume going away, and umount -f only overrides busy, which was never the problem. So a Mac already in this state stayed in it, every retry failing identically. From 0.4.61 the agent falls through to diskutil unmount force, which does override it, on both the shutdown path and the clearing.

Solutions:

  • Upgrade the agent (sudo nobgp upgrade), then sudo nobgp service restart. On 0.4.60 and later that is enough on its own on Linux, and on 0.4.61 and later on macOS too — an already-wedged mount point is cleared on the next mount cycle, within about a minute.
  • On earlier agents, clear a mount already stuck this way by hand — sudo diskutil unmount force /Volumes/nobgp on macOS, sudo umount -l /mnt/nobgp on Linux. Several mounts can be stacked at the same point, so repeat the command until it reports nothing is mounted there. The agent mounts again on its next cycle.
  • On 0.4.59 and later, nobgp status names the reason there is no mount under fs.error, which is the field to read before assuming this is it. From 0.4.61 nobgp show answers the same question in its own output, as fs-backend and fs-error.

Linux: Shared Drive Is Empty and Never Mounts​

Symptoms: On a Linux node, /mnt/nobgp is an empty directory, nobgp status reports fs.mounted: false, and everything else about the node is healthy. On agents before 0.4.51 the only sign in the logs is FUSE mount failed, will retry repeating once a minute, with no cause named.

Cause: A FUSE mount needs both halves — the kernel module and the userspace fusermount helper binary — and only the kernel half was declared by the packages. The kernel half being present is what makes the missing one invisible: /dev/fuse exists, so the agent correctly takes the FUSE path rather than falling back to WebDAV, and then cannot complete the mount. OpenWrt was the platform this actually happened on, because its packages are installed by install.sh rather than resolved by opkg, and neither the module nor the helper was being installed there.

Fixed in 0.4.51: every package declares the helper, the OpenWrt branch of the install script installs it explicitly, and a mount that fails for this reason is reported once as an error naming the package to install instead of as an endless retry warning.

Solutions:

  • Read nobgp status first (agent 0.4.57+). Its fs block names the backend in use, why there is no mount under fs.error, and — under fs.backends — every backend this platform could use with the reason each can or cannot serve. That list is what tells you which package to install, and it may say the node has already mounted over NFSv4 instead, which needs nothing installed. See Which filesystem you get.
  • Install the helper for your platform, then restart the agent:
    # OpenWrt
    opkg install kmod-fuse fuse-utils

    # Debian / Ubuntu
    sudo apt install fuse3

    # Alpine
    apk add fuse3

    # Amazon Linux / RHEL / Oracle / Rocky / AlmaLinux
    sudo dnf install fuse3

    sudo nobgp service restart
  • Installing an NFSv4 client is an alternative to the FUSE helper, not only a fallback behind it — the node then mounts over NFS instead:
    # Debian / Ubuntu
    sudo apt install --no-install-recommends nfs-common

    # Amazon Linux / RHEL / Oracle / Rocky / AlmaLinux
    sudo dnf install nfs-utils

    # Arch Linux
    sudo pacman -S nfs-utils

    # Alpine
    apk add nfs-utils

    sudo nobgp service restart
    From agent 0.4.58 install.sh does this for you on Debian/Ubuntu, RPM hosts and Arch, so a node installed with that script or later already has it. Alpine, OpenWrt and Synology are deliberately left out. From agent 0.4.70 none of these commands is needed on Linux at all: with no mount helper installed the agent makes the NFS mount itself, so a host whose kernel has an NFSv4 client can take the backend with nothing installed. If nobgp status still reports nfs as unavailable on that release, the missing half is the kernel's client, and the message says so.
  • sudo nobgp upgrade does not install the NFS client. No package declares it as a dependency — apt-get install would resolve that without ever running our installer, and it is the installer that confines the rpcbind the client drags in. So a node upgraded in place stays on whichever backend it can already serve; re-run the install script, or install the client by hand, to move it to NFS. On agent 0.4.70+ an in-place upgrade can move a node to NFS without any package, since only the kernel's own client is required there.
  • Upgrade to agent 0.4.75 if a node has been sitting in this state. Until then each failed mount attempt left a background timer and a cleanup task behind it, and the retry runs every 30 seconds — so a node stuck this way accumulated roughly 2,880 of them a day, which on a router or a Pi-class box is memory it does not have. The node still could not mount, but it stopped paying for trying. It was the fuse backend alone; nfs and winfsp already cleaned up after a failed mount.
  • Upgrade the agent (sudo nobgp upgrade) so a future recurrence names itself. On most package managers the upgrade also pulls the FUSE helper in, since the packages now declare it. RPM hosts are the exception: there it is a recommendation rather than a requirement, because EL7-class hosts (Amazon Linux 2, RHEL/Oracle/CentOS 7) have no fuse3 in their default repositories at all and a hard requirement would fail the whole install over an optional filesystem. dnf honours recommendations; yum and rpm ignore them, so those hosts still need the command above.
  • The drive is optional. A node without it is fully functional otherwise — only the shared drive is unavailable, and the nobgp file commands fall back to the API automatically.

Linux: NFS Client Installed and the Drive Still Never Mounts​

Symptoms: On a Linux node — in practice an embedded one, OpenWrt especially — the NFS client package is installed, nobgp status reports nfs: available, and the drive never comes up. fs.type is empty, fs.mounted is false, and the mount fails on every cycle without falling back to FUSE or WebDAV. Agents 0.4.61 and earlier.

Cause: An NFS client is two separate pieces, and on those distributions they are two separate packages: the userspace mount helper (nfs-utils, ~56 KB, which puts /sbin/mount.nfs4 on PATH) and the kernel's own NFSv4 client (kmod-fs-nfs-v4). The probe checked only the helper, so the node advertised a backend it could not mount with — and because a mount failure deliberately never falls through to another backend, it was left with no shared drive at all rather than the FUSE one it could have had.

Fixed in 0.4.62: the probe asks the kernel too, so such a host reports nfs as unavailable and auto moves on to FUSE. nobgp status names which half is missing, since the remedies differ — "install the NFS client" is useless advice to someone who just did.

Solutions:

  • Install the kernel half as well, matching the running kernel, then restart the agent:
    # OpenWrt
    opkg install kmod-fs-nfs-v4

    nobgp service restart
    kmod-fs-nfs without -v4 is not enough — the agent mounts NFSv4 only, and a v2/v3-only kernel cannot serve it.
  • Or upgrade the agent (sudo nobgp upgrade) and let it fall back: on 0.4.62+ a host in this state mounts over FUSE instead of not mounting at all.
  • Installing the module is enough on its own — the agent asks the kernel to load it and retries that every five minutes, so the drive appears on a later mount cycle without waiting for a reboot.
  • On agent 0.4.70+ the kernel module is the only half that matters on Linux: the userspace package is no longer required, so kmod-fs-nfs-v4 (or your distribution's equivalent) is the whole remedy, and a container that cannot load modules is served by the host kernel loading them on demand.

Linux/macOS: The Drive Never Comes Back and the Error Says "Resource busy"​

Symptoms: The drive goes away and never returns. nobgp status shows fs.type empty and fs.mounted: false, fs.backends still says the backend is available, and fs.error reports a mount failure ending in Resource busy or mount point is busy — on every retry, indefinitely. Nothing is listed in the machine's mount table for that path.

Cause: The old mount was torn down, so there is nothing left for the agent to unmount, but some process still holds the mount-point directory open — a shell sitting in it, a Finder window, an editor, a backup agent. The kernel refuses to mount over a directory in that state, and the agent has nothing it can safely do about it: a forced unmount cannot dislodge a live holder, and killing whatever is holding the directory is not a decision the agent takes on your behalf.

Solutions:

  • On agent 0.4.70+, read fs.error — it names the processes and their PIDs (held by bash(4821), Finder(512), which must exit before the mount can recover). End those, and the mount retry that is already running brings the drive back on its next cycle. No restart is needed.
  • On an older agent, find them yourself with lsof /mnt/nobgp (/Volumes/nobgp on macOS) — or fuser -v on Linux — and do the same.
  • If the message says no holder could be named, lsof is either missing or defeated by the wedge. A reboot always clears it.
  • Switching backends is not a way out, because the directory is what is pinned and every backend on the platform mounts on the same one. Through agent 0.4.74 only an nfs mount reported the wedge at all: a Mac pinned to fs: webdav hit the same refusal with fs.error left empty, which is the shape to recognise on those versions — fs.mounted: false, a backend reported as available, and no reason given. From agent 0.4.75 webdav reports it in the same words, and from 0.4.78 so does fuse — which matters most on Linux, where fuse is what the majority of nodes mount with, and where through 0.4.77 the answer was generic advice to go and run fuser -m on the box yourself. A winfsp volume is the one that still says nothing: a refused mount there reports no reason for us to pass on, and Windows has no lsof to name a holder with, so on that backend the shape above is still what you have to recognise.

Linux/macOS: The Drive Path Hangs on a Node That Mounts Nothing​

Symptoms: nobgp status says fs.mounted: false — often with fs.type empty because the node is set to fs: off, or has no mount point at all — and yet reading the mount path hangs or errors instead of showing an empty directory. The machine's own mount table (mount | grep nobgp, or mount on macOS) still has an entry for that path.

Cause: A previous agent was killed while the drive was mounted — a SIGKILL, an out-of-memory kill, a host reset, a crash in the mount driver — so it never got to unmount, and the agent that replaced it was configured not to mount anything. Through agent 0.4.77 those startup paths returned before ever looking at the mount table, so the leftover was never touched again and the agent's report genuinely did not describe the kernel. An orderly stop or restart was never affected.

Solutions:

  • Upgrade the agent (sudo nobgp upgrade) and restart the service. From agent 0.4.78 the leftover is cleared at startup — the kernel is asked what is mounted there and only a mount made by this agent's own backends is removed, so nothing else on that path is disturbed. See A mount left behind when the node is not going to mount.
  • On an older agent, or if the unmount is refused (it is not forced on this path, and is retried at the next start), clear it by hand: sudo umount -l /mnt/nobgp on Linux, sudo diskutil unmount force /Volumes/nobgp on macOS.
  • Windows is not covered by the automatic clearing. A leftover there is a drive mapping — remove it with net use <letter>: /delete.

Linux: The Shared Drive Is Mounted but Bulk Copies Crawl​

Symptoms: On a Linux node mounting over nfs, the drive mounts, lists and reads correctly, and nothing reports an error — but copying anything large onto or off it takes tens of times longer than it should. An 8 MiB write takes seconds where a fuse mount on the same machine takes a fraction of one. Agents 0.4.70 and earlier.

Cause: The kernel's NFS client sizes each read and write from what the server advertises as its maximum, not from the rsize/wsize the mount asked for. The agent's server advertised no maximum, and Linux reads that as an unknown one and falls back to its own 1 KiB floor — so a mount that asked for 128 KiB got 1 KiB, and every transfer was split into ~128× as many requests. macOS honours the requested size either way and was never affected.

Solutions:

  • Upgrade to agent 0.4.71+ (sudo nobgp upgrade) and let the drive remount. The server advertises the size it serves, and the client negotiates 128 KiB. Nothing needs configuring.
  • Confirm what the mount actually settled on — the mount options do not tell you, since these are the negotiated values:
    nfsstat -m # or:
    grep nobgp /proc/mounts # rsize=131072,wsize=131072 once fixed
  • On an agent you cannot upgrade yet, a Linux host with a usable /dev/fuse and fusermount mounts over FUSE by default; if the node is on nfs because fs: nfs is pinned, unpinning it moves the node to a backend this never touched.

Linux: rm -rf on the Shared Drive Deletes Nothing​

Symptoms: On a node mounting over webdav — on Linux, that is a davfs2 mount — rm -rf on a directory in the drive reports one Directory not empty and appears to have removed a file or two. Checking afterwards shows nothing was deleted: every file is still there, from the drive and from every other node. Agents 0.4.78 and earlier.

Cause: davfs2 takes a WebDAV lock before it writes, and hands the lock token back on the DELETE tagged with the URL of the file on the mount. noBGP never saw that URL — the request had already been readdressed to the router by the node's local proxy — so the condition could not be matched, and RFC 4918 makes an unmatched condition a 412. Every delete was refused while holding a lock that was genuinely held, and the only thing the user ever saw was the rmdir behind them failing.

Solutions:

  • Upgrade to agent 0.4.79+ (sudo nobgp upgrade). The proxy now rewrites that condition to name the request's own target, which is the resource the tag named anyway, and deletes go through.
  • Nothing needs cleaning up afterwards — the files were never removed, so a second rm -rf on the upgraded agent finishes the job.
  • The other backends were never affected: they do not send this header at all. davfs2 is where this was measured — and a webdav mount is uncommon on Linux precisely because davfs2 is rarely installed — but the fix is not specific to it, so any WebDAV client that locks a file before deleting it benefits, including a Mac pinned to fs: webdav.

A File on a webdav Drive Cannot Be Written and Nothing Else Can Write It Either​

Symptoms: On a node mounting over webdav, a copy or a save onto the drive fails with a lock error — on Windows, The process cannot access the file because another process has locked a portion of the file (0x80070021). The file then stays unwritable for several minutes, from every node and every backend, not just from the machine that was copying. Agents 0.4.81 and earlier.

Cause: The operating system's WebDAV client takes a lock before it writes, and the answer to that lock names the path the lock was taken on. The node's local proxy presents the share at node/ and networks/<name>/, and it was not rewriting that name onto the mount's own path — so the client got back a lock on a path it had never asked about, could not match it to the file it was writing, and sent the write with no lock token. noBGP refuses a token-less write on a locked path with 423, correctly. The client never sent the matching UNLOCK either, so the lock sat out its timeout — up to 10 minutes on router 0.4.80+, and until a router restarted before that — refusing every other writer in the meantime.

Solutions:

  • Upgrade to agent 0.4.82+ (nobgp upgrade from an Administrator terminal on Windows, sudo nobgp upgrade elsewhere) and restart the service. The lock's answer is rebased onto the mount's own path, so the client matches its own lock and the write carries the token.
  • Wait it out on a path already stuck: on a router at 0.4.80+ the lock expires within 10 minutes on its own and needs nothing done to it.
  • The other backends never sent a LOCK at all and were never the cause — but they were affected while one was stuck, since the refusal applies to the path rather than to the mount that took it.

Windows: Every Copy onto a webdav Drive Fails with 0x80070021​

Symptoms: On a Windows node mounting over webdav, every copy onto the drive fails — Copy-Item, copy from cmd.exe, and dragging a file in Explorer all report 0x80070021 (ERROR_LOCK_VIOLATION), immediately and for every file, however small. Set-Content and anything else that writes without locking first still works, which is what makes it look like a permissions or a path problem rather than a locking one. Nothing is stuck: the same copy fails again a minute later, and no other node is refused that path.

This is a noBGP-side fault, not the node's — nothing on the machine needs changing and no agent version fixes or causes it.

Cause: Windows takes a WebDAV LOCK before it writes, and sends it with no Depth header. RFC 4918 defines an absent Depth as Depth: infinity, so what looks like an ordinary lock on one file arrives as a request to lock a whole tree — and noBGP refused every depth-infinity lock outright, because a lock here names one exact path and cannot cover a subtree. The client got 423, and Windows surfaced it as a lock violation on the copy. On a plain file a depth-infinity lock covers nothing that a Depth: 0 lock does not, so it should always have been granted; the blanket refusal was refusing the default rather than an exotic request. Routers 0.4.80 – 0.4.95.

Solutions:

  • Nothing to do from router 0.4.96. A depth-infinity LOCK on a file, or on a path that does not exist yet — which is the lock-then-create a Windows client opens a write with — is granted. A lock on a directory is still refused with 423, unchanged, and so is a shared lock (501).
  • On a node still meeting it, fs: winfsp is the way round: WinFsp is the default on Windows from agent 0.4.82 and does not send a WebDAV LOCK at all — among the mount backends only webdav was ever affected. Writing through the file tools is unaffected — they take noBGP's own lock rather than a WebDAV one — and so is a plain PUT to the share's own URL, which locks nothing first.
  • Any other WebDAV client that locks before writing could meet the same refusal, on any platform — a client that sends an explicit Depth: 0 never did. Windows is where it was measured, on a fs: webdav node, and it is where the whole Windows fleet would have met it before WinFsp became the default.

Windows: A File Over ~47 MiB on a webdav Drive Reads as "Permission denied"​

Symptoms: On a Windows node mounting over webdav, reading a large file from the drive fails with Permission denied — and so does merely asking for its size. Smaller files on the same drive, in the same folder, read normally. Writing the large file succeeded, and every other node and the share's own URL read it back intact. Retrying later fails identically. Nothing appears in the agent's log or on noBGP's side.

Cause: It is not a permission problem and the message is misleading. Windows' WebClient service — the WebDAV redirector that backs a net use drive — refuses any transfer larger than FileSizeLimitInBytes, which is 50,000,000 bytes (47.68 MiB) on a stock machine, and reports the refusal as an access denial. The redirector decides this client-side and never sends a request, so nothing in noBGP ever sees it and there is no error for the agent to explain. Measured on Windows 11 with agent 0.4.98: 1 MiB and 8 MiB read fine, 64 MiB refused every time.

⚠ The asymmetry is the trap: the write goes through and the bytes are stored correctly, so a file can be written from that machine and then be unreadable on that machine alone.

Solutions:

  • Move the node onto winfsp, which is the better answer and the default on Windows from agent 0.4.82. It is a filesystem driver, so the WebClient service is not in the path at all and the cap does not exist. Install WinFsp (winget install WinFsp.WinFsp, or winfsp.dev) if the machine does not have it, clear any fs: webdav pin from the profile, and restart the agent.
  • Or raise the Windows limit on that machine, if it must stay on WebDAV. Set FileSizeLimitInBytes under HKLM\SYSTEM\CurrentControlSet\Services\WebClient\Parameters (maximum 4 GB) and restart the WebClient service. This is a per-machine Windows setting; the agent neither reads nor changes it.
  • Reach the file another way in the meantime. The file tools and the share's HTTPS URL do not go through the mount and are unaffected, and any other node reads the file normally.

nobgp status Reports a Healthy nfs Mount on an Empty Folder​

Symptoms: On a node mounting over nfs, the mount point is an ordinary empty directory — nothing under node/, no networks/ — while nobgp status reports fs.mounted: true with no error, and the network directory shows the node on a healthy nfs backend. The agent never retries, and a nobgp service restart fixes it until the next time. Agents 0.4.81 and earlier.

Cause: The agent's 30-second mount probe asked only whether the mount point answers, and an ordinary empty directory answers it perfectly well. So a mount that went away while its mount point stayed behind read as healthy for the life of the process: the remount machinery was never broken, it was simply never asked to run.

Solutions:

  • Upgrade to agent 0.4.82+ (sudo nobgp upgrade) and restart the service. The probe now asks the kernel whether anything is mounted at the path as well as whether the path answers, so this state fails within one poll and the node remounts itself in about 90 seconds. Both questions are asked — a mount that is present but wedged still has to fail the probe, which is what the original one was for.
  • On an older agent, sudo nobgp service restart remounts it.

Linux: The Shared Drive Lists Directories as Empty While the Node Cannot Reach noBGP​

Symptoms: On a Linux node mounting over fuse, the drive is mounted and ls on a directory under networks/<name>/ or node/ returns nothing at all — no error, no files — while the same directory is plainly still there in the dashboard and on other nodes. It clears on its own within a listing TTL of the node's connection recovering. Agents 0.4.94 and earlier.

The dangerous version of this has no symptom you would notice: a backup, an rsync --delete or a find -delete pointed at the drive during that window reads the empty directory as the files were deleted and acts on it.

Cause: The mount turned every failed directory listing into an empty directory and cached it for a full TTL — a dropped connection, a 5xx, a timeout, all of them — so the absence of an answer was rendered as the absence of files. noBGP held the data throughout; the mount handed somebody else's tool a deletion signal.

Solutions:

  • Upgrade to agent 0.4.95+ (sudo nobgp upgrade) and restart the service. A listing that could not be obtained now serves the last one the mount held, for as long as the failure lasts; only a genuine no such directory lists as empty, and a directory the mount has never listed answers Input/output error rather than pretending. See a listing noBGP could not be asked for.
  • ⚠ After that upgrade, an Input/output error from ls on a directory the node has never read is the honest answer, not a fault. It is retried every five seconds and clears as soon as the node can reach noBGP; check the agent's connection rather than the mount.
  • On an older agent, keep destructive walkers — anything with --delete — off the drive unless the node's connection is known good, and confirm a directory against the dashboard or another node before acting on an empty ls.

macOS: The Drive Never Comes Back After Changing fs:​

Symptoms: On a Mac, you pinned a different backend — fs: nfs to fs: webdav, or back again; those two are the whole list on macOS — and restarted. Since then /Volumes/nobgp is an ordinary empty directory, nobgp status reports fs.mounted: false, and restarting again changes nothing. The node is otherwise healthy. Agents 0.4.92 and earlier.

Cause: macOS keeps the /Volumes entry that a failed mount leaves behind, and the next attempt reads it as a mount point rather than as debris. The agent undoes a mount point it created and then failed to mount on — but only the one it created in that attempt, and only while it is empty. A directory the previous backend left behind pre-dates the attempt, so the agent does not treat it as its own to remove.

Neither of the two earlier clearings that sound like they should cover this does, which is why it needed a third:

  • The startup clearing added in agent 0.4.78 (a mount left behind) matches what the kernel says is mounted at the path. An ordinary directory is not a mount, so there is nothing there for it to match.
  • The probe fix in 0.4.82 (above) is the opposite case — a mount that has gone away while nobgp status still calls it healthy.

Solutions:

  • Upgrade to agent 0.4.93+ (sudo nobgp upgrade) and restart the service. Before mounting, the nfs backend now removes an empty directory sitting at the mount point that no mount explains, and takes the point itself — so a node stuck this way repairs itself on its next mount cycle instead of waiting for someone at the machine. That is the direction this was measured in: the debris is what a webdav teardown leaves and what an nfs mount then cannot mount over.
  • ⚠ Empty is the whole proof, and it is the only one available. There is no mount-table entry to read and no owner worth checking — the debris is owned by whoever made the previous mount — so the agent cannot show the directory is its own. What it can show is that removing it destroys nothing: rmdir on a directory refuses unless it is empty, so the proof and the removal are the same operation. A directory with anything in it is left exactly where it is, and the mount's own failure is the honest report. A file at the mount point is likewise never touched.
  • On an older agent, do it by hand: sudo rmdir /Volumes/nobgp, then sudo nobgp service restart. The agent creates the mount point itself and mounts on it.
  • ⚠ rmdir, not rm -rf, for the same reason the agent uses it: rmdir refuses a directory that is not empty, which is the guard that stops you deleting a drive that is actually serving. If it refuses, the path is not debris — read fs.error under nobgp status before going further.
  • fs: off never reaches this clearing, deliberately: a node told not to mount does not mount, and does not go tidying a mount point it has no intention of taking.

macOS: The Drive Comes Up After a Restart and Vanishes Seconds Later​

Symptoms: After restarting the agent — most visibly a restart that also changes the fs pin — the drive mounts and then goes away again within seconds. On a Mac, /Volumes/nobgp is gone entirely rather than left as an empty directory. nobgp status reports fs.mounted: false, and the agent's log shows a successful mount immediately followed by an unmount nobody asked for. The node repairs itself about a minute and a half later, when the mount health check notices and remounts. Agents 0.4.105 and earlier.

Cause: A restart runs two agents at once for a moment — the outgoing one is still tearing its mount down while the incoming one is already mounting — and the outgoing one had no way to tell its own mount from its successor's. Every destructive step of a teardown named a path, and a path changes hands, so a slow unmount landed on the mount that had just replaced it. Measured on a Mac at agent 0.4.95: an fs: webdav → fs: auto restart had the outgoing WebDAV unmount take the new nfs mount, and macOS removed the /Volumes entry along with it — about 105 seconds with no filesystem on the node.

Solutions:

  • Upgrade to agent 0.4.106+ (sudo nobgp upgrade) and restart the service. A mount now carries the kernel's identity for the mount that process made, and every teardown step checks it is still what is at the path before unmounting, forcing or removing anything. See A teardown belongs to the process that made the mount.
  • On an older agent, nothing needs doing — the node remounts itself within about 90 seconds. sudo nobgp service restart brings it back sooner, but a restart is what triggers this, so expect it to be a race you may lose again.
  • ⚠ The fix covers more on macOS than on Linux. macOS gives each mount a kernel id, so any two mounts at the same point are told apart; the Linux mount table carries no such field and the match falls back to the mount's source and type, which separates a restart that changes backend but not one nfs mount from the next.

Node Is Online but Every Command and File Call Fails​

Symptoms: The node shows online in the dashboard and nobgp status looks healthy, yet every command, fs_* and net_* call against it fails. The agent's log names a DNS failure for the router's own hostname — lookup router.nobgp.com … server misbehaving — on a machine whose control channel is plainly still up. Agents 0.4.78 and earlier, and in a narrower form through 0.4.80 on a machine whose only other resolver belongs to another overlay.

Cause: Two things had to be true at once, and 0.4.79 fixes both:

  • The host's own resolver was answering and failing. On a machine where another overlay product provides DNS this is easy to reach: Tailscale's MagicDNS was excluded from the agent's forwarding upstreams only at its IPv4 address, so on a host where Tailscale had installed its IPv6 resolver instead the two forwarders could hold each other as upstream — a query loop, answered SERVFAIL, which Go reports as server misbehaving.
  • Only some of the agent's connections could survive that. The agent falls back to public DNS when the host resolver answers and cannot resolve, but several of its connections to the router were assembled separately and without that fallback — the channel that carries command, file and terminal sessions, the event stream, and the node's local MCP proxy. The control channel had it, so the node stayed registered and reported healthy while every new session died on a name that would not resolve.

Solutions:

  • Upgrade to agent 0.4.79+ (sudo nobgp upgrade). Every connection to the router is now made through one dialer with the public-DNS fallback, and Tailscale's whole IPv6 range is excluded from the agent's forwarding upstreams the way its IPv4 one always was.
  • Agent 0.4.80 adds a second layer for the case where public DNS is not reachable either — a network that blocks outbound DNS, which is exactly the one the fallback cannot rescue. The agent remembers the addresses it has actually connected to for the router's name (up to four, in memory only, never written to disk) and tries them when a lookup fails, after the system resolver and before the public resolvers. A fresh lookup still leads on every connection, so an address that legitimately changes is picked up as before and nothing about a healthy node changes; a remembered address is only ever reached once DNS has already failed, so a stale one costs one connection attempt rather than an outage.
  • Agent 0.4.81 adds the two remaining pieces. It stops the node's DNS going dark in the first place on a machine whose only other resolver belongs to an overlay: rather than assuming that resolver is there to fall through to, the agent checks whether it can still resolve the router's name, and only when it cannot does it forward out-of-zone queries to public resolvers itself. While the other resolver works, nothing changes — see Other VPN clients. And when a new session's connection to noBGP does fail, the node now says that is what happened (target_unreachable) instead of returning it as a final error, so noBGP asks the node again rather than handing you a failure from a moment that has already passed.
  • On an older agent, fixing the host's resolver clears it — the agent recovers on its own, with no restart, as soon as names resolve again.
  • If the machine runs Tailscale, see Other VPN clients.

Node Never Reconnects After the Site's IPv6 Prefix Changes​

Symptoms: A node on a dual-stack site goes offline when the ISP hands the site a new IPv6 prefix (a DHCPv6-PD renumber) and does not come back, even though the machine holds a fresh global IPv6 address and a working IPv4 path the whole time. The agent's log shows connection attempts timing out rather than being refused. Agents 0.4.107 and earlier.

Cause: Two things had to line up. The host's own resolver had already failed, so the agent was on one of its fallback paths — the addresses it remembers having connected to, then public DNS. Both of those dial one address at a time, which loses the automatic IPv6→IPv4 fallback the normal path gets for free from the system resolver, and every address shared a single connection budget in the order it arrived. A withdrawn prefix does not refuse connections, it swallows them, so one dead IPv6 address consumed the entire budget and the IPv4 addresses sitting behind it were never dialed at all — and the next attempt spent the same budget in the same order.

Solutions:

  • Upgrade to agent 0.4.108+ (sudo nobgp upgrade). Each candidate address now gets its own share of the budget, so a dead address costs one attempt instead of all of them, and the list is ordered so the two address families alternate — IPv4 first on this path, since it is only ever reached after the resolver has failed and a delegated IPv6 prefix is the thing more likely to have just been withdrawn. On a host with no usable global IPv6 address, IPv6 candidates are dropped rather than dialed.
  • ⚠ A healthy connection is byte-for-byte unchanged: while the host's resolver works, the agent still hands it the router's name and the operating system picks the order, which on a dual-stack host still prefers IPv6.
  • On an older agent, restarting the service (sudo nobgp service restart) clears the remembered addresses, which are held in memory only — but the same ordering applies to the public-DNS answers, so it may not be enough. Fixing the host's resolver takes the agent off the fallback path entirely and is the reliable workaround.

A Machine's Own Uplink Breaks Whenever the Node Sends IPv6​

Symptoms: a node runs on hardware — usually a small appliance or router-class box — whose network drops out under load, and the outages line up with what the agent is doing rather than with anything on the network: an upgrade download, a large file transfer, a busy mesh session. The fault is on the machine itself; other devices on the same link are fine, and the box recovers when its interface is reset. Disabling IPv6 at the operating system does not stick, because the machine's own network manager writes the setting back when it brings the link up.

Cause: a firmware or driver defect on that host that misbehaves on IPv6 traffic, at a rate that grows with volume. noBGP is not the cause, but it can be the biggest IPv6 talker on a quiet appliance — the upgrade package alone is the largest single transfer the agent makes. Through agent 0.4.112 there was no way to ask a node to prefer one family: every connection either raced both families or let the operating system choose, so IPv6 went out on every connect whatever the machine could stand.

Solutions:

  • Upgrade to agent 0.4.113+ and set transport-family to the family that works on that host — v4 on the box above:

    sudo nobgp config --transport-family v4

    The node then dials IPv4 first and on its own and reaches for IPv6 only if IPv4 fails, so under normal operation the other family never leaves the machine.

  • It can also be set from your assistant with node_config_set, which is the point on a box nobody can walk up to. The change applies on the agent's ordinary reload, within about five seconds.

  • ⚠ It is a preference, not a block. If the chosen family stops working the node falls back to the other rather than going offline — that is deliberate, and it means a lasting IPv4 outage will put IPv6 back on the wire.

  • ⚠ The shared drive does not follow the setting yet, and it moves more bytes than anything else here. On a host that must emit no IPv6 at all, pair the setting with fs: off or an interface-level block.

  • Expect a few seconds' extra delay per connect if the preferred family is dark rather than merely slower — a black-holed path is silence, so the leading attempt has to time out before the fallback runs.

Node Restarts Every Time the Machine Wakes​

Symptoms: A laptop or desktop that sleeps shows its node dropping offline and coming back around each wake, and nobgp status reports a small uptime_secs even though the machine itself has been up for hours. On macOS a dead nobgp volume may be left behind by the restart.

Cause: In the first seconds after a wake the resolver has no servers yet and the interface has not been handed its address back, so the agent's first attempt to reach the router fails with a name-resolution or no-route error. Until 0.4.49 that was treated as fatal: the agent exited, the filesystem mount came down with it, and the service manager restarted it with its backoff — for a machine that was seconds away from being fine. Fixed in 0.4.49: registration and the mid-run reconnect retry in place for up to five minutes while the only thing wrong is that the network is absent, logging the network looks absent, retrying in … as they wait.

Solutions:

  • Upgrade the agent (sudo nobgp upgrade), then sudo nobgp service restart.
  • On 0.4.49 and later, read the agent's own line before assuming this is it: only the absent-network errors are retried, so a restart at wake that logs anything else — a rejected credential, for example — is a different problem. See Credentials Rejected / Node Deauthorized.
  • A node that stays unreachable for longer than five minutes after a wake still restarts, which is intended — beyond that, a clean start is more likely to help than more waiting.
  • Agent 0.4.83 removes the other restart in this family. A control connection that drops mid-run — the machine changing network, a router deploy, a few seconds of packet loss — is now re-established inside the running process on both transports. Through 0.4.82 only a QUIC connection was; a node connected over WebSocket, which is every node on a network that blocks UDP and every node before its certificate pin is learned, exited instead and waited for its service manager to bring it back, rebuilding the tunnel, the peer sessions, the DNS configuration and the drive for a blip that was over in seconds. nobgp status counts the in-place recoveries as router.control_reconnects. A node that keeps flapping is still handed back to its service manager rather than reconnecting forever, and a rejected credential still stops the agent as it always did.

Can't Reach Other Nodes​

Symptoms: Node is online but can't communicate with other nodes

Solutions:

  • Verify all nodes are in the same network
  • Check that all nodes show as "online"
  • Ensure encryption settings match across nodes
  • Review network-level firewall rules
  • Run nobgp status on both ends: an environment.tunnels entry means another VPN client on that machine holds an active adapter and may be swallowing overlay traffic — see Security software and VPNs
  • On Linux, if that machine also runs Tailscale, WireGuard or a similar overlay, see the section below

Peers Are Unreachable on a Machine Running Another Overlay VPN​

Symptoms: two shapes, depending on the agent version.

  • Agent 0.4.110 and earlier. The node registers, nobgp status reports it healthy and the control channel stays up, but no peer can be reached from that machine and traffic to overlay addresses never arrives on the noBGP interface. The machine also runs Tailscale, WireGuard or another overlay product.
  • Agent 0.4.111. The agent exits during startup, over and over, and the log reads no free /20 slice available in 100.64.0.0/10 (host is fully claimed).
  • Agent 0.4.112 and later. The node stays online and answers command, file and nobgp status, but has no TUN interface, no overlay DNS and no peer traffic. nobgp status names the reason under network.overlay, and network.dns reads unavailable.

Cause: each node self-assigns a /20 out of 100.64.0.0/10 (Overlay addressing) and picks one that does not collide with a route already on the machine. Through 0.4.110 the Linux side of that read only the main routing table, while products using policy routing keep their routes in a table of their own and reach them with an ip rule ahead of main — so the claim was invisible. The agent took 100.64.0.0/20, the kernel kept sending that range to the other tunnel, and the node looked healthy while nothing on the overlay worked. A second, related fault made a claim on the whole pool look like the agent's own leftover route from a previous run, because 100.64.0.0/10 starts at the same address as the agent's first-choice /20; hosts on any other slice saw such a claim correctly.

From 0.4.111 the agent reads every routing table and treats a route containing its slice from the outside as the conflict it is. That turns a silent failure into a visible one: if the other product claims all of 100.64.0.0/10 there is no free /20 left, and the agent says so instead of taking one it can never use. It deliberately does not carve inside another VPN's range — that would hijack that product's own peers.

From 0.4.112 that verdict no longer ends the process. 0.4.111 turned a silent misconfiguration into a crash loop, and a machine whose agent will not stay running is a machine nobody can reach remotely — so the one defect that could be fixed over command needed physical access to fix. Such a node now starts without an overlay: peers are unreachable, exactly as the condition demands, while the control channel, command, file, the shared drive and nobgp status all keep working. nobgp status is where the reason lives, since a log line on a box you cannot reach is a reason nobody reads.

Solutions:

  • See what is claimed on that machine, and by whom:

    ip route show table all | grep 100.64 # every table, not just main
    ip rule show # which table wins for that range
    nobgp show # the slice this agent is using
  • Free space in 100.64.0.0/10. Narrow the other product's routes so it does not claim the entire range — many VPN clients have a split-tunnel or route-scope setting for exactly this. Once part of the pool is free, the agent picks a /20 there on its next start (sudo nobgp service restart) and needs no further configuration.

  • On 0.4.112+ you can do all of that remotely, since the node is still online and still serves command — including reading nobgp status for the exact reason and restarting the agent afterwards.

  • ⚠ Pinning overlay-cidr is not a way around it. A pinned value that overlaps a live route is refused with a warning and replaced by an auto-pick, which then meets the same fully-claimed pool.

  • ⚠ A carrier-grade NAT uplink is a different case and still works. Where the machine's own default gateway is a 100.64 address (Starlink and similar), the whole pool is the carrier's shared space rather than a per-node claim, and the agent carves a /20 inside it as before — unless a VPN on that machine also holds a 100.64 address, in which case it declines to carve rather than risk hijacking the VPN's peers.

  • On an agent older than 0.4.111 the same machine needs the same fix; upgrading is what makes the condition visible instead of silent.

Credentials Rejected / Node Deauthorized​

Symptoms: The agent logs the router rejected this node's credentials, and the node stays offline.

Cause: The node was deleted, its key was revoked, or the network was removed on the server side, so the router no longer accepts this node's credentials.

Solutions:

  • Re-enroll the node: sudo nobgp register.
  • Or remove the install entirely: sudo nobgp service uninstall.
  • Until you act, the agent paces its retries (roughly every 15 minutes) instead of hammering the router.

Registration Token Expired​

Symptoms: The node is offline and nobgp status reports the profile under unavailable with registration token expired <instant> — the router refuses its connections; re-register with 'nobgp register'. nobgp show prints the same instant on a token-expired line.

Cause: The node's registration token expired while it was offline. A token is normally refreshed over the node's authenticated connection, and an expired one cannot make that connection — so nothing on the node can repair it on its own.

Solutions:

  • Re-enroll the node: sudo nobgp register. On agent 0.4.72+ this replaces the expired token under the node's existing identity key, so the node keeps its place in the network.
  • On 0.4.71 and earlier that command answers Agent is already registered. and does nothing. Upgrade first (sudo nobgp upgrade -f, which needs no registration), then re-register.
  • Before 0.4.72 the reason above is not reported either: the profile reads as an ordinary stopped agent, because the agent is being started, is refused, and is paced fifteen minutes out — so nobgp status almost always catches it between attempts. On such a node, treat start it with 'nobgp service start' that never helps as a reason to check the token.

The Browser Page Says the Sign-In Was Refused and Asks for an Upgrade​

Symptoms: sudo nobgp register or nobgp login prints its link, and the page that opens reads Sign-in refused — No noBGP terminal started this sign-in, or it expired. with If you ran nobgp login or nobgp register, upgrade the agent (nobgp upgrade) and try again. Nothing is registered and nobody is signed in. On router 0.4.184 and earlier, sudo nobgp register on the same machine instead got all the way through the page — sign in, pick a network, approve — and then the terminal waited out its 15 minutes with no token. nobgp login got this page then too.

Cause: The agent is older than 0.4.153, the release that tells noBGP about the terminal before the browser opens. Without that, the page a person opens cannot be tied to the machine in front of them, so noBGP's routers refuse that way in. From router 0.4.185 a registration is refused at the start rather than after the approval.

Solutions:

  • sudo nobgp upgrade -f on the machine (on Windows, nobgp upgrade -f from an Administrator prompt), then run the command again. On a machine too old to upgrade itself, reinstall with the install script.
  • Or register with a key — --key / NOBGP_KEY — which is unaffected, needs no browser, and works on any supported agent.

nobgp register Fails with enrollment failed​

Symptoms: sudo nobgp register on a machine that is already enrolled ends with enrollment failed and no new token, so the node stays offline. Registering a brand-new machine into the same network works.

Cause: The node's name in noBGP is not the machine's hostname. node-name is removed from the config file after the first successful registration, so a later nobgp register presents the hostname — and through router 0.4.115 noBGP matched a re-registration on that name alone. The lookup missed, a second node was attempted under the same identity key, and the write was refused. Measured on a live fleet, 74 of 149 nodes carried a name that differed from their hostname, so this hit most machines that had ever been renamed.

Solutions:

  • On router 0.4.116+ there is nothing to do: the machine is matched on its identity key first, comes back as the same node, and keeps the name noBGP holds for it — along with its labels, role grants, services and storage.
  • On an earlier router, pass the name noBGP has for the node: sudo nobgp register --name "<the name shown in the dashboard>".
  • A refusal that names another node — this machine's key is already registered to node "..." in this network — is a different situation: two registrations presented the same identity key. Remove or deprovision the node named in the message, then register again.

RHEL-Family Host: The Install Script Finished but the Node Never Registered​

Symptoms: On RHEL, Rocky Linux, AlmaLinux, Oracle Linux or CentOS Stream, the install script installs the package and then stops without enrolling the node, or sudo nobgp status answers sudo: nobgp: command not found while nobgp status without sudo runs perfectly well.

Cause: The package puts the binary in /usr/local/bin, and on those distributions sudo's secure_path is /sbin:/bin:/usr/sbin:/usr/bin — it does not include /usr/local/bin. So anything run through sudo by bare name is not found. Through agent 0.4.123 the install script called nobgp register that way: it exited 127, the script stopped there, and nothing reported a failure, so the install looked like it had succeeded while the machine had never reached noBGP. Debian, Ubuntu, Alpine and Amazon Linux 2023 carry /usr/local/bin in secure_path and were never affected.

Solutions:

  • Agent 0.4.124 fixes both halves: the script calls the binary by absolute path and reports a registration failure as one, and the RPM package links /usr/bin/nobgp at the same binary so sudo nobgp … works by name. The link is only created where nothing is already at that path, is removed when the package is removed, and leaves an unrelated /usr/bin/nobgp alone.
  • On a machine already in this state, finish the enrollment by path: sudo /usr/local/bin/nobgp register. Re-running the current install script does the same thing and adds the link.

Node Will Not Start After a Power Cut: not logged in​

Symptoms: A machine that was enrolled and working comes back from an unclean reset — a power cut, a hard reboot, a pulled plug — and the agent refuses to start, saying it is not logged in. nobgp show reports the profile as not registered. The log line above it reads ignoring token with mismatched akid expected=<key id> got="", and the profile's .jwt file is zero bytes.

Cause: The node's token was being written when the machine stopped. Through agent 0.4.120 that write truncated the file first and wrote the new bytes second, so a reset inside that window left an empty file with the previous content already gone — and an empty token is a de-registered node. It is the one piece of node state that cannot be repaired remotely: a node that loses its config reads it again, but a node that loses its token cannot make the authenticated connection every repair path needs. A machine that resets often can hit it more than once.

Solutions:

  • Re-enroll at the machine: sudo nobgp register. The identity key is intact, so the node comes back as itself with its name, labels, grants, services and storage. This needs someone at the device (or console/SSH access to it) — there is no remote route in while the token is empty.
  • Upgrade to agent 0.4.121 to stop it recurring. Both identity files are written atomically now — into a temporary file that is renamed into place, so an interrupted write leaves the previous complete file rather than an empty one — and the token is rewritten only when its value actually changed, so a node reconnecting repeatedly no longer rewrites identical bytes (and no longer wears the flash of a small board) every cycle.
  • ⚠ Check the file's size before reading that log line as this fault. The same got="" also appears for a complete token that carries no key-id claim, which is not damage on the node and which re-registering at the device will not fix.

Agent Version No Longer Supported​

Symptoms: The node never comes online, and the agent logs a rejection reason such as agent version no longer supported (pre-0.3.51): reinstall the current agent.

Cause: The agent build is older than the minimum supported version. Agents below that floor are refused at registration — no node record is created or modified, so nothing on the server side needs cleaning up.

Solutions:

  • Run sudo nobgp upgrade -f on the host. This is the smallest fix: it does not require the node to be registered, and on an unregistered machine it installs the stable channel version.
  • If the agent binary is too old to upgrade itself, reinstall it with the install script.
  • Then confirm the result with nobgp version and nobgp show.

Upgrading​

Minimum Supported Agent Version​

The router refuses registration from agents older than 0.3.51. Those builds predate the current overlay addressing model and cannot operate on the network, so they are rejected with a clear reason (agent version no longer supported) rather than being allowed to connect and fail later.

Any agent on 0.3.51 or newer registers normally. Nodes with auto-upgrade enabled stay well ahead of this floor on their own — this only affects machines that have been offline or pinned to an old build for a long time. Recovering one is a reinstall or a single sudo nobgp upgrade -f.

A version floor can hold a node out of the network​

From router 0.4.168 a release channel can carry a minimum version — a floor a node has to be at before it joins the overlay. It is set only for a change an older agent cannot be allowed onto the network with, so most releases have nothing to do with it and most nodes never meet one.

A node below its channel's floor connects to noBGP as usual, and is then held:

  • It is not in the network. No peers, no relay sessions, no overlay names, and network_directory reports online: false for it while its connection is live, with info.version_status reading below_floor.
  • It is offered its channel's version at once, outside the rollout it would otherwise wait its turn in. With auto-upgrade on it installs that version and joins on its next connection, usually within a minute or two.
  • The remote tools still reach it. command, the file tools, node_logs and node_config_get / node_config_set all work on a held node, so one that will not upgrade itself can be fixed remotely instead of at the machine.
  • It is never disconnected for being held, whatever the reason it is not upgrading.

A node with auto-upgrade turned off stays held until someone upgrades it: run sudo nobgp upgrade on the machine, or through command. Whatever the node itself reported about the offer — automatic upgrades off, an installer error — is on its directory entry as info.upgrade.

You are emailed about a held node that has not taken its upgrade about an hour after it was offered one, once per offered version. It goes to whoever created the node — or to the organization's Owners where that person is no longer a member — and names the node, the version it runs, the floor it is below and the version it was offered. The email asks for nobgp upgrade on the machine and names no cause; what the node itself reported is on info.upgrade.

A node that is already in the network when a floor is raised stays in it until it next reconnects. Only the upgrade offer is immediate.

This is not the minimum supported agent version above. Below that one an agent is refused at registration and never becomes a node at all; a node held below a channel's floor is registered, connected and reachable, and simply is not on the network until it upgrades.

Manual Upgrade​

Run the upgrade command on any registered node:

sudo nobgp upgrade # Upgrade to this node's channel version
sudo nobgp upgrade -y # Same, without confirmation prompts

Without a version, nobgp upgrade installs this node's release-channel version — the same version automatic upgrades install. On a machine that isn't registered, the stable channel is used. A test or pre-release build is never installed unless you name its version explicitly.

To upgrade to a specific version:

sudo nobgp upgrade 0.4.2

Anything off the safe path asks for confirmation first (with a "no" default). A version ahead of the node's channel, or one that can't be verified, is answered by -y. A downgrade or reinstall requires -f — -y alone won't authorize it (so a stray -y in a script can never downgrade a node):

sudo nobgp upgrade 0.4.2 -f # Downgrade or reinstall without prompts

-f covers everything -y does. Run bare sudo nobgp upgrade -f to converge a node that is ahead of its channel back down to the channel version.

warning

Don't downgrade a node below the minimum supported version (0.3.51). The install itself succeeds, but the router then refuses the node's registration and it stays offline until you upgrade it again.

Running agents detect the new binary and restart automatically — no manual service restart needed.

Auto-Upgrade​

The agent supports automatic upgrades driven by the router. When a new version is released, the router signals all connected agents and they upgrade themselves in the background.

Auto-upgrade is enabled by default. To disable it, add this to your configuration file:

auto-upgrade: false

Since agent 0.4.35 the key is already in the file as a commented default (# auto-upgrade: true) — uncommenting that line and setting it to false does the same thing. Either way the agent keeps the key, because its value is now a choice rather than the default.

tip

Auto-upgrade is safe to leave enabled. It never downgrades, never reinstalls the same version, and requires SHA-256 checksum verification. Only packages from the official CDN are accepted.

How Upgrades Work​

Regardless of whether the upgrade is manual or automatic:

  1. The agent downloads the platform-appropriate package from downloads.nobgp.com
  2. The SHA-256 checksum is verified (required — upgrade fails if unavailable)
  3. A lock flag prevents the service from being stopped mid-upgrade
  4. The package is installed via the platform's native package manager
  5. The running agent detects the new binary via mtime polling and restarts

Step 5 waits for step 4 to finish. Since 0.4.48 the watcher holds off while an upgrade is installing, so a slow install — a Raspberry Pi Zero spends about a minute unpacking a package — is not interrupted by the restart it is about to trigger. Before that, such a node came up on the new version but logged auto-upgrade failed and paused for five minutes before it would try again.

How the restart happens depends on how the agent runs:

  • On a machine whose service manager is supervising the agent: the whole service is restarted, so both the background supervisor and the agent it manages come up on the new binary — this prevents the supervisor from being stranded on the old version. Expect a brief reconnect while the service restarts. Through agent 0.4.146 this covered systemd and Windows only; from 0.4.147 it also covers Synology DSM, OpenWrt (procd), OpenRC and sysv init, and from 0.4.148 macOS (launchd) — see Supervisors off systemd below.
  • In containers and in the foreground: the agent re-execs itself in place. The old binary inode stays open until the restart, so in-flight connections are not interrupted.

On Windows the running executable is renamed and the new binary is put into place atomically, with automatic rollback if the replacement fails. From agent 0.4.97 the new binary is staged inside the install directory rather than moved in from a temporary folder, so it inherits that directory's permissions and stays runnable by ordinary accounts — see nobgp says "Access is denied" for what the old shape cost.

On macOS and on a host with no package manager, step 4 instead downloads the raw static binary and atomically swaps it into place. Such a host runs some other platform's build, so it keeps naming that platform and fetches the binary published beside the package rather than the package itself: on Synology DSM the one next to the Alpine .apk, and — from agent 0.4.99 — on a Buildroot-class image the one next to the Debian package.

Resilient downloads​

The package download in step 1 is resumable, so upgrades survive flaky or slow connections:

  • If a transfer is interrupted, the next attempt continues from where it left off instead of starting over — the partially downloaded file is kept between attempts.
  • A stalled connection (no data received for 60 seconds) is aborted and retried automatically rather than hanging.
  • Before writing, the agent checks that the staging filesystem has enough free space and fails early with a clear error if it doesn't.
  • If that filesystem cannot hold the package at all, the download moves somewhere with room and starts again (agent 0.4.117) — see Where the package is staged below.

Because progress is preserved, an auto-upgrade on an unreliable link makes forward progress each time the router re-offers it, rather than repeatedly re-downloading from the beginning. The SHA-256 checksum is still verified against the fully assembled package before anything is installed.

Where the package is staged​

The agent downloads each release into the system temporary directory first — /tmp on Linux, the equivalent on macOS and Windows. On many appliances, routers and small Linux images that directory is a RAM disk of a few megabytes, so a machine with a nearly empty disk could still be unable to take any release at all, and would stay on its old version through every retry.

From agent 0.4.117, a download that runs out of room there is moved and started again:

  • The fallback is the configuration directory — /etc/nobgp on Linux, /usr/local/etc/nobgp on macOS, C:\ProgramData\nobgp on Windows. It exists on every platform the agent installs on, and the agent already writes its log mirror there, so a package is not a new kind of demand on it.
  • The temporary directory is still the first choice, and the move needs evidence. On most appliance firmware /tmp is the larger of the two — the configuration directory is a few megabytes of flash — so the agent moves only when the configuration directory really does have meaningfully more space free. A filesystem whose free space cannot be read is skipped.
  • It moves once, not down a list. If the second directory cannot take the package either, that failure is reported rather than a third location being guessed at.
  • Nothing else changes. Every other download failure keeps its partial file and its retry schedule, and the checksum is verified in the same way wherever the package landed.

The agent's log names both directories when it moves, so nobgp service logs shows which filesystem ran out and where the retry went.

note

OpenWrt: the agent writes a marker file while an upgrade is in progress; without it, the platform's package scripts remove the service during a reinstall. If that marker cannot be written — in practice, when the temporary directory is genuinely full — the upgrade stops with an error instead of proceeding. The node stays on its current version rather than removing its own service.

systemd hosts​

On a systemd host the agent runs step 4 outside its own service, as a transient unit. Step 5 restarts the whole service the moment the new binary lands, and that restart kills everything in the service's control group — a package manager still mid-transaction dies with it, which on Debian and Ubuntu strands dpkg's journal and makes every later apt run fail until it is repaired. Running the install in a unit of its own puts it out of that blast radius while still returning its output and exit code.

From agent 0.4.132 there is a second form for hosts that refuse the first. The wrapper normally hands the agent's own output streams to systemd over D-Bus (systemd-run --pipe), and some hosts refuse exactly that descriptor passing while every other systemd operation on the box works — measured on Oracle Linux 9.8 with dbus-broker, where systemd-run --collect --wait --pipe answers Failed to start transient service unit: Connection reset by peer and nothing is logged anywhere.

  • Untreated it was permanent and silent. The wrapper failed before the package manager ran, so every auto-upgrade failed the same way for ever. One beta node refused four consecutive releases across three days while fully online; from the fleet it looked like a node that was simply not upgrading.
  • The fallback is a transient scope (systemd-run --scope), which keeps both properties that matter: it runs outside the service's control group, so it survives the restart, and it inherits the agent's own pipes, so output and exit code come straight back. The agent logs systemd refused a --pipe transient unit … installing in a transient scope instead, once.
  • It is tried only when systemd refused to create the unit at all, never when the package manager itself failed — a package manager that has already run must not be run a second time under a different wrapper. A host that can create neither form keeps failing, loudly, in the log.

Nothing changes off systemd: in containers, in the foreground, and on macOS, OpenRC and procd the install command runs directly, because those restart by process exit rather than by a control-group kill.

Supervisors off systemd​

The agent runs as two processes: a small supervisor the service manager starts, and the agent it manages. An upgrade has to replace both, or the supervisor keeps running the old program for as long as the machine stays up.

Through agent 0.4.146 the whole-service restart ran only where the agent could recognise systemd, plus Windows. Everywhere else it re-exec'd only itself, so the supervisor stayed behind — measured across the fleet on 22 September 2026, three Synology nodes were running a supervisor 34 to 55 days old and both OpenWrt routers one 31 days old, each still running a program file that had been replaced, while the directory reported the current version of the child beside it. Synology is the case that shows why: DSM 7 is systemd, but a release old enough to lack the marker the agent looked for and with no systemd-run on the box, so it read as a machine with no service manager at all.

From agent 0.4.147 the agent instead asks the machine's own service records whether the process that started it is the service, and restarts the whole service wherever it is. That covered Synology DSM, OpenWrt (procd), OpenRC and sysv init alongside systemd and Windows; from agent 0.4.148 it covers macOS too, where the agent asks launchd for the process its own job runs. A Mac was left out because the first version of the check read Linux's /proc, which macOS does not have, so every Mac answered "no" — measured on 23 September 2026, a Mac running a 0.4.147 agent under a supervisor started six days and several releases earlier.

  • The upgrade that brings the fix still re-execs in place, because the decision is made by the version that is already running — the upgrade to 0.4.147 off systemd, and the upgrade to 0.4.148 on a Mac. The supervisor is replaced at the next upgrade after that — or immediately, with sudo nobgp service restart.
  • A stale supervisor now says so. When the agent starts and finds its supervisor running a program file that has since been replaced, it writes one warning in its log; nothing else on the node reports the supervisor's version. Read it with nobgp service logs or node_logs.
  • A restart the service manager reports as failed no longer leaves the node without a service. On sysv and busybox init the script gives the stop 10 seconds and then refuses to start the service again, while the agent's own orderly shutdown can take about 20; on macOS a launchd job that was unloaded and then not loaded again stays unloaded, and KeepAlive does not bring back a job that is not loaded. In both cases the agent now waits up to 30 seconds for the old supervisor to exit and starts the service itself, trying the start up to three times — see nobgp service restart.
  • OpenRC configured to tear down the service's control group (rc_cgroup_cleanup=YES) keeps the in-place re-exec. Stopping the service there would kill the very helper that is meant to start it again.
  • Containers and a foreground agent still re-exec in place, since there is no service manager whose record could name the agent's parent as the service.

Debian and Ubuntu hosts​

Step 4 normally installs the .deb with apt-get. Two host conditions make apt refuse the install outright, and the agent recovers from both on its own:

  • The host's own dependency tree is broken. If unrelated packages on the machine have unmet dependencies (a common cause is a hand-installed .deb from a different release), apt refuses every install, including ours. The agent falls back to dpkg -i, which enforces only the noBGP package's own dependencies (iptables, procps, ca-certificates — already present on any host running the agent) and leaves the pre-existing breakage exactly as it found it. The agent deliberately does not run apt --fix-broken install, because that removes whichever of your packages conflict — a decision for you to make, not a side effect of an upgrade.
  • A previous package operation was interrupted. If dpkg's journal was left stranded (E: dpkg was interrupted, you must manually run 'dpkg --configure -a'), the agent runs that repair once and retries the install.

Neither fallback touches a package operation you are running yourself: lock contention from a concurrent apt/dpkg is left alone, and the upgrade is simply retried later.

If the fallback also fails, the agent reports both errors and the node stays on its current version. That means your host genuinely needs attention — resolve the broken dependencies with sudo apt --fix-broken install (reviewing what it proposes to remove) and the next upgrade attempt will succeed.

The install script applies the same apt → dpkg fallback, so a first-time install on a host with a broken dependency tree works too.

Uninstallation​

The simplest way to remove noBGP on any platform is the built-in nobgp uninstall command. It stops and removes the service, deletes the binaries, and clears runtime files in one step:

# Keep the node identity so a reinstall re-enrolls as the same node
sudo nobgp uninstall

# Fully remove the agent, including configuration and credentials
sudo nobgp uninstall --purge

By default the command keeps the profile's identity files (.key, .jwt, .yml) so that reinstalling later restores the machine as the same node. Add --purge to remove the configuration directory too, so a subsequent reinstall enrolls a brand-new node. Pass --force (-f) to skip the confirmation prompt. On Windows, run it from an Admin PowerShell session.

If you installed via a package manager and prefer to remove noBGP that way instead, the manual steps below still work.

Linux​

# Stop and uninstall service
sudo nobgp service uninstall

# Remove the agent binary (location varies by distro)
# Ubuntu/Debian:
sudo apt purge nobgp

# Alpine:
sudo apk del nobgp

# Or manually:
sudo rm $(which nobgp)

# Remove configuration
sudo rm -rf /etc/nobgp/
info

apt remove vs apt purge:

  • apt remove leaves configuration files intact
  • apt purge completely removes configuration files
  • If configuration files are present when reinstalling, the device will be automatically restored to its previously configured state

macOS​

# If installed via Homebrew
brew uninstall nobgp

# If installed via install.sh
sudo nobgp service uninstall
sudo rm $(which nobgp)
sudo rm -rf /usr/local/etc/nobgp/

Windows​

Run the following in an Admin PowerShell session:

# Stop and uninstall service
nobgp service uninstall

# Remove the agent binary and configuration
# (The installer places nobgp.exe in C:\Program Files\nobgp\)
Remove-Item -Recurse -Force "C:\Program Files\nobgp"
Remove-Item -Recurse -Force "C:\ProgramData\nobgp"

Next Steps​

Now that your agent is installed:

Support and Resources​

License​

The noBGP agent and noBGP router are proprietary software. All rights reserved. Unauthorized copying, modification, distribution, or use of this software, via any medium, is strictly prohibited. Users of the software must comply with the terms of the license agreement.

For licensing inquiries, please contact licensing@nobgp.com.