Core Concepts
Understanding these core concepts will help you get the most out of noBGP.
Networks
A network is an isolated overlay that connects your infrastructure together. Think of it as a private virtual network that spans across all your machines, clouds, and locations.
Key Features
- Isolation: Each network is completely separate - nodes in one network can't see or access nodes in another
- Secure: All communication within a network is encrypted end-to-end
- Distributed: Networks span across clouds, data centers, and edge locations seamlessly
- Authenticated: Access controlled by registration keys and OAuth permissions
Registration
Nodes join a network through registration. There are two ways to register:
- Interactive (recommended): Run
nobgp register— this opens your browser for OAuth sign-in, then lets you pick which network to join. No keys needed. - Key-based (for automation): Pass a registration key via
nobgp register --keyor theNOBGP_KEYenvironment variable. This is designed for Docker containers, CI/CD pipelines, and other headless environments.
Registration keys are base64-encoded tokens managed through the noBGP web dashboard. Treat them as secrets — don't commit to git.
After registration (either method), the agent receives a JWT token for ongoing authentication. The registration key is removed from the config file.
Default Network
Every noBGP account comes with a default network automatically. When you ask your AI assistant to provision nodes or show networks, it uses your default network unless you specify otherwise.
Nodes
A node is any machine connected to a noBGP network. Nodes can be:
- Cloud instances (AWS, GCP, Azure, etc.)
- Physical servers in your data center
- Raspberry Pis and edge devices
- Docker containers
Node Components
Every node runs the noBGP agent - a lightweight background process that:
- Maintains connection to the noBGP router
- Handles encrypted communication with other nodes
- Executes commands from your AI assistant (with your authorization)
- Reports system status and metadata
Node Identity
Each node has:
- Name: Human-friendly identifier (e.g., "web-server-1", "raspberry-pi"), unique within its network
- Network: Which network it belongs to
- Status: Online or offline — plus deprovisioned, which removes it from the directory while keeping its identity
- Metadata: OS, architecture, IP addresses, etc.
A node's identity is its key, not its name, from router 0.4.110. When an agent reconnects, noBGP matches it on the key it registered with and falls back to its name only if that key is unknown — a machine that reinstalls mints a fresh key and is still recognised by name, which is why both arms exist. Two consequences worth knowing:
- Where the two disagree, the name noBGP holds wins. Editing
node-nameon the machine and reconnecting no longer renames anything and no longer enrols a second node carrying none of the first one's labels, grants, services or storage — the node comes back as itself, under the name noBGP has for it. - A machine that both re-keys and changes its name between connections is genuinely a new node, because nothing links it to the old one. That was already true.
nobgp register follows the same rule from router 0.4.116. Re-registering a machine — the repair path for a node whose token expired while it was offline — used to match on the name alone, and the agent presents the machine's hostname there, so a node that had been renamed could not re-register at all.
Renaming a node needs control of it — its owner, at any role, or an Owner or Admin of its organization. Router 0.4.112 moved it off the Member tier, where any member could rename any node in the organization; router 0.4.179 gave it back to the person who owns the machine, so a Member renames what they provisioned without being promoted. The same rule now covers configuring a node, setting its release channel, labelling it, publishing its services, stopping or removing its compute, and deciding who else may use it. See Roles & Permissions.
It can be done in the web app or, from router 0.4.120, by your assistant with node_rename. The node keeps everything but the name — the same node ID, its labels, its role grants, its published services and their public URLs, and its storage area — which is why renaming is not the same as leaving the network and rejoining, and why it is the only way to change a name without breaking every shared service link. Three things are worth knowing before you do it:
- The machine is never told, and nothing needs restarting. The agent goes on calling itself whatever its own
node-namesays and goes on being recognised by its key. - Peers that are already connected keep resolving the old name until each of them reconnects, after which both spellings work on that peer until it reconnects again. noBGP itself answers with the new name from the moment of the rename, so
network_directoryis right immediately. Nothing is misdirected in the meantime — a session is keyed on the node rather than on the name. - The old name is freed, and noBGP remembers that this node released it. A machine that reinstalls later still presenting the old name is never handed a different node's record: where the name is contested it is admitted under a suffixed name with a new node ID and an empty storage area. Also note that
provision_nodere-adopts by name, so after a rename the old name provisions a new node and the new name brings this one back.
Only three states exist — online, offline and deprovisioned — and from router 0.4.112 there is no fourth: a provisioned machine noBGP stops for its deadline, the compute allowance, an empty credit balance or a lapsed payment leaves an ordinary offline node, with nothing extra to look for and no timer counting down.
A contested name costs the name, not the connection
From router 0.4.113, an agent presenting a name noBGP cannot confidently attribute is enrolled under the next free name in the family — web-1, then web-2 — rather than being turned away. Registration with a valid network key always succeeds; what a contest costs is the name, never the connection. Two situations are contests, and neither is visible from the machine:
- Another node that still exists was renamed off this name and is not answering. A machine that reinstalls mints a fresh key, so an agent presenting
webwith a key noBGP does not know is indistinguishable from that node reinstalling — and handing it the row would hand it another machine's id, labels, role grants, services and storage. - The node holding the name is answering right now on a different key — two machines configured with the same name. Through router 0.4.112 the second was refused outright and could not connect at all, with the reason visible only in its own log; a naming accident became an outage. It now enrols beside the first.
The incumbent is untouched either way — same row, same key, same connection.
"Answering" means the agent itself replied, from router 0.4.153. noBGP puts the question to the node over its own connection and gives the agent two seconds to answer — from router 0.4.154, longer on a link whose measured round trip needs it, up to five seconds, and the question is asked twice inside that wait, so one lost packet on a slow link cannot cost a live node its name. A machine that is genuinely gone waits out the same window it always did, so replacing one is no slower. Through router 0.4.152 the question was put to noBGP's own end of that connection, which stays open for up to 90 seconds after an agent is killed, so a machine whose agent had just been stopped hard — a power cut, a kill -9, a container killed rather than shut down — read as answering, and its own restarting container lost the name to a suffix and came back with a new node ID and an empty storage area. A node that is asked and does not answer is also marked offline there and then, rather than when its connection eventually times out.
A returning workload gets its own suffixed row back. The first name in the family that is not answering is re-adopted rather than a fresh one minted, so a container that loses its key on every start keeps one node id and one storage area across restarts instead of accumulating an empty one per restart. The walk starts after any numeric suffix on the name presented, so a machine configured as web-3 is considered for web-4 upward and never takes web-1, which belongs to whoever asked for it.
⚠ A revoked node is still refused. Revocation is a deliberate act against one identity, and quietly minting a working node beside it would read as working around it.
⚠ Two different unidentifiable machines presenting the same base name land on the same suffixed row, and so share that row's node storage. They are in one organization, so this is mis-attribution rather than a boundary being crossed — but where it matters, give machines distinct names.
⚠ The name is noBGP's, and the machine does not learn it. The agent goes on calling itself whatever its own node-name says, so an unexpected web-1 in network_directory is where a contest becomes visible. Address the node by the name noBGP holds.
provision_node asks the same question before it starts a container, so the two surfaces cannot disagree about when a name is contested.
The browser registration page asks it before you approve, from router 0.4.164. A name a live node answers on is refused there, with a free name from the same family offered instead, and a name an offline node holds is taken over only when you say so — so neither the suffix nor the takeover arrives as a surprise on a machine you enrolled from the browser. See nobgp register.
Node Discovery
Your AI assistant can discover nodes by asking noBGP:
Show me all nodes in my production network
This returns real-time status for all connected nodes.
Node Labels
Beyond what a node reports about itself (platform, OS, architecture), you can attach your own labels — key/value pairs that keep their type (string, number, boolean, or an RFC3339 timestamp):
Label the Raspberry Pis in production with role=edge and site=irvine
Labels do two jobs:
- They remember things across sessions.
last_checked: "2026-07-28T17:04:00Z"orreplicas: 3reads back unchanged next time. - They name a cohort. Anywhere noBGP asks "which nodes?" — watching files, dispatching a command, granting access — you can say
labels: {role: edge}instead of listing machines. Because cohorts are re-evaluated live, a node labelled later joins on its own.
Labels are set by you, never by the agent, so a node cannot tag itself into a cohort. Setting them needs control of the node — its owner, or an Owner/Admin of its organization: a label can move a node in or out of the reach a node_grant handed to another machine, so it sits with that switch rather than with ordinary use. A selector that spans nodes you do not control relabels the ones you do and lists the rest in refused. See node_label.
DNS & Hostname Resolution
Nodes in the same network can reach each other by name. When the agent connects, it configures your system's DNS so that bare hostnames resolve automatically:
# From any node in the same network, reach others by name
ping web-server
curl http://api-server:8080/health
ssh raspberry-pi
Under the hood, each node gets its own overlay zone, <node-id>.nobgp.net, and the agent adds it to your system's resolver configuration as a DNS search domain. This means ping web-server expands to web-server.<node-id>.nobgp.net and resolves to the peer's overlay IP address — no manual /etc/hosts entries or DNS records needed.
The zone is answered by a DNS server built into the agent itself, running locally on the node. Lookups never leave the machine, so overlay names keep resolving even when the node cannot reach the router.
The agent integrates with your platform's native DNS to route the zone to that local server:
- Linux: resolvectl/systemd-resolved, NetworkManager, OpenWrt's dnsmasq, or
/etc/resolv.conf - macOS:
/etc/resolver/directory - Windows: NRPT rule plus the DNS suffix search list
On most hosts this is true split-DNS: only the overlay zone is routed to the
agent and every other name keeps its normal path. Some Linux configurations
offer no way to express that (plain /etc/resolv.conf, Alpine), so the agent
lists itself as the first nameserver and therefore sees every query. There
it forwards anything outside the overlay zone to the machine's own
nameservers, so ordinary internet names resolve exactly as before — including
for applications whose resolver treats a refusal as a fatal error rather than
trying the next nameserver (Node.js and Bun, among others). The host's
nameservers are re-read every minute, so a DHCP lease that swaps them is picked
up without a restart. ⚠ Being first only decides anything where the machine's
resolver asks its nameservers in order, which musl — Alpine Linux and the other
musl-based images — does not: from agent 0.4.137 such a machine takes a
different path, below.
On macOS the search domain your network hands out is kept, from agent 0.4.121. The agent writes the overlay zone into each network service's search list, and that list is the static one macOS layers over the one a DHCP lease supplies — so writing it used to replace the leased domain, and bare LAN names (ping nas, ssh printer) stopped resolving for as long as the agent ran. The agent now reads the lease itself, per interface, and writes it back behind the overlay zone, so both keep working. Stopping the agent restores the list the machine had before it started, rather than re-deriving it — a laptop shut down off the network could not be asked what its lease said, and an entry left behind in the static list would have suppressed every future DHCP search domain on every network it joined, outliving even nobgp uninstall.
On OpenWrt a search suffix something else removed is put back, from agent 0.4.152. There the overlay zone is routed to the agent through the router's own DNS service, and the search suffix that makes a bare name work is written into /tmp/resolv.conf — a file other software on the router rewrites for its own reasons, with no interface event for the agent's hook to see. When that happened the suffix was gone until the next event, so ping web-server stopped resolving while the fully-qualified web-server.<node-id>.nobgp.net kept working, for hours. The agent's DNS monitor now reads that file once a minute and puts its own suffix back, so a bare name recovers on its own within a minute. It only ever restores the suffix this node asked for, and it is not put back after the agent removes it on shutdown.
Hostname resolution works across clouds and locations. A node in AWS can reach a Raspberry Pi at home simply by name, as long as both are in the same network.
To look up a peer's overlay address directly — without going through your system's DNS — use nobgp resolve. It asks the agent for the address straight from the mesh directory, which is handy in scripts (ssh admin@$(nobgp resolve web-server)) or on hosts where overlay DNS is reported as unavailable.
On Alpine and other musl hosts the agent owns the resolver file
From agent 0.4.137 the agent lists itself as the only nameserver on a musl machine, instead of putting itself first. Alpine Linux and the other musl-based images resolve names differently from glibc: the C library asks every nameserver in /etc/resolv.conf at the same time and takes the first reply that is not an outright refusal, so being listed first decided nothing. An upstream nameserver's no such name for an overlay name came back in a millisecond or two while the agent was still asking noBGP for a name it had not seen before, and that answer won the race — so the first lookup of each new peer name failed, with Name does not resolve, and only a second attempt succeeded. With the agent as the sole entry in the file there is nothing left to race.
- Ordinary names resolve exactly as before. The machine's own nameservers are read out of the file and become the servers the agent forwards everything outside the overlay zone to, and they are re-read whenever something rewrites the file.
- The original file is backed up and put back when the agent stops normally and on
nobgp uninstall. - ⚠ While the agent is not running, the machine has no DNS at all — it is the only nameserver in the file. The service manager restarts a crashed agent within seconds; an agent you stop yourself leaves the machine without DNS until you start it again or restore the backup. That is the price of overlay names resolving on a platform that offers no way to route one zone.
- A run that ended without putting the file back is undone at the next startup, from agent 0.4.138. The agent restores the machine's own file before it resolves anything, and takes it over again once its DNS server is up. Before, the restarting agent asked the dead server its predecessor had left in the file: where DNS goes out only through the machine's local resolver — corporate egress rules, many Docker and Kubernetes setups — it could not reach noBGP, never registered, and the machine stayed without DNS until someone restored the backup by hand. It restores only its own profile's takeover, so a file another profile currently owns is left alone, and a machine with no backup to restore to is left as it is.
- Only musl machines move. Whether the machine's own programs use musl is read from
/bin/sh, not from the distribution name, so a glibc host with musl installed alongside is unaffected — glibc asks the nameservers in order, so listing the agent first does decide there. - A NetworkManager-managed machine is left as it was, because NetworkManager rewrites
/etc/resolv.confat every DHCP lease event and would undo the takeover each time. - ⚠
/etc/resolv.confhas one owner, so on a machine running several noBGP profiles only the first takes it over. The others stay on the first-nameserver path, where a cold name's first lookup can still lose the race, and say so in their log.
A name one member reaches on its own LAN
Not every name a network resolves belongs to a node. When a member is asked for a name it does not hold, it can ask noBGP, which asks the other members — and a member that reaches that name on its own LAN answers for it: a switch, a printer, a NAS on the same subnet, none of them running an agent. noBGP records that answer as that member's claim, with a lease on it, and points any other member that asks at the member that vouched.
From router 0.4.136 such a claim is delivered only to the member that made it. Before, it was pushed to every member of the network, so machines that had never asked for the name were handed a handle for it — and because the overlay zone is first in the DNS search list, that handle answered their bare-name lookups ahead of their own LAN. A machine sitting on the same subnet as the device therefore resolved the name across the overlay to the far member instead of reaching it directly, with no way to prefer the local one short of removing the agent.
- The name stays resolvable network-wide. A member that wants it asks for it and noBGP answers, pointing at the member that vouched. What ended is the unasked-for push to machines that never looked the name up. ⚠ Resolvable is not the same as reachable — see below.
- ⚠ Fixed with it: on the member that vouched, the name resolved to that machine itself. Measured 7 September 2026 on agent 0.4.114 — a gateway's own lookup of a switch on its LAN answered with the gateway's address, so
ping switchysucceeded against the gateway while the switch stayed unreachable, and traffic other members sent through the gateway was delivered to the gateway rather than forwarded. The vouching member now resolves the name to the real LAN address — and noBGP never hands a node back its own claim, so a member asking about a name it is itself the answer for is told nothing and settles it on its own LAN. - ⚠ Whether packets sent to that name arrive is a separate question from whether it resolves, and the answer depends on the member that vouched — see Reaching the device through that member directly below. From agent 0.4.128 a member that cannot carry the traffic does not claim the name at all, which turns that gap into a name served by somebody who can, or by nobody.
- A device that answers only over mDNS can be discovered too, from agent 0.4.128. Until then a member vouched from its hosts file and its LAN's DNS servers alone, so anything whose name exists only as
<name>.local— most consumer speakers, printers, smart-home hubs and KVM appliances announce themselves that way and appear in no DNS zone — could not be found by any member. A member that finds neither a hosts-file entry nor a DNS record now asks for<name>.localon each of its own LAN links and vouches for what answers.- It never takes port 5353 from the host's own responder (mDNSResponder, avahi, the Windows DNS client): the query is the ordinary one-off kind, sent from a normal port and answered directly.
- Only real LAN links are asked — those holding a private address — because mDNS is link-local and a query on an uplink reaches nothing.
- It applies to a bare name, or one already ending in
.local. A dotted name in some other domain (nas.lan) belongs to DNS and is not asked for this way. - ⚠ Windows names that exist only over LLMNR or NetBIOS are still not found. Those are different protocols; only mDNS was added.
- A device whose own name is not a host name is served under a label, from agent 0.4.147. Some devices lease a DHCP name nobody can type —
WiiM Mini-9D84, with a space — which the LAN DNS server keeps and publishes no record for, so no member could reach it. A member that runs its LAN's DHCP server now reads its own leases and vouches for the label that name maps to (WiiM-Mini-9D84). It is the last source, after the hosts file, LAN DNS and mDNS, so a real name always wins; only dnsmasq's leases are read, and only on the node serving DHCP. See Devices whose name is not a host name. - The machine's owner decides which LAN names it will serve, from agent 0.4.128, with the
lan-namessetting —nas*,printerto serve two devices,*,!routerto serve everything but one,offto serve none. The default is every name, which is what every earlier release did. A name outside the list is not looked up at all, so the node reveals nothing about whether that host exists;net_reachreports it asresolution.refused: "lan_names". It bounds only what the node answers for the network — names its own applications look up are untouched — and there is no remote path to it: it is the machine owner's say over their own LAN. - Node names are untouched. Every member still learns every other node in the network exactly as before; only these discovered LAN claims are scoped.
- A claim nobody keeps confirming is forgotten about a day later, from router 0.4.150. Every lookup that resolves the name on that member's LAN renews noBGP's record of it, so a name in use stays; once nothing renews it, noBGP deletes the claim about 24 hours later — the sweep runs hourly, so in practice 24 to 25 — and tells the member holding it that the entry is gone. The next lookup of the name, from anywhere in the network, establishes it again if a member can still reach the device. It used to be seven days, so a device that had left the LAN, or a name nobody asked for any more, stayed visible and resolvable network-wide for about a week. ⚠ A day is the horizon for a claim nobody asks about. A name somebody is still looking up goes far sooner once the member that claimed it stops vouching — about two minutes, from router 0.4.167.
- ⚠ The shorter horizon relies on a member renewing a name it is using, which is agent 0.4.125. An older agent goes back to noBGP for a name only once nothing on the machine is addressing it, so a name it keeps busy is never re-confirmed and is withdrawn once a day rather than once a week. Upgrade a member that vouches for names you depend on.
- ⚠ The transition is not instant on a node that already holds one. An agent's directory is add-only and kept across restarts, so a handle a member was pushed before this release is withdrawn when noBGP ages the claim out — a week past its lease — or at once on a fresh install. From router 0.4.141 there is nothing left to wait out: that release cleared every discovered claim noBGP was holding and withdrew each one from the member that held it — see One name, one answer.
- The node owner decides which LAN names the node answers for, from agent 0.4.128. The
lan-namessetting takes patterns such asnas*,printeror*,!router, and a name outside the list is never looked up. From the same release a node also finds devices that have only an mDNS name.
When several members claim one name
From router 0.4.155 noBGP serves the claim most likely to deliver. More than one member can genuinely reach one device — two machines on the same subnet, a laptop that moves between offices — and each of them vouches for the name, so noBGP has to pick which one to point everybody at. It prefers, in this order: a member that is online, then one that is still vouching for the name (router 0.4.161), then the longest-standing claim.
The member's platform does not enter into it, from router 0.4.166. Until that release a member that was not macOS or Windows was preferred, because neither could carry the LAN packet — a difference agent 0.4.131 removed. After that the preference only took names away from members that could serve them: measured 18 September 2026, a container on a Mac read the Mac's own hosts file, claimed a name it could not reach and won it from the Mac, which could — 0 of 11 clients reached the device through the container, and 18 of 18 reached it once the container stopped claiming the name. Whether a member can deliver is the member's own call, expressed by whether it vouches at all, and noBGP does not overrule it from the platform.
- It is a preference, not a filter. A name only an offline member claims still resolves to that member — a worse answer is better than none, and the member's own report is what decides whether it can forward, not a rule noBGP applies over its head.
- ⚠ A member that cannot deliver does not claim the name in the first place, from agent 0.4.128: such a member says nothing, noBGP asks the next one, and where none can deliver the name does not resolve at all. That is deliberate — a name that resolves and answers nowhere is the failure this ends. Between agents 0.4.128 and 0.4.130 that took a device whose only route was through a Mac or a Windows machine out of reach; agent 0.4.131 gives those members a delivery path, so upgrading one is what brings the name back.
- ⚠ Whether a member is still vouching used to be asked after the member's platform, and router 0.4.161 put it first. Until that release a Linux member that had withdrawn a name — its owner narrowed
lan-names, or its delivery leg came down — kept being served ahead of the macOS or Windows member that had just taken the name over and could deliver it, and every lookup landed back on the member that had stopped serving it. Router 0.4.166 dropped the platform preference altogether, so on a current router the question does not arise. - A member that is reconnecting keeps the name, from router 0.4.176. A member counts as online for two minutes after its connection ends, and its claim counts as still vouching for the same stretch — so an agent restart, a Wi-Fi blip, or a noBGP deployment that reconnects every member at once no longer moves a name onto a different member and back. Before it, a lookup landing in that window picked another member's claim, which is a different session key, and the machine that asked held it for a whole lease: two moves for one blip, and two machines that resolved either side of a one-second gap disagreeing about where the name lives. ⚠ The cost is that a member that has really gone keeps the name for those two minutes, so the worst case for moving off it is the five-minute claim lease below plus that grace.
- ⚠ Failing over is not instant. The pick moves on the next lookup, and noBGP only knows the serving member is gone once its connection times out — about a minute, or up to about thirty when the noBGP instance serving it went down with it.
- ⚠ A member already holding the answer keeps it until its lease lapses, which is the bound below — unless the member serving the name tells it otherwise, which from agent 0.4.130 is the shorter path.
A claim answer now carries a five-minute lease (router 0.4.155), where a node name carries an hour. A claim can move between members and a node's identity cannot, so the shorter lease is what bounds how long a member goes on sending traffic to a gateway that has gone offline. ⚠ Five minutes is noBGP's half of that bound, not the whole of it: add the member's own re-check interval and, where a renewal cannot be completed, its retry backoff — so the honest worst case is closer to twenty minutes. The cost is a renewal every five minutes for a LAN name carrying traffic; a name nothing is using is not renewed at all. ⚠ From agent 0.4.130 the lease is not the only bound: a member that has stopped delivering the name says so, which ends the wait in seconds — and covers the case the lease never did, a machine whose traffic kept its own address from ever lapsing.
A claim noBGP can no longer confirm is re-checked with the members before it is answered (router 0.4.160). The record behind a claim is stamped by the vouch that created it and stands for a minute, and nothing else renews it — so at almost every five-minute lease renewal the record noBGP holds has gone stale, even for a name that works perfectly. Until this release noBGP answered from that stale record at once and asked the members behind the answer, so a member that had withdrawn the name went on being served for another whole lease: measured, a machine took 8 minutes 34 seconds to move onto a member that could still deliver it. noBGP now asks the members first — for at most 1.5 seconds, comfortably inside the two seconds the asking machine waits before it gives up — and answers with what comes back.
- A member that vouches during that moment is what the answer names, whether it is the member that had the name or another one, and it carries the usual five-minute lease. A renewal is therefore where a name moves, rather than where it is pinned again to a member that has stopped serving it.
- Where no member answers, the stale record is still served — a known answer beats none, and the device may simply be quiet — but with a 30-second lease instead of five minutes, so the machine that asked comes back promptly and picks up a member that returns. ⚠ From router 0.4.167 that holds while the member holding the claim is offline; one that is online and stays silent loses the claim instead — directly below.
- ⚠ The cost is latency on that one lookup. A member that does not know the name says nothing at all rather than answering no, so the ask usually runs its full second. A claim noBGP can still confirm is answered immediately, exactly as before, and where the last ask found no member noBGP does not ask again for a minute.
A claim whose member is connected and stays silent is dropped (router 0.4.167). Silence is the answer not here: a member that cannot reach the name sends nothing at all rather than saying so, so a member that is online, is asked, and answers none of three asks in a row no longer holds the name. noBGP deletes the claim and tells that member the entry is gone, and where no other member claims the name the lookup is answered not found — never the record it read a moment before the ask. Until this release such a name was served until the claim aged out a day later, so every machine asking for it was sent through a member that could not deliver: measured, more than 90 minutes on one switch.
- A name that has really gone stops resolving after about two minutes. noBGP asks the members for one name at most once a minute, so three asks is the two-minute figure rather than something longer.
- ⚠ An offline member is never counted. Only connected members are asked, so a member that is restarting or briefly disconnected cannot lose the names it serves; that case is still bounded by the day-long horizon above.
- Three asks, not one. A single lost message costs nothing, and any answer puts the count back to zero. The count is held in memory by the noBGP instance that did the asking, so it can start again after a restart — which delays the drop and never causes one.
- Nothing is lost by dropping it. The next lookup of the name establishes it again from whichever member can reach the device, exactly as the first one did.
- ⚠ A member has to answer inside that second, and from agent 0.4.137 it does. Because silence is read as not here, a member that reaches the device but answers slowly loses the name exactly as one that cannot reach it at all. Until 0.4.137 a member asked its LAN DNS servers first and only then went out over mDNS, inside a two-second budget of its own — so a device that answers only a resent mDNS query (a WiiM speaker was measured at 272–689 ms) or a LAN whose first nameserver does not answer took longer than noBGP waits, and a name that was there read as gone. The member now runs both lookups at once and answers within 800 ms, whichever source replies first. Upgrade members that serve mDNS-only devices; one below 0.4.137 can lose a name it genuinely reaches.
To see which member a name is being served from, ask a node. net_reach resolves the name on one node and probes what it got back, naming the member that delivers it; net_reach_many asks every node in the network at once and counts the answers, which is the quickest way to tell the name is wrong from one gateway cannot forward.
When the member serving a name stops delivering it
From agent 0.4.130 that member says so, and the machine sending to it looks the name up again within seconds. A member handed traffic for a name it does not serve — the device left its LAN, its delivery leg came down, or noBGP has since pointed the name at somebody else — answers on the same connection saying it does not deliver that name. The machine holding the address then asks noBGP for the name again and lands on whichever member noBGP names now, on the same overlay address, so nothing running through it has to be pointed anywhere new.
- ⚠ Before this release nothing ended that black hole. A machine looks a name up again only once its address has stopped being confirmed, and an address carrying traffic was never in that state — so a machine that kept sending to a withdrawn gateway went on doing so for as long as the traffic continued. That is the case the twenty-minute figure above does not cover.
- The machine waits a random moment of up to ten seconds first. Every peer of one withdrawn gateway is told at about the same time, so without the spread they would all ask noBGP in the same second.
- An offline member cannot say anything, so that case is read from the absence of an answer: after four connection attempts in a row that nothing answers — about 20 seconds — the sending machine asks noBGP again anyway. One unanswered attempt is a slow noBGP rather than a gone member, and a member can take more than ten seconds to reconnect after a noBGP deploy, which is why it is not the first.
- An address noBGP no longer confirms is renewed too while something is still using it, from the same release. Renewal used to skip such an address, which is precisely the state a withdrawn name leaves it in.
- Asking again is rate-limited per name — at most once a minute at first, doubling to fifteen — so a name nobody serves any more costs one lookup a minute rather than one per packet.
Reaching the device through that member
From agent 0.4.131 every member carries the traffic, on every platform, and there is nothing to configure. The member ends the flow from the other node inside the agent and opens an ordinary connection to the device, so its own kernel picks the machine's LAN address as the source and the reply comes back to the socket that asked. No packet is forwarded, no source is rewritten, and no firewall rule is involved. Measured 16 September 2026 on macOS, Windows and Linux, with nothing configured in any firewall.
- What this replaced was a return-path problem no rule could solve off Linux. Until agent 0.4.129 the member forwarded the packet onto its own LAN with an overlay address as the source, and the device had no route back to it — measured 8 September 2026 on the device itself, 85 pings arrived and not one reply left. Linux could rewrite that source with an
iptablesrule; macOS and Windows had no equivalent, so from agent 0.4.128 a member on those platforms declined to vouch for LAN names at all. Measured 15 September 2026: a macOS and a Windows gateway each got 0 of 4 replies where a Linux gateway on the same LAN got 4 of 4. - ⚠ It carries TCP, UDP and ICMP echo only. Kernel forwarding carried any protocol and a proxy cannot, so anything else addressed to such a name is dropped. A ping is answered only when the real device answers the member — nothing is invented for a device that is silent.
- From agent 0.4.140 the member reaches the device over IPv6 as well as IPv4. It uses every address the name resolves to on its LAN, preferring the IPv4 one where the device has both, so a device that answers with an IPv6 address alone is served rather than refused — see the bullet on a device with no IPv4 address below. Through agent 0.4.139 the delivery leg was IPv4 only, the same limit the forwarding path had.
- At most 512 relayed TCP flows and 256 UDP flows at once on one member, counted separately from the flows it carries to its own services. Past the cap a new flow is refused, with at most one log line a minute.
- A LAN answer of
0.0.0.0, a loopback, a multicast or the broadcast address is refused. Pi-hole and hosts-file blocklists answer0.0.0.0, and connecting to it would reach services the member binds to127.0.0.1. - The Linux firewall rules the old path needed are gone, and an agent that finds rules an earlier release installed removes them — the source-NAT and forward-accept chains named after the overlay interface (
NOBGP_nobgp0_NAT,NOBGP_nobgp0_FWD), the matching pair in OpenWrt'sfw4, and afirewalldbinding of the overlay interface to thetrustedzone. ⚠ That last one is worth knowing about if you ran agent 0.4.118 to 0.4.130 on afirewalldhost: while it was in place your peers could reach ports on the host itself that its own zone would have refused, and the agent now takes it back out. The sweep runs when the agent stops and again when it starts, so rules a killed agent left behind are removed too. - Nothing on the member has to be enabled for this. An ordinary socket needs no rule and no extra privilege, so a member that could not install firewall rules at all — a container without
iptables, or without the privilege to change it — now serves LAN names like any other. - ⚠ A member whose deliverer did not start serves no LAN name and vouches for none. There is no second path to fall back to, so it claims nothing rather than claiming names it would drop.
- The
lan-deliverysetting is retired. Agents 0.4.129 and 0.4.130 chose between the two paths with it; a member upgrading from either lands on this one whatever it was set to, and the agent removes the key from its profile. See How a node delivers to a LAN device. - The name survives a restart of the member that serves it, from agent 0.4.128. noBGP holds a claim of this kind as a short-lived record stamped by that member's own vouch, so a reboot — always longer than the record lives — brought the member back without it, and every packet another member sent was dropped in silence. Nothing repaired it either: a member holding the answer from before has it cached and never asks again, so it took a lookup from some third machine that had never held the name to bring it back for everyone. Measured 8 September 2026. The agent now keeps these names itself and re-resolves each one on its own LAN as it starts.
- ⚠ Fixed with it: a member up for more than a week could lose the name outright. The seven-day clock that ages out an offline peer was being applied to these names too, and nothing ever re-stamps it for a name a machine serves off its own LAN — so on the next restart the name was dropped and its address could be handed to a different name while other machines still held the old answer. That clock no longer applies to them: whether the name still answers on the LAN is what decides it, which is the stricter test of the two.
- A name the member cannot deliver right now is kept rather than forgotten, from agent 0.4.130, and is served again on the same overlay address as soon as the member can deliver it — which from agent 0.4.131 means the next start of the agent, since there are no firewall rules left to come and go underneath it. Before 0.4.130 the member dropped such a name from its own record of it, so it came back as a new address while every peer holding the old one was black-holed until that address stopped being confirmed — minutes on macOS and Windows. Its address stays reserved for it in the meantime, and the name stays out of the directory rather than pointing at something stale.
- ⚠ A name held this way for seven days is released, so a device that has really gone cannot hold an address for ever. That clock counts how long the name has been undeliverable; it is not the seven-day clock for an offline peer, which does not apply to these names at all.
- A name that no longer answers on that LAN is not re-armed, and says so. The address stays reserved for it, the name is left out of the directory rather than pointed somewhere stale, and the agent logs which name it did not restore and why — readable with
nobgp service logsornode_logs. Re-resolving on every boot is deliberately stricter than trusting what was recorded: the device may have moved to another address while the member was down, and delivering to the old one would quietly reach the wrong host. - ⚠ A name the member can no longer reach stops being delivered to while it runs, from agent 0.4.122. The member re-checks each such name on its own LAN at every directory sync and refuses one it cannot serve — but until this release the refusal only stopped it arming a new leg: the leg armed by the last sync that did resolve stayed up, so peer traffic went on being forwarded to an address the member had just established it could not reach, and which may by then belong to a different machine on that LAN. Nothing else took such a leg down, so only a restart of the member cleared it. The leg now comes down at the moment the member refuses, and it writes one line naming the name, the address it stopped delivering to, and why — readable with
nobgp service logsornode_logs. - The name keeps resolving; only delivery stops. Its overlay address stays bound to it — so the address cannot be handed to a different name while other members still hold the old answer — and packets sent to it are dropped rather than delivered somewhere stale. The leg is re-armed by the first sync that resolves the name again, with no restart. ⚠
nobgp statusstill lists the name with its old LAN address: that a leg has been taken down is recorded in the log, not in the status output. - A device with no IPv4 address is served from agent 0.4.140, as long as the member has a route to it over IPv6. The member vouches for exactly what it can deliver: a name answering with an IPv6 address alone is taken where this member can reach it and refused where it cannot, and noBGP then asks the other members — so on a LAN where one member has IPv6 and another does not, the name is served by the one that has it, and where none has it the name does not resolve. ⚠ Through agent 0.4.140 a
pingto such a device was not answered by a Windows member — the interface the agent used there was IPv4 only, while TCP and UDP to it worked normally. A Windows member answers it from 0.4.141.- Through agent 0.4.139 the delivery leg had only an IPv4 address to send from, so such a name was refused at registration and again at restart, with a message naming the two remedies: publish it by its IPv4 address, or give the name an A record. Accepting one used to produce a packet the member could not send at all, which closed its overlay device and took that machine's whole overlay down with it, silently; that is no longer fatal either.
- ⚠ Removing the A record from a name that already had one is the case the bullets above cover: the leg is re-armed on the member's own terms, and before agent 0.4.122 the leg to the old IPv4 address stayed armed behind it.
One name, one answer
From router 0.4.141 a network has a single name space, and a node name wins it. In one network a name means exactly one thing, and where a live node carries that name it is the node — a name discovered on a member's LAN can never sit beside it as a second answer for the same name. From router 0.4.149 a node's own hostname sits between the two.
- A node taking the name displaces the discovered claim. It applies both to a node arriving with the name — registration, provisioning — and to a node renamed onto it. The member that had vouched for the name is told the entry is gone, rather than left holding a handle that now means something else.
- A discovered claim can never take a name a node holds. The lookup already has its answer, so noBGP records nothing and keeps answering with the node. Nothing is refused to you and nothing has to be cleaned up afterwards: the name simply goes on meaning the machine.
- A name is reserved only while a live node carries it. Rename a node off
printer, or delete it, andprinteris free from then on — a device of that name on some member's LAN can be discovered under it again. Precedence is taken when a node acquires a name, never held in reserve against one it gave up. See what a rename frees. - Case is not a distinction.
PRINTERandprinterare one name here, exactly as they are everywhere else noBGP compares names. - ⚠ What this ends is an entry that could shadow a node's name on a member. An agent's directory is keyed by name, so a discovered claim and a node of the same name were two entries competing for one key on any machine holding both. noBGP itself already preferred the node; what could differ was where that member's own lookup of the name landed. The two entries can no longer both exist.
- ⚠ The discovered claims noBGP already held were cleared once, as this release landed. Each is a cache entry with a short lease, so the next lookup of such a name establishes it again — and a member that was offline at the time is told on its next connection.
A node also answers to its own hostname
From router 0.4.149 a name that is the hostname of one of your nodes resolves to that node, even when the node's noBGP name is something else. The common case is a machine you enrolled as ds918 whose hostname on its own LAN is nas: nas now reaches the node directly, rather than being discovered on some member's LAN and relayed through it.
That gives one name space three rungs, tried in this order:
- A node's name. Identity, and it wins outright — a node name answers even while the node is offline.
- A node's short hostname — the first label of the hostname the node reports, so
nas.lanmatchesnas. Case is not a distinction, as everywhere else. - A name a member reaches on its own LAN, which is what answers when neither of the above does.
- ⚠ A hostname two nodes share answers for neither of them. Where more than one live node in the network reports the same short hostname —
raspberrypi,ubuntu— noBGP declines to pick one, and rung 3 answers instead. Rename the machines, or address them by node name. - ⚠ A node stops answering its hostname about an hour after it goes offline, and this is where rung 2 differs from rung 1 deliberately. A node name is identity and keeps resolving while the node is down; a hostname is an alias, and the real host of that LAN name may be a different machine some member can still reach. An hour is far longer than any reconnect or router deploy, so a node that is merely blipping never loses the name — and a laptop closed overnight has stopped answering for it by morning.
- A node taking a name this way also displaces the discovered claim of that name, the same as a node named for it does: the member that had vouched is told the entry is gone, so the network does not end up with two answers for one host. ⚠ Where two nodes share the hostname, nothing is displaced — consistent with the rung-2 rule above.
- ⚠ Nothing is reserved against a hostname. A name noBGP refuses to answer at rung 2 can still be discovered on a LAN and saved as a claim, and a node's hostname changing moves the answer with it.
A name this machine reaches itself
From agent 0.4.116 the overlay gets out of the way for a name the machine can reach on its own LAN. Before a bare name is answered from the overlay, the agent asks whether this machine resolves it off the overlay — through its own hosts file and its own nameservers — and when it does, the lookup gets the device's own LAN address rather than an overlay address routed through another member. One hop on the local segment beats a tunnel to somebody else's, and it costs no overlay address, no lease and no translation.
- noBGP's own zone hands back the LAN address, from agent 0.4.139. Agents 0.4.116 to 0.4.138 answered not found instead and left your system's resolver to fall through to its own search domains — which works only where the resolver then asks the LAN for the name.
systemd-resolveddoes not: by default it sends a single-label name to no ordinary nameserver at all, so after the not found it had nothing left to try, and the name failed on that one machine while every other node in the network resolved it through that very machine. Measured 18 September 2026 on Ubuntu 24.04. Only the machine next to the device needs the upgrade.- The answer carries a 2-second TTL, the life of the LAN check behind it, so nothing downstream keeps pointing at a device that has left the LAN.
- The family the device does not have is answered no address, not not found, so the other family's lookup of the same name still gets the address.
- ⚠ One case still answers not found: a device that has an IPv6 address alone, on a node that has no IPv6 overlay address. No lookup on such a machine could carry that answer, so the zone declines and leaves the resolver to fall through as earlier releases did.
- ⚠ Fixed: a device on your own subnet could be answered with a path through another member. The overlay suffix is first in the DNS search list — deliberately, so a DHCP suffix like
lancannot shadow a node's name — which also means a positive answer from the overlay is final: the resolver stops there, and the fall-through you expect never happens. A plain LAN device was therefore answered with an address routed through whichever member had vouched for the name — and that relayed path carried no traffic at the time, so the one machine that could always reach the device stopped being able to. The reported case ended with a customer removing the agent to reach a switch on their own subnet. - It covers a name this node already holds, not only an unknown one. The check also runs when the name is in this node's directory and the entry is one this node hosts itself — its own node name, or a name it discovered on its LAN and published for the others. That is the shape measured above, where the entry pointed at the gateway's own address and
ping switchysucceeded against the gateway while the real switch stayed unreachable. A lookup that never misses cannot be caught on the discovery path alone, which is why this rung exists as well. - ⚠ A peer's name is never declined, however loudly the LAN claims it. Only entries this node hosts itself are given up; another node keeps its overlay address, which is exactly what the search-list ordering above exists to protect.
- A name this node hosts that its LAN cannot answer right now keeps its overlay address. Declining there would take the name away from everybody with nothing to replace it.
- The check looks off-overlay only — the hosts file and the machine's real nameservers, with the overlay's own suffix dropped from the search list and any answer inside this node's overlay range discarded — so it can never be satisfied by the overlay answer it is deciding about. It sees the same names the vouching mechanism above sees: LAN DNS records and hosts-file entries, and from agent 0.4.128 names that exist only over mDNS as well, asked for as
<name>.localon the machine's own LAN links when nothing else answers. Names that exist only over LLMNR or NetBIOS are still not seen. See Devices that have only an mDNS name. - Both outcomes are remembered for 2 seconds, so a held
pingdoes not re-probe the LAN on every retry, and a device that appears on — or leaves — the LAN is picked up on the lookup after next. The probe also shares the cap on concurrent name discovery: in a burst of distinct unknown names it answers no rather than queueing, and the ordinary path runs. - One lookup gets one answer for both address families, from agent 0.4.133. Your system asks for the IPv4 and the IPv6 address of a name at the same time, and both questions now wait for the single LAN check and take its verdict. On agent 0.4.132 — the release that gave each machine an IPv6 overlay address — the second question was answered no straight away, so one family fell through to the LAN as intended while the other was answered with an overlay address routed through another member. Applications that prefer IPv6, which is most of them on Linux, took that path.
nobgp resolveis deliberately unchanged. It answers from the mesh directory and bypasses the system resolver entirely, so it still reports the overlay address for such a name — it is the tool for what does the overlay say, while your applications go through the resolver this rung feeds.
Seeing peers and their services on this machine
From agent 0.4.128 a node can tell the network what its own host offers, and show the network's peers to this machine's own applications. They are two settings pointing in opposite directions, and the first is on by default while the second is not — see mDNS keys for the full reference.
What this host offers is reported (mdns-report-services, default true). The node looks for a fixed list of services on its own machine — file sharing, screen sharing, remote login, AFP and RDP — either because the host announces them over mDNS or because something is listening on the port, and reports what it finds. Read it back with network_directory as info.host_services, or on the machine itself in nobgp status.
- It adds discovery, not reach. Peers can already reach those ports over the overlay; what they could not do is know they were there. That is why it is on by default.
- A service found by its port alone is marked as such (
source: listener), because a listener proves the port and not the protocol. One the host announces carries the name the host gave it (source: mdns). - A service listening only on
localhostis reported from agent 0.4.140, because a peer reaches loopback on the machine from that release; before it, such a service was left out, since no peer could reach it. From agent 0.4.153 it is left out again on a node whose owner setpeer-loopback: false— the report asks the node path's own question, so it can never claim a peer reaches a port the node itself would refuse. On Windows, file sharing is reported only if the machine really shares a folder — every Windows machine listens on the SMB port whether or not it shares anything. mdns-report-services: falseis the owner's veto: the node looks for nothing, and reports an empty list, which also clears whatever it reported before.
Peers can be shown to this machine's applications (mdns-announce, default false, macOS and Linux). Turn it on and each peer is registered with this machine's own mDNS responder as <name>-nobgp.local, with one entry per service that peer reports — so another node's SMB share turns up in Finder's sidebar, and ssh pi5-nobgp.local works from a terminal that knows nothing about noBGP.
- ⚠ The registrations are local-only: they exist for applications on this machine and nothing goes out on the LAN. The address behind such a name is this node's own handle for that peer, which means something different — or nothing — on every other machine, so announcing it to a LAN would be worse than useless.
- Windows announces nothing, because it offers no way to register a name that local applications resolve and the LAN does not see. The setting is accepted and
nobgp statusreports the state asunsupported. - On Linux it needs
avahi-daemon. With none installed the names are simply held until one appears —nobgp statuscounts them aswaiting— rather than silently dropped.
Overlay Addressing
The overlay uses the 100.64.0.0/10 address space (the CGNAT range, which never collides with home/office LANs or the public internet). Each node self-assigns a small /20 slice of it for its own view of the network; peers appear as individual addresses inside that slice.
Two properties follow:
- Addresses are node-local. The same peer can have a different overlay address on each machine that talks to it. Always address peers by name — names resolve to the right address everywhere; raw overlay IPs are not portable between nodes.
- It coexists with other VPNs. If another interface already holds a 100.64.x.x address (Tailscale, WireGuard, a carrier connection), or another product already routes part of the range, the agent picks its slice to avoid the conflict — reading every routing table on Linux from agent 0.4.111, since products that use policy routing keep their routes in a table of their own. This includes carrier-grade NAT uplinks like Starlink, where the ISP itself uses this range and the agent carves a slice inside it. ⚠ If another overlay claims the whole of 100.64.0.0/10 there is no room to take: the agent reports the pool as fully claimed rather than using a slice it could never route. From agent 0.4.112 such a node still comes up, without an overlay — no peer traffic, but the control channel,
command,file, the shared drive andnobgp statusall keep working, andnobgp statusnames the reason undernetwork.overlayso it is legible from off-box. Agent 0.4.111 alone exited instead, leaving a machine nobody could reach remotely to fix the one thing that was wrong with it. See Other VPN clients.
To pin the slice explicitly (e.g. for firewall rules), set overlay-cidr in the agent config or NOBGP_OVERLAY_CIDR — any /20 within 100.64.0.0/10:
overlay-cidr: "100.64.16.0/20"
A pinned value is a preference, not a hard pin. If it is malformed, outside 100.64.0.0/10, not a /20, or overlaps a route already on the machine, the agent logs a warning and picks a replacement rather than refusing to start — so a bad value is loud in the log instead of taking the node offline, and it is not a way to override a range another product has claimed.
⚠ There is one case where a pin is honoured and an auto-pick is not, from agent 0.4.112: a host whose routing table could not be read at all. The agent picks a slice by dodging the routes already on the machine, so with no view of them it has no evidence that any /20 is free and refuses to invent one — while a slice you pinned, or one this node already recorded on a previous boot, was chosen against evidence and is kept, with a warning that it could not be verified this time. Pinning overlay-cidr to a /20 you know is free is therefore the way to bring such a host back onto the overlay.
nobgp show displays the slice in use; nobgp status lists each reachable peer with its current overlay address.
An IPv6 address beside the IPv4 one
From agent 0.4.132 a node also takes an IPv6 slice, and a peer's name resolves
to both families. Beside its /20 of 100.64.0.0/10 the node picks a random
unique local /64 inside fd00::/8 — its own, private to this machine, the
same shape as the IPv4 slice and subject to the same rules: node-local addresses,
one per peer, held for a name rather than handed around. The node's own overlay
interface takes ::1 in it, and peers are numbered from ::2 up. It is on by
default on Linux, macOS and Windows.
- Your applications choose. The node's resolver answers
AandAAAAfor the same name, with no preference expressed either way, and your system's own address selection picks between them. In practice that differs by platform: a glibc machine (Debian, Ubuntu, Fedora) generally connects over the IPv6 address, a musl one (Alpine) over the IPv4 address. Both reach the same peer. - A peer that cannot take IPv6 is sent IPv4. Through agent 0.4.139 an IPv6 overlay packet was translated to IPv4 before it left the machine and back again on the way in, whoever was on the other end. From agent 0.4.140 it travels as IPv6 to a peer that also runs 0.4.140 or later, and is translated as before for every other peer — one that is older, one in another organization, or any peer at all where the router is too old to tell the two apart. Either way an older node on the other end needs no upgrade, there is no new firewall port, no new endpoint and nothing to configure on the network.
- The
/64is practically unlimited. The IPv4 slice holds 4,093 peer addresses; this one is not a bound you will meet. On a node whose/20is already full, a newly discovered name is given an IPv6 address alone and resolvesAAAAonly. - A LAN name a member serves stays IPv4-only, as does the
node's own name for itself. That delivery path sends from the machine's own
LAN socket, which is IPv4, so those names answer
Aand nothing else. - A host that cannot take it keeps working on IPv4. If IPv6 is turned off in
the kernel, or every
/64the node picks is already routed on that machine, the node runs its overlay on IPv4 exactly as before andnobgp statussays why undernetwork.overlay6. This is never fatal and never affects the control channel. - On OpenWrt the overlay zone is exempted from the router's DNS rebind
protection, from agent 0.4.134. OpenWrt's DNS service counts the address
range these overlay addresses come from as private and discards answers
carrying it, so on agents 0.4.132 and 0.4.133 an OpenWrt node dropped every
AAAAanswer for a peer's name — the name still resolved over IPv4, and the system log gained a rebind warning per lookup. The exemption covers the overlay zone alone, so rebind protection for every other name is untouched, and it is removed again when the agent stops. - ⚠ A reply grows by 20 bytes crossing families. Measured between two Linux
nodes, UDP replies from 500 bytes to 30,000 arrive intact, fragments included,
and TCP is unaffected because the connection's own segment size already allows
for the larger header. It has not been measured on macOS or Windows; if you see
large UDP replies go missing on those platforms, lower
mtuor turn the family off as below.
Turning it off. overlay-ipv6: false — or --overlay-ipv6=false, or
NOBGP_OVERLAY_IPV6=false — takes the node back to an IPv4-only overlay. It is
also settable remotely with
node_config_set, which is deliberate: a node
whose IPv6 path misbehaves keeps its IPv4 overlay and its control channel, so
that surface is the one still working when the datapath is not. The change
applies on the agent's ordinary reload, within about five seconds.
The slice itself is recorded as overlay-cidr6 in the profile, so a restart
reuses it and your peers keep the addresses they had. Like overlay-cidr it is
agent-written state rather than a setting you tune, and it cannot be changed
remotely — changing a slice re-numbers every peer.
Addresses are stable. Once a peer has an address in your node's slice, it
keeps it — across agent restarts, and across the peer going offline and coming
back. If a peer leaves the directory, its address is reserved rather than
handed to someone else: the peer stops appearing in nobgp status (it is no
longer known to be reachable), but the address is held for it for 7 days. If
you look the name up while it is reserved, the agent checks with the router
first: a peer that still exists resolves to its usual address, a deleted peer
reports as not found right away, and if the router is unreachable the reserved
address is returned so a brief control-channel outage never breaks name
resolution.
Names that move between nodes. If a name a peer used to serve is republished on this node — you moved a service, or pointed a name at a different machine — the agent reconciles the two on the next directory update from the router:
- The old host is online and no longer advertises the name: its address is released right away, and lookups return the local one.
- The old host is offline: whether the name really moved is not yet
knowable, so its address stays reserved (and hidden from
nobgp status) while lookups already resolve to the local copy. The peer's return settles it — if it advertises the name again it keeps the address, otherwise the address is released. A reservation held this way still expires after the same 7-day window, so a decommissioned node cannot hold an address forever. - The old host still advertises the name too: nothing is released. The same name served from two nodes is a valid topology, and both keep their addresses.
Moving a name back later simply hands out an address again.
An address noBGP put a lease on stops being served once it is both past that
lease and unused (agent 0.4.103; agent 0.4.104 for what happens to it
then). When a name is discovered by asking noBGP for it — the lookup an agent
makes when a name is not yet in its directory — the answer can carry how long it
stays authoritative. The address minted from such an answer keeps that bound, and
once both conditions hold — the lease has passed and nothing on the machine
has addressed it for 30 minutes — the address moves into exactly the
reserved state described above: hidden from nobgp status, and re-confirmed
with noBGP on the next lookup instead of being answered from cache.
The address is not handed to another name (agent 0.4.104). It stays bound to the name it was minted for, so anything already running through it keeps working, and it can never be re-issued to a different peer while an application is still holding the old answer. Releasing it outright remains the 7-day reservation's job. Agent 0.4.103, where this first shipped, released the address at this point instead — which could silently drop packets on a connection that had merely gone quiet, and then deliver them to the wrong machine once the address was re-issued. Upgrade a node running 0.4.103 to 0.4.104.
Both halves of the condition are load-bearing, and the reason to say so is that either one alone would be visible to you:
- Past-lease alone would stop serving an address an application is still using, costing it a re-resolve for nothing. A lease says how long noBGP's answer stays good, not whether your program has stopped sending to the address it resolved.
- Idle alone would stop serving a live name. A perfectly healthy connection can go quiet for half an hour — an idle SSH session, a pooled database connection — and a name can be the right answer with no traffic on it at all.
Both are elapsed-time measurements, and a clock jump cannot move them (agent 0.4.104). On 0.4.103 they were read from the wall clock, so a large forward step — NTP landing on a machine with no real-time clock, a VM resuming, a laptop waking from sleep — made every address the machine was actually using look idle at once, while addresses nothing had touched were unaffected.
An address noBGP stated no lease for is never touched by this: it stays governed by the 7-day reservation above, which is the behaviour every earlier agent has. In practice that is every address learned from an ordinary directory update, so on most nodes this changes nothing — it is a freshness preference, deciding how quickly a past-lease name goes back to noBGP to be re-confirmed rather than being served from this node's cache.
A name still in use is renewed rather than left to lapse
From agent 0.4.125 an address that is past its lease and still being used is re-confirmed with noBGP in the background, while it goes on being served. This is the other half of the fork above: the sweep demotes an address that is past its lease and idle, so before this release an address that was past its lease and busy was simply never asked about again.
⚠ That could cut a live connection. Only a demoted address makes a node go back to noBGP for the name, so a name carrying traffic stopped renewing noBGP's own record of it — and once that record expired, noBGP's cleanup could withdraw the name from under the connection running through it. Nothing on either side said so.
The renewal is the same lookup an ordinary miss makes, and there are three answers:
- noBGP still knows the name: the address is kept — the same address, and for an unchanged answer the same peer — the new lease is recorded, and nothing is interrupted. This is the ordinary case, and you will not see it happen.
- noBGP no longer knows the name: the address moves to reserved, exactly as the idle sweep would leave it. A connection already running through it keeps running; what stops is answering new lookups from cache.
- noBGP could not be reached, or did not answer in time: nothing changes. The address stays served with its lease still past, and the name is tried again later — after a minute at first, doubling to at most fifteen. A brief control-channel outage therefore costs nothing.
Two bounds worth knowing:
- A renewal does not count as use. The idle clock is not reset by one, so a name that really does go quiet still becomes reserved on the ordinary path rather than keeping itself alive.
- Renewals cannot crowd out ordinary lookups. At most four run at once and one per name, sharing the same capacity as lookups for names not yet in the directory; a name that does not get a place waits for the next sweep.
From agent 0.4.130 an address that is already reserved is renewed too while something is using it, whatever its lease says. Reserved is where a renewal that finds nothing leaves an address, and it is also where an ordinary directory update leaves a name learned by asking — so before this release such a name was never asked about again, and an application holding the address kept sending to whichever member used to serve it. A renewal that finds the name serves it again on the same address; one that finds noBGP no longer knows it changes nothing but is retried, after a minute at first and doubling to fifteen, so a name nobody serves any more does not cost a lookup per packet. See When the member serving a name stops delivering it.
All of this needs the lookup itself — asking noBGP for a name that is not in the node's directory — which is available on any network that is not isolated. Where it is not, an address is governed by the 7-day reservation above and nothing here applies.
What the overlay carries
From agent 0.4.130 a packet is carried whatever protocol it is and whatever is inside it. noBGP reads only what it has to rewrite — the addresses, the ports of TCP and UDP, and the identifier of an ICMP echo — and passes the rest through untouched. Before this release three shapes of ordinary traffic were lost, each of them silently:
- A UDP packet on a well-known port arrived empty, or not at all. Ports such as 53 (DNS), 123 (NTP), 67/68 (DHCP) and 4789 (VXLAN) were interpreted rather than forwarded, so a DNS query or answer between two nodes lost its message, and a body that did not match what the port implies was dropped outright. A resolver, an NTP server or anything else sitting on such a port now answers over noBGP.
- A protocol other than TCP, UDP and ICMP was refused — GRE, ESP, SCTP, IP-in-IP and unassigned protocol numbers among them — and a tunnelled packet lost its inner IP header. ⚠ They are carried now, but only TCP, UDP and ICMP echo have a return path: a protocol that has no ports gives the reply nothing to be matched against, so it reaches the node at the other end and the answer does not come back. Treat such traffic as one-way until that changes.
- An IPv6 packet carrying extension headers was dropped — Hop-by-Hop, Routing, Destination Options, or a fragment header — and so was an ICMPv6 echo, in both directions.
From agent 0.4.132 a UDP service bound to the wildcard address answers. A
daemon left on 0.0.0.0 or :: — the packaged default for most of them — has no
address of its own, so its reply leaves with the machine's overlay interface
address as the source rather than the address the request arrived on. That did
not match the request, and the reply was dropped: the caller saw a service that
was listening, reachable by every other measure, and simply never answered. TCP
was never affected, because an accepted connection carries the exact address it
was accepted on. Binding the service to the machine's own address was the
workaround and is no longer needed.
ICMP errors are carried, in both directions and across families. Destination
unreachable, time exceeded, packet too big / fragmentation needed and, from
agent 0.4.132, parameter problem reach the sender with the quoted packet
rewritten to match — so traceroute, path-MTU discovery and a plain
connection refused behave over the overlay the way they do on a LAN. Before
0.4.132 every parameter problem was dropped, in both directions and even
between two nodes of the same family.
A packet the overlay does refuse is counted, from agent 0.4.133. Each
machine keeps a tally by class — a packet refused on its way out, a reply that
arrived for a flow that had already gone, a reply with nowhere left to go — and
net_metrics reports it as datapath_drops.
Individual drops are logged at debug level, because the machine sending them
decides how often they happen, so these counters are how a machine that is
losing traffic tells itself apart from a quiet one. From agent 0.4.146 the
tally also covers what a peer sent to the machine's own services that the
machine did not deliver — see below — as protocol,
icmp_type, malformed, no_delivery and deliver_failed; those five need a
router on 0.4.174 or later.
This is about traffic between nodes. What a member carries onto its own LAN for a name it serves there is a separate question, with a narrower answer: from agent 0.4.131 that path carries TCP, UDP and ICMP echo only. How a packet addressed to a node is handed to that machine's own services is a third question, and it changed in 0.4.140 — see How another node reaches this machine's own services.
One packet can no longer cost more than itself. A packet a node cannot read, or cannot write back out, is dropped and counted in that node's log while the node carries on. Before, a single UDP-Lite packet stopped the agent outright, and a packet a peer's newer build sent that an older one could not decode ended the connection to that peer rather than being skipped.
How another node reaches this machine's own services
From agent 0.4.140 a connection from another node ends at the agent and is
opened again from this machine, instead of being routed through the overlay
interface to the machine's own address. What arrives at the service is an
ordinary local connection, so a peer reaches exactly what a program running on
this machine reaches: a service bound to 127.0.0.1, to [::1], to the
wildcard address, or to any address the machine holds. Only the machine receiving
the traffic needs this release; nothing changes for the node sending it.
That is the intent of the change, and both sides of it are worth reading before you rely on either:
- ⚠ A service that trusts loopback is now reachable by every node your network
lets reach this machine — a database or Redis on
127.0.0.1with no password, a Docker API on TCP, an admin page bound to loopback because it is "not exposed". Which nodes may reach this one is decided by network membership, not by where a service binds. Put authentication on such a service, stop it, or keep the machine in a network whose members you trust with it.- From agent 0.4.153 the machine can refuse the loopback half outright:
peer-loopback: false, set by the node's owner at the machine, refuses a peer every port bound only on this node's loopback while leaving the ports it binds on a real interface reachable as before.true— the reach described here — stays the default, so a node that has not set it is unchanged. Earlier agents had no such setting.
- From agent 0.4.153 the machine can refuse the loopback half outright:
- The host firewall no longer filters this traffic. Windows Firewall,
ufw,firewalldand macOSpfsee a local connection made by the agent, not traffic arriving from the network, so a rule written against the noBGP address range or the noBGP interface does not match it — includingwindows-firewall: respect-windows, which from this release no longer keeps peers off the machine's TCP and UDP ports. - A service still cannot tell a peer from a local program, and never could: before this release every peer arrived with the same overlay gateway address as its source, so no per-peer rule inside a service is lost.
- The agent's own ports are refused to peers, whatever else is true: the
node's local MCP server, the
shared-drive proxy, the NFS server,
the overlay DNS server, the Windows agent API, and
rpcbindon port 111, which the installer brings in for the drive. A peer connecting to one of those is refused; a program on the machine still reaches them. From agent 0.4.144 the same check covers the other way in: a published service whose target is one of those ports is refused as well. - A ping to a node is answered by the agent itself. A running node answers, whether or not its host would have.
- A Unix domain socket is not reachable. The agent connects to TCP and UDP
ports only, so the Docker socket
/var/run/docker.sock, for example, stays local. A service that must stay local to the machine must listen on a Unix domain socket, not on a TCP or UDP port on loopback. - A port nothing listens on is refused instead of leaving the caller waiting. A TCP connection is reset once every address on the machine has refused it, and on a UDP flow the error the machine itself saw — connection refused, host unreachable, and on Linux message too long — is sent back, so a failure over noBGP reads the way it does on a LAN. ⚠ On Windows, and on macOS from agent 0.4.141, the agent reads the machine's listener table to answer quickly, so a service that starts in the second before a connection arrives can refuse one connection before it is seen. Neither platform answers a closed port promptly on its own — Windows takes about two seconds for each address of the machine, and macOS does not refuse a connection to some of its own addresses at all — so without the table a closed port took seconds to answer and a scan of closed ports could hold off new connections, both to the machine and to the LAN names it serves. Linux refuses at once and needs none of this.
It covers TCP, UDP and ping — including an IP fragment of one, which the agent reassembles, from agent 0.4.144, and a fragmented IPv6 ping from agent 0.4.145: on 0.4.144 a ping larger than the overlay MTU sent to a node over IPv6 got no reply at all. From 0.4.144 nothing else from a peer reaches the machine at all: any other IP protocol, ICMP other than a ping, and a packet the agent cannot read are dropped rather than delivered over the tunnel adapter. Through agent 0.4.143 they arrived there at the machine's primary address, where the host firewall applied to them. For a protocol other than TCP, UDP and ICMP that was a one-way delivery rather than a working protocol — the return path had no case for it, so a peer never got a reply — which is why dropping it takes nothing away that worked.
From agent 0.4.146 each of those drops is counted, by class, and reported by
net_metrics in datapath_drops: protocol,
icmp_type and malformed for what arrived, no_delivery for a machine whose
socket delivery never started — which drops everything a peer sends it —
and deliver_failed for a packet this machine was meant to deliver and could
not, including one for a LAN name it no longer
serves. The counts need a router on
0.4.174 or later; before agent 0.4.146 they appeared only in the machine's
own log.
⚠ So from agent 0.4.144 the host firewall sees nothing at all from peers, on
any platform, and windows-firewall: respect-windows
no longer filters any traffic from another node.
The agent no longer turns IP forwarding on, from agent 0.4.143. Earlier
releases set the machine's forwarding switch — net.ipv4.ip_forward on Linux,
net.inet.ip.forwarding on macOS, IPEnableRouter on Windows — for the routed
path this section replaced. It is a setting for the whole machine rather than for
noBGP, and on a container whose /proc/sys is read-only, writing it failed the
agent's startup. ⚠ A machine where an earlier agent turned it on keeps it on:
the agent cannot tell whether something else on the machine now depends on it, so
turn it back off yourself if it was off before you installed noBGP.
A machine relays at most 1024 TCP and 512 UDP flows to its own services at once, a budget of its own separate from LAN-name delivery. Past the TCP limit a new connection is refused; past the UDP one a new flow takes the slot of the flow that has been quiet longest.
Read which path a machine is on in network.peer_access in
nobgp status, or as the peer-access line of
nobgp show. It says as a local client on a machine
delivering the way this section describes, and — from agent 0.4.144 — none
(socket delivery did not start …) on one whose agent could not start that
delivery, which since that release drops every packet a peer sends it rather
than falling back to the tunnel adapter. Agents 0.4.140 to 0.4.143 said through
the overlay interface, at the primary address only there, which is what they did.
This is different from a published service. A published service has a URL on the internet, and noBGP controls who opens it. A service that a peer reaches by the node's name is reachable only from the nodes in your noBGP network.
Services
Services are the way you expose functionality from your nodes to the internet. noBGP supports two types of services:
Every service, of either type, is served at https://<name>.nobgp.link (router 0.4.172) — a different registrable domain from the nobgp.com the dashboard and the API use, so a published service never shares a browser's cookie scope with them. See Where a service is served.
1. Proxy Services
Proxy services expose HTTP applications running on your nodes to the public internet with a unique URL.
Example use cases:
- Expose a web app running on
localhost:8080 - Share a development server temporarily
- Provide access to an internal dashboard
How it works:
Features:
- Automatic HTTPS with valid certificates
- Optional authentication (require OAuth to access)
- Enable/disable without deleting
2. Terminal Services
Terminal services provide browser-based terminal access to your nodes through a web interface.
Example use cases:
- Remote shell access without SSH
- Shared terminal sessions for collaboration
- Emergency access when SSH is unavailable
- Terminal access for air-gapped systems
How it works:
- Your AI assistant publishes a terminal service
- noBGP generates a unique URL (e.g.,
https://xyz789.nobgp.link) - Opening the URL shows a full-featured web terminal
- WebSocket connection provides real-time interaction
Features:
- Full terminal emulation (colors, control characters, etc.)
- Support for interactive programs (vim, nano, htop, etc.)
- Command history and tab completion
- Resize support
- Optional authentication
Service Authentication
Both service types support optional authentication:
- Auth required: Users must sign in with OAuth before accessing
- Public: Anyone with the URL can access (useful for demos, temporary shares)
Always use authentication for production services or those containing sensitive data.
Sharing Services with Specific People
When a service has authentication enabled, you control exactly who gets in. Visitors who aren't on your approved list see an "Access Required" page where they can request access with one click. You receive an email notification and can approve or deny from a built-in permissions dashboard — no configuration required.
You can also add people proactively by email before they visit, including whole domains with a wildcard (*@yourcompany.com).
See Access Control in the Publishing Services guide for the full details.
Sessions
A session is a running command on a node, managed by the AI assistant via the command MCP tool. Each session runs one specific command and closes automatically when that command exits.
Session Lifecycle
- Start: Your AI assistant calls the
commandtool with asessionobject specifying the node and command to run - Interact: The AI reads output using the returned
command_id; for interactive programs, it sends additional input - End: The session closes automatically when the command exits
Session Patterns
One-shot commands — run a non-interactive command and read its output:
df -h,systemctl status nginx,cat /var/log/app.log- Session closes as soon as the command finishes; no explicit cleanup needed
Interactive shell sessions — start bash (or another shell) and send multiple commands while maintaining state:
cdand exported environment variables persist across inputs within the session- The AI sends each command as stdin input to the running shell
- Close by sending
exitor aSIGTERMsignal
When to Use Each Pattern
Use one-shot commands for:
- Quick status checks (
uptime,df -h) - Non-interactive operations
- Scripted automation
Use interactive shell sessions for:
- Multi-step procedures that share state (navigate directories, set env vars, run related commands)
- Interactive programs like
vim,less,top,psql - Debugging with tools that require back-and-forth input
Subscriptions and Events
A session answers "what is happening on this machine, right now". A subscription answers "tell me when something happens", across as many machines as you like — without holding a connection open or polling.
There are three things you can subscribe to:
| Source | What it reports | Who does the work |
|---|---|---|
| Filesystem | Files and directories changing on the matching nodes | Each node runs a native watcher (inotify, FSEvents, ReadDirectoryChangesW) |
| Command | A command dispatched once to each matching node, and how each run went | Each node runs the command detached |
| Presence | Nodes coming online, going offline, or re-registering | The router — nodes run nothing |
Every subscription follows the same three steps: subscribe (you get a subscription_id), read events, unsubscribe. Reading is what keeps a subscription alive — it expires about 10 idle minutes after the last read — and unsubscribing stops the work on every node, not just the delivery.
Subscriptions are scoped by the same node selector everything else uses, so "watch /etc on every edge node" is one call rather than a loop. A node's owner keeps the last word over what that machine serves — which capability domains, which directories, whether elevation is allowed at all — and when a node won't serve a source it says so with a refused event instead of going quiet.
See the Monitoring & Events guide for the full workflow.
Authentication & Security
noBGP uses multiple layers of security:
OAuth 2.0
All AI assistant access is protected by OAuth:
- Sign in with Google, GitHub, or other providers
- Token-based access control
- Automatic token refresh
- Per-session permissions
Agent Registration
Nodes authenticate through a two-phase process:
- Registration: Either via OAuth browser login (
nobgp register) or with a registration key (nobgp register --key) for automation - Ongoing auth: After enrollment, the agent receives a JWT token signed with per-network Ed25519 keys — no key or login needed for subsequent connections
- Registration keys are unique per network and can be rotated without downtime
Encryption
All communication is encrypted:
- TLS for web traffic
- End-to-end encryption within the overlay
- Certificate pinning for agent connections
Peer key pinning
From agent 0.4.98, a node remembers the public key it was given for each peer the first time the two talked, and tells you when that key changes.
Each node generates its own keypair and never hands over the private half, and every mesh session between two nodes is encrypted end to end with those keys. The public key of the peer, though, reaches each side over the control channel — so pinning is what turns the key I was handed today into the key I have always been handed for this peer. A first contact records the key; every session after that compares against it.
Be precise about what this buys, because the honest version is narrower than "it prevents interception":
- It catches a peer's key changing after the two nodes have already been talking, makes it visible, and attributes it to a named peer.
- It does not catch a key that was substituted from the very first contact — there is nothing yet to compare against, so the first key seen is recorded as the truth. Pinning is a detection layered on top of a session that already works, not a guarantee that a session cannot be intercepted.
The fingerprint is comparable with what the peer says about itself. The value
nobgp peers prints for a peer is the same string that peer reports as
registration.agent_key_id in its own nobgp status. Compare them on the
peer's own console — read back through noBGP's own surfaces, the comparison is
not evidence, because those are the party being checked.
What a node does about a change is one setting,
peer-key-pinning: warn (the default —
report it at ERROR and let the session proceed), enforce (refuse that
session), or off. A refusal under enforce fails one session and never the
node's control channel: a changed key is a claim about one peer, and taking a node
off the network over it would be an outage anybody able to offer a bad key could
cause at will.
The report itself is a line in the node's own log
(nobgp service logs) naming the peer, the fingerprint
offered, the fingerprint pinned, the verification to do at the peer's console, and
the exact nobgp peers --forget command that accepts the new key — plus what the
node just did about it, so it never states a consequence that does not match the
mode it is running in.
warn is the default for an operational reason, not a security one.
Re-provisioning a node gives it a fresh key under the same node ID, so an
ordinary rebuild and a substituted key look identical, and enforce would take a
just-rebuilt node out of reach of every peer that had ever talked to it until a
person accepted the new key on each one. Warn does not mean forget: the pin is
left as it was, so the mismatch is reported on every session until somebody
resolves it.
Resolving one is nobgp peers: it lists what is
pinned, and --forget <node-id-or-name> accepts a rebuilt peer's new key on that
node. The setting is per node and so are the pins — there is no fleet-wide accept,
and the setting cannot be changed remotely, since the check is aimed at the same
control plane a remote config call travels over.
Authorization
Not all users can do everything:
- Node provisioning is open to every account, on every plan — what bounds it is the plan's included compute, not an approval step. The target network may narrow it: a network whose
members_may_add_nodessetting is off takes nodes only from its owner and the organization's Owners and Admins - Creating a network is open to any member, who becomes its owner; managing or deleting one is its owner, or an Owner or Admin of the organization
- Controlling a node — renaming it, configuring it, publishing its services, deciding who else may use it — is its owner (at any role) and the organization's Owners and Admins
- Command execution requires
useof the node, from router 0.4.179: the node is shared with its organization, or you are named on its access list. Organization membership alone lets you see a node - Running as root is a further grant, held per node and given explicitly
The whole model is Roles & Permissions.
File Sharing
noBGP provides a shared virtual drive scoped to each network. This allows you to share files across all nodes and manage them through a web interface.
How It Works
Virtual Mount Point:
- Each network gets a shared drive mounted on every connected node
- All mounts are backed by S3 storage
- Files written on one node are visible on all other nodes in the same network
- Default mount points:
/mnt/nobgp(Linux),/Volumes/nobgp(macOS), drive letter (Windows)
Mount Technology: see Which backend mounts the drive below. Whichever backend the node lands on, the agent runs a small proxy on 127.0.0.1:19840 that adds the node's credentials to the requests it forwards, so the operating system's own client needs no noBGP configuration of its own. On Linux and macOS you can restrict which local accounts may reach it — see Shared drive proxy keys.
Multi-Profile Mount Points:
When running multiple profiles (see Multi-Profile Support), each profile gets its own mount point with a -{profile} suffix (e.g., /mnt/nobgp-work, /mnt/nobgp-staging). The default profile always keeps the bare name (/mnt/nobgp). On Windows, each profile gets a separate drive letter.
File Names:
-
Names are stored exactly as written, and are unique without regard to case (router 0.4.71).
Foo.txtkeeps its capitals forever — nothing is lowercased, folded or rewritten, and a listing gives back the bytes that were written — butfoo.txtis the same name, so creating it besideFoo.txtis refused asalready_exists, naming the spelling that holds the name. Over the share's URL that refusal is a 409. Overwriting the stored spelling is ordinary and always allowed, and renaming a name to another spelling of itself (foo.txt→Foo.txt) is how a stored casing gets corrected. -
Lookups are still exact.
README.mddoes not findreadme.mdthrough the file tools or the share URL; only the mount folds, on the platforms whose own filesystems do. The rule exists so a share can be copied onto a Windows or macOS filesystem at all — neither can hold two names differing only in case, so a directory that did lost a file silently on the way off the share. It is not retroactive: a directory that already holds both spellings keeps both, both readable and both writable, and only a new second spelling is refused. Between router 0.4.53 and 0.4.71 either spelling was a separate file; before 0.4.53 both reached the same one. See the changelog. -
%, spaces,#and?are ordinary characters in a name and survive the round trip through the mount, the share URLs and the file tools unchanged. On the mount this needs agent 0.4.56: before it, a file created on the drive asa%20b.txtwas stored asa b.txt, and a#or?cut the path short so the write landed somewhere else entirely. -
.nobgp-upload-…and.nobgp-trashare reserved. The share writes an upload under that prefix and renames it into place when the body is complete, and a deletion is moved into.nobgp-trashat the top of the tree until its bytes are reclaimed — so a name of your own using either is refused rather than stored. The upload prefix is reserved at the end of a path, the trash name at every level of one. ⚠ One door accepted them until router 0.4.177: a copy onto one of these trees, which published a file under the reserved name that no listing then showed. It is refused there now too, on any segment of the destination path. -
Four OS presentation files are accepted and discarded —
.DS_Store,desktop.ini,Thumbs.dbandautorun.infnever take up space and never appear in a listing, so copying a folder from Finder or Explorer does not litter the share. Each of them describes one machine's view of a folder — icon positions, a thumbnail cache — and is read by no other machine. From router 0.4.77 deleting one succeeds instead of failing, and every read of one gives the same version tag rather than a fresh one each time. -
macOS
._namesidecars are stored like any other file, from router 0.4.128. They are not presentation state: macOS has no other way to carry a file's extended attributes over a network drive, so._report.pdfis wherereport.pdf's tags, quarantine flag and resource fork live. They are stored, listed, read back and deleted with no special handling — and they follow their file, so deleting or renamingreport.pdftakes._report.pdfwith it. A copy does not carry one; the client writes the attributes itself on the far side. ⚠ A folder a Mac has written to therefore lists about twice the entries you expect, which is what every other server a Mac copies to does. Storing them also means a file of your own genuinely named._notesis kept rather than eaten.Through router 0.4.127 these were discarded with the four above, and the discard was not quiet:
cpfrom a Mac onto the share printed could not copy extended attributes and exited 1 while the file itself landed correctly, because the sidecar it wrote read back as zero bytes. If you have a script that treats that exit code as a failed copy, it will stop firing.
Web Dashboard:
- Browse a network's files in the web dashboard: open Networks, choose your network, then Files
- Signed in with your noBGP account — the same access your account already has to that network
- Upload and download files through your browser
- Manage directories and organize your shared files
What the mount contains
One mount, two folders (agent 0.4.54+):
/mnt/nobgp/
├── node/ this machine's own storage area
└── networks/
└── production/ one folder per network this node belongs to
-
node/is the node's own storage area — the same tree ashttps://files.nobgp.com/nodes/<node-id>/, reachable from the machine itself with no URL and no token. It sits at that exact path on every node, never named after the machine and never carrying its id, so the same script deployed across a fleet reaches each machine's own storage by the same path. -
networks/<name>/is a network's shared drive, one folder per network the node belongs to, carrying the network's name. Everything above about the shared drive applies inside it. -
The root and
networks/are a view, not storage. Creating, deleting or renaming directly in either is refused as a read-only directory — a file written there would have nowhere to live. Everything belownode/andnetworks/<name>/is read-write exactly as before. -
A file cannot be moved between the two trees in one step.
node/and each network are separate filesystems, so a move across them cannot be a rename. On Linux the mount says so in the waymv, Finder and Explorer understand, and they copy and then delete for you. On macOS and Windows the move is refused (agent 0.4.56+) — copy the file across and delete the original. -
The layout is the same whichever backend mounts the drive. On macOS and Windows, run agent 0.4.56 or later: 0.4.54 presented the two folders correctly but dropped entries from the listings inside them, so a folder holding six files could show four with nothing reporting an error, and 0.4.54–0.4.55 turned a rename in Finder or Explorer into a wrongly-placed new file inside the share. Linux was unaffected by both.
-
Renaming a network reaches the mount right away, from agent 0.4.75. noBGP pushes the node's whole network list down the connection it already has, so
networks/<old>/becomesnetworks/<new>/on every node in the network at once — no reconnect, no restart. What can still lag is the operating system's own directory cache, bounded byfs-cache-ttl, which is 30 seconds unless you have changed it.Through agent 0.4.74 the mount learned its network names only when the agent connected, so the old folder name stayed until that node next reconnected — up to a day, and by a different amount on every node depending on when each last connected. Nothing was lost or misfiled in the meantime, since the folder still pointed at the same network, but two machines could show different names for it, and every rename also put a spurious
received unknown messagewarning in each node's log.
Before 0.4.54 the mount root was the network's shared drive, so a file that used to be at /mnt/nobgp/report.pdf is now at /mnt/nobgp/networks/production/report.pdf. Nothing moved in storage; only what the mount shows you changed. The networks/ level exists so that a network named node cannot shadow the one path that has to mean the same thing on every machine.
What travels with a file
A file is its name and its bytes, and the share keeps nothing beside them. Every file is one stored object, on every backend and on the share's URLs alike — this is a property of the storage, not of the driver a node happens to mount with.
- Metadata a format keeps inside its own bytes travels intact, because it is the bytes: an Office or
.msiproperty block, EXIF in a photo, ID3 in an MP3. A copy onto the share is byte-identical, so anything read out of the file itself reads the same on the other side. - Metadata your platform keeps beside a file does not travel — NTFS alternate data streams on Windows, and extended attributes on any platform. The share has one stored object per file and nowhere to hang a second stream off it.
- macOS is the exception, from router 0.4.128. A Mac carries a file's extended attributes in a hidden
._namecompanion beside it, and the share now stores that companion as an ordinary second file rather than discarding it — so tags, quarantine flags and resource forks written from a Mac survive a copy onto the share and come back on the next Mac. It is a file, not metadata: it appears in listings, counts against your storage limit, and is only understood by a Mac.
Windows records that a file came from the internet in a Zone.Identifier stream beside the file's data — the thing SmartScreen reads before running an executable, and what puts a downloaded Office document into Protected View. The share has nowhere to keep a stream, so the mark is gone on the copy and cannot be written back afterwards.
It is silent. A file whose only extra metadata is the mark copies with no prompt of any kind. So a file downloaded on one machine, staged on the share and picked up on another arrives unmarked, and will open or run with none of those defences engaged — and moving files between machines that cannot reach each other is exactly what the share is for. Treat anything you collect from the share as you would treat the original download.
Measured on a Windows machine on 2026-08-17, agent 0.4.81, on a winfsp volume: the copy carries no stream, enumerating its streams fails with The parameter is incorrect, and writing one onto the share afterwards is refused. The webdav drive on the same machine was measured on 2026-08-22, agent 0.4.95, and answers the same way: enumerating the copy's streams fails, and writing a Zone.Identifier onto the share is refused. So the share holds no named streams on either backend. The one thing not yet seen on the WebDAV drive is the silence — an unrelated defect made that day's copy fail outright rather than complete quietly — so expect the loss there to be as silent as it is on winfsp, and treat the mark as gone on both.
Separately, an Explorer copy onto a winfsp volume can ask whether to copy the file without its properties, and it warns about more than it takes: Windows writes a file's property store after the copy, the volume has nowhere to put one, and Explorer reports that refusal in the words of a failure. Measured on the same machine and day: for the installer that reliably raises the prompt, every property the dialog named — Title, Subject, Tags, Comments, Authors, Revision number, Program name — was still there when the copy was read back on the share, because those live inside the file's own bytes. A file with no property store raises no prompt at all. Do not read that as nothing is ever lost: what it costs is a click and a scare, and the streams above are the case where something really is gone.
When the agent cannot read its own mount
On some nodes the drive works normally for everything on the machine — Finder, a login shell, sudo from Terminal, and any app you open — while the one process that cannot read it is the agent's own background service. The file tools and command reach a node's drive through the agent, so where that happens a path under the mount point is refused: on macOS with a bare Operation not permitted.
The node reports whether it is in that state, and network_directory publishes it per node as info.mount_readable (router 0.4.83+, agent 0.4.85+ — the node sends it, so an older agent reports nothing here either way), beside info.fs_mount — where the drive is — and info.writable_roots — which directories under it accept writes. true means the agent reads its own mount and those paths work; false means they are refused at every identity; absent means the agent never probed, usually because nothing is mounted. Read the field rather than assuming either way.
From agent 0.4.93 the node says it on the box as well. nobgp status adds an fs.unreadable line to the fs block whenever the drive is mounted and the agent cannot read through it, naming the remedy for whichever of the two causes below it is in. Until that release the block reported mounted: true and named no condition, so the person standing at the machine — the only one who can grant macOS consent — had no signal at all, while a tool caller had been told since router 0.4.83. The line appears only when something is wrong, so a healthy node prints nothing here. ⚠ A probe that cannot answer within two seconds reports that it could not tell and says plainly that this is not a permission verdict: a wedged mount and a refused one call for opposite remedies, and granting Full Disk Access on the strength of a timeout costs you the one clue you had.
Two causes are known, and they share no remedy:
- macOS — the consent layer (TCC) gating access to a network volume on the process doing the reading. A background daemon has no consent and cannot ask for one, so no retry, no execution identity and no backend change makes a difference — measured across eleven variations on 2026-08-12, including mounting as the console user, mounting outside
/Volumes, andwebdavin place ofnfs. A Full Disk Access grant ends it, which is why this is a per-node state and not a property of the platform. - Windows on the
webdavbackend — a mapped drive belongs to the logon session that created it, and the agent runs as a SYSTEM service, which has no view of it. There is no consent to grant here; awinfspmount is not session-scoped and does not have the problem.
Linux nodes have no equivalent restriction.
What it costs, and what it does not:
- Nothing about the drive itself. Reading, writing, locking and syncing all work for the machine's own users and programs. The mount is healthy; the daemon is blind to it.
- The file tools reach that content a different way. Ask the router for it instead of the node —
network_namealone for a network's shared drive, orstorage: truefor a node's own area. Those serve the same bytes, and the node need not even be online. See Addressing storage instead of a node. - The share's HTTPS endpoint is unaffected, from a Mac or anywhere else.
- Granting the
nobgpdaemon Full Disk Access does lift it, for every tool at once. On that Mac: System Settings → Privacy & Security → Full Disk Access →+, then press ⌘⇧G (the picker hides/usr/local) and enter/usr/local/bin/nobgp— the daemon's own binary. It takes effect on the running daemon immediately, with no restart and no remount. It takes a person at that machine, or an MDM profile: nothing sent over noBGP can grant macOS consent, which is why the file tools address the remedy to the node's owner from router 0.4.76. Since agent 0.4.72 the grant survives upgrades. The macOS binaries are signed with a Developer ID certificate, so macOS keys the consent to the signing identity rather than to the exact bytes, and a later build satisfies it. Under 0.4.71 and earlier the binaries carried an ad-hoc signature, which tied the grant to that one binary and let the next auto-upgrade void it silently — so on those versions the grant was worth doing by hand on a machine you were debugging and not worth rolling out. A node upgrading from such a version needs the grant given once more, against the signed binary; from then on it holds. - From agent 0.4.144 the Mac asks for it while you are there.
sudo nobgp registerrun at a terminal on the machine has the agent open that pane itself and names the row to switch on, so a node enrolled by hand arrives granted. Running the command again on a Mac that is already enrolled makes the same offer, which is the way to reach the pane on a machine enrolled before this. From agent 0.4.145 a registration with a key asks as well, the install script included, as long as the command's output is a terminal; an enrolment that writes to a log or a pipe asks nothing.
Which backend mounts the drive
The drive is presented to the operating system by one of four backends, and the agent mounts with the first one this host can actually serve:
| Backend | What it is | Available when |
|---|---|---|
nfs | The agent's own NFSv4 server on loopback, mounted by the kernel's own NFS client. Runs as an ordinary user on a high port — no rpcbind, no privileged port, nothing to install on macOS. From agent 0.4.69 its locks hold between nodes too | The kernel can mount NFSv4 — and on Linux, from agent 0.4.70, that is the whole question: with no mount helper installed the agent makes the mount itself, so a host with no NFS packages at all can serve this backend (see below). On macOS the helper is mount_nfs, part of the base system. Agents 0.4.62–0.4.69 required a helper on Linux too, which is nfs-common/nfs-utils; from agent 0.4.58 install.sh installs it on Debian/Ubuntu, RPM hosts and Arch |
fuse | An in-process filesystem over /dev/fuse, with on-disk caching — files are cached locally for fast reads and uploaded asynchronously on writes. From agent 0.4.64 its locks hold between nodes — the first backend where they did, joined by nfs in 0.4.69 | /dev/fuse is usable. Linux only today. The userspace fusermount helper (fuse3, or fuse-utils on OpenWrt) is needed as well, but that only shows up when the mount is attempted — see Shared drive is empty and never mounts |
winfsp | An in-process filesystem over the WinFsp driver — the same idea as fuse, with a Windows driver underneath, and from agent 0.4.67 with the same on-disk caching and asynchronous write-back. Read-write from agent 0.4.63 (read-only in 0.4.62), though a real mount only accepts writes from 0.4.65. ⚠ Opt-in through agent 0.4.81; from 0.4.82 auto chooses it on Windows wherever the driver is present (see below). Agent 0.4.62+ | WinFsp is installed and its library loads for this machine's architecture — and from agent 0.4.73 the probe checks the version too, so an install older than 1.10 (2022) reports as unavailable with an upgrade remedy instead of being chosen and failing the mount. Windows only |
webdav | The operating system's own WebDAV client against the agent's local proxy — mount_webdav on macOS, net use on Windows, davfs2 on Linux | The proxy is running and the OS mount helper exists (nothing to check on Windows) |
Whether a backend can serve is a question about the host, and from agent 0.4.57 it is always answered by probing that host rather than assuming from the platform — is there an NFSv4-capable mount helper and a kernel that can mount NFSv4, is /dev/fuse usable and its helper on the path, does the WinFsp library load. Which of the capable backends wins when you have not pinned one is a different question, and from agent 0.4.65 it has a different answer per platform:
| Platform | Preference order |
|---|---|
| Linux | fuse → nfs → webdav |
| macOS | nfs → webdav (agent 0.4.65 alone had these the other way round — see below) |
| Windows | winfsp → webdav |
| Synology DSM | fuse alone — DSM's only package manager has neither davfs2 nor an NFSv4 client, so there is no second backend to fall to |
| OpenWrt | as Linux, but kmod-fuse + fuse-utils come from the installer and NFSv4 and davfs2 only if you install them by hand |
| Linux with no package manager | as Linux in principle, but the installer can add nothing — a Buildroot-class image serves whichever of these its own build already supports, and often none |
| Containers | as Linux: privilege to mount at all, then /dev/fuse for fuse and nothing for nfs |
webdav is the fallback on every platform and is never removed from a chain — but it is probed like the others, so a Linux box without davfs2 reports it as unavailable rather than pretending it could mount. On Linux that is the common case rather than the exception, and worth knowing before you plan around it: install.sh tries fuse3 first and only falls back to davfs2 if that fails, which on a Debian-family host it almost never does, and on OpenWrt, Arch and Synology DSM it does not try davfs2 at all. So a Linux node's chain is usually two backends deep, not three, and on Synology — whose only package manager has neither davfs2 nor an NFSv4 client — it is one. If a host can serve none of them, nothing is mounted and fs.error says no filesystem backend is available on this host; installing davfs2 by hand adds the fallback where the platform has it. macOS has mount_webdav in the base system, and Windows needs nothing.
From agent 0.4.83 a node whose chain is one deep says so, without your having to count the backends lines: nobgp status carries an fs.no_fallback line naming the one backend that can serve, and its presence is the warning — if that mount fails there is nothing to switch to. Where the condition is fixable it names the remedy: a Windows node without WinFsp is one deep on webdav and is told that installing WinFsp ends it, while a Synology is told plainly that there is no second backend and nothing to install. It is a statement about the host, so a perfectly healthy node prints it — the point being to know before a mount fails rather than after. It is absent on a node that can run more than one, and absent as well when the node is running no local WebDAV proxy (fs: off, no mount point, or a proxy that failed to bind), because webdav then reads as unavailable for a reason that is not the host's and the line would blame the machine for it.
A backend that cannot serve a platform at all is not in that platform's list, and nobgp status no longer names it: there is no FUSE build for macOS or Windows, no WinFsp anywhere but Windows, and no NFSv4 client on Windows (it ships v2/v3 only, which cannot mount the agent's v4 server). The line it used to print was misleading anyway — telling a Mac owner that /dev/fuse is not usable invites installing something that will not help. The name still parses everywhere, so fs: fuse in a config carried to a Mac is refused by name rather than read as a typo.
Falling through happens on availability, never on failure. If the chosen backend is available and its mount then fails, the agent retries that same backend rather than dropping to the next one — cascading down the list on a failure is how a node ends up flapping between two broken mounts. Once mounted, the agent probes the mount every 30 seconds and remounts after three consecutive failures, about a minute. A healthy mount is left alone, so installing a backend the node would have preferred changes nothing until the next nobgp service restart.
A mount attempt that never finishes is stopped after two minutes, from agent 0.4.147. Every backend's mount helper can hang — a mount_webdav, mount.davfs, mount_nfs or net use waiting on something that never answers — and until this release such an attempt simply never concluded: the node went on publishing the backend it had before, with no drive under it, for as long as the helper hung. The attempt is now given two minutes, which is long enough for a first mount on the smallest node over a slow link; past that it is stopped, its helper is killed, the node reports fs_backend: error like any other failed attempt, and the retry cycle carries on. ⚠ No new attempt starts over a helper that is still running. A process stuck in the kernel does not die from being killed, so the agent waits for it to exit and clears up after it — an empty mount point, an orphaned helper — before it mounts again. Each cycle in between concludes as error rather than stacking a second helper on the same mount point.
A mount that was attempted and failed is reported as a failure, on every backend from agent 0.4.151. A backend can also decline, which is a different answer: nothing will change until the host does — no davfs2 or mount_webdav to mount with, no free drive letter on Windows. A decline is said once and then retried quietly, because repeating it every minute buries the one line worth reading. Through agent 0.4.150 a mount_webdav or mount.davfs that ran and failed was reported as a decline as well, on macOS and Linux: fs.error read webdav declined to mount at … (reason logged in the agent log) and the log said once that the retry could not succeed until the host changed — about a failure that is usually temporary, and which the next cycle often clears on its own. Such a failure now names itself in fs.error and warns on each cycle that it will retry, exactly as fuse, nfs and a Windows net use mount already did. ⚠ The trade is that a mount which keeps failing now writes a warning per cycle rather than one line.
Pinning or changing fs needs that restart as well, and from agent 0.4.87 the agent no longer reloads on that key at all: it records the edit as pending, keeps serving on the backend it started with, and holds back every other config edit until it restarts. Swapping a backend under a live agent is what the rule exists to prevent — see Filesystem keys.
From agent 0.4.82 that probe asks two questions rather than one on an nfs mount, and the second one is what was missing: the kernel is asked whether anything is mounted at the path, as well as whether the path answers. An empty directory answers a probe perfectly well, so an nfs mount that vanished while its mount point stayed behind read as healthy for the life of the agent — nothing retried, and nobgp status, the backend reported to noBGP and the network directory all described a healthy drive on a node with no drive at all. Reproduced on a Mac, where removing the leftover directory failed the probe within one poll and the node remounted itself in 90 seconds — the recovery machinery was never broken, it was simply never asked to run. Both questions are asked, never one: a mount that is present but wedged still has to fail, which is what the original probe was for.
Limitations worth knowing before you rely on one
These are properties of the platform, not bugs waiting on a release:
- ⚠ The agent's own service may not be able to read the node's mount. File operations and commands the agent runs for you against the mount then fail at every identity including root — on macOS with
Operation not permitted. It is not a permission or privilege problem, and your own account reads the mount normally; only the agent's service is blind. Two causes, with different fixes: macOS TCC consent, which a Full Disk Access grant ends, and a Windowswebdavmapped drive, which belongs to a logon session the agent's SYSTEM service cannot see. It is a per-node state rather than a platform rule —network_directoryreports it asinfo.mount_readable(router 0.4.83+). See When the agent cannot read its own mount. - ⚠ On macOS the two backends trade against each other, and the default takes the honest lock. Both serve — an earlier result where WebDAV did not finish a 60-file read in 328 seconds did not reproduce when both were re-measured on one Mac on agent 0.4.77. NFS answers questions about files and folders locally where macOS's WebDAV client goes back over the network for each one, and WebDAV is quicker at creating, renaming and deleting them.
nfsstays first because its locks are real: on a macOS WebDAV mount two processes can each take an exclusiveflockon one file and both be told they hold it. See what each backend measures at on a Mac. - ⚠ WebDAV is not available on much of the Linux fleet. It needs
davfs2, which the installer only reaches for when FUSE cannot be installed — which almost never happens on Debian, Alpine or Amazon Linux — while Arch, OpenWrt and Synology have no automatic fallback to it at all. Synology cannot supply it at any price:synopkg, DSM's only package manager, has no WebDAV package, and DSM registers no NFSv4 client either, so a Synology node has exactly one backend and nothing behind it. - ⚠ Nothing beside a file's own bytes travels with it, on any backend. A file is one stored object, so NTFS alternate data streams and extended attributes are dropped. The one exception is macOS, whose
._namesidecars are stored as ordinary files from router 0.4.128, so attributes written from a Mac do survive. The one that matters is mark-of-the-web, which Windows keeps in a stream and which is dropped silently — see What travels with a file. - ⚠ On Windows,
autoselects WinFsp from agent 0.4.82 where the driver is installed and is version 1.10 (2022) or newer; WebDAV is the fallback there, as it is everywhere else. WinFsp is the better drive on Windows — machine-wide rather than confined to one logon session, and not labelledDavWWWroot. ⚠ It is a shared driver, installed by rclone, sshfs-win, Cygwin and MSYS2 among others, so a machine can have it without anyone choosing it for noBGP — which means a node can change backend at an upgrade without you doing anything.fs: webdavpins it back. Seewinfspon Windows. - ⚠ On Windows, a
webdavdrive cannot read a file larger than about 47 MiB, and saysPermission deniedinstead. Windows' own WebDAV redirector caps a transfer at 50,000,000 bytes on a stock machine and refuses anything over it before a request is ever sent, so nothing on noBGP's side sees it and there is no error to enrich. The write succeeds, which is what makes it a trap: the bytes are stored correctly and every other node reads them back, while the machine that wrote the file cannot read it or evenstatit.winfspis a driver rather than a service in the path and has no such cap — a further reason it is the right Windows default — or the node's owner can raise the Windows limit by hand. See the troubleshooting entry. - On WebDAV, the guard against overwriting another node's change covers less. The other backends condition every upload on the version the handle read and refuse when it has moved. The operating system's own WebDAV client names no version, so from agent 0.4.147 the agent's proxy supplies one for the files it has served — and for everything it cannot vouch for, two nodes writing one file is still last-writer-wins with neither told. Below 0.4.147 that was every file. See The mounted drive uses it too.
- On FUSE, a write can fail after your program was told it succeeded. Uploads are queued, so
write,fsyncandcloseall return success and the upload can be refused seconds later. The bytes are not lost — the rejected copy is kept beside the cached file under a.rejectedname and the agent logs where — but nothing tells the program that wrote it, at the time. NFS reports a refused write properly, atwrite,fsyncorclose. - Reads and writes move whole files, on every backend. Reading 4 KB out of a 1 GB object fetches 1 GB; changing 4 KB inside a 10 MB file uploads 10 MB.
What changed in 0.4.65 and 0.4.66
Through 0.4.64 one fleet-wide order put nfs first on every platform. Nothing about the drive's contents or layout changes when the backend changes, and pinning fs holds any node exactly where it is.
-
Most Linux nodes move from
nfsback tofuse(agent 0.4.65). FUSE caches file contents and directory entries locally, so re-reading a tree it has already seen costs nothing, and it was then the only backend whose locks hold between nodes —nfscaught up in 0.4.69, so that half of the reason has since expired.nfskeeps second place as the fallback for a Linux box with no usable/dev/fuseor nofusermounthelper, which is a real population — such a host is unaffected and stays on NFS. Pinfs: nfsto keep NFS on a node that has both. -
Macs moved to
webdavin 0.4.65 and are back onnfsin 0.4.66. The move was a safety measure, not a performance one: two macOS kernel panics had been recorded inside the system's own NFS client, and being wrong about the cause costs the whole machine rather than one operation. What the fleet showed within the release was the price — the Macs became the only nodes reporting that a lock on their drive is not honestly enforced (see what each backend does with a lock below), because macOS WebDAV tells two processes they both hold the same exclusiveflockwhile NFS refuses out loud, and no WebDAV write-back was conditional at the time either. The panics have since been attributed to the agent's own NFS test suite rather than to a mount — the long-lived mounts on those machines have never panicked one — so 0.4.66 putsnfsback in front, and it lands together with the shared cache that makes an NFS mount cheap to re-read.The reason is the locking, and it holds whatever the speeds turn out to be — which is worth saying plainly, because this page previously said WebDAV on a Mac did not serve at all and that turned out to be wrong. See the numbers below.
-
Windows was unchanged in practice at the time.
winfspled the list, butautodeclined it, so a Windows node kept mounting over WebDAV exactly as it did — through agent 0.4.81. From 0.4.82autotakes it, and the preference order finally means on Windows what it means everywhere else (see below).
Which NFS client a Linux node has is largely a matter of when it was installed. From agent 0.4.58 the install script installs an NFSv4 client on Debian/Ubuntu, RPM hosts and Arch — worth having as the fallback a host without FUSE lands on. sudo nobgp upgrade does not install it — no package declares the client as a dependency, deliberately, because the installer is also what confines the rpcbind it pulls in — so a node upgraded in place keeps the backends it could already serve. Alpine, OpenWrt and Synology are left out of that on purpose; from agent 0.4.70 a Linux host needs no package at all to mount over NFS, only a kernel that has the v4 client, so those three can now take this backend as well.
The two macOS backends, measured
Both macOS backends were run against the same 60-file fixture on two Macs on agent 0.4.77 — one on Ethernet and idle, one on Wi-Fi and busy. This replaces an earlier reading on this page that WebDAV on a Mac "did not serve": it completed every phase on both machines.
| NFS (Ethernet, idle) | WebDAV (Ethernet, idle) | NFS (Wi-Fi, busy) | WebDAV (Wi-Fi, busy) | |
|---|---|---|---|---|
| Walk a folder it has never seen | 0.5 s | 3.9 s | 0.5 s | 7.5 s |
| Ask about 60 files it has seen | 0.02 s | 3.8 s | 0.01 s | 7.2 s |
| Write 8 MB | 0.8 s | 0.6 s | 14.4 s | 6.6 s |
| Create 60 files | 40 s | 15 s | 250 s | 57 s |
| Rename 60 files | 43 s | 12 s | 168 s | 48 s |
| Delete 60 files | 32 s | 9 s | 50 s | 18 s |
Read the ratio, not the seconds. The same trade holds on both machines — NFS is far quicker at looking at a folder, WebDAV 2.6–4.4× quicker at changing one — but the absolute cost is dominated by the machine and its network rather than by the backend: the same touch of the same file on the same build cost 0.4 s on one of these Macs and 3.4 s on the other. If creating files on the drive feels slow from a Mac, look at the link before you look at the backend.
What that means in practice:
- Work that reads and searches — a build, a backup,
git statusover the drive — is where NFS answers locally and WebDAV asks the network for every question. - Work that creates and renames a lot of files pays on either backend, and most on NFS. Part of that cost is inherent to a Mac: it writes a hidden
._namecompanion beside every file it creates on a network volume, so one create is two writes. Those companions are stored rather than discarded from router 0.4.128, which is what stoppedcpreporting a failure on a copy that had in fact landed — the numbers above were measured before that change. The rest of the gap is NFS confirming each change is stored before it says the change happened, and will remain. - The lock is why the default does not follow the create numbers. WebDAV's price for being quicker at changes is that it cannot lock, and on macOS an exclusive
flockreturns success to both holders. Its version check is narrower than the other backends' too, so several ordinary shapes stay last-writer-wins. See what each backend does with a lock. fs: webdavis still not a proven escape hatch on a Mac. What was measured is the WebDAV mount, made by hand beside a live NFS one. The agent falling back to it on its own — the probe, the order, and the switch a failing NFS mount would take — has not been exercised on macOS. If NFS cannot mount on a Mac,fs: offis the answer that is known to behave.
What nfs needs on a host
An NFS client is two separate pieces: the userspace mount helper (/sbin/mount.nfs4 on Linux, mount_nfs on macOS) and the kernel's own NFSv4 client. On Linux, from agent 0.4.70, only the second one is required. With no helper on the box the agent performs the mount itself, with the same options the helper would have passed — there is nothing for a helper to resolve on a loopback mount — so a container with no NFS packages installed mounts fine. That is worth having on Alpine above all, where the client package costs 34 packages and 57 MiB and pulls in rpcbind and python3, neither of which a single-port NFSv4 loopback mount uses. A helper that is present is still used in preference, so nothing changes on a node that already had one. On macOS the helper is part of the base system, and its absence there means a broken install rather than a missing package.
Through agent 0.4.61 only the helper was probed, so a box that had the helper and a kernel without NFSv4 reported nfs: available, was selected best-first, failed the mount — and, because a failed mount deliberately never falls through to another backend, was left with no shared drive at all. Embedded distributions are where the two come apart: on OpenWrt nfs-utils is the helper alone and kmod-fs-nfs-v4 is the kernel half.
From agent 0.4.62 the probe asks the kernel too, and nobgp status names what is missing, because the two halves have different remedies:
backends:
- "nfs: unavailable (this kernel has no NFSv4 client (no `nfs4` in /proc/filesystems after a module load; OpenWrt: `opkg install kmod-fs-nfs-v4` matching the running kernel))"
A kernel that has the modules on disk but has not loaded them still counts as capable: the agent asks the kernel to load them before deciding, and retries that every five minutes, so a node where you install kmod-fs-nfs-v4 picks NFS up on a later mount cycle without waiting for a restart. Inside a container there are no modules to load and no privilege to load them, yet the host kernel loads its own filesystem modules on demand — so from agent 0.4.70 the probe also settles the question by attempting a throwaway mount against a closed loopback port, which fails in microseconds on a capable kernel and registers the client on the way. The verdict is the kernel's registered filesystem list either way. On macOS nothing changes — the client is in the kernel and mount_nfs is part of the base system, so the helper is the whole question there.
winfsp on Windows
⚠ From agent 0.4.82 this is what auto mounts on Windows, wherever the driver is installed and is 1.10 or newer. Read the "opt-in" history below as history: it explains what held the promotion back and how each reason went, which is worth keeping because each will be re-proposed. What matters if you run a Windows node is the upgrade below.
Through agent 0.4.81 Windows was the one platform still mounting over WebDAV by default, and two long-standing quirks of the Windows WebDAV redirector came with it: the drive letter is visible in exactly one logon session, and Explorer labels the volume DavWWWroot. A WinFsp volume has neither — it registers a single machine-wide drive letter through the Mount Manager, and from agent 0.4.76 Explorer shows it as noBGP.
Through agent 0.4.75 it carried no label at all — vol N: answered "Volume in drive N has no label" — and Explorer fills that gap with its generic name for an unnamed local volume, so the drive read as Local Disk, indistinguishable from a real disk on a machine with several letters in use. The label is deliberately not per profile: the mount point already carries the profile name, and on Windows that means two profiles land on two drive letters, which is what a path uses and what tells them apart.
It was read-only in agent 0.4.62 — every write refused, so writing to the share from that node had to go through the file tools or the share's URL instead. From agent 0.4.63 it writes: create, write, truncate, rename, delete and mkdir/rmdir all work, and every write-back carries the same version check the nfs and fuse mounts use.
Writing on a real mount needs agent 0.4.65, though. On 0.4.63 and 0.4.64 the volume mounted, listed and read perfectly while every attempt to create a file was refused with Access is denied — from every account on the machine, including the one the agent itself runs as. Windows decides that before any of the agent's code is reached, from a security descriptor WinFsp builds out of what the filesystem reports about ownership, and what the volume reported granted full access to an account identifier no token on the machine holds, leaving read-and-execute as the only thing any real caller matched. From 0.4.65 the mount is given an explicit descriptor — full access for LocalSystem, Administrators and Authenticated Users, which covers the agent, the desktop account and the node's configured user — so creates, writes and renames work. The same release stops the volume reporting zero free space, which Explorer refuses a copy into before it starts. This needs WinFsp 1.10 (2022) or newer, because an older WinFsp rejects the option that carries the descriptor and refuses the whole mount rather than the one option it does not know.
From agent 0.4.73 the version is part of the probe, so an older WinFsp is reported as unavailable and never selected, rather than being pinned or chosen and then failing every mount cycle. It is a third state with its own remedy — "not installed" says install it, this one says upgrade it — and it names what it found and where:
backends:
- "winfsp: unavailable (WinFsp 1.7 is installed at C:\Program Files (x86)\WinFsp\ but this backend needs 1.10 or newer (older WinFsp refuses the whole mount rather than the one option it does not know) — upgrade it: install WinFsp (`winget install WinFsp.WinFsp`, or https://winfsp.dev))"
A DLL whose version cannot be read is treated the same way, deliberately: the node stays on the backend it already had rather than guessing.
From agent 0.4.67 a winfsp volume reads and writes through the node's own disk cache, the one the fuse and nfs backends already use. Until then every open downloaded the file again into a temporary of its own, which made it the most expensive backend per operation and gave two handles on one path two private copies — a write through one was invisible to the other. Now there is one cached copy per path: opening a file the node has already downloaded and nothing has changed costs no request at all, both handles read and write the same bytes, and the copy outlives the handle rather than being deleted at its last close, so the next open is free too. It is bounded like the others — an hour since last use, 1 GiB, 100 MiB of free disk kept back — and the versions it remembers do not survive an agent restart.
The same release makes the upload asynchronous, the shape a fuse mount has always had. Through 0.4.66 the write-back happened inside the CloseHandle that triggered it, so a refused save failed that call; now the bytes are queued and carried up behind the caller, and a refusal arrives after the close has already returned success — see the mounted drive's version check for where it surfaces instead. fsync jumps the queue's coalescing delay, so a program that wants its save sent now should call it. Deleting or renaming a file whose upload is still queued takes the queued bytes with it, so a deleted file is not recreated seconds later. Through agent 0.4.92 the one gap was a file inside a renamed directory that had no handle open on it and an upload still pending, which went up under the directory's old name — closed in 0.4.93.
Through agent 0.4.78 a delete that raced its own upload could bring the file back. Taking the queued bytes off the queue is only half of stopping an in-flight write-back; the other half is removing the node's cached copy, and on Windows that failed with a sharing violation for as long as the upload held the copy open — a file opened by Go is not shareable for delete, even to the process that opened it. So an upload already past the queue completed after noBGP had removed the file, and the object reappeared moments after a delete that had reported success. The same share mode is what a refused write's set-aside copy and a rename of a cached file need, so both could fail in the same window. Fixed in agent 0.4.79 on both the volume and the shared cache underneath it; no other platform was affected, because renaming and unlinking an open file is ordinary there.
What is different about the drive when a node moves onto it
Everything about the share's contents is the same — same folders, same files, same layout, and pinning fs holds any node exactly where it is. What changes is the drive:
- Better: the drive letter is machine-wide. A service, a scheduled task and an elevated Administrator prompt all see it, where the WebDAV drive existed in exactly one logon session. Explorer names it noBGP instead of
DavWWWroot. - Better: a losing cross-node write is refused rather than blended. Every write-back is conditional on the version the handle read, and the bytes it could not save are kept on the node. The Windows WebDAV drive's check is narrower — it is per file rather than per open, and only for files that node read through the agent — so several ordinary shapes there are still last-writer-wins with neither writer told, and below agent 0.4.147 every write was.
- ⚠ Worse: locks stop at that machine. This backend has no way to forward a lock to noBGP, so an exclusive open holds between programs on that machine and means nothing against a writer on another node — and because it is enforced locally, testing on one machine will suggest locks work across the network. The WebDAV drive it replaces is worse rather than better here: its lock request does leave the machine, so it looks cross-node while protecting nothing.
nobgp statusprints the contract this node actually has underfs.locking; read it before putting a database or a lockfile on the drive. See what each backend does with a lock. - ⚠ Renaming can be slower than on the WebDAV drive — the one ordinary operation this backend can lose at, and the one you feel when moving a folder of files between directories; how much slower is not settled, and one later run put the two at parity. Walking, listing, creating and deleting are all quicker. See what each operation costs.
- ⚠ Neither drive carries anything beside a file's own bytes, so mark-of-the-web is dropped on both, and both have now been measured. See What travels with a file.
- The free-space figure is invented. The share reports no quota, so the number Explorer shows is not one to plan against.
- The top-level
nodeandnetworksfolders cannot be created or deleted, exactly as on every backend; Windows reports access denied.
A leftover WebDAV mapping in someone's logon session is cleared before the volume mounts, from agent 0.4.94. A machine-wide drive letter is still shadowed, for one user, by a mapping that user's own session holds on the same letter — and a WebDAV drive that the agent had migrated into the console session leaves exactly that behind when the agent dies without unmounting (a crash, a kill, a power cut). The result reads as a contradiction: the file tools and nobgp status see a healthy volume while the person at the keyboard opens a dead share. The agent now walks the active logon sessions before it mounts and deletes its own stale mapping there. A mapping that is not noBGP's — the user's NAS on the same letter — is never touched and is reported instead, on nobgp status as fs.note, because that is what that user will see; clear it with net use <letter>: /delete in that session, or move the node to another letter with mount:.
To move a node back, set fs: webdav in the profile and restart the agent — or fs: off for a node that should mount nothing at all.
auto takes it from agent 0.4.82
⚠⚠ If you run a Windows node, this is the part to read. Through agent 0.4.81 auto declined this backend and a Windows node mounted over WebDAV. From 0.4.82 auto selects it wherever the driver is present and is 1.10 or newer — so a node that upgrades can change backend without you asking it to, and the only thing that says so is the backend named in nobgp status.
That reaches machines that never chose it for noBGP. WinFsp is a shared driver — rclone, sshfs-win, Cygwin and MSYS2 all install it — so "has WinFsp" is not the same as "wanted WinFsp here".
Two ways to hold a node where it is, and they are not the same:
fs: webdavkeeps the WebDAV drive the node has today. This is the way back.fs: offmounts nothing at all. It is not a fallback; it is no drive.
An explicit fs is honoured or refused, never quietly replaced, so either one holds through upgrades.
What it costs, and it is one operation: renaming a file was measured about 1.7× slower than WebDAV — the only phase this backend lost. Agent 0.4.82 removes a wasted network round trip from every rename, create and folder creation, so the gap should now be smaller; it has not been re-measured, which is why the older figure is the one quoted.
How it stopped being opt-in
Through agent 0.4.72 the reason given was that the one real mount it had run on lost data; that was fixed in 0.4.69 and confirmed on hardware on 2026-08-11, and from agent 0.4.73 the printed reason said what was true then instead. Two reasons remained, and each is worth recording because each will be re-proposed:
- Locks are node-local, and this is the one that came off without being fixed. This backend has no way to forward a lock to noBGP, so an exclusive open holds between processes on that machine and not against a writer on another node — see what each backend does with a lock. It stopped being a reason because it was never the safer half: it held the fleet on WebDAV, whose Windows client reports a lock it does not enforce across nodes. An honest node-local lock beats a dishonest cross-node one, and a losing cross-node write is refused rather than blended either way, because every write-back is conditional on the version the handle read. ⚠ The limitation itself has not changed and still applies to every Windows node.
- Creating a file is slow — this one was measured, and the objection lost: 402 ms/op against WebDAV's 600 on the same machine, a 1.49× win for WinFsp. The clause had been written against a 1025 ms figure with no WebDAV number beside it, so it had never actually been priced.
A third reason, case-sensitivity, was answered in agent 0.4.74 (see below) and came off the printed reason in agent 0.4.75. Agent 0.4.74 alone printed all three.
On agent 0.4.62–0.4.81, to try it: install WinFsp (winget install WinFsp.WinFsp, or winfsp.dev), pin fs: winfsp in the profile, and restart the agent. On those releases, leaving fs alone kept a Windows node on WebDAV.
On a node that does not have the driver, install WinFsp (winget install WinFsp.WinFsp, or winfsp.dev) and restart the agent; from 0.4.82 nothing else is needed.
What each operation costs
Measured on 2026-08-22, agent 0.4.95, on one Windows machine over a real link, with both backends run on that same machine on the same day so the two columns are comparable. Read the ratio rather than the seconds: a different machine and a different link give different numbers, and it is the shape between them that holds.
| Operation | Cost per file on winfsp | Against webdav on the same machine |
|---|---|---|
| Walk a folder | 2.4× quicker | |
| Ask about a file it has seen | 7× quicker | |
| Create | 1.5× quicker | |
| Delete | 2.4× quicker | |
| Rename | ~0.56 s | 1.75× slower (WebDAV ~0.32 s) |
- Renaming is the one to watch, and the one ordinary operation this backend can lose at — so moving a folder of files from one directory to another is the slowest ordinary thing you can do with the volume. It is not the network: noBGP's own side of a rename takes about 64 ms.
- ⚠ How big that gap is, is not settled. A second record from the same machine and the same script — 2026-08-24, agent 0.4.98, WebDAV alone over 200 files — averaged 541.9 ms per rename with a 21.9% spread, which against the 560.5 ms WinFsp figure above is parity rather than a loss. So the honest answer today is somewhere between nothing and about 1.75×. An earlier 1.7× this page used to give (2026-08-14, agent 0.4.80) came from a single WebDAV run and is superseded by the two above.
- Two releases took round trips off the rename before those runs. Agent 0.4.80 made a rename and a delete state their one change in the directory listing the node already holds rather than forgetting the listing and buying it back — renaming twenty files inside one folder used to pay twenty re-listings of a folder that was not even changing size, and an
rm -rfdid the same per file. Agent 0.4.82 removed another: creating a file, making a directory and a rename's destination all name something that is not there yet, and the volume was spending a lookup discovering that on top of the one the folder listing had already answered. Both are included in the figures above. - Deleting is quick, including large files this machine has never read — a 32 MiB file the node had never opened was removed in 0.1 s, because from agent 0.4.79 the volume no longer downloads a file in order to delete it (see below).
Opening a file costs nothing, from agent 0.4.79
Windows deletes and renames files through a handle — CreateFile, then set the disposition, then close — and Explorer, its shell extensions, the search indexer and antivirus all open files whose contents they never read. Through agent 0.4.78 a winfsp volume downloaded the file at the open, so the drive fetched a file in order to delete it, and browsing a folder could pull down everything in it.
From agent 0.4.79 the open transfers nothing at all. The download happens at the first operation that actually needs the bytes — a read, a write, or a truncate to a non-zero size — and a truncate to zero fetches nothing, because the content is being replaced anyway. What that changes for you:
- Deletes, renames and folder browsing on cold files cost no transfer. A warm file was already free, so the win is exactly the bulk case.
- A network failure now surfaces at the first read rather than at the open. Whether the file exists, and whether it is a directory, are still answered before the open, as they always were — those come from the same lookup that makes the open free.
- A handle that never read anything writes unconditionally, and one that read first is still conditioned on the version it read: the check moves to the first write rather than being dropped.
Two spellings of one name, from agent 0.4.74
From agent 0.4.74 a winfsp volume is case-insensitive and case-preserving — what every Windows program assumes, and what the WebDAV drive this backend would replace already gave them.
- A mis-cased path resolves.
Get-Item CASE.TXTfindscase.txt, and so do open, read, write, truncate,mkdir, delete and rename: every operation that takes a path resolves it for itself, so the file's contents are found under the folded name as well as its size and timestamps. Through 0.4.73 a wrong-case read failed outright. - The stored spelling is what comes back. Listings, Explorer's address bar and a shell's tab completion show the name as the share stores it, not the case you typed.
- An exact spelling costs nothing. A path that came out of a listing — which is nearly all of them — matches byte for byte and never folds. Folding is a fallback, applied one path component at a time against listings the node already holds, so a path under a directory it has seen recently costs no extra requests.
- It is safe because the share refuses collisions. Router 0.4.71 made names on the storage trees unique without regard to case, so a fold can never have two files to choose between. A directory that already holds
Foo.txtandfoo.txt— from before that rule, or written straight to the bucket behind noBGP's back — keeps both, both listed and each reachable by its own exact spelling. Only a third spelling that is neither of theirs is folded, and it resolves to the same one of them on every lookup rather than to whichever came up first. - Only this backend folds. A Linux
fuseornfsmount still presents the share exactly as it is stored, because that is what those platforms' own filesystems do and what every script on them assumes. The rule is Unicode simple case folding, the same one the share enforces:K(U+212A) is the same name ask,İ(U+0130) is not the same name asi, and composed and decomposedéstay two names. - A router it cannot reach costs the fold, not the operation. Resolving is best effort and stops at the first component it cannot answer, leaving the rest of the path as you wrote it — so a blip degrades to the pre-0.4.74 behaviour instead of failing the call.
Creating a file, from agent 0.4.73
Every create on this volume used to ask noBGP whether the path existed before claiming it, and that question could never be answered from the node's cache: a directory listing tells you what is there, never that a name is absent, so the check was a request guaranteed to come back "not found" in front of every single create. It is gone. The existence question is asked of the directory listing the node already holds, and the claim itself is what settles it — a create takes the path exclusively, so being wrong about the answer costs a refusal rather than someone else's file.
Two behaviours change with it, both on a path this mount has no warm listing for:
- A create onto a file that already exists answers
EEXISTwhere it used to empty the file and carry on. For a caller that asked to create a file that is the right answer, and Windows applies the caller's own disposition to it — an overwrite-minded program retries as an open-and-resize and gets what it wanted. - A create onto an existing directory answers
EISDIRrather than queueing a whole-file write over it.
Agent 0.4.82 takes a second round trip off the same shape. Since agent 0.4.74 the volume resolves a mis-cased path by asking about each part of it in turn — free for a name the node has a listing for, a request for one it does not. A create, a mkdir and a rename's destination all name something that is expected not to be there, so that lookup was buying a second "not found" on top of the one the folder listing had already implied. Those three now resolve the last part of the path from the listing the node already holds and from nothing else, so they can never cost a request; the folders above are folded exactly as before. On a folder the node has no listing for, the last part is left spelled as you wrote it — and if that spelling turns out to collide with a stored one, the create is refused as already existing, which is the answer it would have got anyway.
Appending to a file the volume just wrote
Agent 0.4.69 fixes a data loss on a winfsp volume, found on the one real mount this backend has run on (2026-08-10). Windows asks a file for its size by name as well as through an open handle — that is what every dir, Get-Item and Explorer size column does — and the by-name question fell through to the cached listing, which still held the size before the write. Windows then computes an append's write offset from exactly that number, so Add-Content on a file this volume had just created landed at offset 0 and destroyed what was there, while reporting success and uploading the bytes perfectly.
There were two doors to it, and 0.4.69 closes both:
- While a handle is open, a by-name question is now answered from that handle's unsaved bytes, the same answer the handle itself already gave.
- After the last handle closes, it is answered from the pending write-back for as long as the upload is queued or in flight — nominally about five seconds, and up to five minutes when the upload is retrying under backoff. The answer has to outlive the handle because the upload does:
Set-Content f; Add-Content fcloses the file between the two commands and lands in exactly that window.
The fix was confirmed on that same hardware on 2026-08-11 — the size correct immediately after a write, Add-Content landing at the right offset, the share's copy matching byte for byte — so from agent 0.4.73 this is no longer among the reasons auto declines the backend. Agents 0.4.69 through 0.4.72 went on reporting it in nobgp status after it had been fixed. On agent 0.4.67 and 0.4.68 — the releases where the volume both wrote and deferred its uploads — do not append to a file on a winfsp mount; write it whole, or use the file tools or the share's URL. webdav — which through agent 0.4.81 was what every Windows node mounted with unless you pinned otherwise — was never affected.
A directory listing kept the pre-write size until agent 0.4.146, which is the same gap at a third door. Both answers above are given to a question asked about one file, and a dir, an Explorer window or a backup tool that compares sizes reads a whole folder instead — served from the listing the node already holds, which still carried the row as it stood before the write. On a winfsp volume that row is the empty file a create publishes, so a program that wrote a new 150 KB file and closed it saw 0 KB in dir and in Explorer while Get-Item on the same name answered correctly, for as long as the upload was queued or retrying. From agent 0.4.146 a listing served from the node's cache carries the size and modification time of that machine's own unsent writes, on a winfsp or nfs mount. Nothing else in the row changes — a file the listing does not hold is not added to it, and which files a folder contains is answered exactly as before. A fuse mount was never affected: it keeps its own listing and already answered from the open handle.
How much a Linux nfs mount moves per request
Through agent 0.4.70 a Linux nfs mount moved bytes 1 KiB at a time. The mount has always asked for 128 KiB reads and writes, and macOS gives it exactly that — but the Linux kernel client sizes its transfers from what the server says it can serve, and the server said nothing. Linux reads a missing maximum as an unknown one and falls back to its own 1 KiB floor, whatever the mount asked for, with nothing logged and no error anywhere: the mount succeeded, the drive listed and read correctly, and every byte crossed in 1 KiB pieces. An 8 MiB write became roughly 8192 separate write requests, measured at 8.8 seconds against 0.1 s for the same bytes on a fuse mount on the same machine.
From agent 0.4.71 the server advertises the 128 KiB it will really serve, so the client negotiates 128 KiB and that same write costs 64 requests. Nothing about mounting changes — same port, same options, same drive — and there is nothing to configure: upgrading the agent and letting the drive remount is the whole of it. What the mount settled on lives in the kernel's own view of it rather than in the mount options, so read it there:
nfsstat -m # or:
grep nobgp /proc/mounts # rsize=131072,wsize=131072 from 0.4.71
macOS was never affected — it honoured the requested size with or without the advertisement — and no other backend was. Which Linux nodes carried it is worth knowing, because auto prefers fuse there: it bound the hosts with no usable /dev/fuse and the nodes pinned to fs: nfs.
The nfs server was replaced in 0.4.69
From agent 0.4.69 the NFSv4.0 server behind the nfs backend is a different implementation. Nothing about mounting it changes — same loopback port, same mount options, same client, same drive — but the old library had no NFSv4 state machine at all, which is what put a ceiling on several behaviours at once:
- Locks are forwarded to noBGP, the headline of the change and the reason for it. An
nfsmount joinsfuseas a backend whereflockandfcntlmean something on a second node. fsyncreports a refused write-back. The old server discarded the error from bothfsyncandclose, so a save the router refused could only be discovered from the nextwriteon the same handle. A program that callsfsyncand checks its return value now learns straight away.rmdiron a directory that still has something in it is refused, with "directory not empty". Before this it deleted the whole subtree and reported success — the share's own delete is recursive, and the emptiness check the old server made could never fire against a remote tree.rm -ris unaffected either way: it empties the directory through the mount before it callsrmdir.- Two writes from different nodes inside the same second are told apart. The marker the kernel uses to decide a file has changed was derived from the modification time, which is second-granularity, so a peer's write landing in the same second as the one before it could be missed. It now comes from the share's own version identifier.
- Bytes a client had buffered but not committed survive an agent restart. The client re-sends them when it notices the server restarted, where before they died with the agent.
truncateon a file nobody has open takes effect immediately, rather than waiting for a close that is never coming.
Under nfs, a file is fetched whole the first time something actually reads or writes its contents, and uploaded whole when it is closed, and only if it was modified — the share has no partial write, so a small edit to a large file is a full upload however it is made. Directory and attribute lookups are cached by the kernel, which is what keeps an ls -l of a deep tree from becoming a request per entry — for a fixed minute through agent 0.4.66, and from 0.4.67 for as long as you tell it to, the same window the agent's own listing cache behind it uses.
Re-reading a file the kernel has already cached costs nothing from agent 0.4.65. Through 0.4.64 the bytes were fetched when the file was opened, so a second pass over a tree spent the same whole-file downloads as the first even though the kernel's own cache meant it never asked for a single byte — on a 60-file tree, opening every file without reading any of them took longer than reading the whole tree cold. The fetch now happens on the first operation that needs the contents, so a warm read touches noBGP zero times. A file another node has changed is still re-read: the kernel notices the change on its next open and asks for the bytes, which is what triggers the fetch.
And from agent 0.4.66 the fetch itself is served from the node's own disk cache — the same one a fuse mount has always had, which until then no other backend did. When the kernel does drop a file from its cache and asks for the bytes again, an unchanged file is copied from the local copy instead of downloaded (measured at ~57 MB/s against a ~90 ms download for a small file on a Pi CM4), and the version check that decides "unchanged" costs no request of its own. Two limits worth knowing, because they are what the win is not: the cache is bounded — an hour since last use, 1 GiB, and 100 MiB of free disk kept back — and the versions it remembers do not survive an agent restart, so the first read of each file after one downloads as before. Writes are unaffected: the share has no partial write, so a small edit to a large file is still a whole-file upload.
And from agent 0.4.68 the open is free too. Through 0.4.67 a warm re-read still spent about one request per file — not for the bytes, which were already local, but for the file's parent directory: every open stats it, and the agent looked for it only in the listing of the directory above it, which reading a folder's files never has a reason to fetch. So each open fell through to a lookup of a directory that had just been listed, measured at roughly 113 ms each on a real mount, which was nearly all of what a second pass over a tree still cost. A directory's own entry now comes back with its listing — the same request always carried it — so that stat is answered from what the node already holds. It expires with the listing it came from, on the same fs-cache-ttl clock, so nothing is served for longer than before, and a directory not covered by a warm listing is still asked about. A winfsp volume looks paths up through the same cache and gains the same thing; a fuse mount keeps its own directory cache and never paid this.
Appending to a file on an nfs mount needs agent 0.4.63 on Linux. Through 0.4.62 the shell's >> redirect emptied the file it appended to: the NFS server library the agent uses marks every ordinary create as truncating, and on a path that already existed the agent honoured it, so echo line two >> notes.md left the share holding line two alone — with every syscall reporting success. Agent 0.4.62 fixed one half of this (it stopped publishing an empty file the moment such a handle was opened) and 0.4.63 fixes the other. macOS was never affected, because its client checks the file exists before opening it and never asks for the truncation; Linux, where most nfs-backed nodes are, was. A client that genuinely does want the file emptied still gets that — it restates it as a separate resize, which the mount serves. If you are on an older agent, append over the share's URL or with the file tools until you can upgrade.
That upload happening on close is why agent 0.4.61 answers an attribute read from the file's unsaved changes while it is still open. Between a write and the upload, the router still holds the previous version, and the node used to answer from it — so a program that truncated or wrote through an open descriptor was told the size it had just changed, and the kernel's one-minute attribute cache then held that stale answer while the content was already correct. On 0.4.60 and earlier, truncate -s 8 file on the drive is the shape that shows it; close the file, or wait out the cache, and the size corrects itself.
A writer killed before it closes a file no longer strands what it wrote (agent 0.4.64). Because the upload happens on close, a kill -9 of a program writing to an nfs mount left the agent holding bytes it would never send: the mount went on answering that path's size and modification time from the abandoned copy — for every process on the node, outranking anything another node or the file tools wrote afterwards — until the agent restarted. The agent now saves those bytes itself once nothing has touched the handle for a minute, which is the POSIX answer: writes the program completed are what the file contains. It logs a warning naming the file when it does. A handle that is merely quiet is unharmed — it is flushed, not closed, and a writer that comes back can carry on writing to it — and if the save is refused because another node wrote the file in the meantime, the unsaved copy is set aside exactly as it would be on close. Unmounting the drive or dropping the connection already healed this case and still does.
From agent 0.4.59 an nfs mount also survives a restart of the agent. The server binds the same loopback port it bound last time, recorded per profile as the agent-written nfs-port key: the kernel remembers the port a mount was made against, so an agent that came back on a different one left the existing mount pointing at a port nothing was listening on, and every operation on it answered Operation timed out. If the remembered port is taken by something else, a new one is chosen, recorded, and the drive remounted. On 0.4.57 and 0.4.58, unmount the drive by hand after a restart that leaves it timing out.
Recording that port used to restart the agent, and agent 0.4.78 stops it. The agent watches its own profile file so that an edit you make is picked up without a restart — and through 0.4.77 it could not tell its own write from yours. Anything it recorded after startup therefore read as an edit and unwound the process for a reload: measured, the log said config changed, exiting for reload about five seconds after a mount came up, tearing down the drive that had just succeeded and costing the node a restart. It affected the values written from the mount and MCP paths — nfs-port, mcp-port, and last-mount-point in this release. The agent now recognises the write it just made and carries on; an edit of yours lands afterwards and is still picked up exactly as before.
The same release clears a mount left behind by a previous agent at the node's own mount point before mounting there. Mounting over a leftover stacks a second mount on top rather than replacing it — one extra volume per restart — and a wedged leftover holding the mount point made every new mount fail, leaving the node at mounted: false retrying indefinitely. Through agent 0.4.77 that was the only mount the agent would remove, because it is the one it can prove is its own — which is why fs: off left a stray mount alone until agent 0.4.78. Several layers can be stacked at once, so the clearing unmounts until the mount point is free, up to four layers.
On 0.4.59 that clearing ran after the mount point was created, which left the worst case unreachable: a wedged mount makes its own mount point unanswerable, so creating it failed first and the clearing never ran — fs.error read nfs: mount point /Volumes/nobgp: mkdir /Volumes/nobgp: file exists with fs.type empty, indefinitely. From 0.4.60 the clearing runs before the mount point is touched, so a node already stuck this way repairs itself on its next mount cycle. A failed mount that created the mount point also removes it again on the way out, since on macOS a leftover /Volumes entry reads to the next attempt as a mount point rather than as debris.
On macOS, agent 0.4.61 escalates an unmount that is refused. A macOS unmount answers Operation not permitted — to root, with nothing holding the volume — when something has objected to the volume going away, typically a Finder window on it or the "Server connections interrupted" dialog. That is not the same as busy, so umount -f does not override it and every retry failed identically, wedging the node with no way back. The agent now falls through to diskutil unmount force, which does override it, on both the shutdown path and the clearing above — so a node in this state clears the mount itself instead of needing a hand. Linux has no equivalent step and needs none: the lazy detach it already uses is the last escalation there.
A mount point that is busy but is no longer a mount says what is holding it, from agent 0.4.70. There is one shape the clearing above cannot help with: the old mount has already been unmounted, so there is nothing in the mount table to clear, and yet some process still has the directory open — so every remount answers Resource busy and the node retries that forever, reporting only those two words. The agent now names the holders in fs.error and in its log:
fs:
type: ""
mounted: false
error: "nfs: mount point /Volumes/nobgp is busy but not in the mount table — held by bash(4821), Finder(512), which must exit before the mount can recover: ..."
It names rather than reclaims, deliberately: a forced unmount cannot dislodge a live holder, and killing whatever is holding your mount point is not a decision the agent gets to make. Ending those processes is enough — the mount retry that is already running picks the drive back up on its next cycle with no restart. Where the holders cannot be identified (the lsof it uses is not installed, or the wedge defeats it) the message says so and still tells you the shape of the problem.
From agent 0.4.75 you get that answer whichever backend is mounting. The pin is on the mount point, which every backend on the platform shares, so switching backends was never an escape from it — and through 0.4.74 only an nfs mount said so. A Mac pinned to fs: webdav met the identical Resource busy with no diagnosis at all, and worse: a webdav mount failure was reported as a decline rather than an error — which it no longer is, from agent 0.4.151 — and a decline then cleared fs.error, so such a node showed mounted: false with an empty reason indefinitely while the one sentence naming the process that had to exit went only to the local log. A busy point is now reported as the failure it is, in the same words, prefixed with the backend that hit it.
fuse joins them at agent 0.4.78. It leads the chain on Linux, so it is the backend most Linux nodes actually mount with — and through 0.4.77 a busy point there answered with generic advice to go and run fuser -m yourself, on a machine whose log you are usually reading from somewhere else. It now names the holders in the same words as the other two, both when a remount finds the point busy and when the unmount of a leftover is refused. The same release fixes a fuse mount point whose path contains a space: the kernel escapes it in its own mount table, and the agent's reading of that table did not undo the escaping, so such a point read as not mounted — which silently skipped both the leftover clearing and this diagnosis, on precisely the paths that are hardest to sort out by hand.
⚠ winfsp still reports nothing here, and not for want of trying: a refused WinFsp mount hands back no reason at all for the agent to pass on, and Windows has no lsof to name a holder with. On that backend a mount that will not come up still shows mounted: false with an empty fs.error.
Two machines on the same platform running the same agent can land on different backends, because the answer depends on what is installed. nobgp status reports which one was chosen and why the others were not:
fs:
type: fuse
mount: /mnt/nobgp
mounted: true
selected: auto
locking: "Locks reach other nodes: the first lock on a file takes an exclusive lease at the router, at whole-file grain between nodes …"
backends:
- "fuse: available"
- "nfs: unavailable (this kernel has no NFSv4 client (no `nfs4` in /proc/filesystems after a module load; OpenWrt: `opkg install kmod-fs-nfs-v4` matching the running kernel))"
- "webdav: available"
The list is that platform's preference order, best first, and holds only the backends that platform can run at all — a Mac lists nfs and webdav, a Windows node winfsp and webdav.
locking says what a lock on this mount is worth, from agent 0.4.79. It is one sentence about the backend the node is actually mounted with — whether locks reach other nodes and, where they do not, what still holds between programs on this machine — and it is the line to read before putting a database or a lockfile on the drive. It is reported because the mount will never tell you: a lock the drive cannot enforce still returns success, so the program that took it carries on believing it and nothing anywhere fails. Every backend answers, including the two whose answer is good news, so silence never has to be read as either "fine" or "no guarantee". It is absent when nothing is mounted, and it is never reported for the other entries in backends — those are advice about what could serve, and a locking line beside one would read as a promise about the mount this node has. See what each backend does with a lock for the same information per backend, and network_directory's info.fs_locking for the fleet-wide view.
The backends list is reported even when nothing is mounted, which is when it is most useful — it is the answer to "which backend was this node even trying, and what is stopping it". From agent 0.4.59 an fs.error line sits beside it saying why there is no mount, when there is none — a pinned backend the host cannot serve, a failed mount, a backend that declined:
fs:
type: ""
mount: /Volumes/nobgp
mounted: false
selected: nfs
backends:
- "nfs: available"
- "webdav: available"
error: "nfs: mount_nfs -o port=61601,vers=4.0,...: exit status 1"
It is empty while the drive is mounted, and empty before the first attempt — absence means nothing has gone wrong yet, never that nothing is wrong.
Pinning a backend is the fs config key (auto, off, nfs, fuse, winfsp, webdav; default auto). It has no command-line flag — set it in the profile's YAML file (see Configuration file):
fs: nfs
- A pinned backend is honoured or refused, never substituted. Setting
fs: nfson a host with no NFS client fails out loud and retries the same backend — quietly mounting WebDAV instead would defeat both reasons to pin one (holding a backend steady, or diagnosing it). - A value that isn't one of the six is a typo, so it warns and falls back to
auto. A misspelling must not take a node's filesystem away. - A backend that
autoskips can still be pinned. Naming a backend the probe found is honoured — pinning is how one gets used beforeautowill take it, which is howwinfspwas reached through agent 0.4.81, beforeautotook it in 0.4.82. - Leaving the key out means
auto, so a node configured before the key existed keeps mounting exactly what it mounted. offmounts nothing — and serves nothing: the local shared drive proxy is not started either, so an operator who turns the drive off does not leave a loopback listener holding the node's credentials behind. It is also the escape hatch if the mount machinery itself misbehaves on a host. From agent 0.4.78 it clears a drive left behind by an agent that was killed, on Linux and macOS — see A mount left behind when the node is not going to mount; through 0.4.77 such a mount stayed until someone unmounted it by hand. A restart after an orderly stop is clean either way.- A mount failure never falls through to another backend. Selection falls back on a backend being unavailable, which is a property of the host; a failed mount is usually specific and transient, so the same backend is retried on the next cycle.
Changing fs takes effect on the next agent restart.
A mount left behind when the node is not going to mount
There are two ways a node starts up and never mounts anything: fs: off, and no mount point at all (a Windows profile that found no free drive letter). Through agent 0.4.77 both returned before looking at the machine's mount table — so an agent that had been killed while the drive was mounted (a SIGKILL, an out-of-memory kill, a host reset, a crash in the mount driver) and then came back configured not to mount left that mount in place, serving nothing. nobgp status said mounted: false beside a path the kernel still had a mount on, every read into it hung or errored, and the agent never touched it again — the way out was a person working out that the agent's report was not describing the kernel, and unmounting by hand.
From agent 0.4.78 the agent clears that mount at startup, on Linux and macOS:
- It asks the kernel, and matches the mount's source rather than its location. A mount is removed only when what the kernel says is mounted there is something this agent makes — its own FUSE filesystem, its own loopback NFS export, or its own local shared-drive proxy. Anything else at that path is somebody else's mount and is left where it is.
- Two paths are checked: the configured mount point, and where a mount was last actually up. The second is recorded in the profile as the agent-written
last-mount-pointkey, and it is what covers a mount point that was edited or blanked since the drive came up. It is never mounted on, and a stale value is harmless, because the source still has to match. - It does not escalate. Unlike the clearing on the mounting path, a refused unmount is not forced here — nothing is waiting on it, so it is left for the next start rather than pushed through against whatever is holding it. What it did, or could not do, goes to the agent's log.
- Windows is not covered yet. A leftover there is a drive mapping rather than a row in a kernel mount table; a Windows node running
fs: offstays silent about it rather than warning about a question that cannot be asked on that platform. Remove one withnet use <letter>: /delete.
An orderly stop — nobgp service stop, sudo nobgp service restart, a normal shutdown — has always unmounted cleanly on the way out, so this only ever mattered after a kill.
Which is why fs: off is worth setting in the right order. fs: off stops the agent mounting; it is not a way to take down a mount that is already up, and least of all one that is already stuck — the clearing above runs at startup, matches the mount's source, and, as it says, does not force a refused unmount. So on a node whose drive is currently mounted, sudo nobgp service restart first, and set fs: off after. A healthy mount goes down on the orderly stop; a stuck one is not cleared by that stop either, but the restart still comes back into the mounting path, whose own clearing does force an unmount past whatever is refusing it. Setting the key first removes both: no orderly unmount and no mounting path, so the wedge stays exactly where it was and the agent has no reason to look at it again.
A teardown belongs to the process that made the mount
A restart runs two agents at once for a moment: the outgoing one is still tearing its mount down while the incoming one is already mounting. Through agent 0.4.105 the outgoing one had no way to tell those apart. Every destructive step of a teardown named a path — umount <point>, diskutil unmount force <point>, and the "is anything mounted here" check in front of them — and a path changes hands, so a slow unmount could land on the mount its successor had just made.
Measured on a Mac at agent 0.4.95: an fs: webdav → fs: auto restart had the outgoing WebDAV unmount take the nfs mount the incoming process had already made, and macOS removed /Volumes/nobgp with it. The node had no filesystem for about 105 seconds, and came back only because the mount health check noticed and remounted.
From agent 0.4.106 a mount carries the kernel's own identity for the mount that process made, captured the moment it succeeded:
- Every step re-asks, not just the first one. Each of
umount, the forced escalation behind it, and the removal of the mount point is a subprocess that can run for seconds, and the measured handover happened inside one of them. - It stops only on positive evidence — there is a mount at the path and it is not the one that was captured. A point holding no mount, or a mount table that cannot be read, tears down exactly as it always did: a mount left behind is its own problem, and on macOS a stale
/Volumes/nobgpblocks the next mount outright, so the fallback is always to carry on. fuse,nfsandwebdavall take it, on Linux and macOS. Awinfspvolume does not, because Windows keeps no mount table of this kind.- Clearing a leftover before mounting is deliberately unaffected. A mount sitting at the node's own mount point, at the moment the node is about to mount there, is provably the node's own to remove — which is exactly the claim a teardown cannot make, and why the two paths ask different questions.
- On macOS the identity is the kernel's per-mount id, so it tells one mount from the next mount of the same backend at the same point. On Linux the mount table carries no such field and the match falls back to the mount's source and type, which still separates a restart that changes backend — the shape this was measured on — but cannot separate one
nfsmount from the next.
An orderly stop that is not immediately followed by another agent was never affected, and neither is a node whose backend does not change across a restart on macOS.
How long a directory listing is cached
The mount caches the listing of a directory it has read, so walking a tree does not ask noBGP about the same directory once per file in it. From agent 0.4.66 how long it may serve that copy is yours to set, as the fs-cache-ttl config key (or NOBGP_FS_CACHE_TTL, or --fs-cache-ttl on a foreground nobgp agent). The default is 30s, and the value is clamped to [0, 10m]:
fs-cache-ttl: 2m
- It is a staleness window, not a performance dial. It is exactly how long this node may keep showing a directory as another node left it before that node's newer write — and the same number is how often every mount re-asks about a directory nothing changed. Both directions cost something.
0is a real setting, not "off". It re-asks every time, while file contents are still served from the local cache when the version has not moved — the freshness answer for a tree several machines write to at once.- Write a unit.
fs-cache-ttl: 30is thirty nanoseconds, not thirty seconds, so a value under a second is treated as a missing unit and ignored with a warning rather than obeyed. Write30s. Anything that is not a duration at all, and anything past the ten-minute ceiling, is likewise reported and clamped rather than taking the node somewhere nobody asked for. - From agent 0.4.67 it is the whole answer on
nfs,fuseandwinfsp. In 0.4.66 it reached thenfsbackend alone, whose own listing cache was a fixed 60 s before that; afusemount kept a separate hard-coded 30 seconds and awinfspone cached nothing. All three now read this key, and where the kernel caches in front of the agent — annfsmount's attribute cache, afusemount's directory entries — that window comes from the same number, so the two caches in series cannot disagree. When they did, the longer one won while everything claimed the shorter: annfsnode set to5sstill let the Linux kernel serve a directory attribute up to 30 seconds old, because the floor it mounts with defaults to 30. Awebdavmount is the operating system's own client and still caches on its own terms. - The kernel's half is fixed when the drive is mounted. On
nfsandfusethe timeouts handed to the kernel are mount options, and a mount option cannot be changed under a live mount — so a changedfs-cache-ttlreaches the agent's own cache on the config reload and the kernel's on the next mount cycle. Restart the agent if you want both at once. - It does not govern file contents. Whether a cached copy of a file may be served is decided by comparing versions with the router, not by this clock, so raising it does not make a mount serve stale bytes for longer.
Unlike the other flags, --fs-cache-ttl is not written to the config file by nobgp config — to make a value durable, put it in the profile's YAML (the agent reloads on the edit) or in the environment.
One related behaviour that is not tunable: if the refresh of an expired listing fails — a router blip, a moment without a network — the copy the node already had is served for up to 30 seconds past its TTL rather than reporting an I/O error or, worse, an empty directory. Past that the error is real.
On a Linux fuse mount that grace stops being a clock, from agent 0.4.95. While the refresh keeps failing, the listing the node already holds is served for as long as the failure lasts rather than for a fixed grace, and the only directory that reports an error is one this mount has never listed and therefore has nothing to say about. What it must never do — and did through 0.4.94 — is answer an unreachable noBGP with an empty directory: see a listing noBGP could not be asked for.
A momentary busy answer from noBGP no longer reaches the program as an I/O error, from agent 0.4.77. A short burst of 503s — a gateway with no healthy target for an instant, one refusing new work, one that answered too slowly, or a rate limit (429, 502, 503, 504) — used to surface on the mount as ls: fts_read: Input/output error, which reads as corruption for a condition that means come back shortly. Measured on a Mac: 35 of them in 16 seconds. The mount now retries such a request itself, and because the retry lives in the one client every backend shares, nfs, fuse, winfsp and webdav all get it.
- Bounded, deliberately short. All attempts of one operation share about 5 seconds, and never outlive whatever deadline the calling program already had. It does not paper over a real outage — holding a filesystem call open for the length of one would trade an honest error for a process you cannot kill — so a longer outage still reaches you as an error, with the node's log naming the status and how many attempts it spent.
- Giving up names the status, wherever the clock ran out — from agent 0.4.83. The point of the error is the
503:context deadline exceededin its place sends the next reader looking at the network instead of at the gateway. That held only when the budget expired between attempts; expiring inside one let the timeout displace the status the earlier attempts had already established, and which of the two you got depended on how loaded the machine was — so the busier the box, the worse the diagnosis. The last transient status a run saw is now reported either way. - The pause is the server's if it says so. A
Retry-Afterthat fits the budget is honoured; one that does not ends the retrying rather than being shortened to something noBGP never agreed to. Otherwise the wait backs off from 100 ms with jitter, so the many files a directory listing opens at once do not come back in lockstep. 500is never retried. That one is a failure handling this particular request, and repeating it only turns a bad request into a loop.- Uploads are not retried here, because a request body cannot be replayed. Writes are already covered where it counts: a queued write-back retries on its own schedule and survives a restart. What still reaches you is an upload a program is waiting on — a
close()onnfs, a flush onwinfsp— and that is the honest answer rather than a gap.
Writing a lot of small files stopped throwing the listing away, in agent 0.4.70. Every completed upload used to discard the cached listing of the file's directory, so a burst of writes left the mount with no warm listing at all and every subsequent open paid its own round trip to ask about one file. Measured on a winfsp volume: 100 reads of 4 KB files took 60 seconds, against 1.1 s for the same files through the Windows WebDAV drive. Two changes fix it, both in the cache the nfs, fuse and winfsp backends share. An upload now edits its own row into the listing — the name it wrote, the size it sent, the version the router answered with — instead of forgetting the whole directory; and a question about a file the node has no answer for lists the parent directory, one request, which warms every file beside it for the TTL rather than answering about one. Nothing is served for longer than before: the edited row expires with the listing it joined, on the same clock, and a name a listing does not hold is still asked about individually rather than reported as missing. On 0.4.70 the nfs backend got the second half but not the first — its own uploads still dropped the parent listing, so a write-heavy burst there cost what it always had; from agent 0.4.73 the edit happens inside the upload itself, which is the one door every backend's writes go through, so nfs gains it too. Writing 60 files into a directory no longer throws that directory's listing away 60 times, and reading them back afterwards costs no request per file.
A fuse mount also stops inventing a change every time it fetches a file (agent 0.4.70). A file's size and modification time were answered from the node's local copy whenever one existed, and a local copy's timestamp is when the node downloaded the file rather than when anyone wrote it — so an unchanged file looked freshly modified after every fetch, which defeats the attribute cache the Linux kernel keeps in front of the mount and the re-download suppression that rides on it. The local copy now answers only while it is the only place the bytes exist: an open handle holding writes, or an upload still queued or in flight. Everything else is answered from what noBGP holds. It is the same rule nfs has followed since 0.4.61 and winfsp since 0.4.69, and a rename carries it with the file so a renamed file with unsaved bytes still reports its own size. Agent 0.4.79 closes the last gap in that hand-over: a rename landing in the instant a program's first write was being recorded left the record attached to a name nothing was writing, so the file being written fell back to the size noBGP last saw — the same wrong-offset shape, in a window a few instructions wide.
A renamed directory takes its cached copies with it, from agent 0.4.93. The cache is keyed on each file's path, so before that release moving a directory left every cached copy beneath it filed under the name it had moved off, with nothing able to name them again. Reading one of those files under its new name was never wrong — the lookup missed and fetched — but two things rode on the stranding: the orphaned copies held cache space against the node's quota until they aged out an hour later, which is real on a small machine; and a queued write-back for a file inside that directory went up under the old name, which is a save landing somewhere you did not put it. A rename now moves the path and everything beneath it — cached bytes and queued uploads together — and drops any entry it meets that no longer has either, so a rename is also when the index is swept.
Moving the cache is not the same as moving the rename, and a fuse mount had only half of it until 0.4.94. Renaming a directory on a Linux mount left the mount's own idea of where each file below it lives pointing at the path it had moved off — only the directory's immediate children were re-pointed — so the first read of a file further down asked noBGP for a path it no longer had, and the mount answered Input/output error for the whole window before the kernel looked the name up again: 5 to 10 seconds, measured on a Pi CM4 and a Pi 5. The same release drops the stale directory listings at both ends of the rename, so creating a new directory with the name you just moved away from no longer shows the old one's contents for the length of the listing cache. An nfs or winfsp mount already moved its whole subtree; a webdav mount has no such table.
On an nfs mount that applies to a renamed file too, and that is the half worth knowing about: through agent 0.4.92 that backend moved no cached state on a rename at all. It re-pointed its own handles and dropped the affected listings, and left the bytes and any queue entry under the old key — so a renamed file, which is the rename every program actually makes (write a temporary, rename it over the real name), could send its pending save back up under the temporary's name. A fuse or winfsp mount already moved the file's own copy and gains only the subtree half; a webdav mount keeps no cache of its own and was never affected.
fuse or nfs mount enforces a lock between two nodesFrom agent 0.4.64 a Linux fuse mount forwards the locks taken on it to noBGP, so flock and fcntl there hold between nodes and not only between the processes on one machine; from agent 0.4.69 an nfs mount does the same, on Linux and macOS both — see Locks on a fuse or nfs mount below. On the remaining backends a lock your program takes is at best local to the machine holding it, so a lock that works perfectly while you test on one box buys nothing the moment a second node touches the file, which is the way this misleads hardest. The one thing that does cross nodes on a webdav mount is the WebDAV LOCK its client sends before a write, which noBGP arbitrates from router 0.4.80 — that is a different mechanism from your program's flock, and it is described further down this box.
Which of these a node lands on moved twice in as many releases (see the preference order): from 0.4.65 a Linux node with FUSE lands on a backend where a lock does hold, and a Mac spent that release on webdav — where flock is told it succeeded and enforces nothing — before 0.4.66 put it back on nfs, which refused a lock out loud then and enforces it between nodes from 0.4.69. A Mac pinned to fs: webdav stays on the backend that lies. Windows moved in agent 0.4.82, from webdav to winfsp on any machine with the driver: neither reaches a second node, but the row that now describes most Windows nodes is winfsp's.
Even where the lock does hold between nodes, keep SQLite and anything else that keeps a file open and writes in place off the share. The drive uploads a whole file when it is closed and serves reads from a local copy, so the data path is wrong for that shape of program whatever the locking says.
What each backend does, as measured:
| Mount | What a lock does |
|---|---|
nfs, agent 0.4.69+ | Enforced between processes on the machine and between nodes — noBGP arbitrates it, through the same lease a fuse mount takes. Linux and macOS both. A lock it cannot forward is refused rather than granted. See below. |
nfs, agent 0.4.68 and earlier | Not implemented, and fails loudly rather than pretending. Nothing that locks before writing is protected — and nothing is misled either. On macOS that was a straight improvement on WebDAV. |
fuse, agent 0.4.64+ | Enforced between processes on the machine and between nodes — noBGP arbitrates it. A lock it cannot forward is refused rather than granted. See below. |
fuse, agent 0.4.63 and earlier | Enforced between processes on that machine and not between nodes. Measured on two nodes holding the same file at once, 2026-08-07. |
webdav on macOS | Reports success and enforces nothing. Two processes each took an exclusive flock on one file and both were told they held it; fcntl refuses cleanly, flock is the one that lies. Measured on /Volumes/nobgp, 2026-08-07. |
webdav on Windows | An exclusive open is enforced between local processes and reported honestly — and that is the trap. It does not exclude a writer on another node, and a held handle's write-back on close overwrites theirs, producing a file that is neither version with success reported to both. Measured 2026-08-08. From router 0.4.78 network_directory flags this as the dangerous case (fs_locking.honest: false), where it had been reporting true off the local half alone. |
webdav on Linux | Unmeasured. Assume nothing is enforced and that a failure may be silent. |
winfsp — what a Windows machine with the driver mounts from agent 0.4.82 | An exclusive open is enforced between processes on the machine and reported honestly — and that is the trap: it does not exclude a writer on another node. A losing cross-node write is refused rather than blended into the other version — every write-back is conditional on the version the handle read, and the rejected bytes are kept. This platform's WebDAV mount checks the same way from agent 0.4.147, but per file rather than per open and only for what that node read, so more of it stays last-writer-wins. Measured 2026-08-11. |
On a webdav mount the operating system's own lock request does leave the machine, and from router 0.4.80 it is a real lock. Unlike flock, a WebDAV LOCK is a request on the wire: the OS client sends it to the agent's local proxy, the proxy forwards it, and noBGP answers it. Through router 0.4.79 that answer came out of one router instance's in-memory table — shared with no other instance, no other node, and with none of the leases a fuse or nfs mount takes — so two nodes asking for the same exclusive lock were both told they held it, and the lock was invisible to file with op: "lock". From router 0.4.80 a LOCK takes the same lease every other lock on the share takes: one holder at a time across every router and every node, released when the node holding it disconnects rather than only when it expires, and contending with file with op: "lock" and with a fuse or nfs mount's lease on the same path.
What that means in practice, on a webdav mount and on the share's own URL alike:
- The holder is the node, not the program. Every program behind one node's mount is one holder upstream, exactly as on a
fuseornfsmount. ALOCKon a path this node already holds is refused with423— one holder means one hold, and a second token over it would let eitherUNLOCKdrop both — while that node's own writes are still served, so a mount cannot lock itself out of a path it locked. - A write that carries no lock token is refused with
423when someone else holds the path. That is what makes the lock worth anything against a second node: aPUT,DELETE,MOVEorCOPYfrom another holder is turned away rather than quietly winning. - A lock lasts at most 10 minutes whatever the client asks for. A request with no
Timeoutheader, or one asking forInfinite, is granted 10 minutes; a shorter request is honoured exactly as sent. Through router 0.4.79 those two forms produced a lock that never expired — a cancelled copy on Windows left one file unwritable from every node and every backend until a router restarted, measured still refusing 34 minutes later, with the mount reporting nothing more useful than an I/O error. - Locking a whole directory tree is refused with
423: a lease names one exact path, so a lock covering a directory and everything under it is not offered. Lock the paths you are about to write. - ⚠ A
LOCKon a file with noDepthheader is granted, from router 0.4.96 — and through router 0.4.95 it was not. RFC 4918 makes an absentDepthmeaninfinity, and sending none is what a Windows client does when it locks an ordinary file: on a plain file such a lock covers nothing but the file itself, so it is the same lock asDepth: 0and is granted as one. Routers 0.4.80 – 0.4.95 refused every depth-infinity lock without asking what the path was, which refused the default rather than an exotic request — on a Windows node mounting overwebdavthat failed every copy onto the drive with0x80070021, since the client locks before it writes. A lock on a directory is still refused, as above. See the troubleshooting entry if you have been meeting that. - A noBGP release no longer drops the lock, from router 0.4.91. Releasing a lock the moment its node disconnects is what stops a machine that died holding a file from holding it out to the timeout — but a routine deploy moved every node's connection to another instance, and each of those momentary drops read as the node going away, so every lock every node held was released. A replica going away and a node going away are now told apart: only the second releases anything, and a node that genuinely disconnects still frees its locks immediately.
- Shared (read) locks are not offered either — a
LOCKhere is exclusive, and asking for a shared one answers501.
Through agent 0.4.81 a webdav mount could lock itself out of its own file. A LOCK's answer names the path the lock was taken on, and the node's local proxy — which is the layer that presents the share at node/ and networks/<name>/ — was not rewriting that name onto the mount's own path the way it does for a directory listing. So the client got back a lock on a path it had never asked about, could not match it to the file it was about to write, and sent the write with no lock token at all — which noBGP correctly refuses with 423 (0x80070021 ERROR_LOCK_VIOLATION on Windows). It never sent the UNLOCK either, so the lock sat out its timeout, and while it did every node and every backend was refused that path, not just the one that took it. Fixed in agent 0.4.82; nothing else about the lock changes, and the other backends never sent a LOCK at all.
From agent 0.4.79 a node also logs a warning the first time a program on that machine takes one, since nothing fails and the event leaves no other trace. That warning describes the behaviour above as it was through router 0.4.79 — it overstates the risk against a router on 0.4.80 or later — and nobgp status answers the same question in advance under fs.locking.
From agent 0.4.79 you can also read this on the box, for the backend the node actually mounted with: nobgp status reports it as fs.locking. From router 0.4.62 you can read it per node without going to the machine at all: network_directory reports it as info.fs_locking, derived from the node's backend, platform and — from router 0.4.68 — its agent version. That last part is what makes the table above readable per node: an nfs mount reads cluster from agent 0.4.69 and none below it, a fuse mount reads cluster from 0.4.64 and node below it, so a node that has not upgraded yet is described as what it actually does rather than as what your newest node does. A node whose agent version the router cannot read is described as pre-forwarding, never as cluster. A winfsp node reads node from router 0.4.71 — before that it was the one backend with no row, so the field was absent there and silence read as "no answer" rather than "no guarantee". Its row is not version-aware like the other two: the backend has nothing to forward a lock with, so node is the floor rather than a claim that improves with the agent.
If you need a lock on a backend that does not hold one, take it through the file tools rather than on the mount: file with op: "lock" locks a path on the share itself, where noBGP arbitrates it. It is the same lease a fuse mount on agent 0.4.64+ and an nfs mount on agent 0.4.69+ take, and the same one a WebDAV LOCK takes from router 0.4.80, so all of them arbitrate against each other on the same path. It is mostly advisory — it coordinates callers that also lock, and the file tools themselves do not consult it before writing — but from router 0.4.80 a write that reaches noBGP over WebDAV, which is every write a mounted drive makes and every PUT on the share's URL, is refused with 423 while someone else holds the path.
If you write over HTTPS rather than through the mount, you can also make the write itself conditional and have a stale one refused rather than accepted — see Conditional writes.
When noBGP cannot be reached
A mounted drive is a cache in front of noBGP, so every operation on it has two halves: the local one, which happens at once, and the one that has to reach noBGP. Agent 0.4.95 settles what happens to the second half when it does not arrive, and it changes three answers a Linux fuse mount used to give. Agent 0.4.96 then stops a node that already knows it is offline paying to rediscover that on every request.
A rename, mkdir or delete that noBGP refuses now fails the call
Through agent 0.4.94 a fuse mount treated noBGP said no and noBGP could not be asked as the same thing at four places — mv, mkdir, rm and rmdir — and carried on locally in both cases, reporting success to the program.
The rename is the one that cost bytes. mv a b where b collides with a name that already exists under a different case is refused by the share's case-uniqueness rule; the call returned 0 anyway, and the queued upload of a was re-pointed onto the name noBGP had just said it would never hold, where it retried every 20–80 seconds for the life of the process — absent from the mount, absent from noBGP and absent from nobgp status alike.
From agent 0.4.95 a refusal belongs to the caller:
- A refusal is noBGP judging the request — a name it will not hold, a permission that is not held, a storage allowance that is full — and the call now fails with the error that matches it: permission denied for a refused permission, file exists where the destination is already taken, no space left on device for a full allowance (so a full organization reads as a full disk in every program that has ever handled one), and an ordinary I/O error for the rest.
- A failed link is not a refusal, and never fails the call. That is the case the record below exists for.
nfsandwinfspwere never affected. Those two return every server error from those four operations already, which is why they have no offline tolerance to speak of; awebdavmount is the operating system's own client and answers for itself.
An operation noBGP could not be asked about is recorded and delivered
Proceed locally and let the write-back queue reconcile was only ever true for writes. That queue carries uploads; nothing anywhere re-sent a rename, a mkdir or a delete. So on a node whose link had blinked, mv photos archive returned 0 and then, measured at the default listing TTL: archive stayed readable by full path only until something listed its parent, ls archive came back empty from the first refetch, and the name reverted to photos within one TTL — with any file written under the new name refused for the life of the process. A local success followed by a silent revert is not tolerance.
From agent 0.4.95 a Linux fuse mount records such an operation instead of dropping it:
- It is applied locally at once, exactly as before, so nothing the program sees changes at the moment it makes the call.
- It is written into the node's cache directory and replayed after a restart, so a laptop closed on a plane still delivers what it accepted.
- It is delivered in order. Renames,
mkdirs and deletes are a diff —mkdir b; mv b c; rm c/xmeans nothing in any other order — so they are replayed strictly oldest-first, one at a time, behind the same backoff the write-back queue uses. One failed link stalls the whole backlog, which is the point: an offline node's operations wait together. - Uploads wait for the operations they would race. A queued write under a recorded operation's path is held back until that operation lands, so a save into a directory created offline follows the directory's creation instead of being refused for having no parent, and a save to a renamed file follows the rename instead of being overwritten by it.
- The mount shows the namespace you asked for, not the one noBGP still holds. Every listing and every read is resolved through the record until it lands:
ls archiveshows what noBGP holds underphotos, plus what was created offline, minus what was deleted offline, plus files whose bytes are still queued for upload. A path the record says is gone — under a pending delete, or renamed away — reads as no such file, never as an empty directory. nobgp statusshows the backlog underfs.pending, one line per operation, oldest first.
A recorded operation that noBGP then refuses is abandoned at once — the opposite of a refused upload, and deliberately. An operation has no bytes that could deliver themselves later, the namespace corrects itself at the next listing, and it sits at the head of an ordered backlog where ageing it for an hour would stall everything behind it. It is reported under fs.conflicts with a line saying the operation was not applied and that the mount now follows noBGP's name — never as a loss of bytes, because none were at stake — and the record clears when a later operation lands at that name.
⚠ Only a Linux fuse mount records. nfs and winfsp still fail those four operations outright when noBGP cannot be reached. An nfs mount that inherits a fuse session's cache directory — the Linux order is fuse then nfs — reads through that session's record but does not add to it, so a recorded name is refused there until the backlog drains.
A listing noBGP could not be asked for is never an empty directory
This is the half to upgrade for. Through agent 0.4.94 a fuse mount turned every failed directory listing into an empty directory and cached it for a full TTL — so a node that could not reach noBGP listed every directory as successfully empty, and went on doing so for up to a TTL after the link came back. ls showing nothing is an annoyance; rsync --delete, a backup computing a delta, find -delete — anything that walks the drive and acts on what it finds — reads an empty directory as the files were deleted and propagates it. noBGP held the data throughout.
From agent 0.4.95 the answer depends on what the failure actually was:
- noBGP answered that the directory is not there — or this node's own record says it was deleted or renamed away offline — and it lists as empty, cached normally, carrying whatever this mount still owns in it: a file whose bytes are queued for upload, or one behind a handle that has written to it. That is the locally-created directory noBGP has not heard of yet, and it is the only case that lists as empty.
- The link failed, or noBGP answered a
5xx, a timeout or a permission refusal — none of which is a statement about the directory — and the mount serves the listing it already holds, re-asking once per TTL for as long as the failure lasts. Last-known-good beats emptiness, and it is what keeps a shell able to resolve the directory it is sitting in across an outage. - A directory this mount has never listed has no last-known-good and nothing true to say, so it answers
Input/output error— onls, on a name lookup inside it, and onrmdir— and re-asks every 5 seconds rather than once per TTL, because what it is serving meanwhile is an error and has to stop the moment noBGP is back. Nothing can be created in such a directory until it answers.
⚠ Being refused is not an answer about the directory. A permission refusal is noBGP declining to be asked, so it takes the second arm above rather than listing as empty. Only a genuine no such directory lists as empty.
A node that already knows it is offline stops rediscovering it
From agent 0.4.96 a request the mount makes is short-circuited once the node's control channel has been down long enough that the outcome is not in doubt, instead of waiting out a network timeout to learn the same thing again. On a node measured with all outbound HTTPS dropped, an mv issued 150 seconds into the outage — long past every keepalive deadline, on an agent that had already given up reconnecting — still spent 36 seconds rediscovering that before completing locally; earlier in the outage it spent 60. That whole time was rediscovery, once per call, for as long as the outage lasted.
Three things are worth being precise about:
- It changes how long you wait, never what you get. Every request it declines to attempt would have failed anyway, and it fails in exactly the shape the offline machinery already handles — the record, the upload backlog and the listing rules above all engage unchanged. The error says the router has been unreachable, for how long, and that the operation will be retried when the node reconnects.
- It waits for the link to be provably gone. The control channel is only marked down a full keepalive window after the last successful exchange, the agent then fails several reconnect attempts, and this waits one more window on top. On a node connected over
wssthat is roughly two minutes of real unreachability before anything is short-circuited. ⚠ On a node connected over QUIC it is roughly three minutes from agent 0.4.97, because that release doubled the QUIC idle timeout to stop healthy connections being closed under load, and the first of those two windows is the transport's. That is longer waiting, never a different answer — the requests in that stretch would have failed either way.router.transportsays which window a given node is on. - It does not help at the start of an outage. A request in the first minute or so is issued before the link has been observed to be down, and still pays its own timeout. Narrowing that would mean guessing about a connection nothing has yet seen fail.
The residual: a router that is refusing control connections while its storage endpoint still serves would put the drive on this path for as long as that lasts. Nothing is lost — that is what the record is for — but reads block where they would have been answered.
Locks on a fuse or nfs mount
From agent 0.4.64 a Linux fuse mount stops letting the kernel arbitrate locks by itself and hands them to noBGP instead, which is what makes them mean something on a second node. From agent 0.4.69 an nfs mount does the same, on Linux and macOS alike. flock and fcntl byte-range locks both go through it; nothing in a program has to change.
The two backends share one lease manager inside the agent deliberately, so they cannot drift into two renewal cadences or two ideas of when a lock has been lost. Everything below holds on both unless a bullet says otherwise.
Two layers do the work, and they are deliberately different shapes:
- On the machine, the agent keeps a real byte-range table — per path, per lock owner, shared and exclusive distinguished — so everything the kernel used to do locally still holds, including disjoint ranges and shared read locks.
- Between nodes, noBGP holds one whole-file exclusive lease per path, taken while this node holds any lock on the file and dropped when it holds none. That is coarser than what the program asked for, always in the safe direction: a shared read lock excludes a remote reader it need not have, and two disjoint ranges on two machines contend. It never grants where a local lock would refuse.
Things worth knowing before relying on it:
- The node is the holder, not the process. noBGP names the lease from the node's connection, so every process on one machine is one holder upstream, and re-taking a path this node already holds renews rather than contends. It is the local table that keeps two processes on the same box apart.
- The lease is the same one the file tools take. A path locked on a mount is contended for a caller using
filewithop: "lock", and the other way round. - A blocking
fcntl(F_SETLKW)behaves differently on the two backends. Onfuseit waits up to 30 seconds and then returnsEINTR, which means "ask again" — a caller that wants to keep waiting reissues the call, and the agent polls with a backoff on your behalf. Onnfsthe kernel's own client does the retrying, so a blocking wait simply blocks. noBGP holds no queue of waiters either way, and a signal interrupts the wait immediately, as it would on a local filesystem. F_GETLKasks noBGP too, so it can report a file another node holds. A remote holder is described as a whole-file exclusive lock belonging to no process on your machine — onfusethat is spelled pid 0, and onnfsthe holder carries no client identity at all. The lease has no range and no process behind it, so neither answer can be mistaken for a local one.- A lock that cannot be forwarded is refused, never granted. On
fusethe refusal isENOLCK("no locks available") rather thanEAGAIN; onnfsa contended path is denied outright, a momentary rate limit is retried by the kernel client without your program seeing it, and anything permanent surfaces as an I/O error with the agent logging the path. A mount never grants a lock it cannot enforce, so a program that checks its return value stops rather than proceeding unprotected. - A lease that is lost surfaces on the next write. The agent renews every 20 seconds; if renewal is refused because another node took the path, or fails for longer than a minute, the lock is gone — and since neither mount has a way to announce it,
write,fsyncandcloseon the handles that held it answerEIO. Reads keep working. The agent logs the file it happened to. - Locks are released when the file is closed and when the drive unmounts, so a process that exits holding one does not leave the path locked for the rest of its two-minute expiry.
- On
nfs, a released lease lingers about five seconds, from agent 0.4.80. When the last local lock on a path goes away, the lease behind it is dropped a few seconds later rather than at once. Both halves of that are real. What it buys: a program that takes and drops the same lock in a loop — a build, an indexer, SQLite, anything using a lockfile as a mutex, which is most of what locks at all — used to pay a round trip to noBGP and back on every iteration, and now pays one per linger however tight the loop is, because re-locking the same path inside the window keeps the lease it already has. What it costs: another node asking for that path inside the window is refused where it would just have been granted — a blocking waiter gets it a moment later, a non-blocking one sees the refusal. Renewal keeps running while a lease lingers, so it never quietly expires underneath you, andF_GETLK, which takes and drops a lease to answer, still drops it immediately. Afusemount releases as it always did, at once.
Reaching a share directly
Every network's share also answers over HTTPS as a WebDAV endpoint, keyed on the network's id (router 0.4.54+):
https://files.nobgp.com/networks/<network-id>/
PROPFIND lists a directory, GET downloads a file, PUT uploads one, MKCOL, MOVE, COPY and DELETE manage the tree, and LOCK / UNLOCK take and release an exclusive lock on a path. Authenticate with your noBGP bearer token; the URL carries exactly the access your account already has to that network, so there is nothing extra to grant. You never have to assemble it by hand — ask your AI assistant for the network directory and each network comes back with its own URL as files_url.
It is keyed on the network id rather than its name deliberately: the same name can exist in more than one organization you belong to, so a name would not say which share you meant.
A DELETE answers as soon as the name is gone (router 0.4.77). The entry is moved out of its directory in one step and its bytes are reclaimed in the background, so removing a large file or a directory full of them no longer holds the request open while the underlying storage works through it — and a directory costs the same one step whatever is inside it. What you get back is unchanged in every way that matters: the moment the 204 arrives the path is gone for every client, every node's mounted drive and every router, the name is immediately free to create again, and metering has already stopped counting those bytes. This is the same change behind fs_delete on the storage trees. Deleting a directory is still one request per file for clients that walk the tree themselves — that is the client's doing, not the share's.
A LOCK here is the same lock everything else on the share takes (router 0.4.80). It is exclusive, held across every router and every node, and contends with file with op: "lock" and with a fuse or nfs mount's lease on the same path — so a client of your own, a person's assistant and a program on a node's drive can no longer each believe they hold the same file. It lasts at most 10 minutes: ask for less and you get what you asked for, ask for Infinite or send no Timeout header at all and you get 10 minutes rather than a lock that never expires. A lock covering a directory and everything under it is refused with 423, and shared locks with 501; a LOCK on a file that carries no Depth header — which RFC 4918 reads as Depth: infinity, and which is what a Windows client sends — is granted from router 0.4.96, because on a plain file it covers nothing the depth-0 lock does not. Routers 0.4.80 – 0.4.95 refused it, which is what broke Windows copies onto a webdav drive. A write that carries no lock token is refused with 423 while another holder has the path. Everything in what a lock on a webdav mount is worth applies here too — it is the same endpoint.
Conditional writes
Every file on the share carries an ETag, and from router 0.4.63 you can make a PUT conditional on it so a write that would clobber someone else's is refused instead of silently winning.
# Read the file and its validator
curl -si -H "Authorization: Bearer $NOBGP_TOKEN" \
https://files.nobgp.com/networks/<network-id>/notes.md
# ...
# ETag: "1f4a2-17c9d3e8b2a4c00-3e8"
# Write it back only if nobody changed it in the meantime
curl -X PUT -H "Authorization: Bearer $NOBGP_TOKEN" \
-H 'If-Match: "1f4a2-17c9d3e8b2a4c00-3e8"' \
--data-binary @notes.md \
https://files.nobgp.com/networks/<network-id>/notes.md
- The same
ETagcomes back fromGET,HEAD,PROPFINDand thePUTresponse, so whichever way you read a file, the validator you hold is the one a laterPUTis compared against. It is an opaque string — never parse it — and it changes whenever the file's bytes change. If-Match: "…"refuses a stale write with412 Precondition Failed. Read the file again, merge, and retry.If-Match: *requires the file to exist — useful when you mean to replace something and not to create it.If-None-Match: *is an exclusive create: thePUTsucceeds only if nothing is there, and comes back412if there is.- A
PUTthat sends neither header behaves exactly as it always has, and is accepted whatever the file looks like now. Conditional writing is something a client of your own opts into. - The condition is checked twice, and the second check is the one that decides (router 0.4.76). It is read once when the request arrives, which refuses a doomed write before you spend the upload on it, and again immediately before the upload is published. An upload takes as long as its body — seconds on a large file — and through router 0.4.75 anything that landed inside that window was overwritten anyway, with success reported to both writers. That upload is refused with
412now, andIf-None-Match: *claims the path against other uploads in flight rather than only against the requests that had already finished. - It narrows the window rather than closing it. What is left is the gap between that last check and the moment the file becomes visible — microseconds rather than the length of a body. Closing it entirely would need a form of rename the operating system does not offer.
This is not the WebDAV If: header, which carries lock tokens; these are the ordinary HTTP conditional headers.
The mounted drive uses it too
From agent 0.4.62 the mounted drive conditions its own write-backs the same way, so the lost update above is refused on the mount rather than only in a client you write yourself. It needs a router on 0.4.63 or later; against an older one the write goes out unconditional exactly as it did before. The winfsp backend joins from agent 0.4.63, which is the release where it started writing at all.
The shape it protects against is the everyday one: the mount uploads a whole file when it is closed, so an editor that keeps a file open and saves its whole cached copy each time would otherwise destroy anything another node wrote in between, with success reported to both.
nfs— from agent 0.4.63 the refusal reaches the program on itswrite, and that is the fix that makes the rest of this true on Linux. Through 0.4.62 the bytes were correctly refused but the error was raised atfsyncorclose, and the NFS server library the agent used then discarded the error from both — so the losing writer was told its save had landed after all. The handle now re-checks the version on its way to its first write (one extra round trip per handle, and only for a handle that read the file and is now writing it), latches the refusal, and hands it back from every laterwriteon that handle. It narrows the window rather than closing it: a write that lands in the milliseconds between that check and the upload is still refused by the router, and that refusal still arrives where nothing can carry it. The file stays unsaved on the node, and its unsaved bytes are kept, not deleted — set aside next to the agent's cache, with the log naming the path (agent 0.4.63; before that a refusednfswrite-back discarded its local copy). From agent 0.4.69 a refusal also reachesfsync, which the replaced server is what makes possible — until then only a laterwriteon the same handle could tell you. From agent 0.4.70 the copy is set aside in the same place every other backend uses — beside the cached file, with.rejectedon the end — so hunting for a rejected save needs no knowledge of which backend the node mounted with.fuse— the upload is asynchronous, sowritehas usually already returned by the time the router refuses. The nextfsyncorcloseon that file gets a hard write error instead of another silent success, the mount goes back to showing the router's version, and the rejected copy is kept, not deleted — the agent log names the file it was set aside as.webdav— from agent 0.4.147 the agent's local proxy supplies the missing version itself. The operating system's own WebDAV client reads a file whole when it is opened, writes its whole cached copy back when it is closed, and names no version — so the proxy remembers the version of each file whose bytes it handed the client, and conditions that file's next save on it. A save the router refuses comes back to the client as412 Precondition Failed, the other node's version stays on the share, and your bytes are kept beside the node's cache with the usual note, listed bynobgp statusunderfs.conflicts. Open the file again, merge, and save: that save goes through and the kept copy is removed. Through agent 0.4.146 the proxy forwarded anIf-Matchits caller sent and never added one, so two nodes writing one file was last-writer-wins with neither told. ⚠ It covers less than the other backends do — see what thewebdavcheck does not cover.winfsp— from agent 0.4.63, when the backend writes at all, and in practice from 0.4.65, when a real mount started accepting writes. Through 0.4.66 the write-back happened inside theCloseHandlethat triggered it, so a refused save failed that call. From agent 0.4.67 the upload is asynchronous likefuse's, so the close has usually already reported success by the time the router refuses: the refusal is latched onto the handle and handed back from the next flush,fsyncor close on that file, and the unsaved bytes are set aside beside the cache with the log naming the path. In 0.4.62 the backend was read-only and nothing was written from the mount.
A queued upload survives the agent restarting, from agent 0.4.75. On a fuse or winfsp mount the upload is deferred by a few seconds so several quick saves cost one transfer — and through agent 0.4.74 a restart inside that window dropped the save, silently, after the program that made it had already been told it succeeded. Restarts are routine: an auto-upgrade or a config reload is enough. It got worse from there, because the node's cache dates a copy it adopts from a previous process by the file's modification time, so an hour-old unsaved save was already past the cache's one-hour age limit the moment the agent came back — and the cleaner deleted the only copy of it.
The queue now records each pending upload on disk beside the bytes, at the moment the write is queued rather than at shutdown, so a kill -9 loses no more than an orderly stop does. The next start re-queues them and logs how many. Until the bytes are upstream the copy holding them is exempt from the cleaner and from the cache's size limit — an unsaved save is never what gets evicted to make room, and the limit is exceeded and reported instead. A resumed upload goes out unconditionally, like any copy adopted from a previous process. An nfs mount uploads on close rather than queueing, so it was never exposed to this.
Deleting a file whose upload is still queued now reaches noBGP too, from agent 0.4.154. Through agent 0.4.153 a Linux fuse mount skipped the delete whenever an upload for that path was queued, on the reasoning that noBGP had never been told about the file — which is only true of a file created on this node and never sent. An overwrite queues an upload for a file noBGP already holds, and so does a rename the queued upload followed, so rm after either one removed the local copy and left the earlier version on the share. shred -u is exactly that sequence — overwrite, rename, unlink — and it left the file readable through noBGP with nothing on the machine to show it. The delete is now skipped only where all of these hold: the copy is a create noBGP has never held, its upload is queued under that same name, none is in flight, and no recorded operation covers the name. That is the .git lock file written and removed inside the write-back window, and nothing else. A delete noBGP refuses still fails the call with the file present on the mount and its queued bytes still queued, so a refusal cannot lose the last write.
Some writes still go out unconditionally on every backend, because there is no version to demand: a file opened with O_TRUNC (the caller named an existing file to replace), a file cached before the agent last restarted (the validators do not survive a restart), and — on nfs from agent 0.4.65, where the contents are only fetched when something asks for them — a handle that reached its first write without ever having read the file. It is one rule in three shapes: a caller that read nothing cannot be overwriting anything it read.
None of this makes a mount safe for two nodes editing one file — it turns a silent loss into a visible error. For coordination that actually holds between nodes, lock the path through the file tools, or lock it on the mount itself from a Linux fuse node on agent 0.4.64+ or an nfs node on agent 0.4.69+, both of which take the same lease.
What the webdav check does not cover
The other backends hold a version per open handle, because the agent is the filesystem. On a webdav mount the filesystem is the operating system's own client and the agent only sees the requests it makes, so the check the proxy adds from agent 0.4.147 is narrower. Every gap below leaves that save exactly as it was before — unconditional, last-writer-wins — and never turns into a refusal of a write that raced nobody:
- Only files this node read through the agent since the agent last started. A file the client had already cached, and a file this node is creating, have no version to condition on.
- One version per file, not one per program. Every program on the machine reaches the share through the one OS client, and nothing in a request says which open it belongs to — so any later read moves the version forward. If a preview pane, an antivirus scan or an indexer reads another node's newer version while your editor still holds the older one, that editor's save is conditioned on the newer version, passes, and replaces it.
- A save made by writing a temporary file and renaming it over the document is not checked. The condition would name the temporary file, not the document. The version does follow the rename, so a later in-place save of that name is checked again.
- A weak validator is never used.
If-Matchcompares strongly, so conditioning on one would refuse every save of that file for ever. - Your own client's condition wins.
davfs2sends its ownIf-Matchand handles its own412; the proxy adds nothing to a request that already carries a precondition. - ⚠ A
412is not proof of another writer when the client also sent anIf:header — that header's lock condition can fail it too. The log line and thefs.conflictsentry say so rather than naming a change nobody made. - ⚠ Only a complete save is kept. If the client stops sending, or the node cannot stage the bytes, the save is still refused — the other writer's acknowledged bytes are never destroyed — but nothing is set aside and the log says so.
- ⚠ What the Windows and macOS clients show you when a save is refused has not been measured.
What a node could not save
A refusal above that got as far as accepting your bytes keeps them, and through agent 0.4.79 finding them again was the difficulty: the kept copy sits beside the node's cache under a name derived from a hash, so the file it belongs to was named exactly once, in the log line the refusal wrote. Nothing counted those copies against anything, and nothing ever removed one. Agent 0.4.80 gives them a lifecycle:
nobgp statuslists them. Thefsblock gains aconflictsline per unresolved refusal, in path order: the file, where its unsaved bytes are, when the write was refused, and why. It is absent on a mount holding nothing. The index is rebuilt from disk at startup, so it survives a restart or an upgrade. Awebdavmount joins from agent 0.4.147, when its proxy started keeping a refused save: that mount keeps no cache, so the list is read from the notes the proxy wrote beside the kept bytes. Below 0.4.147 awebdavmount refused nothing and so listed nothing.- Each kept copy carries a note naming its document, written beside the bytes — the file it came from, why it was refused and when — so a cache directory found filling up a week later reads back to files and reasons without hunting through old logs.
- They are counted but never evicted. Refused bytes are exempt from the cache cleaner, because they may be the only copy of your document — so a node that keeps losing races could grow its cache directory past its 1 GiB limit while the amount the agent was accounting for stayed comfortably under it. The agent now says so in its log, naming how much of the directory is unsaved bytes, and still deletes none of it.
- One thing reclaims a kept copy: a later write of the same file reaching noBGP. There is no timer, deliberately — a lifetime on a hidden file destroys the only copy of somebody's document on a schedule nobody is watching. What makes the deletion honest is that the program has had its whole chance (refused, logged, listed) and this node has since written that path successfully. It is not a claim that the write which landed contains the refused bytes.
- A refusal stands until then rather than being consumed by whoever reads it first. For the handle whose copy was refused that means the error is reported on every later flush,
fsyncand close rather than only on the first — that handle genuinely cannot save, because its bytes have been moved aside. A handle opened after the refusal holds the version that superseded it and is not refused for it. - A refusal that kept nothing is recorded too, and says so. On an
nfsmount the check a handle makes on its way to its first write refuses before any of your bytes have been accepted, so there is nothing to set aside — and nothing was lost, since thewritecall itself failed and the data is still your program's. The status line spells that out rather than describing a loss that did not happen, which is what the backend that refuses soonest deserves. - An upload noBGP keeps refusing stops being invisible, from agent 0.4.95. A refused write-back is retried on its own schedule, and for the common case that is exactly right — remove the colliding name and the upload delivers itself intact, with nobody ever learning anything went wrong. Past two minutes of refusals the silence is the defect: six such uploads were measured piling up on one node, each retrying every 20–80 seconds indefinitely, while every surface said the file simply did not exist. Such an upload now gets a
conflictsline while it is still queued and still being retried, and the line says exactly that — remove the cause and it delivers itself — rather than pointing at a set-aside copy that does not exist. The next flush,fsyncor close on that file reports it too. - After an hour of refusals it stops retrying, and its bytes are set aside beside the cache with the usual note naming the document, instead of cycling against a wall for the life of the process. ⚠ Both clocks count refusals only, never a failed link. A laptop off the network is the case the queue exists for, and ageing its backlog into set-aside copies — where nothing ever delivers itself again — would be far worse than the unbounded retry the ceiling exists to bound. The hour is measured within one run of the agent: a restart re-queues the upload with that clock at zero, while the two-minute record survives, so such a node is never silent — only unbounded. Both apply where the upload is queued, so on a
fuseorwinfspmount; annfsmount uploads on close. - On an
nfsmount the refusal now outlives the file being closed (agent 0.4.82), and it had to. A refused write-back does not clean the page the kernel was refused from, so the client comes back and sends those same bytes again after the close — and through agent 0.4.81 the refusal went away with the handle, so the retry was rebased onto the version that had won and accepted. The result was the winner's file destroyed silently, at the right length with the wrong contents. The refusal now sticks to the file, and lifts for the two things that cannot be carrying the refused bytes: a read, which is a program looking at what noBGP actually holds, and a truncate to zero, which replaces the whole content — and only while nothing else has the file open, so an unrelated reader (Spotlight, the Windows indexer,grep -r) cannot lift it on a writer's behalf.fuseandwinfspwere not affected: their refusal is latched onto the handle, and their write-back is queued rather than being retried by the kernel.
Two nodes creating one file
A create used to be on that list of unconditional writes, and from agent 0.4.70 it is not. The old reasoning was the same rule — a caller that read nothing cannot be overwriting anything it read — and it is wrong for exactly one shape, measured on two nodes on 2026-08-11: both create the same path inside the write-back window, neither read anything because nothing existed in either one's view, and the loser silently destroyed bytes the winner had already been told were saved.
A create now claims the path exclusively (If-None-Match: *, the same conditional write a client of your own can make), so the second one loses out loud instead of quietly winning:
- On a
fuseorwinfspmount, and for any file the node's cache is managing, the refusal takes the same path a stale version does: the copy is set aside next to the cache with the log naming it, the refusal is latched onto the handle, and the nextfsync, flush or close reports it rather than another silent success. Onwinfspa create that loses answersEEXIST, which Windows turns into whatever the caller asked for next — an overwrite-minded program retries as an open-and-truncate, which is a different question and is answered normally. - On an
nfsmount the placeholder a create publishes is exclusive too.O_EXCL— the guarded create — now fails across nodes rather than only against the processes on this machine, which is whatO_EXCLis for; an ordinaryopen(O_CREAT)that finds the peer got there first simply opens the peer's file, which is what that call means on a path that exists. - Renaming a freshly created file keeps the claim with it, so the write-a-temp-then-rename pattern most editors use is covered.
It needs a router on 0.4.63 or later, the release that honours the header. Against an older one the create degrades to the unconditional write it was before, silently — the same compatibility rule the version check itself follows. Router 0.4.76 is what makes the claim hold between two uploads in flight: before it, the exclusive claim was settled when the request arrived, so two nodes that both started a create while neither had published still ended with one overwriting the other. The claim is now re-checked as the upload is published, which is the moment it has to be true. One narrow case stays uncovered: if something else on the same node reads the path in the window between the create and its upload, the write is conditioned on the version that read found rather than on the path being empty, so it composes with the peer's file instead of refusing. That is still strictly better than the clobber it replaced.
Each node also has a storage area of its own, separate from every network's share:
https://files.nobgp.com/nodes/<node-id>/
It follows the machine rather than a network — a node that joins several networks still has exactly one area, belonging to none of them — and it is not a network's shared drive. On the machine itself it is the node/ folder on the mount (agent 0.4.54+, see What the mount contains), which needs no URL and no token. Reaching it over HTTPS needs the Owner or Admin role in the organization that owns the node, a step above the network share, because nothing the node's owner sets on the box filters writes that arrive this way. A node always reaches its own area; from router 0.4.73 it reaches a peer's only when it holds the manage tier and the peer is in its own network.
It outlives the node, for a month. Deleting a node does not delete its storage there and then — the area and everything in it stay, reachable by the URL above, so a machine that is rebuilt or replaced does not silently take its files with it. From router 0.4.73 that grace has an end: 30 days after the node is deleted the area is queued for deletion, and the bytes go about a week after that — so nothing is removed sooner than roughly five weeks from the deletion, and after that the area is gone for good. Copy out anything you want to keep inside that window; delete the contents yourself sooner if you are done with them. A deleted node's area stops counting toward your storage limit from the moment the node goes, so the retained files never hold space against your allowance.
Bringing the node back inside the window calls the deletion off, and the area comes back with the node's identity, contents intact. Past the window the identity, labels and role grants still return and the area is empty.
A deleted network's shared drive is swept on the same clock — 30 days, then queued — but the resemblance ends there, and the difference is what to plan around. A network has no way back once deleted, and neither does its drive: https://files.nobgp.com/networks/<network-id>/ answers 404 from the moment the network goes, for everyone, so there is nothing to copy out during the retention window. Unlike a node's area, that window is bookkeeping rather than a grace period — copy out what you want to keep before you delete the network.
A node's area counts toward your storage allowance while the node exists (router 0.4.61+), against the organization that owns the node and pooled with that organization's network shares — see Storage. Storage is a plan limit, not a metered charge: reaching it stops new bytes going in, and never produces a bill.
⚠ A node that is rebuilt from scratch gets a NEW area, not its old one. The area follows the node's identity, which survives a reinstall that re-uses the existing registration but not a fresh enrolment — so a replacement machine starts empty, and the predecessor's files remain where they were, reachable by the URL above.
Your AI assistant can reach both without a WebDAV client (router 0.4.56+). The same file tools it uses on a node's own disk now address these two places directly: name a network and no node for its shared drive, or a node with storage: true for that node's area. The router answers from its own storage, so nothing has to be online — you can read a machine's files while it is switched off, and put files there for it to find when it comes back. Each call moves at most 1 MiB: a larger file can be read a page at a time (router 0.4.59+, in 32 KiB pages by default from 0.4.84), but writing one still means a PUT to the URLs above. Details in the MCP reference.
⚠ To move a file from one machine to another, the shared drive is not the shortest way round (router 0.4.84+). Ask for it directly — "copy /tmp/build.tar.gz from pi5 to cm4" — and fs_copy streams the bytes between the two ends in a single call: neither machine needs a route to the other, nothing is stood up or cleaned up, and the file never passes through the conversation. From router 0.4.170 a whole directory goes the same way — name the files, or ask for every file under a directory, in one call. Staging on a shared drive is for bytes that several machines will collect, or that should be there before the target comes back.
Using File Sharing with Your AI Assistant
Your AI assistant can:
- Point you to your network's files in the dashboard (Networks → your network → Files)
- List files and directories currently stored
- Guide you on uploading or downloading files
- Help you organize and manage your shared files
The nobgp file CLI commands (upload, download, ls, rm) let you transfer files directly from any node. Your AI assistant can run these on your behalf using the command tool. You can also manage files in the web dashboard under Networks → your network → Files.
Use Cases
- Share configuration files across nodes - Upload once, access from all nodes
- Distribute deployment artifacts - Share builds or packages with your infrastructure
- Collect logs from multiple machines - Centralized storage for log files
- Transfer files between disconnected systems - Bridge air-gapped environments
Example Workflow
You: I need to share a config file with all my production servers
AI: I can help you with that. You can upload your config file to the shared drive
for your production network.
Open app.nobgp.com, go to Networks, choose "production", then Files.
Once uploaded, the file will be available at /mnt/nobgp/networks/production/
on all nodes in your production network.
You: How do I access it from my servers?
AI: On any node in the production network, the file will be at:
/mnt/nobgp/networks/production/your-config-file.conf
You can copy it to the appropriate location, for example:
cp /mnt/nobgp/networks/production/your-config-file.conf /etc/app/config.conf
Web Dashboard
The noBGP web dashboard provides a graphical interface for managing your account:
- Network management: Create and configure networks
- Node overview: See all connected nodes and their status
- Registration keys: Generate and manage keys for automated agent enrollment
- Service management: View and configure published services
- File sharing: Upload and download a network's files under Networks → your network → Files
The web dashboard complements your AI assistant — use it for account management tasks and visual overviews, while using the AI assistant for operational tasks like running commands and troubleshooting.
Putting It Together
Here's how these concepts work together in a typical workflow:
The flow:
- You authenticate to noBGP via your AI assistant (OAuth)
- Your request is authorized based on your account permissions
- Operations execute within your networks (isolated from other users)
- Nodes authenticate via registration (OAuth or key) and then JWT tokens
- All communication is encrypted and secured
- Results stream back to your AI assistant in real-time
Next Steps
Now that you understand the core concepts:
- Try Your First Steps - Hands-on walkthrough of common tasks
- Install an Agent - Connect your first machine
- Provisioning Guide - Create nodes on-demand
- Service Publishing Guide - Expose your applications
- Monitoring & Events - Subscribe to file changes, fleet-wide commands, and node presence
Questions?
Q: How many networks can I have? A: As many as your plan allows — see Plans & Billing. Create separate networks for different environments (dev, staging, prod).
Q: Can nodes be in multiple networks? A: Yes! A node can join multiple networks simultaneously using multiple agent profiles (see Multi-Profile Support)
Q: What happens if I regenerate my registration key? A: Already-registered nodes are unaffected — they use JWT tokens. Only new registrations using the old key will stop working.
Q: Are services publicly accessible? A: Only if you make them public. By default, all services require OAuth authentication.
Q: How long do sessions stay alive? A: Sessions have a 1-hour idle timeout. After the command exits, there is a 30-second grace period to retrieve the final exit code.