Skip to content

Case study

Remote access without a hosted control plane

Design and implementation · Multi-site overlay, roaming clients, self-hosted control plane
WireGuard Overlay mesh VPN Self-hosted control plane Subnet routing Split-horizon DNS IPsec
Problem

Remote administration usually costs either a management port published to the internet or a hosted control plane that can change its terms, and the networks you work from fight the tunnel either way.

Solution

An encrypted overlay with a control plane on my own hardware, one subnet router into the LAN, split-horizon DNS inside the tunnel, and direct, port-mapped then relayed connections tried in that sequence.

Result

Administration that survives a hostile internet with no management port published anywhere, a control plane nobody else operates, peers added as routine, and one load-bearing node named rather than discovered.

How I found it

A name that resolved correctly on the LAN came back with the wrong address over the tunnel, so remote clients reached the public edge and were refused by the address filter instead of reaching the service. Querying each authoritative server directly showed the two of them disagreed about one record — and that the tunnel clients were using neither, because the control plane was pushing a public resolver and overriding the local one. Two causes in two places, which is exactly why it looked intermittent from the client end.

An estate with no management port published to the internet, administered entirely over an encrypted overlay, on a control plane nobody else operates. It is my own infrastructure, described separately in the self-hosted production infrastructure case study, which is why the decisions behind it can be set out here in full rather than in general terms.

The problem

Remote administration is normally bought in one of two ways, and both have a price. Publishing a management port — RDP, SSH, a hypervisor console, a firewall’s web interface — means being scanned continuously by people who are patient, automated and uninterested in how small you are. Buying a hosted overlay instead means peer coordination depends on a third party that can change its terms, its price or its availability, and that holds a list of everything you own.

Then there is the network you happen to be sitting on when something breaks, which is rarely one you control. Symmetric and carrier-grade NAT, proxies with deep packet inspection, captive portals, IPv6-only segments: the tunnel has to be established outward, from somewhere restrictive, to somewhere that publishes nothing. And the far end contains switches, printers and appliances that will never run a client of any kind.

A remote-access and site-to-site overlay with a control plane the client ownsRoaming clients and sites reaching one estate over encrypted peer-to-peer tunnels: a self-hosted control plane, a single subnet router into the LAN, and split-horizon DNS inside the tunnel.
Roaming clients and sites reaching one estate over encrypted peer-to-peer tunnels: a self-hosted control plane, a single subnet router into the LAN, and split-horizon DNS inside the tunnel.

What I did

  • A control plane on my own hardware. Peers coordinate through a server inside the estate rather than a service somebody else runs, so the membership list and the ability to admit a new device stay in one place, with one operator, and cannot be repriced or withdrawn.
  • One subnet router instead of an agent on every host. A single node advertises a route into the whole LAN, so the equipment that will never run a client is reachable over the overlay without installing anything on it or renumbering anything around it.
  • Split-horizon DNS inside the tunnel. A name resolves to the internal address over the overlay and to the public one outside it. That removes the hairpin consumer routers handle badly, and it keeps internal structure private from anybody resolving the name from outside.
  • A small public edge server, joined back by an encrypted tunnel. Services are published from the edge, so nothing in the building is directly reachable. The list of published ports is deliberately short, and everything not on it is refused there rather than inside.
  • IPsec or WireGuard chosen on evidence, not habit. WireGuard where both ends are mine and a small, fast, readable configuration is worth more than options; IPsec where the far end is somebody else’s firewall and interoperability decides the argument before performance gets a vote.
  • A ladder for hostile networks, tried in sequence. A direct connection first; then a port mapping asked of the router; and only then a relay that routes by identifier and cannot read the session. In that order rather than all at once, so a direct path is used whenever there is one and the relay stays the last resort rather than the default.
  • Keys and peers as routine, not as events. Each key is generated on the device that keeps it and never travels. Peers are added and removed as ordinary operations, so nobody is tempted to leave an old one in place. The configuration is backed up in a way that has been restored from.

What it costs

Two trade-offs are worth stating plainly, because a design presented without any is a design nobody has run.

The subnet router is load-bearing. Giving the overlay one route into a whole LAN means one node carries remote access to everything behind it, and restarting that node cuts the path you would use to fix whatever made you restart it. So it is treated accordingly: changed in a planned window, worked on from its own console rather than through the tunnel it provides, and named as a dependency in the documentation instead of being discovered as one during an incident.

Hosting the control plane yourself means nobody can change its terms, and equally that nobody else will restore it at three in the morning. That is the trade. It is why the rebuild path for it is written down and rehearsed rather than assumed, and why its configuration is part of the same backup discipline as everything else.

The result

Remote administration that survives the internet being hostile. No management port is published anywhere, so the scanning that never stops has nothing to find. The control plane is operated by nobody but me. Names resolve correctly from inside the tunnel and from outside it, so one set of documentation works in both places. Equipment that cannot run a client is still reachable through the subnet router, and adding or removing a peer is a routine change rather than a project.

The single load-bearing node is known, documented and handled as such, and the path to rebuild the overlay and its control plane exists in writing. That is the honest version of the outcome: not that nothing can fail, but that what would hurt has been identified, written down, and made recoverable.

Leave a response

Every comment is read before it appears. Yours will not show up straight away, and that is not a fault.

Not published, and not used for anything else.