Skip to content

Case study

Self-hosted production infrastructure

Problem

Production mail, DNS and web have to be reachable from the internet without anything in the building being directly reachable, trusted by large receivers, and recoverable after the machine they run on is gone.

Solution

One hypervisor of single-purpose containers behind a small public edge server joined by an encrypted tunnel, split-horizon DNS, certificates renewed with the reload tested, detection on every host and blocking at the edge.

Result

This site, its mail and its DNS run on that estate in production, nine ports are published and the rest refused at the edge, and two documented scripts rebuild the estate from bare metal.

How I found it

The symptom was a container with a working LAN and no route out, on an address every check said was free: nothing answered a ping, nothing was in the ARP table, nothing held a lease. The cause was the site router, still mapping that address to another device's MAC from an old reservation. I found it by watching the gateway's own traffic with tcpdump and reading the destination MAC on the frames coming back — they were addressed to a machine that was not this one. Every new guest here is verified that way now, before anything is built on it.

This website, its mail, its DNS and its tools all run on infrastructure I built and operate myself. It is not a demonstration environment. It carries real mail, real traffic and real consequences when something breaks, which is what makes it worth describing.

The problem

Running production services out of a building means solving three separate problems at the same time. The services have to be reachable from the internet without anything in the building being directly reachable. Mail has to be trusted by large receivers, which is a different job from making mail work. And the whole estate has to come back after the machine it runs on is gone, which is the part that is usually assumed rather than tested.

The alternative is to rent all three from somebody else. That is a reasonable answer for many organisations, and a poor one for anybody who has to know how the thing works when it breaks at an inconvenient hour. The cost of getting it wrong is concrete: an exposed management interface, mail that is silently discarded by the receivers that matter, or a backup that turns out to be a hypothesis.

The self-hosted estate, from the public edge to the containers behind itA small public server joined back by an encrypted tunnel, a hypervisor of single-purpose containers behind it, detection on every host and blocking where blocking is cheap.
A small public server joined back by an encrypted tunnel, a hypervisor of single-purpose containers behind it, detection on every host and blocking where blocking is cheap.

What I did

A single hypervisor running seven isolated containers, each with one job: a reverse proxy terminating TLS for every published service, two authoritative DNS servers, a full mail stack, a control server for the private overlay network, the web server this page is served from, and a local agent holding a copy of the documentation and the recovery kit.

In front of it sits a small public server connected back over an encrypted tunnel. Nothing in the building is directly reachable from the internet. Nine ports are published deliberately; everything else is refused at the edge.

Split-horizon DNS. Internal clients resolve services to internal addresses, external clients to the public one. It removes a hairpin that consumer routers handle badly, and it keeps internal structure private.

Certificate automation. Every service presents a valid certificate, renewed automatically, including the ones that never appear in a browser — mail transport, submission and IMAP. Renewal that does not also reload the service is the most common way this fails, so the reload is part of the process and is tested.

Mail that is actually delivered. Running your own mail server is straightforward; being trusted by large receivers is not. Forward and reverse DNS that agree, SPF, DKIM and DMARC with reports that are read, MTA-STS with TLS reporting, and monitoring for queue depth and blacklist status.

Intrusion detection with enforcement at the edge. Detection runs on every host — hypervisor, every container and the public server — reporting to a single decision store. Blocking happens on the public server, before unwanted traffic crosses the tunnel and consumes the link. Detection and enforcement in different places, because the place that understands the traffic and the place where blocking is cheap are not the same place.

Recovery, tested rather than assumed. Every container is backed up nightly and the host’s own configuration separately — network, storage definitions, users, firewall, scheduled jobs — because container backups restore onto a machine that must already exist. Two documented scripts rebuild the whole estate from bare metal, and both have a mode that reports what they would do without changing anything. Copies live inside every container, so any single archive carries the instructions for rebuilding the rest.

That last detail was not a design decision. It came from checking a claim I had already written down, finding that four of the six archives I checked did not contain what I had said they did, and fixing it. A backup nobody has restored is a hypothesis.

The result

This site, its mail and its DNS run on that estate in production. Nothing in the building is published directly, nine ports are open at the edge and the rest are refused there. Certificates renew and the services that use them are reloaded, because the reload is tested rather than hoped for. Mail is delivered, and the reports that prove it are read. Two documented scripts rebuild the estate from bare metal, each with a mode that shows what it would do before it does it.

Operating something you depend on also teaches what building it does not. Certificates expire at inconvenient times. Logs fill disks. A service that has run for a year turns out never to have been enabled at boot, and you find out during an unrelated reboot. Every one of those has happened here, and each produced a specific, transferable habit: verify rather than assume, decide how a change undoes itself before making it, and write down why rather than what.

The tools published on this site came out of the same environment. They exist because I wanted the answer to something and the available options were either slow, full of advertising, or wanted an account. None of this needs to be taken on trust: the estate is serving the page you are reading, and the tools run on it.

Leave a response

Every comment is read before it appears. Yours will not show up straight away, and that is not a fault.

Not published, and not used for anything else.