Skip to content

TechProValley

Services

Engagements usually fall into one of six categories. All of them start with the same thing: finding out what is really there.

Network architecture & troubleshooting

Two problems arrive under this heading. The first is a network that has outgrown the design it was built on — it still works, in the sense that packets arrive, but nobody can explain it and every change carries a risk nobody can size. The second is a fault that will not go away: intermittent, hard to reproduce, and already the subject of several attempts to fix it.

The method is the same either way. Every running configuration is collected and compared against its neighbours, so the variations are catalogued rather than discovered mid-change. That comparison is mechanical enough to hand to a tool, so it is: ConfigAudit reads the configurations offline and reports what disagrees, what overlaps and what was left at its default. Then the design: campus and multi-site wide-area networks, OSPF and BGP routing policy, VLAN and VRF segmentation, high availability tested by switching the primary off, and quality of service for voice and video. For a fault it is measurement rather than opinion — isolate the layer, prove the cause, change one thing with a rollback, verify it, write it down.

The diagnostic method, from a reported symptom to a verified and documented fixMeasure before changing anything, isolate the layer, prove the cause — which is usually not where the symptom appeared.
Measure before changing anything, isolate the layer, prove the cause — which is usually not where the symptom appeared.

You get: an audit of what is actually running rather than what the documentation claims, a design document, a configuration set and a migration plan with a rollback — not a diagram and good wishes. For a fault: the cause, the evidence for it, and the change that removes it.

Typically: Cisco IOS and IOS-XE, MikroTik RouterOS, Juniper Junos, and the mixed estates where all three meet.

The boundary: a diagnosis cannot be quoted like a build. I will tell you what I expect to find and what it costs to look, and sometimes the finding is that the design is sound and one interface has been quietly discarding frames for a year.

VPN & secure tunnelling

Most VPN problems are not encryption problems. The tunnel comes up, both ends can ping each other, and the thing it was built for still does not work — because the routing inside the tunnel was never designed, the names resolve to the wrong side of it, or the path crosses a network that interferes with anything it does not recognise. That is usually what I am called about, once the tunnel itself has been declared healthy and the problem has not moved.

The work covers both shapes of it: site-to-site, where two or more estates have to behave like one network without becoming one failure domain, and remote access, where staff and administrators need to get in from wherever they happen to be, over a connection you do not control.

Before anything is designed, the path is measured. TunnelCheck, run at both ends, reports the largest packet the path carries in each direction, how the NAT in the middle behaves, which ports cross and how long a quiet connection survives — the numbers the tunnel’s MTU, MSS clamp and keepalive have to be set from.

  • IPsec or WireGuard, decided on evidence. What each one buys, where each is the wrong answer, and the fact that an interoperability requirement usually settles it before preference gets a say. Vendor kit already in place at one end is a design input, not something to argue with.
  • Overlay and mesh, with a control plane you own. Peers that find each other and connect directly, coordinated by a control plane running on your infrastructure rather than a third-party service that can change its terms, its price or its availability.
  • Tunnels that survive networks which fight back. Symmetric and carrier-grade NAT, proxies that inspect packets, captive portals, IPv6-only segments. The ladder is a direct connection first, then a port mapping asked of the router, then a relay that routes by identifier and cannot read the session — tried in sequence, not all at once, because each rung costs something the one before it does not.
  • Routing and DNS inside the tunnel. The part that is usually what actually broke. A subnet router so the overlay reaches a whole LAN instead of needing an agent on every host, and split-horizon DNS so a name resolves to the internal address over the tunnel and the public one outside it, which removes a hairpin that consumer routers handle badly and keeps internal structure private.
  • Key and peer lifecycle. Keys generated on the device that keeps them, adding and removing a peer treated as a normal operation rather than an event, and configuration backed up in a way that has been restored from.
A remote-access and site-to-site overlay with a control plane the client ownsRoaming clients and sites reaching one estate over encrypted peer-to-peer tunnels: a self-hosted control plane, a single subnet router into the LAN, and split-horizon DNS inside the tunnel.
Roaming clients and sites reaching one estate over encrypted peer-to-peer tunnels: a self-hosted control plane, a single subnet router into the LAN, and split-horizon DNS inside the tunnel.

You get: a tunnel design with the routing, the name resolution and the failure behaviour written down, the configuration for both ends, and a peer and key procedure your team can run without me. Where a management interface has been published to the internet to make administration possible, it comes off and the administration moves inside the tunnel.

The boundary: a VPN cannot repair the link underneath it. If the circuit loses packets the tunnel will too, and the honest first step is proving which of the two you have. Where one node ends up load-bearing — a subnet router is the usual example, and restarting it cuts remote access — I will say so in writing rather than leave it to be discovered.

Worked example: Remote access without a hosted control plane — an estate administered entirely over an encrypted overlay, with no management port published anywhere, and the trade-offs of that choice set out rather than glossed over.

Server & service infrastructure

The situation is usually an estate that grew by accident: services on whichever machine had room that year, a backup regime nobody has restored from, and one or two machines everybody is afraid to reboot. It works until the day it is asked a question nobody has asked it before.

Web, mail, DNS, FTP, application and database servers — built properly, secured, monitored and documented. Virtualisation and container platforms underneath. Mail with the forward and reverse DNS, SPF, DKIM and DMARC records that decide whether anyone receives what you send, and certificate renewal that also reloads the service, which is the most common way that fails. Backup and disaster recovery rehearsed rather than assumed, and the alerting to go with it. Including consolidation work: taking an estate that grew by accident and giving it a shape.

You get: built, hardened and documented services, a recovery procedure that has been run rather than only written, and a record of what depends on what.

The boundary: I build and hand over, and I am glad to stay on for operations, but a service with no owner after handover will drift. If nobody is going to own it, that belongs in the plan before the build.

Monitoring & custom tooling

This is the last category for a reason: it starts where configuring what you already have has run out. Either the platform is too large for a single network team, or too shallow to answer the question you actually have, or it will not look at the ping-only devices that make up a good part of a real estate.

So: monitoring and reporting that raises a problem while it is still small, configuration automation for work that is too repetitive to do by hand and too risky to do carelessly, integrations between systems that were never meant to talk to each other, and occasionally a full application. Three of them are published on this site — NetMon for monitoring mixed estates, McastTool for proving whether a multicast group crosses a path, RmDesk for reaching a machine across a network that fights back. Each exists because the available options stopped short of the job: too large for a single team, too shallow once the detail mattered, impossible to install on a machine in a comms room I did not control, or giving up in exactly the network conditions the work was in.

You get: a tool built to your environment and handed over with source and documentation, written to be picked up and understood by someone else rather than to be clever.

The boundary: software is the last resort, not the offer. Most of what gets asked for here is better served by configuring what you already have, and I will say so before quoting to build anything.

Unified communications & voice

The complaints here are specific and hard to argue with: calls break up on some routes and not others, a site goes silent when its wide-area link drops, or adding a branch means touching translation patterns in four places. None of them is really a complaint about the phone system.

CUCM deployment, migration and repair. Cluster design with redundancy tested by failing the publisher deliberately, a dial plan structured to survive growth, voice gateways with local breakout so a WAN failure degrades a branch to local and external calling rather than silence, SIP trunks with the codec and transcoding policy decided explicitly rather than left to negotiate, and the QoS path from handset to provider. This includes the unglamorous part: fixing the call-quality complaint that turns out to be a queue, or a trust boundary that was never applied, on a switch three hops from anything anyone would call a voice system.

You get: a cluster and dial plan documented well enough to extend without restructuring, and a QoS policy that new switches inherit rather than miss.

The boundary: the voice platform and the network are one conversation. Asked to fix call quality without being allowed to look at the switching and the wide-area path, I will say no rather than guess.

Satellite & remote-site connectivity

Sites where a visit costs a day, the round trip is measured in hundreds of milliseconds, the weather is part of the topology, and often there is no second route to fail over to. Six years of hub station operations sit behind this: administering, configuring and managing a VSAT hub and the remote terminals working through it.

Hub and remote configuration and management, commissioning new remote sites, carrier and bandwidth planning across remotes competing for the same space segment, link budgets built with margin because rain fade is a design input and not an incident, and performance work on paths where latency and loss are permanent conditions rather than faults. Also the terrestrial side of hybrid designs, where satellite is the backup path and the failover has to be believable — a backup path that has never carried production traffic is not a backup path.

You get: a remote site whose configuration was validated before the terminal shipped and whose monitoring was installed before handover, a bandwidth and prioritisation policy, and a failover tested by taking the primary path away.

The boundary: the ground segment, the terminals and the network around them are what I work on. Buying capacity and putting hardware on a roof at the far end still needs someone local, and I will tell you what to ask them for.


How an engagement runs

  1. Conversation. What is happening, what it costs you, and what fixed looks like. Free — and it sometimes ends with me saying you do not need me.
  2. Audit. A documented picture of what is actually running, including the risks nobody has written down. Yours to keep either way.
  3. Proposal. Scope, sequence and cost, with the risky steps named before they are taken rather than discovered during them.
  4. Execution. Changes in planned windows, each with a rehearsed rollback, and progress reported as it happens.
  5. Handover. Documentation, configurations and a walkthrough, so your team owns the result instead of depending on me.

If your network or VPN setup is complex, let’s discuss an architecture that actually works — before you spend. Discuss your architecture.

Leave a response

Every comment is read before it appears. Yours will not show up straight away, and that is not a fault.

Not published, and not used for anything else.