Engineering Journal
Real decisions. Real outcomes.
Concise narratives showing how evidence changed architecture, cost, risk and operational outcomes.

Every infrastructure engineering case study is based on real experience across networks, platforms, SQL Server, backup, identity, storage and virtualisation. Client identities and identifying details are protected, while the decisions, trade-offs and measured outcomes remain genuine.
The replacement wasn't the fix
A wireless refresh had been approved. The complaints survived the new access points.
4 min readThe cheapest tender became the most expensive decision
A discounted storage platform won the procurement. The architecture still had to work afterwards.
5 min readPerformance wasn't the problem
A switching refresh was proposed. Packet analysis showed the hardware was not the constraint.
5 min readThe upgrade wasn't the objective
A platform upgrade preserved the client outcome while permanently changing the effort required to deliver it.
5 min readTechnology isn't capability
Buying the platform is not the same as becoming able to operate it.
4 min readThe vendor has left. The work has not.
The right answer should still be the right answer when the vendor leaves the room.
4 min readThe platform wasn't the decision
A vendor deadline created urgency. A working alternative was proven in the client environment within two hours.
4 min readThe contingency became the production platform
A temporary remote access path proved more dependable than the service it had been introduced to protect.
4 min readThe database wasn't the bottleneck
An inventory search fell from fifteen minutes to fifteen seconds when the application, SQL and infrastructure were treated as one platform.
5 min readOne source of truth mattered more than another system
A defence environment needed every operator to act from the same sonar tracking picture without conflicting information.
4 min readThe thirty-day project that took four
A network transformation estimated at thirty days was completed in four because the scope was understood before delivery began.
4 min readThe biggest vendor is still a single point of failure
CrowdStrike and Entra outages proved the same point: market leadership does not remove the need for independent recovery paths.
5 min readThe working solution was replaced before its purpose was understood
The backups continued to run, but the recovery design, retention model and compliance position had been lost.
5 min readThe certificate did not expire unexpectedly
The expiry date was known years in advance. The outage happened because renewal ownership and deployment knowledge were not.
5 min readThe array was healthy. The disk was not.
The platform-level status remained reassuring while physical media telemetry showed that replacement planning should already have started.
5 min readThe missing agent exposed the missing baseline
A Proxmox telemetry warning revealed that a guest had been deployed outside the intended operational standard.
5 min readThe backup succeeded. The restore had never been tested.
A green backup report confirmed that data had been copied. It did not prove that the service could be recovered.
5 min readThe dashboard was green because collection had stopped.
The last known state looked healthy. The monitoring system had quietly stopped learning anything new.
4 min readEverything was monitored. Nobody owned the alert.
The condition was detected correctly. The operating model gave nobody a dependable reason to act on it.
5 min readThe failover worked. The service still failed.
The clustered platform moved exactly as designed. A dependency outside the cluster prevented the application from recovering.
5 min readThe firewall change was correct. The dependency map was not.
The rule matched the approved request. The request did not describe every service that depended on the traffic path.
5 min read