Engineering Journal

Real decisions. Real outcomes.

Concise narratives showing how evidence changed architecture, cost, risk and operational outcomes.

Every infrastructure engineering case study is based on real experience across networks, platforms, SQL Server, backup, identity, storage and virtualisation. Client identities and identifying details are protected, while the decisions, trade-offs and measured outcomes remain genuine.

EN-001 · Architecture

The replacement wasn't the fix

A wireless refresh had been approved. The complaints survived the new access points.

4 min read
EN-002 · Procurement

The cheapest tender became the most expensive decision

A discounted storage platform won the procurement. The architecture still had to work afterwards.

5 min read
EN-003 · Performance

Performance wasn't the problem

A switching refresh was proposed. Packet analysis showed the hardware was not the constraint.

5 min read
EN-004 · Applications

The upgrade wasn't the objective

A platform upgrade preserved the client outcome while permanently changing the effort required to deliver it.

5 min read
EN-005 · Architecture

Technology isn't capability

Buying the platform is not the same as becoming able to operate it.

4 min read
EN-006 · Operations

The vendor has left. The work has not.

The right answer should still be the right answer when the vendor leaves the room.

4 min read
EN-007 · Security & Identity

The platform wasn't the decision

A vendor deadline created urgency. A working alternative was proven in the client environment within two hours.

4 min read
EN-008 · Security & Identity

The contingency became the production platform

A temporary remote access path proved more dependable than the service it had been introduced to protect.

4 min read
EN-009 · Applications & Data

The database wasn't the bottleneck

An inventory search fell from fifteen minutes to fifteen seconds when the application, SQL and infrastructure were treated as one platform.

5 min read
EN-010 · Operational Assurance

One source of truth mattered more than another system

A defence environment needed every operator to act from the same sonar tracking picture without conflicting information.

4 min read
EN-011 · Infrastructure & Networks

The thirty-day project that took four

A network transformation estimated at thirty days was completed in four because the scope was understood before delivery began.

4 min read
EN-013 · Architecture & Resilience

The biggest vendor is still a single point of failure

CrowdStrike and Entra outages proved the same point: market leadership does not remove the need for independent recovery paths.

5 min read
EN-014 · Data Protection & Recovery

The working solution was replaced before its purpose was understood

The backups continued to run, but the recovery design, retention model and compliance position had been lost.

5 min read
EN-015 · Security & Identity

The certificate did not expire unexpectedly

The expiry date was known years in advance. The outage happened because renewal ownership and deployment knowledge were not.

5 min read
EN-016 · Storage & Operational Intelligence

The array was healthy. The disk was not.

The platform-level status remained reassuring while physical media telemetry showed that replacement planning should already have started.

5 min read
EN-017 · Virtualisation & Operational Assurance

The missing agent exposed the missing baseline

A Proxmox telemetry warning revealed that a guest had been deployed outside the intended operational standard.

5 min read
EN-018 · Data Protection & Recovery

The backup succeeded. The restore had never been tested.

A green backup report confirmed that data had been copied. It did not prove that the service could be recovered.

5 min read
EN-019 · Monitoring & Operational Intelligence

The dashboard was green because collection had stopped.

The last known state looked healthy. The monitoring system had quietly stopped learning anything new.

4 min read
EN-020 · Operations & Service Assurance

Everything was monitored. Nobody owned the alert.

The condition was detected correctly. The operating model gave nobody a dependable reason to act on it.

5 min read
EN-021 · Resilience & Service Architecture

The failover worked. The service still failed.

The clustered platform moved exactly as designed. A dependency outside the cluster prevented the application from recovering.

5 min read
EN-022 · Network Architecture & Change

The firewall change was correct. The dependency map was not.

The rule matched the approved request. The request did not describe every service that depended on the traffic path.

5 min read