Services

Engagements are scoped in weeks, billed monthly, and cancellable with 30 days' notice. We would rather leave early than be needed forever.

Practices

Pick one, or let the assessment tell you which two matter most.

Platform build-out

Cluster topology, ingress, secrets management and GitOps delivery. Terraform or Pulumi, your choice — we write it the way your team already writes it.

Migration

Data-centre to cloud, cloud to cloud, or VM to container. Each workload moves behind a documented rollback, and we keep the old path warm until you say otherwise.

Observability

Prometheus, OpenTelemetry and structured logging, with SLOs agreed by the people who carry the pager rather than the people who buy the tooling.

Reliability engineering

Load and failure testing, capacity planning, backup restore drills, and blameless incident reviews that produce actual backlog items.

Security hardening

Network segmentation, least-privilege IAM, supply-chain controls and secret rotation. We report findings in writing, ranked by exploitability.

Cost engineering

Usage attribution per team and per service, commitment and spot strategy, plus a monthly review so savings do not quietly evaporate.

Engagement models

Three shapes cover nearly everything we are asked for.

ModelCommitmentBest for
AssessmentFixed, 1–2 weeksYou suspect something is wrong but cannot name it
ProjectFixed scope, 4–16 weeksA migration or platform build with a defined finish line
RetainedMonthly, rollingOngoing operations alongside your own engineers

Tooling we work with

Not an exhaustive list, and not a religious one. We will happily work inside whatever you have already standardised on.

Orchestration

Kubernetes, Nomad, ECS, systemd where it is genuinely simpler.

Infrastructure as code

Terraform, OpenTofu, Pulumi, Ansible, Packer.

Delivery

GitHub Actions, GitLab CI, Argo CD, Flux, Buildkite.

Telemetry

Prometheus, Grafana, Loki, Tempo, OpenTelemetry, Vector.