Upgrades and lifecycle
Version planning, upgrade rehearsal, and add-on lifecycle management, so the platform tracks supported Kubernetes versions without drama.
The platform launched, but upgrades, cost, capacity, support, and ownership now need a routine the team can keep.
The routines that keep a live platform healthy: upgrades, capacity, cost, ownership, runbooks, support rhythm, and the backlog after launch.
Launch creates the platform; day two proves whether it can be lived with. Kubernetes versions move, add-ons age, teams need support, capacity shifts, and cloud bills start telling a story. We put the operating loop around the platform and either hand it over or run it with you.
Scope
Every engagement is scoped to the pressure in front of you. These are the areas we usually need to make reliable for the change to stick.
Upgrades and lifecycle
Version planning, upgrade rehearsal, and add-on lifecycle management, so the platform tracks supported Kubernetes versions without drama.
Cost and capacity
Per-team cost allocation, utilisation reviews, and autoscaling tuned to real usage. Visibility first, then optimisation.
Incident readiness
Runbooks for the likely failures, recovery procedures that get tested, and clear escalation, so incidents are handled rather than improvised.
Typical operating model
A day-two operations engagement usually strengthens how the platform is maintained after launch: runtime health, upgrade handling, cost posture, and operational routines.
reliability cues
change management
operating rhythm
Engagement shape
Day-two engagements range from a structured handover to ongoing shared operations. Most teams start with one pressure point and grow from there.
Platform handover
Documentation, runbooks, and a working transition into your team, operating habits included.
Run-with-you support
We stay involved after the buildout: upgrades, incident support, and platform evolution alongside your engineers.
Cost and health reset
A focused pass over spend, capacity, upgrades, and alerts when a live platform has drifted into reactive mode.
Outcomes we are aiming for
Kubernetes and add-on upgrades on a rehearsed path instead of a yearly leap
Cost and capacity visible per team and per workload, with rightsizing built into the routine
Runbooks, ownership boundaries, and a support rhythm your team can actually keep
From the blog
Deep dives from the engineering blog covering the tools and patterns this service is built on.
Kubernetes
Surviving Kubernetes node drains: PDBs and topology spread
Node drains are ordinary Kubernetes maintenance, but they only stay boring when replicas are spread out, disruptions are budgeted, and pods shut down cleanly.
10 min read
Observability
Kubernetes cost allocation with OpenCost
Kubernetes cost allocation with OpenCost: install it with Helm, query per-namespace showback from the Allocation API, and account for idle spend honestly.
8 min read
Kubernetes
Kubernetes autoscaling and scheduled scaling with HPA
Scale Kubernetes on demand and on a schedule, so idle nodes stop costing you overnight.
4 min read
Start with the problem
An honest picture of the upgrade backlog, on-call load, or the cloud bill is the best starting brief we could ask for.