About the role
You own the backbone: scale, reliability, observability, and how fast everyone else can ship. An agent people trust with their working life cannot be down, and cannot be slow, and the reason it is neither will be work you did. This is a small team, so you will set the SLOs rather than inherit them.
What we need to see
- Deep distributed-systems experience running production at scale, including being on call for it
- Fluency with cloud, containers, CI/CD, and real observability rather than dashboards nobody reads
- A reliability mindset you can evidence: SLOs, error budgets, blameless postmortems
- You improve deploy frequency and recovery time, not just uptime
Nice to have
- Google Cloud specifically, since that is where we run
- Edge or hybrid infrastructure, where some compute is on hardware we do not own
- You have taken an organisation through a compliance regime such as SOC 2 or FedRAMP
What winning looks like
- Uptime against SLOs and a shrinking error budget burn
- p99 latency and cost-to-serve
- Deploy frequency and mean-time-to-recovery
Where and how we work
In the office together five days a week, in any of these cities. Remote-friendly around your family, arranged one person at a time.