Job Description
- Own cloud cost reporting across GCP and AWS — weekly, monthly, and ad hoc dashboards for engineering and business stakeholders.
- Drive cost optimisation: rightsizing, committed use discounts (CUDs / RIs), idle resource cleanup, storage lifecycle policies, and network egress reduction.
- Design and enforce a cloud tagging strategy — define taxonomy, automate compliance checks, and resolve untagged resources.
- Build showback and chargeback models so teams understand and are accountable for their spend.
- Lead monthly FinOps review meetings — present spend trends, highlight anomalies, explain variances, and recommend actions.
- Set up and maintain budget alerts, anomaly detection, and forecasting models.
- Partner with product and engineering to embed cost awareness early in architecture and feature design.
CloudOps & SRE — Secondary
- Support day-to-day cloud operations on GCP and AWS: incident response, monitoring hygiene, and infrastructure health reviews.
- Contribute to Datadog observability — dashboards, monitors, and alert tuning, particularly for cost-correlated signals.
- Participate in on-call rotations and post-incident reviews with a cost lens on RCA.
- Work with platform and SRE teams on Kubernetes cluster efficiency — node utilisation, autoscaling, and namespace-level cost attribution.
TECH STACK
Required
- GCP — primary cloud platform (Billing API, BigQuery cost export, CUD analyser, GCP Recommender)
- AWS — secondary cloud platform (Cost Explorer, Savings Plans, Trusted Advisor)
- Kubernetes — node pools, resource requests/limits, namespace cost attribution
- Datadog — monitoring and cost dashboards
- Terraform — infrastructure as code for cost-related configurations
Good to have
- CloudZero — unit economics and cost allocation
- Istio — service mesh and traffic cost observability
- Python or SQL — cost data analysis and automation
- Looker or Grafana — cost visualisation