Kubernetes
Run EKS across five AWS accounts: version upgrades, node groups, custom AMIs and launch templates.
I'm an infrastructure engineer working on smart-metering and IoT platforms for utilities. I run Kubernetes across several AWS accounts, look after the Kafka and MQTT brokers that carry device traffic, and plan migrations between AWS and OCI. I like systems that are easy to operate: alerts that mean something, deploys that only touch what changed, and upgrades with a written plan.
how device data moves through the platform
Twelve areas I work in every week, across AWS and OCI.
Run EKS across five AWS accounts: version upgrades, node groups, custom AMIs and launch templates.
Operate the MQTT and Kafka brokers that carry device traffic, and move managed Kafka onto Strimzi.
Design AWS VPCs and OCI VCNs, load balancers, route tables and hub-and-spoke ingress.
Tune Grafana, CloudWatch and Alertmanager so the right alerts reach people and monitoring cost stays in check.
Build pipelines that only rebuild what changed, and deploy across accounts with shared IAM roles.
Move large MySQL datasets between RDS and EC2 with pipelines that resume after a failure.
Layer Karpenter, KEDA, HPA and VPA so bursty Kafka consumers scale on lag, request services scale on load, and nodes follow the pods.
Terraform modules with remote state, CloudFormation for the MQTT clusters, and Helm with Argo CD for application rollout.
Own cost across a 24-account AWS estate: right-sizing, reserved and Savings Plan coverage, cleanups, and a monthly cost-per-meter report.
Account guardrails with SCPs and tag policies, SSO permission sets in place of IAM users, and WAF rate limits on auth endpoints.
Carry on-call for the head-end and meter-data systems, lead the debugging, and write the postmortem so the fix outlives the incident.
Lead a five-person team: three DevOps engineers, a support engineer and a security engineer. Sprint planning, reviews, on-call rota and sign-off on shared production changes.
Recent work on the platform, with the numbers that came out of it.
Every commit used to rebuild and redeploy the whole platform, which took about 30 minutes even for a one-line change. I reworked the pipeline to compare the changed paths against a map of services, build only what changed, and deploy across AWS accounts with shared IAM roles.
Moved EKS from 1.33 to 1.35 across five AWS accounts before extended-support pricing applied, working through launch template, custom AMI and IMDS issues.
Designed the move of an EMQX cluster from AWS to OCI, which meant reworking TLS termination and spreading nodes across fault domains.
Moving managed MSK clusters to Strimzi. Root-caused an outage where the operator's 30-day reconcile deleted SCRAM users created outside it.
Traced a $13/day CloudWatch charge to one wildcard alert rule, and found routing that sent almost every alert to a null receiver.
Moved three months of data off RDS with a checksum-verified pipeline that resumes cleanly when the 60-minute session limit kills a run.
Built a capacity model from 14 days of production data (Kafka offsets, load-balancer logs, database load) to size a 10-million-meter rollout instead of multiplying today's numbers. Pods and Kafka partitions scale out; the database writer doesn't. Its load is 85–100% writes, so read replicas can't help. Checked against a second week of data, I/O-bound consumers drained 7–35× slower than a CPU-only model predicted.
Moved the head-end services from static node groups to Karpenter, with KEDA scaling Kafka consumers on lag, HPA on request services and Goldilocks-guided VPA for sizing. Some services stay pinned on purpose: a decode fleet sized one-to-one with its partitions beats one that rebalances all day.
Cut one production cluster over from MSK to Strimzi on KRaft and from ElastiCache to Valkey 9. I measured MirrorMaker 2 first: it was I/O-bound at about 2k messages a second, far too slow to copy history, so we switched at the live tail with a small reprocess and kept the managed services warm as the rollback.
Ran a full failover round trip for the head-end and meter-data workloads on two production clusters: out to a second availability zone and back. The drill found two wrong assumptions in the plan: a cluster autoscaler kept reviving the source node group, and meter-data pods were pinned to a node-group name.
Designed the tag taxonomy, tag policy and SCP model for a flat 24-account organisation. Backfilled canonical tags on a production account, and kept cost tags report-only on purpose: blocking launches on a missing tag would have stopped about 44% of node launches mid-scale-up.
Measured VPC Flow Logs on two production VPCs before a rollout. Logging all traffic would grow to about $15.4K a month at 1.2 million meters. Reject-only logging costs $0.93 a month today and stays at a few dollars at that scale, because rejects are almost all internet port scans and don't grow with meter count.
Production problems I traced to a root cause, and what changed afterwards.
Every service on a head-end cluster kept losing its database connection with "unexpected EOF". The database host ran strict memory overcommit with no swap, so Postgres had about 140 MB left to fork new backends. The fix was one kernel setting, with no restart.
About sixty pods kept failing health checks. Day-old orphaned queries had pinned the vacuum horizon, tables bloated, and the database sat at its disk IOPS cap. I cancelled them, set statement timeouts, vacuumed, and added two missing indexes.
A Kafka consumer group was stuck in an endless rebalance. VPA had shrunk decode pods from 500m to about 53m of CPU, so each batch outran the poll timeout and offsets never committed. I pinned the resources and kept VPA as recommend-only.
A boot script's cluster-join check found peers by EC2 tag and, when none of them answered, assumed the join had failed and stopped the broker. With every node down at once, no node could ever go first. I seeded one node by hand, joined the rest and restored the config from backup. The boot log, not the broker log, showed it, and it also showed the same check had quietly stopped the backups.
Clients couldn't reach a healthy MQTT cluster. A CloudFormation update had rewritten the Auto Scaling group's target groups and dropped one I had attached by hand weeks earlier. I put that target group in the template for both production stacks, including the one that hadn't broken yet.
A meter-data Aurora writer was being restarted every few minutes. A table had just moved to 1,219 daily partitions, and a hot UPDATE by id didn't filter on the partition key, so each call opened every partition across about 200 connections. I diagnosed it from RDS events, Performance Insights and logs, and the change was rolled back the same day.
I dropped old TimescaleDB chunks table by table, because one big drop exhausts the lock table. Then I shrank the root volume by cloning onto a smaller disk with the same filesystem UUID. I boot-tested the clone before the swap; the IPs stayed the same and downtime was about ten minutes.
Three years at Polaris Smart Metering, from intern to leading the infrastructure team.
The open-source path a meter reading takes: MQTT on EMQX, Kafka on Strimzi, microservices on Kubernetes, watched by Prometheus, Grafana and Loki. And what it costs to run one meter.
Layered autoscaling with KEDA, HPA, VPA and Karpenter, with cost per meter as the number that settles each scaling decision.
Four current certifications across Kubernetes, AWS and Terraform.
Open to infrastructure, platform and SRE roles. Email is the fastest way to reach me.
Looking for my next infrastructure role