Senior DevOps Engineer
Boats Group
Senior DevOps Engineer — Miami (Hybrid)
(Must be local to Miami area or commutable distance. If willing to relocate, please specify on your application.)
About The Role
As a Senior DevOps Engineer you will be part of the Infrastructure Team responsible for Boats Group’s internal developer platform. The platform provides self-service infrastructure for 500+ repositories across 13 international marketplace brands, running on AWS with EKS as the primary compute target.
You’ll own and evolve the platform that engineering teams use every day: Kubernetes clusters, GitOps delivery pipelines, CI/CD automation, observability, Cloudflare edge infrastructure, and Terraform/Atlantis-based infrastructure provisioning. You will also drive active migrations from legacy infrastructure (Jenkins, ECS, CloudFormation) to the modern platform.
A successful candidate is someone who thinks in systems — you understand how a Helm chart, an ArgoCD ApplicationSet, a Terraform module, and a GitHub Actions workflow connect to deliver a service from PR to production. You’re comfortable operating at scale (200+ deployed services, multi-cluster EKS) and you know when to automate and when to simplify. Boats Group is a heavy AI adopter — LLMs and AI coding tools are part of everyday engineering work, and we expect you to bring real experience using them.
Every day is different: one day you’re debugging a Karpenter node scaling issue, the next you’re writing Terraform modules with Atlantis PR workflows, the next you’re migrating a Jenkins pipeline to GitHub Actions. You won’t be bored.
What You’ll Do
- Operate and evolve our multi-cluster EKS platform, ArgoCD GitOps (200+ ApplicationSets), and Helm charts
- Write Terraform modules and manage infrastructure changes via Atlantis PR workflows
- Maintain and extend GitHub Actions reusable workflows and composite actions
- Manage the Grafana LGTM observability stack (Mimir, Loki, Tempo, OpenTelemetry)
- Manage Cloudflare edge infrastructure — 70+ domains, WAF, DNS, Workers, Bot Management
- Migrate legacy services and pipelines from ECS/Jenkins/CloudFormation to the modern platform
- Use AI coding tools (Claude Code, MCP integrations) daily to accelerate infrastructure work
- Improve deployment safety with progressive delivery (Argo Rollouts, Kargo)
- Monitor and optimize AWS spending, security policies (Kyverno, RBAC), and IAM trust chains
- Participate in on-call coverage for production outages
What You Must Have
- Kubernetes — EKS operations, pod/node troubleshooting, RBAC, networking
- GitOps — ArgoCD or Flux; reconciliation model, sync failure debugging
- IaC — Terraform/Terragrunt at scale; Atlantis PR-based workflows
- CI/CD — GitHub Actions (or similar) reusable workflows, not just consuming pipelines
- AWS — Multi-account, IAM/OIDC, EKS, ECR, RDS, S3, SQS, VPC, CloudWatch
- Helm — Writing and maintaining production charts, values layering, template debugging
- Observability — Prometheus/Grafana stack; OpenTelemetry, Mimir, or Loki a plus
- SRE practices — SLOs/SLIs, error budgets, incident response, blameless postmortems
- AI/LLM fluency — Daily use of AI coding tools; prompt engineering, code generation, critical evaluation of LLM output
What You Should Have
- K8s Ecosystem: Karpenter, Cilium, KEDA, Kyverno, External Secrets, cert-manager, Argo Rollouts/Kargo
- IaC & CI/CD: Terraform/Terragrunt module authoring, Atlantis, GitHub Actions (composite actions + reusable workflows), ArgoCD ApplicationSets, ECR image lifecycle
- Migration: Jenkins → GHA, ECS → EKS, CloudFormation → Terraform, monolith decomposition
- Languages: TypeScript/Node.js (primary ecosystem), Python, Go (bonus), Bash
- AWS: EKS, ECR, Lambda, RDS/Aurora, DocumentDB, Elasticsearch, S3, SQS, EventBridge, IAM/OIDC, Cognito, SSM, VPC, Route53, CloudFront, CloudWatch, Athena, Organizations
- Observability: Grafana, Mimir, Loki, Tempo, OpenTelemetry (Collector + SDK), Prometheus, PagerDuty
- Cloudflare: CDN, DNS, WAF, Bot Management, Workers (TypeScript/Hono, KV, R2), Terraform at scale, Logpush
- AI & Automation: Claude Code / Cursor / Copilot, MCP server integrations, prompt engineering for infrastructure tasks
- Bonus: Envoy Gateway, FinOps, Snowflake, prior software development experience
What You'll Receive
- Hybrid Work Flexibility: Embrace a balanced work model with remote work on Mondays and Fridays and in-office collaboration from Tuesday to Thursday.
- Generous Time Off: With a strong focus on work/life balance, we offer all employees paid time off starting on day one, multiple paid holidays throughout the year, your birthday off, and a winter break at the end of the year
- Volunteering Time: Participate in our volunteer program with 4 paid days annually to contribute to your community.
- Modern Office Perks: Our vibrant Miami office features cutting-edge amenities, such as an electric sit/stand desk, dual monitors, a gym, and a variety of snacks and beverages.
- Comprehensive Benefits Package: Enjoy top-tier Medical, Dental, Vision, and Life insurance, along with a 401(k) plan featuring a 4% match.
- Commuter Benefits: Park conveniently in our building’s garage at no charge to you. For train commuters, we subsidize most, if not all, of your monthly pass expenses.
- Professional Development: Take advantage of online training, live courses, and additional funds for courses, seminars, and certifications to enhance your skills.
- Team-Centric Atmosphere: Be part of a close-knit team that prioritizes relationship-building and personal connections.