The role
We run a multi-tenant email marketing platform: ~24 Spring Boot services on EKS, React/TypeScript frontend, Kafka, Aurora Postgres, ClickHouse at 9B+ rows, and real MTA infrastructure (GreenArrow, Kumo MTA, Postfix warmup fleets).
The application layer is mature. The ops layer isn't:
- No IaC
- Zero autoscalers across 35 hand-pinned deployments
- Deploys that skip tests
- A half-finished GitHub Actions setup
You ship product features in the same codebase you operate — Java services, React frontend — and you fix the tech debt we keep stumbling over instead of working around it. Decide, build, operate, answer for it. No infra team behind you, no ticket queue to hand off to.
We are agentic-first. Claude is our entire coding and infra-management surface — services, manifests, migrations, runbooks, and incident debugging all run through it. The expectation is an engineer directing agents across several workstreams at once, reviewing and owning everything they produce. "The agent wrote it" is never an excuse; your name is on the output. If you haven't made AI-assisted development a core part of how you ship, this won't be a fit.
You will own
-
Kubernetes & compute
the EKS cluster, Karpenter node pools, deployment manifests. Replace hand-edited replica counts with real autoscaling you can defend on cost.
-
IaC
none exists today. Pick the tooling, migrate incrementally, kill the drift between git and the live cluster.
-
CI/CD
GitHub Actions on self-hosted runners, You own the runners and the pipelines.
-
Reliability & observability
Prometheus/Sentry are wired up; SLOs, dashboards-as-code, and real alerting are not. You own incidents, postmortems, and closing the action items.
-
Data ops
ClickHouse migrations and query performance at billion-row scale, Postgres, Redis.
-
Disaster recovery
the MTA backup posture needs work. Turn it into tested, boring recovery.
-
Email infrastructure
MTAs, DKIM, IP warmup, deliverability. Mail arriving at scale is the product.
Ownership, concretely
- Merge is not the finish line — you watch it in production and you're the first to know it broke.
- "I'll handle it" means handled. Nobody follows up with you.
- You own the outcome, not the ticket.
- You write the postmortem and close its action items.
- You push back early — including on us, when the priority is wrong.
- You leave runbooks behind. Knowledge-hoarding is disqualifying.
You are primary on-call. In exchange: broad authority to fix what pages you, and reducing your own pager load is part of the job.
Requirements
- 7+ years operating production systems — you've written the code and carried the pager.
- Java/Spring at feature-delivery depth, plus JVM ops: heap dumps, GC tuning, thread stalls. Backend is Java 8 / Spring Boot 2.3; modernization is on the table.
- Deep production Kubernetes and AWS (EKS, EC2, Aurora, IAM, networking, cost).
- IaC fluency, including introducing it to a live system without a big-bang rewrite.
- Kafka under load: lag, rebalances, partitioning, ordering.
- SQL depth beyond CRUD; columnar/OLAP experience a plus.
- Fluent agentic development — you already work through Claude (or equivalent) daily, and you can review and stand behind agent output at production quality
Rare and valuable
Email infrastructure (GreenArrow, PowerMTA, Postfix, DKIM/SPF/DMARC, warmup, deliverability). If you have this, talk to us regardless of the rest.
What this is not
- Not greenfield.
- Not hands-off architecture.
- Not buffered from production.
- Not activity theater — stalled migrations and unread dashboards are worse than nothing here.
Process
- 1 Intro call 30m
- 2 Technical deep dive on our real architecture 90m
- 3 Incident/systems exercise, coding problems (Claude is allowed and encouraged) 60m
- 4 Team conversations 2-3×45m
- 5 References
