The role
We run a multi-tenant email marketing platform: ~24 Spring Boot services on EKS, React/TypeScript frontend, Kafka, Aurora Postgres, ClickHouse at 9B+ rows, and real sending infrastructure (GreenArrow, PowerMTA).
We need another senior engineer in that codebase. Most of your time is Java/Spring — features, fixes, and the refactors a mature system actually needs. Some of your time is making deploys, CI, and the cluster less painful than they are today. You will work in production. You will not be the lone SRE, DBA, and email-infra owner.
What you'll do
- Ship features and fixes across the Spring Boot services (and the React apps when that's the shortest path).
- Take work from design through production: deploy it, watch it, debug it. Kafka lag, a bad query, a failed rollout — you'll have access and the authority to fix what you find.
- Improve how we ship, in the places that slow everyone down. Highest leverage first: tests in CI, autoscaling on busy services, infra that lives in git. We'll pick those together.
- Write things down. The next person should not have to reverse-engineer what you did.
How we work
We use Claude as our main coding surface. You direct agents, review what they produce, and put your name on the output. Comfort with that workflow matters here.
Follow-through matters more than theatrics. Merge isn't the finish line. "I'll handle it" means handled. Push back early — including on us — when the priority is wrong.
First 90 days
-
30
days
ship a small backend change end to end; deploy and roll back a service unaided.
-
60
days
a meaningful feature in production; one high-leverage ops improvement (tests gating a pipeline, HPA on a busy service, or similar).
-
90
days
contributing independently on backend work; a second ops improvement in flight, with a realistic plan for the rest.
Requirements
- 5+ years shipping production backend. Java/Spring is the daily language. We run Java 8 / Spring Boot 2.3; modernization is on the table.
- You've operated what you wrote: deploys, logs, a production issue you helped close. Not a dedicated SRE background — just not someone who stops at the PR.
- Enough Kubernetes and AWS to deploy, debug, and improve things (EKS, IAM, the usual). Deep infra expertise is a plus, not the hiring bar.
- Kafka and SQL beyond CRUD. ClickHouse / OLAP is a HUGE plus.
- Comfortable working through AI coding tools daily.
Nice to have
Terraform or equivalent introduced to a live system; JVM performance (heap dumps, GC); email infrastructure (GreenArrow, PowerMTA, DKIM/SPF/DMARC, warmup). If you have the email piece, talk to us regardless of the rest.
What this is not
- Not a greenfield rewrite.
- Not a dedicated SRE/on-call hire.
- Not "own every ops gap by day 90."
- Not a ticket factory — you'll have a say in what gets prioritized.
Process
- 1 Intro call 30m
- 2 Technical conversation on our real architecture 90m
- 3 Practical systems/debugging exercise, no algorithm puzzles 60m
- 4 Team conversations 2×45m
- 5 References
