Senior Platform Engineer (DevOps / SRE) — Remote in Canada | tinyEmail Careers

Engineering · tinyEmail / tinyAlbert / tinyRelay

Senior Platform Engineer (Eng / DevOps / SRE)

Team
Engineering — tinyEmail / tinyAlbert / tinyRelay
Reports to
CTO
Location
Canada - Remote

The role

We run a multi-tenant email marketing platform: ~24 Spring Boot services on EKS, React/TypeScript frontend, Kafka, Aurora Postgres, ClickHouse at 9B+ rows, and real MTA infrastructure (GreenArrow, Kumo MTA, Postfix warmup fleets).

  • ~24 Spring Boot services
  • EKS · Karpenter
  • React / TypeScript
  • Kafka
  • Aurora Postgres
  • ClickHouse · 9B+ rows
  • GreenArrow
  • Kumo MTA
  • Postfix warmup fleets
  • The application layer is mature. The ops layer isn't:

    • No IaC
    • Zero autoscalers across 35 hand-pinned deployments
    • Deploys that skip tests
    • A half-finished GitHub Actions setup

    You ship product features in the same codebase you operate — Java services, React frontend — and you fix the tech debt we keep stumbling over instead of working around it. Decide, build, operate, answer for it. No infra team behind you, no ticket queue to hand off to.

    We are agentic-first. Claude is our entire coding and infra-management surface — services, manifests, migrations, runbooks, and incident debugging all run through it. The expectation is an engineer directing agents across several workstreams at once, reviewing and owning everything they produce. "The agent wrote it" is never an excuse; your name is on the output. If you haven't made AI-assisted development a core part of how you ship, this won't be a fit.

    You will own

    • Kubernetes & compute

      the EKS cluster, Karpenter node pools, deployment manifests. Replace hand-edited replica counts with real autoscaling you can defend on cost.

    • IaC

      none exists today. Pick the tooling, migrate incrementally, kill the drift between git and the live cluster.

    • CI/CD

      GitHub Actions on self-hosted runners, You own the runners and the pipelines.

    • Reliability & observability

      Prometheus/Sentry are wired up; SLOs, dashboards-as-code, and real alerting are not. You own incidents, postmortems, and closing the action items.

    • Data ops

      ClickHouse migrations and query performance at billion-row scale, Postgres, Redis.

    • Disaster recovery

      the MTA backup posture needs work. Turn it into tested, boring recovery.

    • Email infrastructure

      MTAs, DKIM, IP warmup, deliverability. Mail arriving at scale is the product.

    Ownership, concretely

    • Merge is not the finish line — you watch it in production and you're the first to know it broke.
    • "I'll handle it" means handled. Nobody follows up with you.
    • You own the outcome, not the ticket.
    • You write the postmortem and close its action items.
    • You push back early — including on us, when the priority is wrong.
    • You leave runbooks behind. Knowledge-hoarding is disqualifying.

    You are primary on-call. In exchange: broad authority to fix what pages you, and reducing your own pager load is part of the job.

    Requirements

    • 7+ years operating production systems — you've written the code and carried the pager.
    • Java/Spring at feature-delivery depth, plus JVM ops: heap dumps, GC tuning, thread stalls. Backend is Java 8 / Spring Boot 2.3; modernization is on the table.
    • Deep production Kubernetes and AWS (EKS, EC2, Aurora, IAM, networking, cost).
    • IaC fluency, including introducing it to a live system without a big-bang rewrite.
    • Kafka under load: lag, rebalances, partitioning, ordering.
    • SQL depth beyond CRUD; columnar/OLAP experience a plus.
    • Fluent agentic development — you already work through Claude (or equivalent) daily, and you can review and stand behind agent output at production quality

    Rare and valuable

    Email infrastructure (GreenArrow, PowerMTA, Postfix, DKIM/SPF/DMARC, warmup, deliverability). If you have this, talk to us regardless of the rest.

    What this is not

    • Not greenfield.
    • Not hands-off architecture.
    • Not buffered from production.
    • Not activity theater — stalled migrations and unread dashboards are worse than nothing here.

    Process

    1. 1 Intro call 30m
    2. 2 Technical deep dive on our real architecture 90m
    3. 3 Incident/systems exercise, coding problems (Claude is allowed and encouraged) 60m
    4. 4 Team conversations 2-3×45m
    5. 5 References

    Want the pager and the authority?

    Tell us the ops gap you'd close first and how you'd know it worked. We read every application.

    Apply for this role