Skip to content

Job Title: Lead Platform Engineer
Location: Remote (Guatemala/Mexico/Colombia)
Reports To: Director of Engineering
Employment Type: Contractor


Who we are:

We are a rapidly growing, high-impact startup in the procurement space, expanding at 50% YoY. We process massive volumes of sensitive invoice and supplier data using advanced OCR, structured data extraction, matching, validation, and anomaly detection. Because we handle critical financial data for our clients, security, compliance, and trust are foundational to everything we build. As we rapidly expand our product capabilities into Agentic AI workflows, we are overhauling our infrastructure to support our next massive phase of growth securely and efficiently.

The role:

We are looking for a Lead Platform Engineer to spearhead the modernization of our cloud infrastructure. Today, we operate heavily in an AWS Serverless environment (Lambdas, Step Functions, DynamoDB). Tomorrow, with your leadership, we are migrating to a robust, highly scalable, and compliant platform built on Kubernetes, Temporal.io, and a mix of modern databases (PostgreSQL, MySQL, NoSQL).

In this role, you will own the transition, lay down the foundational infrastructure using Terraform CDK, and build out modern CI/CD pipelines from scratch. A major mandate for this role is to design the new platform to be as cloud-agnostic as possible, preventing vendor lock-in while maintaining high performance. You will ensure the platform is secure by design—adhering to frameworks like SOC 2 and GDPR—while designing the infrastructure required to run cutting-edge Agentic AI workloads.

What you’ll do:

  • Architect the Future: Lead the migration from our current AWS Serverless stack to a modern, cloud-agnostic Kubernetes architecture capable of supporting diverse database workloads (PostgreSQL, MySQL, NoSQL).
  • Security & Compliance by Design: Build the new infrastructure to strictly adhere to compliance frameworks (e.g., SOC 2, ISO 27001, GDPR). Implement DevSecOps principles, zero-trust networking, secure secrets management, and robust IAM policies.
  • Workflow Orchestration: Implement and manage Temporal.io to replace AWS Step Functions, ensuring highly reliable, scalable, and resilient distributed workflows.
  • Advanced Observability: Design and implement comprehensive observability using OpenTelemetry and leading log management/APM platforms (e.g., Datadog, Splunk, Elastic, Grafana Loki). You will ensure we have deep, distributed tracing and visibility across microservices, Temporal workflows, and AI agents.
  • Infrastructure as Code (IaC): Fully automate our cloud infrastructure provisioning and management using Terraform CDK, ensuring everything is version-controlled, auditable, repeatable, and easily portable across cloud providers.
  • CI/CD Pipeline Design: Build and maintain fast, automated deployment pipelines using [GitHub Actions / Azure DevOps] with automated security, testing, and compliance guardrails built-in.
  • AI Infrastructure: Collaborate closely with our Data and AI engineering teams to provision and optimize secure infrastructure for Agentic AI workflows (e.g., managing GPU compute, vector databases, and AI model deployments).
  • Cloud FinOps: Actively monitor, analyze, and optimize infrastructure costs without sacrificing performance, security, or reliability.

About you:

  • Experience: 7+ years of experience in DevOps, Cloud Infrastructure, or Platform Engineering, with a proven track record of leading large-scale architectural migrations.
  • Cloud & Agnosticism: Deep expertise in AWS (EKS, networking, IAM), paired with a strong architectural mindset for building portable, cloud-agnostic systems that abstract away underlying provider dependencies.
  • Security & Compliance: Proven experience designing and managing cloud environments subject to rigorous compliance audits (SOC 2, GDPR, etc.). You know how to secure a Kubernetes cluster and manage sensitive data at scale.
  • Kubernetes Expert: Extensive hands-on experience designing, deploying, securing, and managing production Kubernetes clusters.
  • Modern IaC: Strong proficiency in Infrastructure as Code. Experience with Terraform CDK (using TypeScript or Python) is highly preferred.
  • Observability Champions: Deep understanding of distributed tracing, metrics, and logging. Hands-on experience with OpenTelemetry and major observability platforms.
  • Programming Skills: You are a strong coder—not just a scripter. Proficiency in TypeScript, Python, or Go is required to effectively use CDKTF and support our engineering teams.
  • Database Knowledge: Experience with database migrations and optimizing a mix of relational (PostgreSQL, MySQL) and modern NoSQL databases for high-throughput, data-intensive workloads.
  • Bonus Points
  • Hands-on experience with Temporal.io or similar workflow orchestration engines (e.g., Airflow, Cadence).
  • Experience supporting AI/ML infrastructure, MLOps, or running LLM agents in production.
  • Background in B2B SaaS, fintech, or procurement data processing.

What we offer:

  • Relish offers a supportive and inclusive work culture where your contributions matter. 
  • Remote first
  • Annual Company Summit 

Salary offering: $60-$75k  Base Salary

This salary range represents a good faith and reasonable estimate of the range of possible compensation for this role at the time of posting, and Relish may ultimately pay more or less than the posted range.  The final salary for this position will be determined in Relish’s sole discretion, consistent with applicable law, and based on a variety of factors, including but not limited to the employee’s work experience, skills, and qualifications for the role, as well as the needs of Relish’s business and other operational considerations.

Conclusion:

Relish is an equal opportunity employer who encourages applications from all qualified applicants.   We thank all applicants for their interest; however, only short-listed candidates will be contacted.