$75–110/hr · Mercor · Full time, 40 hours a week
An infrastructure engineer with hands-on Kubernetes production experience who writes and grades assessments for cloud infrastructure training.
What you would do
- Author troubleshooting tasks for cluster failures, service integration problems, and CI/CD pipeline failures
- Grade and evaluate infrastructure solutions, flagging weak assumptions and architectural oversights
- Design rubrics assessing Terraform/CDK code quality, state management practices, and deployment risk
- Teach guidance to teams on closing knowledge gaps in cloud infrastructure reasoning
- Collaborate with peers to validate assessment consistency and real-world relevance
Who they want
- 4+ years production experience in Kubernetes operations, DevOps, platform engineering, or SRE at large companies
- Hands-on cluster operations: diagnosed and repaired failures, not just written manifests or used managed services
- Production Terraform or AWS CDK experience and live infrastructure-as-code ownership
- Direct AWS service integration: Lambda, API Gateway, DynamoDB production exposure required
- Built and maintained CI/CD pipelines; demonstrate career progression and technical depth
What the interview asks about
1.Cluster failure root cause analysis
Production clusters fail silently in ways Kubernetes docs don't explain; separating symptom from cause tests operator maturity.
For example: “A pod restarts every 90 seconds but logs show no errors. Describe the first three diagnostic steps you'd include in an assessment rubric for identifying whether this is CPU throttling, memory pressure, or a liveness probe configuration issue.”
2.IaC state management pitfalls
Terraform and CDK state files hide real infrastructure problems until deployment; rubrics must capture where engineers underestimate risk.
For example: “A Terraform module drifts after manual changes to a load balancer security group. How would you structure a rubric to assess whether a candidate understands drift detection, remote state locking, and when to manually reconcile versus regenerate?”
3.AWS service integration failure modes
Lambda cold starts, API Gateway rate limits, and DynamoDB hot partitions interact in ways that only surface at scale.
For example: “An API with Lambda, API Gateway, and DynamoDB throttles after load increases. How would you assess whether a candidate systematically checks throughput limits and partition keys?”
4.CI/CD pipeline resilience and idempotence
Broken CI/CD pipelines can hide infrastructure issues for weeks; assessment must test whether engineers understand failure isolation and retry logic.
For example: “A CI/CD pipeline succeeds inconsistently on the same commit. Design a rubric item that evaluates how a candidate would troubleshoot whether the problem is race conditions, external service timeouts, or state pollution from prior runs.”
5.Trade-offs in architecture choices
Production engineers know that cost, latency, and reliability trade against each other; beginners miss the costs of their choices.
For example: “For a new service, would you use a dedicated RDS instance or Aurora Serverless? Write two rubric items: one checking cost estimation and one checking when serverless cold starts become unacceptable.”
A task you may get
Given a cluster symptom (e.g., persistent pod eviction), an IaC code snippet, and a CI/CD failure, author a diagnostic rubric with criteria for what a candidate must check to rule out root causes systematically.
How to prepare
- Review a recent Kubernetes incident from your production experience; list the diagnostic steps that would catch it in a rubric
- Audit a Terraform or CDK module you've written for state management, idempotence, and failure recovery edge cases
- Document a past CI/CD pipeline failure and the checks that would surface the root cause before it reaches production
The facts
- Pay
- $75–110/hr
- Hours
- Full time, 40 hours a week
- Where
- Remote · United States (Remote)
- Open to
- USA
- Field
- Software Engineering
- Posted
- 8/24/2026
- Places left
- 10
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.