$70–90/hr · Mercor · Hourly, 40 hours a week
DevOps engineer who audits Kubernetes scenario quality and provides detailed feedback for AI model training.
What you would do
- Review Kubernetes manifests, Helm charts, and cluster configuration scenarios for correctness and production readiness
- Assess task scenarios for authentic failure modes and troubleshooting challenges from real clusters
- Evaluate model outputs and proposed solutions against cluster best practices and reliability patterns
- Provide rubric-based written feedback identifying gaps in scenario design or model reasoning
- Contribute realistic scenario ideas grounded in your own incident response and platform-engineering experience
Who they want
- 3+ years hands-on production Kubernetes experience with EKS, GKE, AKS, or self-managed clusters
- Deep understanding of cluster internals: container networking, persistent storage, RBAC, and common failure modes like CrashLoopBackOff and resource eviction
- Demonstrated experience authoring and reviewing manifests and Helm charts, plus debugging live cluster incidents
- Proficiency in Go, Python, or TypeScript for understanding infrastructure-as-code and evaluation logic
- CKA or CKAD certification, service-mesh experience, and observability tools (Prometheus/Grafana) strongly preferred
Main skills
What the interview asks about
1.Evaluating manifest quality
Many manifest issues are silent until production stress: missing resource limits, incorrect RBAC, bad networking assumptions. You must spot these and explain why they matter.
For example: “Review a database deployment manifest with no resource limits, missing anti-affinity rules, and generic service selector. Identify three production risks and explain how an AI model should evaluate this.”
2.Assessing failure-mode authenticity
Generic 'pod fails to start' scenarios don't teach real debugging. You need to recognize which failure modes matter and how they appear in actual clusters.
For example: “A task has a pod in CrashLoopBackOff; the model correctly finds a missing env var. Is this enough for SRE? What additional diagnostics or reasoning would make this task valuable for real-world work?”
3.Cluster networking understanding
Networking problems are common and subtle. You must evaluate whether scenarios test genuine understanding of CNI, DNS, service routing, and ingress.
For example: “Design a networking task where a service works internally but external traffic fails. What issues (routing, policy, DNS) would you embed? How do you force network-layer reasoning, not just syntax recitation?”
4.Storage and data integrity scenarios
Data persistence is mission-critical. Scenarios must realistically test understanding of PV/PVC, stateful sets, and failure recovery without data loss.
For example: “Create a task where a pod using a PVC must be rescheduled to a different node, and the model must reason about volume availability, persistent storage claims, and whether data survives. What edge cases and failure modes should this task probe?”
5.Rubric design and assessment clarity
Your judgment about task quality must be transparent and reproducible so researchers can improve systematically.
For example: “Write a rubric for evaluating whether an AI model understands RBAC sufficiently. What specific behaviors, scenarios, or reasoning patterns would indicate mastery? What gaps or misconceptions should the rubric detect?”
6.Observability and troubleshooting depth
Real cluster operators use logs, metrics, and events to debug. Tasks must test whether the model can extract and interpret observability signals to troubleshoot.
For example: “A task has partial metrics, logs, and events from a degraded cluster. What signals must an AI model interpret to determine if the cause is resource exhaustion, networking, or application? Design for multi-signal reasoning.”
A task you may get
Review three Kubernetes scenarios (manifest review, diagnosis, security audit). Identify what makes each realistic or unrealistic, what the AI model needs to know, and how to improve them.
How to prepare
- Gather 3-4 recent Kubernetes incidents from your own experience and document the root cause, debugging steps, and the manifest/configuration issues involved
- Review a Helm chart or complex Kubernetes YAML and write down the potential production failure modes and best-practice violations you notice
- Prepare specific examples of cluster failure modes you've seen (pod evictions, DNS issues, storage mount failures) and what the correct diagnosis and fix looks like
- Read the Kubernetes documentation on failure modes and cluster internals to refresh your mental model of what real operators need to know
The facts
- Pay
- $70–90/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote · Remote — United States
- Open to
- USA
- Field
- Software Engineering
- Posted
- 8/28/2026
- Places left
- 3
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.