$70–150/hr · Mercor · Part time
You create infrastructure scenarios and evaluate how AI models reason about deployment, cloud operations, and containerization challenges.
What you would do
- Design realistic infrastructure and deployment challenges that test whether AI models understand CI/CD pipelines and containerization at a reasoning level
- Evaluate AI-generated solutions against your expertise in cloud architecture, Kubernetes, and operational best practices
- Create detailed problem specifications with pass-fail criteria for infrastructure scenarios across different cloud platforms
- Provide structured feedback on AI reasoning about orchestration, scaling, and system reliability issues
- Contribute domain insights to ongoing task refinement and benchmark development
Who they want
- Professional experience with CI/CD pipelines, cloud infrastructure on AWS/GCP/Azure, and containerization using Docker and Kubernetes
- Strong technical communication skills and ability to explain infrastructure concepts clearly
- Comfort working independently and managing time across multiple AI evaluation projects
- Experience writing code and infrastructure specifications in Python or similar languages
- Ability to design problems you cannot shortcut but can judge rigorously
Main skills
What the interview asks about
1.Infrastructure reasoning vs. memorization
You must distinguish between AI systems that retrieve infrastructure facts and those that genuinely reason through deployment trade-offs and failure scenarios.
For example: “You design a task about a microservice outage cascading through dependent services. What questions would you ask the AI model to verify it understands root cause analysis vs. just listing common failure patterns?”
2.Realistic problem scenario design
Training data quality depends on your scenarios reflecting actual engineering challenges; contrived examples teach AI systems to pattern-match instead of reason.
For example: “Design a problem where a Kubernetes cluster needs to handle a traffic spike but the CI/CD pipeline has a bottleneck. How do you specify this so an AI model must reason about the interaction between orchestration and deployment speed?”
3.Cloud architecture and trade-offs
Cloud infrastructure involves constant trade-offs between cost, latency, and reliability; your scenarios must surface these tensions so AI learns to reason about them.
For example: “Create a task about choosing between three AWS deployment strategies for a stateful service. What measurable criteria would you use to evaluate whether an AI model understood the trade-offs vs. just guessed?”
4.Specification precision and testability
Vague problem specifications let AI systems generate plausible-sounding but untestable solutions; your specs must be concrete enough to definitively score correct answers.
For example: “Write a problem specification for container image optimization that includes both the constraints and the metrics you would use to score an AI model proposal.”
5.Independent evaluation judgment
You work alone on these tasks; you need to confidently judge whether an AI solution is sound without a team to verify your scoring decisions.
For example: “An AI model proposes a solution to a containerization problem that works but uses an unconventional approach you have never seen. How would you evaluate whether this is valid reasoning or just lucky pattern-matching?”
A task you may get
Design a CI/CD challenge where an AI model must identify why deployments are failing; include the problem statement, success criteria, and at least two example scenarios.
How to prepare
- Document a recent infrastructure incident from your practice; analyze what reasoning steps an AI model would need to reach the same diagnosis.
- Gather 3-4 container orchestration or deployment scenarios you have encountered; identify the core reasoning patterns that distinguish good solutions from bad ones.
- Prepare technical specifications for a Kubernetes scaling problem; practice writing constraints and success criteria that can be objectively scored.
The facts
- Pay
- $70–150/hr
- Hours
- Part time
- Where
- Remote
- Field
- Software Engineering
- Role type
- Talent network
- Posted
- 2/20/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.