$90–120/hr · Mercor · Hourly, 40 hours a week
An MLOps engineer designs and solves complex problems in LLM infrastructure, focusing on GPU kernels, performance profiling, distributed debugging, and inference serving.
What you would do
- Design challenging, technically rigorous problems in GPU kernel optimization, performance profiling, debugging, and inference serving
- Write clear, correct solutions that demonstrate best practices and performance considerations
- Evaluate engineering tasks and solutions, providing detailed feedback on correctness and efficiency
- Develop evaluation rubrics and guidelines covering kernel optimization, profiler interpretation, and serving trade-offs
- Collaborate with research and engineering teams to improve ML systems and model performance
Who they want
- 2+ years hands-on work with ML infrastructure, systems engineering, LLM serving, or GPU acceleration
- Practical experience with GPU kernels (CUDA, Triton), performance profiling (Kineto, torch.profiler, Nsight), or LLM serving frameworks
- Production experience with PyTorch or JAX, ideally with framework-level depth in distributed training or custom operators
- Familiarity with modern accelerators (A100, H100, B200, TPU) and ability to reason about throughput and memory trade-offs
- Full-time availability, 40 hours per week, no conflicting engagements
Main skills
What the interview asks about
1.GPU kernel optimization
Writing efficient CUDA or Triton kernels requires understanding hardware constraints, memory hierarchies, and occupancy; evaluators need to verify both functional correctness and performance awareness.
For example: “A correct CUDA matrix multiply kernel on A100 achieves 50% of peak FLOPS. What profiler metrics would you check to identify the performance bottleneck?”
2.Performance profiling interpretation
Many engineers can run profilers but few correctly interpret the output to identify root causes; this skill separates competent infrastructure engineers from expert ones.
For example: “Training throughput drops 30% with 4-GPU distributed data parallel, showing high all-reduce overhead. How do you determine if this is expected or indicates an implementation bug?”
3.Distributed systems reasoning
Debugging issues in distributed training or serving involves reasoning about synchronization, failure modes, and communication patterns that don't appear in single-machine code.
For example: “A candidate proposes a continuous batching strategy for serving that processes requests with different sequence lengths by padding to the maximum. How would you evaluate the KV cache impact and memory efficiency of this approach compared to ragged batching?”
4.LLM serving architecture
Modern serving systems optimize for throughput and latency through techniques like KV cache, paged attention, and continuous batching; evaluators assess whether candidates understand these trade-offs.
For example: “You're comparing two serving approaches: one that batches all available requests every 10ms, and another that processes requests individually as they arrive. What metrics matter most when evaluating which approach the company should use?”
5.Framework-level ML systems knowledge
Evaluating infrastructure solutions requires understanding how PyTorch or JAX executes graphs, manages memory, and schedules operations across devices.
For example: “A candidate optimizes a model's backward pass by using gradient accumulation across multiple steps before all-reduce. What should you verify about their understanding of gradient computation order and communication patterns?”
6.Technical communication and rubric design
Creating evaluation frameworks requires translating vague performance goals into measurable criteria that others can apply consistently.
For example: “You're writing a rubric for evaluating submissions on CUDA kernel optimization. What specific performance metrics, code quality criteria, and edge case handling would you include to distinguish a correct basic solution from an expert-level optimized kernel?”
A task you may get
Design 2-3 MLOps infrastructure problems covering GPU kernels, profiling, and serving; write reference solutions; and create a detailed evaluation rubric with concrete performance expectations.
How to prepare
- Study GPU architecture and CUDA optimization techniques, focusing on occupancy, memory bandwidth, and common bottlenecks on A100 or H100
- Practice interpreting profiler output from PyTorch (torch.profiler) and JAX; run performance traces on real model training and serving workloads
- Research LLM serving systems like vLLM and SGLang; understand KV cache management, continuous batching, and paged attention
- Review distributed training frameworks (FSDP, DeepSpeed) and understand communication patterns, gradient synchronization, and scaling efficiency
The facts
- Pay
- $90–120/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote · Canada, UK, US
- Open to
- CAN, GBR, USA, USA
- Field
- Software Engineering
- Posted
- 9/15/2026
- Places left
- 10
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.