Training Turk

Trainium (NKI) Kernel Expert

$70–90/hr · Mercor · Hourly, 40 hours a week

You evaluate NKI kernel implementations for AWS Trainium, assessing migration fidelity, performance optimization, and numerical correctness.

What you would do

  • Review developer-written NKI kernels for correctness and hardware-appropriate design choices
  • Assess quality of CUDA-to-NKI code migrations and identify optimization opportunities
  • Validate cross-platform numerical equivalence under different accumulation and precision scenarios
  • Provide detailed written feedback explaining architectural issues and performance bottlenecks

Who they want

  • 2+ years of hands-on NKI kernel development targeting AWS Trainium or Inferentia2 hardware
  • Deep familiarity with NKI design patterns including tile-based execution models, SBUF/PSUM/HBM resource management, and partition-dimension constraints
  • Demonstrated experience evaluating CUDA migrations and profiling Trainium performance characteristics
  • Familiarity with numerical accuracy standards across mixed-precision and platform differences

Main skills

Nki kernel developmentCuda to nki migrationAWS trainium optimization

What the interview asks about

  1. 1.CUDA migration fidelity assessment

    Evaluating whether developers correctly adapted GPU algorithms to Trainium's tile-based architecture while preserving numerical results is core to your role.

    For example: “You review a developer's NKI port of a CUDA matrix multiplication kernel. Output matches GPU results on small tensors but diverges on large tensors. What hardware or numerical precision issues would you investigate first?”

  2. 2.Performance bottleneck identification

    Identifying whether kernel code makes efficient use of Trainium's compute hierarchy and memory systems directly affects model training speed and cost.

    For example: “A developer's NKI kernel achieves 40% of theoretical tensor-engine throughput. What NeuronCore pipeline metrics would you profile to pinpoint whether the bottleneck is memory bandwidth, DMA orchestration, or compute scheduling?”

  3. 3.Memory hierarchy optimization

    Ensuring kernels properly allocate SBUF, PSUM, and HBM resources within Trainium's partition constraints is essential for fitting workloads on hardware.

    For example: “A developer is porting a convolution that needs 2x the available SBUF. What Trainium-specific design changes would you suggest, and how would you weigh tradeoffs between tile size and memory access patterns?”

  4. 4.Numerical correctness validation

    Verifying implementations account for GPU-Trainium differences in accumulation order, rounding modes, and mixed-precision semantics ensures training stability.

    For example: “A developer's BF16 accumulation kernel passes unit tests on GPU but fails on Trainium with 1e-3 weight differences. What precision-specific factors in Trainium architecture might explain this divergence?”

  5. 5.Rubric-based technical assessment

    Communicating findings clearly through structured feedback helps developers understand immediate fixes and underlying architectural principles.

    For example: “You identify three kernel issues: a memory allocation problem, suboptimal DMA usage, and partition constraint violations. How would you structure your assessment to help developers prioritize and understand the architectural concepts?”

A task you may get

Analyze a provided CUDA kernel and NKI port, then write a detailed assessment covering migration correctness, hardware utilization, numerical accuracy, and specific optimization recommendations with profiling rationale.

How to prepare

  • Review AWS Trainium architecture docs, NeuronCore-v2 specifications, and NKI programming model patterns
  • Practice analyzing CUDA and NKI code side-by-side to identify common migration pitfalls and optimization opportunities
  • Study accumulation order, rounding behavior, and mixed-precision semantics differences between GPU and Trainium
  • Prepare examples of performance bottlenecks and optimization strategies for typical ML training workloads

The facts

Pay
$70–90/hr
Hours
Hourly, 40 hours a week
Where
Remote · Remote — United States
Open to
USA
Field
Software Engineering
Posted
8/25/2026
Places left
3

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.