Training Turk

LLM Research Scientist (Pre-training & Computer Vision & Adversarial Robustness)

$100–120/hr · Mercor · Hourly, 40 hours a week

An ML research scientist who trains and improves deep learning models across vision and language, with expertise in robustness, compression, and efficient training.

What you would do

  • Train image classifiers and generative models (diffusion, GANs, VAEs) from scratch using PyTorch or JAX
  • Fine-tune open-weight language models and implement post-training methods like supervised fine-tuning and preference optimization
  • Apply adversarial training and robustness evaluation techniques to harden models against adversarial inputs and adversarial conversations
  • Compress models through quantization, pruning, and knowledge distillation to deploy under hard size and latency constraints
  • Diagnose training issues and optimize learning efficiency to achieve target performance within resource budgets

Who they want

  • 3+ years of machine learning research experience, with PhD research or publications demonstrating research output
  • Strong hands-on expertise with PyTorch, JAX, TensorFlow, or similar deep learning frameworks
  • Domain depth in at least one area: adversarial robustness, efficient vision, generative modeling, or LLM post-training
  • Graduation from a top-100 institution, background with FAANG or comparable AI lab, or equivalent research track record
  • Project-based engagement with flexible hours, bringing end-to-end problem-solving skills and research judgment

Main skills

Adversarial robustness and adversarial trainingModel compression and quantizationGenerative image models and diffusion

What the interview asks about

  1. 1.Adversarial training and robustness

    Robust training involves non-trivial trade-offs and pitfalls (gradient masking, over-regularization); evaluating robustness correctly is essential to avoid false confidence.

    For example: “PGD adversarial training shows 85% robust accuracy, but AutoAttack evaluation shows 71%. Explain what might be happening and how you'd verify whether this is gradient masking.”

  2. 2.Model compression and accuracy trade-off

    Compression techniques (quantization, pruning, distillation) each have different accuracy-efficiency profiles; choosing the right approach and depth requires understanding the model and deployment target.

    For example: “You have a 500M parameter vision model that must run on-device with 50MB memory and 100ms latency. You could quantize to int8, prune 60% of weights, or distill to a 100M student. How would you approach this decision, and what empirical steps would you take?”

  3. 3.Generative model training and quality

    Training from scratch requires managing training instability, mode collapse, and quality metrics; iterating efficiently depends on choosing the right evaluation metric and debugging workflow.

    For example: “Training a diffusion model, FID plateaus at 15 after 100k steps with missing fine-grained details. Increasing capacity or training length doesn't help. What diagnostics would you run and what might change?”

  4. 4.Language model post-training behavior

    Fine-tuning and preference optimization can amplify unintended behaviors (sycophancy, overconfidence, false refusals); shaping behavior precisely while preserving capability requires careful design.

    For example: “You fine-tune a model with DPO to be more helpful, but it now accepts obvious false premises (e.g., 'I know X is true because you said so'). How would you diagnose this regression, and what changes to data or training would you try?”

  5. 5.Training diagnostics and resource optimization

    When performance is suboptimal, diagnosing the bottleneck (data quality, learning rate, model capacity, training length) saves compute and accelerates iteration.

    For example: “Classifier on 10k examples plateaus at 78% accuracy. Budget for 10 experiments. Prioritize: (1) more data, (2) bigger model, (3) tune learning rate, (4) data augmentation. Justify your ranking.”

A task you may get

Implement or outline a training pipeline for either an adversarially robust image classifier or a fine-tuned language model. Specify the loss function, key hyperparameters, evaluation metrics, and diagnostic steps you'd use to identify and fix training issues.

How to prepare

  • Review recent adversarial robustness literature (AutoAttack, trade-off analyses) and hands-on experience with PGD or TRADES training if that's a target area
  • Study model compression approaches (quantization methods, pruning strategies, distillation) and when each applies based on deployment constraints
  • If applicable, review language model fine-tuning and preference optimization methods (DPO, RLHF, RLAIF) with focus on behavioral evaluation
  • Prepare 2-3 detailed case studies from your research: a training challenge you solved, a failure mode you discovered, and what you'd do differently

The facts

Pay
$100–120/hr
Hours
Hourly, 40 hours a week
Where
Remote
Field
Data Analysis
Posted
8/5/2026
Places left
15

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.