$70–90/hr · Mercor · Hourly, 40 hours a week
Applied ML specialist who evaluates the correctness and methodological rigor of machine learning tasks developed to train advanced AI models.
What you would do
- Evaluate applied ML experiment quality, design rigor, and methodological soundness.
- Evaluate the design of experiments, model selection rationale, and assessment methodology
- Identify flaws in data quality practices, including leakage, metric gaming, and validation errors.
- Provide clear, rubric-based written feedback explaining technical issues and their impact.
- Reproduce results and verify technical claims against available evidence.
Who they want
- 3+ years hands-on applied and experimental ML work with strong experiment design skills.
- Experience with popular ML frameworks including PyTorch, TensorFlow, scikit-learn, XGBoost
- Deep comprehension of data quality standards including leakage detection, metric gaming, train-test-CV best practices
- Capability to evaluate ML assertions against evidence and reproduce experimental findings
- Preferred: Kaggle or competition experience, graduate research or publications in applied ML.
Main skills
What the interview asks about
1.Detecting data leakage in experiments
Leakage produces false performance estimates and renders conclusions invalid; specialists must spot it immediately.
For example: “An author reports 94% accuracy on a time-series prediction task using a model trained on data from 2020-2023, validated on 2024 data, but they performed feature scaling on the entire dataset before splitting. Is this approach sound?”
2.Evaluating train-test split validity
Improper splits are among the most common ML errors; they invalidate reported performance.
For example: “A researcher uses stratified k-fold cross-validation with k=5, but applies hyperparameter tuning on the full dataset before splitting. What's the flaw and what's the impact?”
3.Judging evaluation metrics appropriateness
Metrics can mask real problems; specialists must assess whether reported metrics actually measure model capability.
For example: “For an imbalanced classification task with 2% positive class, the model achieves 98% accuracy but 0.3 AUC-ROC. How do you interpret this and what would you recommend?”
4.Assessing statistical rigor in results
Results need statistical validation; a good model on one fold might fail under proper cross-validation.
For example: “An experiment shows improvement of 3.2% on one test set. How would you determine whether this is a real finding or random variation?”
5.Critiquing model architecture choices
Authors must justify framework and design choices; unjustified complexity often hides overfitting or engineering errors.
For example: “An author uses a deep neural network with 15 layers for a tabular dataset with 30 features and 5000 samples. What would you ask them to justify?”
A task you may get
Review a short experimental ML report (2-3 pages) with intentional methodological flaws in experiment design, data handling, or evaluation, and write a rubric-based assessment identifying issues and explaining their implications for result validity.
How to prepare
- Review a recent ML paper in your specialty area and think through what methodological checks you would perform.
- Practice reproducing results from a published ML experiment to understand where leakage and validation errors commonly hide.
- Prepare examples from Kaggle competitions or your own work where you caught (or made) a subtle data leakage or validation error.
- Study common ML pitfalls in experiment design: temporal leakage, information leakage from preprocessing, and incorrect cross-validation schemes.
The facts
- Pay
- $70–90/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote · Remote — United States
- Open to
- USA
- Field
- Data Analysis
- Posted
- 8/28/2026
- Places left
- 3
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.