$140–200/hr · Mercor · Hourly, 40 hours a week
You design and build data pipelines and infrastructure for a short-term AI research initiative, delivering production-quality systems that support model training and evaluation.
What you would do
- Build and optimize production data pipelines handling ingestion, transformation, and delivery at scale
- Design data warehouse or lakehouse architecture appropriate for research and model-training workloads
- Implement orchestration, scheduling, and monitoring so pipelines run reliably with minimal manual intervention
- Write efficient SQL transformations and Python logic for complex data processing tasks
- Troubleshoot and optimize pipeline performance, data quality, and system reliability under production conditions
Who they want
- Hands-on data engineering experience building and owning production pipelines and systems
- Strong proficiency in SQL and Python with production-code quality standards
- Experience with ETL, ELT, orchestration (Airflow), streaming (Kafka), and data warehousing (Snowflake, BigQuery, Redshift)
- Track record with cloud platforms: AWS, GCP, or Azure in production environments
- Based in the United Kingdom with legal right to work and access during UK business hours, September-October 2026
Main skills
What the interview asks about
1.Pipeline architecture decisions
The right architectural pattern balances speed, reliability, and maintainability; your reasoning about trade-offs reveals engineering maturity and project-fit judgment.
For example: “For a research project needing daily-refreshed feature sets from raw event data, would you recommend ELT (load raw, transform in warehouse) or ETL (transform before load)? Walk through the trade-offs and what project factors would influence your choice.”
2.Production reliability and troubleshooting
Pipelines fail; your approach to monitoring, alerting, and debugging reveals whether you can keep systems running under pressure in a compressed project window.
For example: “You deploy a pipeline that runs smoothly for 48 hours then fails during overnight processing. Describe your approach to diagnosing the root cause and your strategy for preventing similar failures.”
3.Performance optimization
Efficiency matters in research; slow pipelines delay downstream work and waste compute budget; optimizing is an ongoing engineering responsibility, not an afterthought.
For example: “Your SQL transformation on a billion-row fact table takes 45 minutes but research timelines need it in under 10. Walk through how you'd investigate and what optimization strategies you'd consider.”
4.Tool selection and configuration
Different problems call for different tools; your reasoning about whether to use Kafka, Airflow, or dbt for a given task demonstrates hands-on judgment.
For example: “The research team needs to ingest real-time event streams and produce fresh features for model evaluation. What tools and architecture would you propose, and why that combination?”
5.Communication and documentation
Handoff happens in a short project; your ability to explain architecture and maintain documentation ensures the system remains understandable and operable.
For example: “Document your approach to a complex multi-stage transformation pipeline in one page, targeting an audience of data scientists rather than engineers. Emphasize what they need to know about reliability and data freshness.”
A task you may get
You receive specifications for a data pipeline supporting model evaluation, produce an architecture document outlining tool choices, orchestration strategy, and data flow, then walk through how you'd implement and monitor a critical stage.
How to prepare
- Review your own production pipeline work; prepare 2-3 examples where you solved a scale, reliability, or performance problem and what you learned
- Study the cloud platform you use most heavily: refresh your knowledge of serverless options, managed services, and cost optimization patterns
- Research recent advances in data orchestration and modern ELT tools like dbt; understand how they differ from traditional Airflow-based ETL
- Practice articulating architectural trade-offs: when you choose one tool over another, be ready to explain what you're optimizing for
The facts
- Pay
- $140–200/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote
- Open to
- GBR
- Posted
- 9/4/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.