$50–70/hr · micro1
Software engineer who evaluates AI-generated developer workflows by testing them in authentic development environments.
What you would do
- Execute structured developer tasks simulating real workflows with version control and CI/CD
- Assess correctness of AI outputs by hands-on testing in development tools
- Design reproducible test contexts with realistic commit histories and build pipelines
- Configure tool integrations, authentication flows, and API connectors
- Investigate undocumented behaviors through systematic testing
Who they want
- Minimum 3 years professional software engineering with demonstrated Git and code review mastery
- Strong proficiency with CI/CD systems including writing and debugging custom workflows
- Advanced scripting in Python or Bash for automation and environment setup
- Fluency with REST APIs, OAuth, webhooks, and enterprise tools
- Experience with technical QA demonstrating consistency in applying rubrics
Main skills
What the interview asks about
1.Diagnosing CI workflow failures
Tests pass locally but fail in CI - identifying root cause requires systematic log reading and reproducing exact context.
For example: “A GitHub Actions workflow fails on push with a permission error 30 minutes into matrix build. How would you isolate whether the issue is environment caching, token expiry, or service limits?”
2.Evaluating code quality hands-on
AI code might be syntactically correct but subtly wrong - requires running tests and checking for race conditions, leaks, or incomplete error handling.
For example: “An AI script processes CSV files with concurrent workers. How would you test for resource leaks, race conditions on shared files, or silently-ignored exceptions?”
3.Configuring authentication flows
Integrations fail silently due to OAuth scope mismatches or webhook retries - evaluators must trace requests and verify contracts.
For example: “A webhook integration receives only 3 of 10 expected events. How would you investigate endpoint configuration, delivery retries, and platform event filters?”
4.Building reproducible test contexts
Evaluation validity depends on test environment reflecting real conditions - history, merges, and configuration matter.
For example: “Design a Git repository for evaluating an AI's merge conflict resolution. What elements would you include: branches, conflicts, commit messages?”
5.Consistency across evaluations
Applying grading rubrics consistently across many code samples prevents subjective judgment from degrading quality.
For example: “You notice variations in how you evaluate similar code samples across sessions. How would you document and standardize your evaluation approach?”
A task you may get
Set up a GitHub repository with a Python application, configure a custom Actions workflow, commit a subtle bug, then assess whether test output correctly identifies it.
How to prepare
- Study Git internals: branches, refs, and object model; practice explaining diffs and rebasing.
- Write and debug a custom GitHub Actions workflow with environment setup and caching.
- Create a Bash script to set up and tear down test environments idempotently.
- Trace a real API request-response cycle to understand headers, tokens, and error codes.
The facts
- Pay
- $50–70/hr
- Open to
- Bangladesh, Hong Kong, India, Indonesia, Japan, Kazakhstan, Kyrgyzstan, Malaysia, Pakistan, Philippines, Singapore, Sri Lanka, Taiwan, Thailand, Uzbekistan, Vietnam, Austria, Belarus, Belgium, Denmark, France, Germany, Greece, Italy, Netherlands, Portugal, Russia, Spain, Switzerland, United Kingdom, Argentina, Brazil, Chile, Colombia, Mexico, Peru, Algeria, Bahrain, Egypt, Iraq, Jordan, Kuwait, Lebanon, Libya, Morocco, Oman, Palestine, Qatar, Saudi Arabia, Tunisia, United Arab Emirates, United States, Canada, Nigeria, Kenya, South Africa, Ghana, Ethiopia
- Field
- Software Engineering
- Role type
- Generalist
- Posted
- 8/18/2026
- Places left
- 4
We wrote this page from the public micro1 listing. It may be out of date, so read the full posting before you apply.