$50–100/hr · micro1
You create complex coding problems that test AI systems on realistic software engineering tasks like debugging, feature building, and performance optimization.
What you would do
- Design problems that demand skills like isolating bugs in legacy systems, adding features to existing codebases, refactoring for readability, or optimizing algorithmic performance
- Set up reproducible environments including starter code, test cases, and build configurations so problems are consistent across repeated attempts
- Write deterministic verifiers that accept any correct solution strategy while catching incomplete implementations and hard-coded answers
- Provide gold-standard reference solutions demonstrating the intended approach and explaining key design decisions
- Iterate on problem design based on feedback, adjusting difficulty and verification logic to ensure problems remain challenging but fair
Who they want
- Strong practical experience in at least one programming language such as Python3, Java, Rust, Go, C++, or TypeScript, with demonstrated ability solving complex problems
- Solid grasp of algorithms, data structures, and software engineering fundamentals including refactoring patterns and performance analysis
- Proven track record debugging complex systems, resolving subtle bugs, and delivering effective performance improvements in production codebases
- Experience modernizing legacy systems and refactoring large codebases to improve maintainability and design quality
- Strong documentation and communication skills for articulating technical rationale in writing
Main skills
What the interview asks about
1.Crafting problems with multiple valid solutions
AI models learn depth when problems admit diverse correct approaches; a single-solution problem encourages pattern memorization rather than genuine reasoning.
For example: “You design a sorting optimization problem with two valid approaches: quicksort with custom pivoting or radix sort. How do test cases verify both solutions are truly optimized, not hardcoded?”
2.Distinguishing between correct and hard-coded approaches
A clever AI can pass weak test suites by memorizing patterns or outputting hard-coded values; your verifier must make this impossible.
For example: “Your problem asks to find the longest increasing subsequence. A contestant submits code that returns 5 only when the input array has length 10 and sum 55. How does your verifier catch this without being brittlely specific?”
3.Handling environment reproducibility
Non-deterministic test failures or setup-dependent bugs break AI training; the problem must work identically across runs and configurations.
For example: “Your problem builds a codebase with a makefile. When run on different machines or with different compiler versions, the behavior varies slightly. How do you debug and fix the environment setup?”
4.Writing useful reference solutions
The golden implementation teaches reviewers and the development team what correctness looks like; a sloppy reference is worse than useless.
For example: “You write a reference solution to a refactoring problem that makes code more readable while preserving behavior. How do you document your refactoring decisions so reviewers understand your approach was intentional?”
5.Assessing problem difficulty balance
A trivial problem wastes training resources; an impossible one derails learning. You must test and adjust to find the sweet spot.
For example: “You set a time limit of 1 second. Your reference solution runs in 0.2 seconds, but contestant submissions average 2 seconds timeout rate. What do you investigate and change?”
A task you may get
Design a small software engineering problem (e.g., debug a race condition, add a feature to a codebase, optimize an algorithm), build a reproducible starter environment, and write a verifier that distinguishes correct from incorrect approaches.
How to prepare
- Review 2-3 recent Codeforces editorial solutions to competitive problems; study how they document approach, complexity analysis, and potential pitfalls
- Practice writing test suites for your own code; focus on edge cases, performance boundaries, and detecting common mistake patterns
- Study examples of legacy code or performance bottlenecks in open-source projects to understand realistic engineering problems worth testing
The facts
- Pay
- $50–100/hr
- Open to
- Bangladesh, Hong Kong, India, Indonesia, Japan, Kazakhstan, Kyrgyzstan, Malaysia, Pakistan, Philippines, Singapore, Sri Lanka, Taiwan, Thailand, Uzbekistan, Vietnam, Austria, Belarus, Belgium, Denmark, France, Germany, Greece, Italy, Netherlands, Portugal, Russia, Spain, Switzerland, United Kingdom, Argentina, Brazil, Chile, Colombia, Mexico, Peru, Algeria, Bahrain, Egypt, Iraq, Jordan, Kuwait, Lebanon, Libya, Morocco, Oman, Palestine, Qatar, Saudi Arabia, Tunisia, United Arab Emirates, United States, Canada, Nigeria, Kenya, South Africa, Ghana, Ethiopia
- Field
- Software Engineering
- Role type
- Expert
- Posted
- 9/7/2026
- Places left
- 100
We wrote this page from the public micro1 listing. It may be out of date, so read the full posting before you apply.