$60–95/hr · micro1
You design and optimize GPU kernels and shaders to accelerate AI model training, balancing performance, architecture constraints, and code maintainability in low-level GPU software.
What you would do
- Implement GPU kernels in CUDA, WebGPU, or GLSL that execute AI training operations with high throughput
- Profile GPU code with profiling tools to identify bottlenecks and measure optimization impact
- Refactor kernels to improve performance by tuning memory access patterns, warp configurations, and synchronization
- Develop C++ host code that manages GPU memory, coordinates kernel launches, and handles error conditions
- Test and validate GPU solutions for correctness and performance across different hardware architectures
Who they want
- Advanced experience with GPU programming languages (CUDA, WebGPU, GLSL) on NVIDIA or compatible hardware
- Background in graphics programming, machine learning acceleration, scientific computing, or high-performance computing with GPUs
- Strong proficiency in C++ and understanding of systems-level programming, memory models, and concurrency
- Demonstrated ability to profile, debug, and optimize GPU code for production performance requirements
- Willingness to start within 24-48 hours of onboarding and commit to steady task output
Main skills
What the interview asks about
1.Kernel performance profiling and optimization
Profilers reveal where GPU time actually goes; optimizing the wrong target wastes effort, so your skill at interpreting profiler data and iterating on hotspots directly impacts training speedup.
For example: “You profile a CUDA kernel and notice memory bandwidth is saturated but compute utilization is only 60%. What architectural changes would you investigate, and why might low occupancy still be causing the bottleneck?”
2.Warp divergence and GPU execution model
Warp divergence silently degrades throughput without errors, so understanding conditional branching impact is essential to avoid shipping slow-looking-correct code.
For example: “You have a kernel where threadIdx.x threads take one code path and threadIdx.x threads take another. Explain the performance penalty this causes and how you'd restructure the code to avoid it.”
3.GPU memory hierarchy and data movement
Memory bottlenecks often dominate GPU performance for AI workloads, so decisions about coalescing, shared memory, and host-device transfers directly determine speedup achievable.
For example: “Your kernel loads a 2D matrix row-by-row into a 1D thread grid, leading to uncoalesced memory reads. Redesign the memory access pattern to improve cache efficiency and explain the tradeoff.”
4.Host-device synchronization and error handling
Silent GPU errors or subtle synchronization bugs can corrupt training data or cause deadlocks, making robust host integration code critical for model reliability.
For example: “Your C++ code launches 100 kernels in a loop without checking return codes. What could go wrong, and how would you add proper error checking and synchronization?”
5.Cross-architecture GPU optimization
Optimal kernel parameters and strategies vary across GPU architectures, so generalizing across NVIDIA and non-NVIDIA hardware requires architectural awareness.
For example: “A kernel optimized for NVIDIA A100 with large shared memory achieves poor occupancy on a smaller GPU. How would you restructure the kernel to scale to different compute capability levels?”
A task you may get
Write a CUDA kernel that multiplies two 1024x1024 matrices using shared memory tiling, then profile it against the naive approach and document the performance gain with optimization rationale.
How to prepare
- Implement 2-3 small CUDA kernels from scratch (e.g., reduction, matrix transpose) and profile them to identify optimization opportunities
- Study NVIDIA GPU architecture documentation for a specific generation and understand the compute capability, shared memory, warp scheduler behavior, and cache hierarchy
- Review 1-2 open-source GPU optimization examples (e.g., from CUTLASS or similar libraries) and trace through how memory access patterns and thread hierarchies are tuned
- Practice writing C++ GPU host code that handles memory allocation, kernel launches, synchronization, and error checking for multiple kernel sequences
The facts
- Pay
- $60–95/hr
- Open to
- Bangladesh, Hong Kong, India, Indonesia, Japan, Kazakhstan, Kyrgyzstan, Malaysia, Pakistan, Philippines, Singapore, Sri Lanka, Taiwan, Thailand, Uzbekistan, Vietnam, Austria, Belarus, Belgium, Denmark, France, Germany, Greece, Italy, Netherlands, Portugal, Russia, Spain, Switzerland, United Kingdom, Argentina, Brazil, Chile, Colombia, Mexico, Peru, Algeria, Bahrain, Egypt, Iraq, Jordan, Kuwait, Lebanon, Libya, Morocco, Oman, Palestine, Qatar, Saudi Arabia, Tunisia, United Arab Emirates, United States, Canada, Nigeria, Kenya, South Africa, Ghana, Ethiopia
- Field
- Software Engineering
- Role type
- Specialist
- Posted
- 7/29/2026
- Places left
- 25
We wrote this page from the public micro1 listing. It may be out of date, so read the full posting before you apply.