$60–100/hr · micro1
Expert GPU programmer optimizing CUDA kernels and high-performance computing systems for maximum computational efficiency.
What you would do
- Profile GPU kernels using CUDA and profiling tools to identify performance bottlenecks
- Refactor CUDA and C++ codebases to improve efficiency and adapt to diverse GPU architectures
- Implement graphics and compute workflows using GLSL and WebGPU
- Collaborate with stakeholders to evaluate GPU-based approaches and performance optimization strategies
Who they want
- Demonstrated track record optimizing GPU kernels and tuning CUDA code for performance
- Advanced C++ expertise in high-performance computing environments
- Hands-on experience with GLSL, WebGPU, and graphics compute shaders
- Advanced skill with performance profiling tools like Nsight and Visual Profiler for kernel analysis
- Strong analytical ability to evaluate kernel performance across hardware generations
Main skills
What the interview asks about
1.Kernel bottleneck identification
Effective optimization requires pinpointing whether limitations are memory bandwidth, arithmetic throughput, latency, or other factors.
For example: “You're optimizing a CUDA kernel that achieves only 40 percent of theoretical peak throughput. Describe your approach to determine what specific resource is limiting performance.”
2.GPU memory hierarchy optimization
Memory efficiency often matters more than algorithm cleverness for achieving actual speedup on real hardware.
For example: “A kernel reads 1000 elements per block with scattered access patterns. Explain the memory optimization strategy you'd employ to reduce bandwidth waste.”
3.Cross-GPU architecture adaptation
Code that performs well on one GPU generation may suffer significant slowdowns on different architectures without careful redesign.
For example: “Your highly optimized kernel for A100 GPUs shows poor performance on H100 hardware. What GPU-specific factors would you investigate to improve portability?”
4.Graphics API integration with compute
GLSL and WebGPU integration requires understanding GPU state and scheduling implications for mixed graphics-compute workloads.
For example: “You need to integrate a GLSL compute shader into an existing graphics pipeline while minimizing synchronization overhead. Describe your design approach.”
5.Performance measurement and reporting
Credible optimization claims require rigorous measurement methodology and clear communication of tradeoffs made.
For example: “You achieved 3x speedup on one kernel but increased register pressure on another. How would you document these changes for stakeholder review?”
A task you may get
Optimize a sample CUDA kernel that has obvious inefficiencies, documenting profiling data, changes made, and performance improvements achieved.
How to prepare
- Review NVIDIA GPU architecture details and memory hierarchy optimization
- Study CUDA profiler output interpretation and performance metric analysis
- Practice identifying bottlenecks in example kernels using profiling tools
- Explore optimization techniques across multiple GPU generations
The facts
- Pay
- $60–100/hr
- Open to
- Bangladesh, Hong Kong, India, Indonesia, Japan, Kazakhstan, Kyrgyzstan, Malaysia, Pakistan, Philippines, Singapore, Sri Lanka, Taiwan, Thailand, Uzbekistan, Vietnam, Austria, Belarus, Belgium, Denmark, France, Germany, Greece, Italy, Netherlands, Portugal, Russia, Spain, Switzerland, United Kingdom, Argentina, Brazil, Chile, Colombia, Mexico, Peru, Algeria, Bahrain, Egypt, Iraq, Jordan, Kuwait, Lebanon, Libya, Morocco, Oman, Palestine, Qatar, Saudi Arabia, Tunisia, United Arab Emirates, United States, Canada, Nigeria, Kenya, South Africa, Ghana, Ethiopia
- Field
- Software Engineering
- Role type
- Specialist
- Posted
- 8/11/2026
- Places left
- 50
We wrote this page from the public micro1 listing. It may be out of date, so read the full posting before you apply.