
Low Level Optimization Engineer
We are looking for exceptional engineers with deep low-level expertise to join our team
in optimizing and scaling next-generation AI systems.
This role sits at the intersection of machine learning, systems, and hardware, and
involves working on highly performance-critical components of our video generation
stack. You will play a key role in driving efficiency, throughput, and latency
improvements across the full software–hardware stack.
What You’ll Do
Design and implement high-performance components for large-scale generative models
Optimize compute-intensive workloads across CPUs, GPUs, and AI accelerators
Develop custom kernels and low-level primitives for critical model paths
Analyze and improve system-level performance (latency, throughput, memory utilization)
Collaborate across ML, systems, and infrastructure teams to co-design efficient solutions
Contribute to scaling real-time video generation pipelines to production environments
Requirements (Must Have)
Strong passion for hardware–software co-optimization
Proven experience in at least one of the following:
Low-level programming (C/C++/CUDA/assembly or similar)
Performance optimization and profiling of complex systems
Hardware-aware software development
SW-aware ASIC / architecture designDeep understanding of system architecture (memory hierarchy, parallelism, compute bottlenecks)
Ability to reason about performance at multiple levels (kernel, model, system)
Nice to Have
Experience optimizing machine learning or deep learning workloads
Familiarity with video generation, diffusion models, or related domains
Experience with compilers, runtime systems, or ML frameworks internals
Background working with AI accelerators (e.g., GPUs, or custom silicon)
