
Senior Machine Learning Engineer - Research Optimisation
Canva
Senior Machine Learning Engineer - Research Optimisation
Canva is seeking a Senior Machine Learning Engineer to bridge research and production, focusing on research enablement and performance. The role involves turning experiments into scalable features, optimizing PyTorch training and inference, and building shared libraries and tooling. Candidates need strong Python skills, hands-on experience with ML systems in production, and expertise in PyTorch optimization and Kubernetes.
Senior Machine Learning Engineer - Research Optimisation
Canva is seeking a Senior Machine Learning Engineer to bridge research and production, focusing on research enablement and performance. The role involves turning experiments into scalable features, optimizing PyTorch training and inference, and building shared libraries and tooling. Candidates need strong Python skills, hands-on experience with ML systems in production, and expertise in PyTorch optimization and Kubernetes.
Salary
Core Qualifications
Technical (Must-have)
Soft Skills
Preferred Qualifications
Technical (Nice-to-have)
Key Responsibilities
- Productionise research models: refactor, test, containerise, and integrate them into the monorepo for scalable reuse.
- Profile and optimise PyTorch training jobs, identifying bottlenecks across compute, memory, I/O, and networking.
- Improve distributed training setups (multi-GPU, multi-node) and help teams pick the right parallelism strategy.
- Build and maintain inference services, SDKs, and shared libraries that standardise pre/post-processing and interfaces across variants.
- Create multi-variant runners and rollout frameworks (feature flags, canaries, A/B testing, automated rollbacks).
- Establish CI/CD workflows, artifact management, and reproducible builds for ML services and model assets.
- Add robust observability (dashboards, alerts) and reliability practices across training and inference workloads.
- Optimise inference (batching, caching, quantisation/compilation, hardware utilisation) to reduce latency and cost.
- Work across the broader training stack, including Kubernetes orchestration, storage, and data pipelines.
- Partner with researchers and product engineers via code reviews, pair sessions, and documentation.
- Drive good engineering hygiene: testing strategy, dependency management, and de-duplication across model variants.