OUR SECTORS
At Tech Recruit, our sectors cover a wide range of industries within the field of technology.
At European Recruitment, our sectors cover a wide range of industries within the field of technology
At European Recruitment, our sectors cover a wide
range of industries within the field of technology
At European Recruitment, our sectors cover a wide
range of industries within the field of technology
Client services
Learn about the range of client services we offer at Tech Recruit, and browse through our case sudies.
At European Recruitment, our sectors cover a wide range of industries within the field of technology
About us
Learn about Tech Recruit's mission, values, our team, and our commitment to DE&I.
At European Recruitment, our sectors cover a wide range of industries within the field of technology
Senior Distributed ML Engineer
Our client is a non-profit organization committed to advancing research and creating technical solutions that enable safe-by-design AI systems.
We are seeking a Senior Distributed Machine Learning Engineer to join a highly technical team working on large-scale machine learning systems.
In this role, you will work closely with ML research scientists and engineers to solve challenging training and inference problems involving large models and distributed computing environments.
You will focus on building and optimising the infrastructure required to train and run advanced machine learning models at scale, investigating performance bottlenecks and developing tools that make distributed computing resources more accessible and efficient for research and engineering teams.
This is an opportunity to work at the intersection of machine learning, distributed systems, high-performance computing and GPU optimisation.
Key Responsibilities
- Collaborate with ML researchers and engineers to accelerate model development, training and inference.
- Enable efficient use of large-scale machine learning models across distributed computing environments.
- Investigate performance bottlenecks across distributed training and inference workloads.
- Profile research and production code to identify opportunities for performance improvements.
- Debug complex issues across distributed ML systems and develop robust solutions.
- Optimise the utilisation of GPU, CPU, memory, networking and other computing resources.
- Develop tools, libraries and infrastructure to simplify and orchestrate distributed computing resources for ML workloads.
- Improve the scalability, reliability and efficiency of distributed training and inference workflows.
- Establish, document and maintain best practices for large-scale distributed ML development.
- Collaborate with research, engineering and infrastructure teams to support evolving technical requirements.
- Evaluate emerging distributed ML technologies and incorporate them where they provide meaningful technical benefits.
Background & Experience
- Degree in a relevant field such as Computer Science, Computer Engineering, Software Engineering or a related discipline.
- Master’s or PhD in Machine Learning, Distributed Systems or a related field is preferred, but not required for candidates with exceptional experience.
- 3+ years of experience designing and implementing distributed machine learning training systems or frameworks.
- Strong experience working with large-scale ML workloads and distributed computing environments.
- Recent hands-on experience with one or more technologies such as:
- Megatron
- DeepSpeed
- Hugging Face Accelerate
- FSDP
- vLLM
- verl
Distributed Systems & Infrastructure
- Experience with cloud platforms such as AWS, GCP or Azure.
- Experience with workload managers and distributed computing frameworks such as Ray or SLURM.
- Strong understanding of distributed training architectures and large-scale ML infrastructure.
- Experience working with containerisation technologies such as Docker and Kubernetes.
- Familiarity with distributed communication technologies and frameworks such as gRPC.
- Familiarity with modern data infrastructure and platforms, including technologies such as vector databases.
GPU Performance & Optimisation
- Hands-on experience profiling and optimising GPU workloads.
- Experience with GPU profiling and performance-analysis tools such as:
- PyTorch Profiler
- PyProf
- NVIDIA Nsight
- Strong understanding of GPU utilisation, memory behaviour, communication overhead and distributed compute performance.
- Experience identifying and resolving performance bottlenecks across large-scale training and inference workloads.
Collaboration & Communication
- Ability to collaborate effectively with research, engineering and infrastructure teams.
- Strong communication and documentation skills.
- Ability to establish and communicate engineering best practices across technical teams.
- Strong problem-solving skills and the ability to debug complex systems.
- Demonstrated ability to stay current with developments in distributed machine learning, GPU computing and software engineering.
- Track record of contributing to high-quality machine learning or deep learning research projects.
What You’ll Bring
We are looking for an engineer with strong expertise in distributed machine learning and high-performance computing who enjoys solving difficult systems-level problems.
You should be comfortable working with large-scale model training and inference, profiling complex workloads and developing infrastructure that enables researchers and engineers to make effective use of distributed computing resources.
Apply Now
By applying to this role, you acknowledge that we may collect, store, and process your personal data on our systems.
For more information, please refer to our
Privacy
Notice