OUR SECTORS
At Tech Recruit, our sectors cover a wide range of industries within the field of technology.
At European Recruitment, our sectors cover a wide range of industries within the field of technology
At European Recruitment, our sectors cover a wide
range of industries within the field of technology
At European Recruitment, our sectors cover a wide
range of industries within the field of technology
Client services
Learn about the range of client services we offer at Tech Recruit, and browse through our case sudies.
At European Recruitment, our sectors cover a wide range of industries within the field of technology
About us
Learn about Tech Recruit's mission, values, our team, and our commitment to DE&I.
At European Recruitment, our sectors cover a wide range of industries within the field of technology
Senior ML Data Processing Engineer
Our client is a non-profit organization committed to advancing research and creating technical solutions that enable safe-by-design AI systems.
We are seeking a Senior ML Data Processing Developer to contribute to the development, curation and scaling of large-scale machine learning data pipelines.
Working at the intersection of data engineering, data curation and machine learning, you will own and develop end-to-end pipelines that transform raw, web-scale data into high-quality datasets used to train next-generation AI models.
In this role, you will go beyond traditional data engineering by actively engineering data quality. You will design algorithmic filtering systems, develop model-based scoring mechanisms, implement rigorous data-quality controls and build novel data transformations for emerging machine learning requirements.
As AI models and research requirements evolve, you will work on problems where established approaches may not yet exist, helping to define new methods for processing, evaluating and improving training data at scale.
Key Responsibilities
- Partner with Research and Engineering teams to define, build, automate, scale and maintain data pipelines that transform web-scale data into high-quality training datasets.
- Build and maintain large-scale data processing pipelines covering deduplication, model-based quality scoring, heuristic filtering, toxicity removal, PII scrubbing, metadata extraction and custom data transformations.
- Ensure datasets have robust versioning, lineage and provenance tracking while optimising pipelines for throughput and cost.
- Ensure ingested and processed data meets relevant compliance requirements, internal data governance policies and legal obligations.
- Develop and refine data-quality tooling, including:
- Heuristic filtering systems
- LLM-as-a-judge evaluators
- Machine learning classifiers
- Metadata extraction modules
- Human-in-the-loop review workflows
- Instrument pipelines with data-quality monitoring, guardrails and alerting to identify regressions before they propagate downstream.
- Work with Research and Engineering teams to understand evolving data requirements and identify or acquire large-scale text corpora that meet those requirements.
- Conduct systematic coverage analyses to identify gaps within existing datasets and develop targeted acquisition strategies.
- Work with Legal and Governance teams where required to assess and license new data sources.
- Design and maintain robust leakage-detection mechanisms to prevent evaluation contamination throughout the data processing pipeline.
- Develop internal tools and interfaces that enable researchers to efficiently explore, query and understand available datasets.
- Identify opportunities to improve the scalability, reliability, quality and efficiency of data processing workflows.
Skills & Qualifications
- Degree in computer science, software engineering, or a related field.
- Proven track record of handling massive unstructured text datasets (trillion-token scale), with 5+ years of experience in data processing, machine learning engineering or Natural Language Processing (NLP).
- Hands-on experience with distributed processing frameworks (e.g., Spark, Ray, Flink), designing and optimizing high-throughput pipelines.
- Experience with data privacy implementation (PII scrubbing), content-safety filtering (toxicity, bias), and evaluation-contamination prevention.
- Demonstrated ability to work across Research, Engineering, and/or Legal/Governance teams, translating varied requirements into concrete pipeline work.
- Strong Python proficiency, including experience writing production-grade data-processing code.
- Experience with pipeline orchestration frameworks (e.g., Airflow, Prefect, Dagster).
Nice to have
- Experience training, fine-tuning, or deploying ML models for data-quality tasks (classifiers, LLM-based evaluators) and familiarity with LLM inference optimization (e.g. vLLM, SGLang).
- Familiarity with containerized deployment (Docker, Kubernetes) and infrastructure-as-code practices.
- Familiarity with ML experiment tracking tools (e.g. Weights and Biases).
- Experience with data licensing workflows or web-scale data acquisition.
- Contributions to open-source data processing or NLP tooling.
Apply Now
By applying to this role, you acknowledge that we may collect, store, and process your personal data on our systems.
For more information, please refer to our
Privacy
Notice