The work

You own the model side of a project: the data it learns from, the evaluation that says whether it is good enough, and the serving that keeps it fast. Some models rank or classify, some read documents and images, some are language models adapted to a client's field.

Our head of engineering has served models at multi-million scale with millisecond latency, and that is the standard we build to: cost, latency and drift are watched for every model we ship.

What you'll do

  • Turn a business question into a modelling problem with a target you can measure.
  • Build datasets and labelling processes, and keep them clean as the data changes.
  • Train and fine-tune models: classical machine learning, ranking, vision and language.
  • Build evaluation sets and the offline and online tests that decide what ships.
  • Serve models with monitoring for latency, cost, drift and quality.

What you bring

  • Models you have trained and put into production, and what they changed.
  • Strong Python and its data tools: PyTorch, scikit-learn, pandas or Polars.
  • Solid statistics: you can tell when an improvement is real.
  • Experience serving models behind APIs, with batching, caching and monitoring.
  • Working knowledge of language models: retrieval, fine-tuning, distillation and evaluation.

What success looks like

  • The model beats a simple baseline on a test the business agrees with.
  • Latency and cost stay within the budget we set together.
  • Quality is monitored, and you notice drift before the client does.

Tools you'll use

  • Python
  • PyTorch
  • scikit-learn
  • Hugging Face
  • SQL
  • Vector search
  • Docker
  • Google Cloud and Vertex AI