Role Description
A senior data scientist who also engineers on a delivery team building an outcome-optimisation and scoring platform for an independent SSP. The role builds, evaluates, and operates models across three families: a real-time curation model, look-alike modelling with segment augmentation, and contextual segmentation. Python is the delivery language.
The two data scientists on the team are peers. Ownership of the three model families is split between them at kickoff, so this role must be credible on all three, including the latency-bound one. Model choice is not the hardest part of the work. Conversion labels arrive hours to weeks after exposure, advertiser data can be pooled only as far as each advertiser permits, and the finished system transfers to a client team with no data scientists.
An AdTech background on the programmatic supply side is required. Depth in CTV, audio, or social is an advantage rather than a gate.
About the Project
Client: an independent supply-side platform with curation, identity, and managed-service products, engaged directly. The first use case is outcome-driven campaigns for financial-services advertisers, who pay for real-world outcomes, account openings and deposits, rather than CPM or CTR.
Product: a scoring and optimisation platform that runs on the client’s infrastructure in two planes. The central plane on Google Cloud holds the data and feature foundation, model training, MLOps and experimentation, and control and reporting. The edge plane is a portable scoring container deployed inside third-party SSP runtimes.
Channels in scope: the initial production release covers one agreed channel set. Rollout follows the client’s priority order: CTV first, then web and app video, display, audio, native, and DOOH. Social is a priority for segment creation and optimisation rather than for campaign optimisation.
Modelling: three model families defined by the client. The curation model uses gradient-boosted trees on tabular features, and scores bid requests in real time. The look-alike and contextual models run in batch. Gradient-boosted trees are the baseline elsewhere, but not a constraint. The approach for the batch families is selected against the data in Phase 1 and signed off with the architecture in Phase 2.
Facts that shape the role
Only the curation model is latency-bound. Its budget is under 5 ms total round trip at p99+, and model execution is only a subset of that. The look-alike and contextual models are not real-time.
Two execution modes run at the same time. Dynamic mode scores inside third-party SSP auctions. Static mode is batch and human-governed. It produces outputs no more than once a day, which the client activates through its curation and audience tools.
Two feedback clocks. Bidstream data returns within two to three hours and drives fast adaptation. Deterministic conversions return in hours to weeks, and up to six months for one industry feed, and drive learning, calibration, and evaluation.
Advertiser data must not be commingled. The design is a general baseline model plus isolated client-specific instances, with isolation enforced in datasets, pipelines, model artefacts, and scoring instances.
Data rights are runtime configuration. Each advertiser’s position is discovered at onboarding, after the system is live, so a new policy must not require re-engineering.
Pooled historical training across advertisers is largely unavailable, so cold start is a design problem from day one.
Initial volume is around 10,000 QPS per campaign. All eligible traffic is scored, but logging is sampled within what each SSP permits to leave its environment.
The client standardises input fields across SSPs, so the models see a consistent request schema.
All datasets are delivered in Google Cloud, from roughly a dozen sources. Establishing schemas, join keys, historical coverage, and label latency is the first phase of work.
The client has no in-house data scientists. The mode



