We are currently looking for a Senior AI Platform and Model Operations Engineer to join a professional German-based product company team that designs innovative eHealth solutions for hospitals.
You will work on a stable healthcare product.
You'd own the models and the platform around them, serving every channel we put in front of a patient.
The phone assistant is first because it's already underway, but it isn't the point of the job. Next is a chat assistant in the patient portal that can answer questions and walk someone through booking an
appointment, and after that whatever channel comes along.
The thing you're actually building is one internal API that every channel calls, so that the rules about what a patient can do live in one place instead of being written three times, slightly differently. The phone assistant and the web chat are two front ends onto the same thing.
Alongside that: moving off the vendor. Right now we pay per minute for speech recognition and language models, and within about a year we want to be running them ourselves — because we're dealing with
patient data, and because we don't want a vendor swapping a model underneath us without telling us.
To be clear about scope: we mean serving open-weight models and adapting them to our domain. We don't mean training a model from scratch.
One thing we want to be direct about. We are not looking for someone to build what we've already decided on. We don't yet know whether running these models ourselves is the right call — we don't have an accuracy figure on real German phone audio, and we don't have a proper cost comparison. Your first job is to produce both, and if they say we should keep paying the vendor for another year, we need you to tell us that plainly. Someone who quietly builds the thing because we asked for it is worse than no hire at all. If being the person who says "the numbers don't support this" makes you uncomfortable, this isn't the job.
You need
To have put a language model behind an API that more than one product called, and kept it running.
Streaming responses, timeouts, retries, and errors that tell the caller something useful.
To have run a speech recognition or language model in production on hardware you were responsible for. Not a notebook, not someone else's API.
To be able to size and cost it: how many people one GPU will serve at once, what an hour of audio or a million tokens costs you, where it starts to fall over.
Tool calling in production, and a clear view on where the boundary sits. The assistant will guide patients through booking, which means the model proposes an action and our code decides whether it's allowed. If you've built something where the model was effectively trusted to act, we'd want to talk about what you'd do differently.
To have dealt with untrusted text reaching a model that can trigger actions. A chat box on a patient portal is a much easier thing to attack than a phone call, and someone will try.
Experience grounding a model on your own content so it stops inventing things — and honesty about where that still fails.
Experience fine-tuning or adapting an open-weight model, LoRA or PEFT or something equivalent, and the honesty to say whether it was worth the effort.
A habit of being suspicious of your own benchmark numbers, especially the good ones.
To understand that using patient conversations to train a model is a different legal question from using them to answer one. If that distinction is new to you, this role probably isn't the right fit.
Python and Linux, containers, and a serving stack like vLLM or TGI.
Enough experience to work without an ML team around you, and to tell a team lead something he'd rather not hear.
We are considering candidates who work through Ukrainian FOP.
Nice to have
Quantization, and a view on what it costs you in accuracy.
Audio processing: sample rates, codecs, what a phone network does to a recording.
GPU scheduling, Kubernetes.
Ha


