About the role
We're building an on-premise AI platform that processes large volumes of multilingual documents and runs autonomous agents that gather and analyse information from open sources. Everything runs on our own hardware: self-hosted models, no external APIs.
You'll own the agentic layer: the LangGraph applications that plan, retrieve and reason, the prompt engineering behind them, and the evaluation that keeps their quality measurable.
What you'll work on
Multi-step agents on LangGraph: planning loops, tool use, checkpointing, crash recovery, human-in-the-loop
Prompt engineering with schema-constrained decoding, where every response is validated against a contract
RAG pipelines: chunking, hybrid dense + sparse retrieval, citation-grounded generation
Information extraction from documents into structured records
Evaluation: gold sets, regression gates, scenario tests over full agent trajectories
Prompt-injection defence: policy layers over tool calls, adversarial test suites
Stack
Python, LangGraph, FastAPI, vLLM, Qdrant, OpenSearch, Postgres, Kubernetes, MLflow, inspect_ai
What we're looking for
4+ years in Python, with at least 1 year building LLM-based systems in production
Hands-on experience with an agent framework (LangGraph, Pydantic AI, or similar), something you've shipped rather than prototyped
Solid RAG experience: you understand why retrieval quality, not model choice, is usually the bottleneck
Practical prompt engineering: structured outputs, evaluation-driven iteration
Experience with self-hosted inference (vLLM, TGI, Ollama)
Strong engineering discipline: typed code, tests, reproducibility
Nice to have
Document processing at scale (OCR, layout extraction) · information extraction or knowledge graphs · LLM evaluation tooling · air-gapped environments · web crawling with robots compliance and rate limiting


