Who we are
Xenoss is an AI engineering and integration services company, helping medium to large enterprises run AI transformation end-to-end, from situation analysis and goals framing to data discovery and preparation, pipeline building, model development, retraining pipeline design, solution deployment, and support.
We build a broad spectrum of AI solutions such as user behaviour prediction, content generation, NLP, audience segmentation, pathfinding solutions, AI assistants, edge computer vision, fraud detection, and others.
We work with prominent companies such as Microsoft, Toshiba, AstraZeneca, Activision Blizzard, Verve Group, Voodoo Games, and Telefonica, among others.
We’re included in the top 100 software companies on the Inc. 5000 list.
What is the project
We’re hiring a Senior LLM Engineer to join a long-term In-Call Assistant initiative for a world-leading financial services company.
The project focuses on building a real-time conversational AI system that supports front-office employees during live customer conversations. The system identifies customer needs, objections, buying signals, and required process steps, and provides concise, context-aware recommendations.
You will primarily work on the recommendation generation layer: training and specializing instruction-tuned language models to produce grounded, policy-aligned next-best-action recommendations based on the live conversation and prepared customer context.
The broader solution combines low-latency signal detection, context preparation, specialist recommendation generation, RAG over approved product and policy knowledge, and compliance guardrails.
What will you do
You’ll own implementation and continuous improvement of the recommendation generation models, working closely with the AI Solution Architect and the rest of the AI team.
Core work includes:
Building and fine-tuning specialist recommendation models for objection handling, discovery, and product guidance
Designing training datasets from historical conversations, outcomes, SME input, and generated supervision
Applying SFT, preference optimization, and parameter-efficient fine-tuning approaches such as LoRA / QLoRA
Evaluating DPO, KTO, and other post-training methods where appropriate
Building RAG capabilities over approved product, policy, and knowledge sources
Designing grounding, abstention, and fallback behavior
Developing evaluation frameworks for recommendation quality, factual accuracy, relevance, and policy alignment
Running systematic error analysis and model improvement cycles
Optimizing model serving for latency, throughput, and infrastructure constraints
Working with AI, data, MLOps, and client teams to move models from experimentation to production
You’re expected to be deeply hands-on in LLM training, fine-tuning, evaluation, and optimization.
Technology landscape
You’ll operate across the modern LLM and applied AI ecosystem, including:
Python
PyTorch and Hugging Face
Instruction-tuned language models
SFT and preference optimization
LoRA / QLoRA and PEFT
DPO, KTO, and related post-training approaches
RAG and knowledge-grounded generation
Embeddings and retrieval
Prompting and context construction
Generative model evaluation
Guardrails, grounding, and hallucination control
Low-latency LLM inference and serving
MLOps, monitoring, and feedback loops
We optimize for measurable recommendation quality, factual grounding, low latency, and enterprise constraints.
Scope of ownership and delivery context
Core ownership
Implement and improve the Recommendation Generation model
Train and maintain specialist adapters for defined conversation scenarios
Build training and evaluation pipelines for generative models
Define and test fine-tuning and post-training strategies
Develop RAG and grounding mechanisms for approved knowledge sources
Establish measurable recommendation quality and factual accuracy
Improve abstention, fallback, and policy-compliance



