← Усі вакансії

AI Developer (LLM Evaluation)

Джерело:
djinni.co
Відгукнутись на вакансію →

Що робити

  • Design and implement automated end-to-end and regression testing infrastructure for an LLM-based assistant;
  • Build an evaluation harness that runs a structured set of prompt scenarios against the deployed service;
  • Implement automated checks for response correctness, tool usage, and guardrail behavior;
  • Generate clear pass/fail results based on agreed evaluation baselines;
  • Collaborate with the internal DevOps team to integrate evaluation runs into CI/CD processes and identify quality and behavior regressions before and after deployments;

Що очікуємо

  • Hands-on experience with LLM evaluation tools such as Promptfoo, Langfuse, Arize, or comparable solutions;
  • Practical experience designing or implementing automated evaluation and regression testing for LLM-based applications;
  • Understanding of approaches used to evaluate non-deterministic LLM responses, tool calls, and guardrail behavior;
  • Solid experience with AWS and CI/CD;
  • Ability to design a solution from scratch and deliver it independently;

Схожі вакансії

З блогу Trackr

Усі статті →

Знайдено через trackr.help/jobs · Канал: @trackrhelp · Бот для персональних сповіщень: @trackrhelpBot