AI Developer (LLM Evaluation)
- Джерело:
- djinni.co
Що робити
- Design and implement automated end-to-end and regression testing infrastructure for an LLM-based assistant;
- Build an evaluation harness that runs a structured set of prompt scenarios against the deployed service;
- Implement automated checks for response correctness, tool usage, and guardrail behavior;
- Generate clear pass/fail results based on agreed evaluation baselines;
- Collaborate with the internal DevOps team to integrate evaluation runs into CI/CD processes and identify quality and behavior regressions before and after deployments;
Що очікуємо
- Hands-on experience with LLM evaluation tools such as Promptfoo, Langfuse, Arize, or comparable solutions;
- Practical experience designing or implementing automated evaluation and regression testing for LLM-based applications;
- Understanding of approaches used to evaluate non-deterministic LLM responses, tool calls, and guardrail behavior;
- Solid experience with AWS and CI/CD;
- Ability to design a solution from scratch and deliver it independently;
Схожі вакансії
З блогу Trackr
Усі статті →Знайдено через trackr.help/jobs · Канал: @trackrhelp · Бот для персональних сповіщень: @trackrhelpBot


