Raiffeisen Bank is the largest Ukrainian bank with foreign capital. For more than 30 years, we have been creating and building the banking system of our country.
Raiffeisen employs more than 5,000 employees, including one of the largest product IT teams, which includes 900+ specialists. Every day, we work side by side so that more than 2.5 million of our clients can receive quality service, use the bank’s products and services, and develop their business, because we are #TogetherWithUkraine.
We are looking for a Middle SRE to join our team to help ensure the stability, scalability, and predictable performance of our services in production. In this role, you will work at the intersection of software development, infrastructure, monitoring, and reliability engineering practices.
Your future responsibilities:
Maintain and evolve production infrastructure
Ensure high availability, reliability, and performance of services
Automate routine operations and minimize manual effort (toil reduction)
Set up and maintain monitoring, alerting, and observability systems
Analyze incidents, actively participate in troubleshooting, and conduct Root Cause Analysis (RCA)
Conduct post-incident reviews and implement preventive measures to avoid recurring issues
Participate in capacity planning, system scaling, and resource utilization optimization
Enhance disaster recovery (DR) processes, backup strategies, and operational readiness
Implement and support a GitOps approach to deployment using Argo CD; manage declarative application configurations in Kubernetes
Maintain comprehensive documentation of standard operating procedures, runbooks, and architectural decisions
Collaborate closely with software development, QA, security, and IT support teams
Participate in on-call rotations according to an agreed schedule
Adopt and drive AIOps practices: leverage ML/AI for anomaly detection across metrics and logs, alert correlation, and automated root-cause hypothesizing
Integrate LLM-powered tools into daily operational workflows: incident diagnostics, reporting, log analysis, and technical knowledge base search
Tech stack:
OS: Linux
Containers & Orchestration: Kubernetes, Docker
Cloud Platform: AWS
Infrastructure as Code (IaC): Terraform
CI/CD: GitHub Actions, Jenkins
Monitoring & Observability: Prometheus, Grafana, Alertmanager, and related tooling
Logging: OpenSearch, Elasticsearch
Scripting & Automation: Python
Version Control & Quality: Git, code review best practices
Your skills and experience:
2+ years of experience in SRE, DevOps, Platform Engineering, or a related role
Solid knowledge of Linux and networking fundamentals: TCP/IP, DNS, HTTP, TLS
Hands-on experience with Kubernetes and containerization
Proven track record of building and maintaining CI/CD pipelines
Experience with Terraform or other Infrastructure as Code (IaC) tools
Strong understanding of observability principles: monitoring, logging, metrics collection, and distributed tracing
Ability to independently diagnose and troubleshoot production issues
Clear understanding of core SRE practices: SLIs, SLOs, SLAs, and error budgets
Proficiency in Python
Ability to read, interpret, and analyze application logs and metrics
Strong and clear communication skills during incident management
Nice to have:
Hands-on experience with or a strong interest in adopting AI/ML-driven approaches within SRE practices
Experience using LLM tools to accelerate diagnostics, script writing, and operational automation
Understanding of AI agent architectures and principles
Experience applying AI for anomaly detection, alert correlation, or incident root cause analysis (RCA)
Experience designing, developing, or deploying AI agents to automate routine SRE workflows and reduce toil
We offer what matters most to you:
Competitive salary: we guarantee a stable income and annual bonuses for your personal contribution. Additionally, we have a referral reward program for attracti



