We’re building a next-generation payment infrastructure designed for scale, speed, and resilience.
Our system dynamically routes transactions across multiple providers in real time — optimizing for performance, reliability, and approval rates.
As we enter a scaling phase, we’re focused on building robust, high-performance systems that can handle complex traffic flows and continuous growth.
This is not about maintaining infrastructure — it’s about engineering a system that powers a new standard in modern payment ecosystems.
Primary Objective of the Role
We are looking for a DevOps Engineer who can maintain and develop our production infrastructure in AWS, ensure high system availability, enable safe zero-downtime releases, automate infrastructure management, and establish effective monitoring.
Mandatory Requirements
- Hands-on experience with AWS. - Understanding of High Availability principles and fault-tolerant infrastructure design. - Experience implementing zero-downtime deployments. - Hands-on experience with Infrastructure as Code, preferably Terraform. - Experience with configuration management tools, such as Ansible or Puppet. - Experience building and maintaining CI/CD pipelines. - Strong understanding of Docker and application containerization. - Experience with load balancers, health checks, and horizontal scaling. - Understanding of network infrastructure: VPCs, subnets, routing, security groups, and firewall rules. - Good understanding of DNS, SSL/TLS, and domain and certificate management. - Experience with managed and self-hosted databases. - Understanding of backup, restore, disaster recovery, and backup restoration testing. - Experience setting up monitoring, centralized logging, and alerting. - Experience with tools such as Prometheus, Grafana, Loki, ELK, Datadog, or CloudWatch. - Good knowledge of Linux and confidence working with the command line. - Experience troubleshooting production incidents. - Understanding of secrets management and secure access control.
Expectations for the Candidate
The candidate should be able to:
- design a production architecture appropriate for the actual workload without unnecessary overengineering; - identify single points of failure; - explain trade-offs between cost, complexity, and reliability; - build infrastructure that can be reproduced automatically; - automate server and environment configuration using Ansible, Puppet, or similar tools; - minimize manual operations and deployments via SSH; - organize a safe rollback after a failed release; - define critical metrics, logs, and alerts; - propose a system scaling plan; - independently analyze and resolve production issues.


