On behalf of our Client, Scalors is looking for a DevOps Engineer to join a remote team for a full-time position.
About Client: Our Client delivers cutting-edge software solutions for the cruise and hospitality industries, driving efficiency and reliability across mission-critical systems. Our Infrastructure & Edge Operations team ensures that our global platforms run smoothly, securely, and at scale.
Role summary:
Remote first, work from anywhere, with occasional travel to customer sites or ships as required for onboarding or incident response. Candidate must be available for on call rotations.
You will own build, release, operations, observability and deployment automation for cloud and shipboard environments.
You will design and operate secure, resilient networks and runtime platforms that support microservices, event streaming and offline sync for guest facing systems such as POS, gangway, dining and guest services.
You will be customer facing, and able to translate shipboard constraints into reliable deployment and support practices.
You will also use AI tools and techniques to design, automate and improve our infrastructure, observability and incident operations.
Core responsibilities:
CI CD and release automation, design and maintain CI CD pipelines for multi environment deployments, including cloud, shipboard on premise and air gapped scenarios.
Infrastructure as code, author and maintain Terraform, Ansible or Pulumi templates for cloud and on premise platforms.
Kubernetes and container platforms, operate and tune Kubernetes clusters in cloud and shipboard edge nodes, manage Helm charts or Kustomize manifests and GitOps flows.
Shipboard network design and troubleshooting, configure IP networks, VLANs, routing, NAT, port forwarding, DNS, NTP and firewall rules for shipboard deployments, including satellite link and VSAT awareness.
Edge constraints and offline design, design resilient services for intermittent connectivity, offline first clients, local stores such as SQLite or Couchbase Lite, CRDT or other reconciliation strategies.
Black box microservice handling, integrate and operate against third party or opaque services using robust contract testing, timeouts and circuit breaker patterns, and observability wrappers to detect failures without requiring changes to the black box.
Observability and diagnostics, deploy and operate Prometheus, Grafana, OpenTelemetry, Jaeger, Elastic or OpenSearch, Loki and tracing to achieve end to end visibility across cloud and shipboard.
Messaging and data platforms, operate Kafka and topic management, ensure broker resiliency and correct retention and consumer behavior for shipboard replication scenarios.
Security and compliance, implement mTLS, certificate lifecycle, Vault secrets management, vulnerability scanning, container image signing, and operate according to PCI and GDPR requirements on payment and guest data flows.
Hardening and incident response, lead troubleshooting for production incidents, runbooks, on call rotations, post incident reviews and automated rollback strategies.
Automation and developer enablement, provide tools and templates that enable developers to produce production quality artifacts, and self service infrastructure for test and demo environments.
Customer engagement, collaborate with customer IT to validate IP designs, test failover scenarios and document network and deployment runbooks.
Backups and disaster recovery, implement backup and restore for stateful components, test DR in constrained network conditions.
Cost and capacity planning, monitor capacity, plan cluster sizing and network bandwidth usage, and recommend changes to reduce cost and risk.
AI driven design and operations, apply AI and machine learning tools to improve infrastructure design, incident management, observability and automation, while maintaining strict security and audit controls.
Required skills and experience:
3 plus years operating productio


