Senior DataOps Engineer (#5858)
- Рівень:
- senior
- Джерело:
- djinni.co
Що робити
- Design, implement, and maintain the AWS Data Platform foundation, including Amazon S3 lake layout, Apache Iceberg table format standardization, and AWS Glue Data Catalog integration.
- Build and optimize scalable batch and streaming data pipelines using Amazon EMR (Serverless and EMR-on-EKS), Apache Spark/PySpark, Apache Flink, and Amazon Athena with workgroup cost caps.
- Implement data tokenization and de-identification pipelines on the data side (PA-T) using Protegrity/Spark UDFs for batch, Spark Streaming for micro-batch, and Kafka Connect SMT for streaming PII masking prior to cloud transit.
- Build and manage real-time streaming architectures with Amazon MSK (Managed Streaming for Apache Kafka), MSK Replicator/MirrorMaker2, and Schema Registry integration.
- Implement Medallion Architecture (Bronze, Silver, Gold layers) for lakehouse data modeling, automating Iceberg table registration and backfill frameworks.
Що очікуємо
- Mandatory Technical Skills:
- 4+ years of hands-on experience as a Data Engineer or DataOps Engineer building enterprise-grade data platforms and pipelines.
- Strong expertise with AWS Data Analytics stack: Amazon S3, AWS Glue Data Catalog, AWS Lake Formation, Amazon EMR (EMR-on-EKS / Serverless), Amazon Athena, AWS DataSync, and Amazon MSK.
- Deep experience with Apache Iceberg table format, cataloging, compaction, and schema evolution.
- Proficient in Apache Spark / PySpark and Spark Streaming for batch, micro-batch, and real-time data processing.
Схожі вакансії
З блогу Trackr
Усі статті →Знайдено через trackr.help/jobs · Канал: @trackrhelp · Бот для персональних сповіщень: @trackrhelpBot


