The Data/Ingestion Engineer will design and implement scalable ingestion pipelines for high-volume structured and unstructured insurance documents, including PDFs, scans, emails, Word, Excel and PowerPoint files.
The role includes OCR/document extraction, text cleaning and normalization, semantic chunking, metadata tagging, vector storage/RAG preparation, validation, monitoring and integration with downstream AI solutions.
At least 7 years of experience as Data Engineer with Python and SQL
At least 5 with AWS, the rest should be more or less but should be real experience that the candidate is able to talk about and give a concrete example/
Must to have :
Python
SQL
AWS
S3
Step Functions
CloudWatch
Data/document ingestion pipelines
Processing of unstructured documents:
scanned documents
Word
Excel
PowerPoint
emails
OCR / document extraction technologies
AWS Textract or equivalent
Connectors/integrations with sources such as:
SharePoint
Text cleaning and normalization
Metadata extraction/tagging
Automated error monitoring
Validation of OCR/extraction results
Git
CI/CD
Testing
Public-cloud data processing/document ingestion experience
Good to have :
Vector databases
RAG
Semantic chunking
Vector-storage schema design
Retrieval mechanisms
Azure
Databricks
AWS Textract specifically
Experience within insurance / financial services
Experience handling enterprise security and low-latency SLA requirements
English Level C1
Duration of the mission : 1 year


