Outstaff Hiring: Remote — Contract (Full-time, long-term)
About the Role
We’re looking for a Senior Data Analyst to support the enterprise Knowledge Graph initiative.
The project connects large-scale company, individual, and relationship data to support compliance, KYC, credit risk, sanctions screening, beneficial ownership analysis, and corporate structure research.
In this role, you’ll work directly with data engineers, product managers, and analysts to validate data, verify query results, investigate discrepancies, and assess data quality across multiple database technologies.
This is a hands-on role focused on delivering clear evidence, including validated results, documented defects, root-cause analysis, and statistical assessments.
Location: Remote (EU) Engagement: Full-time, long-term contract Start Date: ASAP Language: English Time Zone: Ability to overlap with Eastern US hours
What You’ll Do
Validate large datasets loaded into Databricks, PostgreSQL, and graph databases
Confirm data completeness, accuracy, consistency, and structural integrity
Compare query results across different database technologies
Ensure the same queries produce correct and consistent results across platforms
Identify and investigate data discrepancies and quality issues
Determine whether issues come from ETL pipelines, schema mapping, source data, or database-specific behaviour
Develop validation test cases and define expected query results
Maintain data-quality reports, defect logs, and resolution tracking
Support the assessment of compliance and business use cases
Help determine whether each use case requires a graph database or can be handled using a traditional relational database
Apply statistical methods to assess datasets and benchmark results
Analyse distributions, variance, outliers, sampling quality, and measurement reliability
Document findings clearly for engineering and product teams
Work independently with minimal supervision as part of a cross-functional engineering team
Must-Have Requirements
Strong experience in data analysis, data validation, or data quality roles
Experience validating large and complex datasets
Strong SQL skills
Experience working with PostgreSQL or another relational database
Ability to identify discrepancies and perform root-cause analysis
Experience with ETL pipelines and data transformation processes
Understanding of common data-quality issues during data loading and migration
Experience working with large-scale datasets where manual validation is not sufficient
Practical knowledge of statistical analysis, including:
Distribution analysis
Outlier detection
Variance analysis
Sampling validation
Ability to analyse benchmark results and separate real differences from normal performance variation
Experience working across multiple databases or query technologies
Experience with Databricks, Spark, Delta Lake, or similar distributed data platforms
Strong written communication and documentation skills
Ability to work independently in a remote environment
Experience working within a cross-functional engineering team
Nice-to-Have
Hands-on experience with graph databases such as:
Neo4j
TigerGraph
NebulaGraph
ArangoDB
Apache AGE
Understanding of graph data models, nodes, edges, and relationship structures
Experience with Cypher or other graph query languages
Knowledge of Knowledge Graphs or ontology concepts
Experience with MongoDB
Experience with data visualisation or graph analysis tools
Experience in financial services, compliance, KYC, AML, or risk
Understanding of beneficial ownership, sanctions screening, PEP data, or corporate ownership structures
Experience with large company and entity relationship datasets
Familiarity with Bureau van Dijk, Orbis, or similar data sources
Technical Stack
Databricks
PostgreSQL
SQL
Spark / Delta Lake
Neo4j
TigerGraph
MongoDB
Graph databases
Statistical analysis
ETL pipelines
Data validation and quality testin



