Expected, Snowflake Data Cloud, Databricks, Python, SQL, Apache Spark, PyTorch
Optional, Apache Flink, Terraform, Kubernetes, Docker, Apache Hive, AWS, Microsoft Azure, Apache Airflow, Kafka, Debian
About the project, The project migrates data and reporting from a legacy estate onto a modern platform built on Databricks and Microsoft Fabric. Quality engineering is central to the project. The business needs confidence that migrated data reconciles with the source, that pipelines behave as specified, and that reports produce consistent results.
Team size, 1-10
This is how we work, you focus on a single project at a time, you focus on product development, agile
Team members, backend developer, technical leader, devOps, data scientist, automated test programmer
Your responsibilities, Develop, operate, optimize, test, and maintain the data warehouse, including ETL/ELT process development, cube development, database/performance administration, and dimensional table design, Drive the full life-cycle of back-end development for the data warehouse, Identify, design, and implement internal process improvements - redesigning infrastructure for scalability, optimizing data delivery, and automating manual processes, Define data retention policies, Build analytical tools that leverage the data pipeline to deliver actionable insight into key business metrics (operational efficiency, customer acquisition, etc.), Select and integrate tools for monitoring, managing, alerting on, and improving database performance, Develop and implement automated processes to ensure uninterrupted database updates and correction of vulnerabilities, Assemble large, complex datasets that meet functional and non-functional business requirements
5+ years of experience or 5+ completed projects, Advanced SQL and query optimization, with proficiency across popular database variations, Python for data engineering and automation, Cloud data platforms: Snowflake and/or Databricks, Cloud services: AWS and/or Azure, Data transformation and modeling with DBT, ETL/ELT pipeline design, development, and maintenance, Apache Spark (PySpark preferred), Workflow orchestration using Apache Airflow, Dagster, or another widely adopted orchestrator, Relational database design and performance tuning (PostgreSQL, MySQL, SQL Server, Oracle, etc.), Data warehousing concepts and dimensional modeling, Data management fundamentals: data modeling, data quality, metadata management, data warehouse/lake patterns, distributed systems, Version control using Git and CI/CD practices, Data governance, data quality, lineage, and observability practices, Security and access control implementation in cloud data platforms
Optional, Apache Kafka, Apache Flink, Apache Beam, Terraform, Kubernetes, Docker, Apache Iceberg, Delta Lake, or Apache Hudi, Real-time and event-driven architectures, AWS Glue, Amazon MWAA, Azure Data Factory, Data Mesh and Data Product concepts, Machine learning data pipelines, feature stores, or AI development experience, Streaming analytics and Change Data Capture (CDC) solutions (e.g., Debezium)
This is how we work on a project, Continuous Integration
Development opportunities we offer, mentoring, technical knowledge exchange within the company
DataArt Poland sp. z o.o., Our client is a UK financial services group operating in a highly regulated environment, currently rebuilding its data and analytics capability as part of a wider platform modernisation programme.