Data Engineer

Scale.jobsSeattle, United States
Full TimeOn-siteMidLimited info disclosed
57 views0 applications

Description

About The Role The role is responsible for designing, building, and optimizing the core data platform and infrastructure that powers business-critical analytics and machine learning applications. This includes developing robust pipelines that ingest, transform, and load large-scale structured and unstructured datasets. The data engineer will collaborate closely with data scientists, software engineers, and product stakeholders to ensure high data quality, availability, and low latency across the entire data ecosystem. Key Responsibilities Design, implement, and maintain scalable batch and streaming data pipelines using Apache Spark, Kafka, and Airflow. Architect and optimize data models in the cloud data warehouse (Snowflake or BigQuery) to support high-performance analytical queries. Implement data quality monitoring, anomaly detection, and automated alerting frameworks to guarantee the integrity of critical data assets. Collaborate with backend engineering teams to integrate transactional databases and external APIs into the centralized lakehouse architecture. Write clean, reusable, and thoroughly tested infrastructure-as-code and data pipeline configurations in Python, SQL, and Terraform. Optimize query performance, storage costs, and resource utilization across the entire cloud data stack. What We Are Looking For 3-6 years of experience in data engineering, backend software engineering, or database administration. Expert-level proficiency in Python and advanced SQL for complex data manipulation and transformation. Hands-on experience with modern data orchestrators (e.g., Apache Airflow, Prefect, or Dagster) and distributed computing frameworks (e.g., Spark or Flink). Proven track record of managing and modeling data within cloud-native data warehouses like Snowflake, Redshift, or BigQuery. Familiarity with containerization (Docker, Kubernetes) and CI/CD pipelines for automating data deployments. Bonus: Experience with dbt (data build tool), streaming architectures (Kafka, Kinesis), or AWS/GCP data engineering certifications. Show more Show less