Senior Data Engineer

VeevaBoston, United States (Fully Remote)
InternshipFully RemoteSeniorSome info disclosed
0 views0 applications

Description

Veeva Systems is a mission-driven organization and pioneer in industry cloud, helping life sciences companies bring therapies to patients faster. As one of the fastest-growing SaaS companies in history, we surpassed $3B in revenue in our last fiscal year with extensive growth potential ahead. At the heart of Veeva are our values: Do the Right Thing, Customer Success, Employee Success, and Speed. We're not just any public company – we made history in 2021 by becoming a public benefit corporation (PBC), legally bound to balancing the interests of customers, employees, society, and investors. As a Work Anywhere company, we support your flexibility to work from home or in the office, so you can thrive in your ideal environment. Join us in transforming the life sciences industry, committed to making a positive impact on its customers, employees, and communities. The Role The NitroAI team is seeking a Senior Data Engineer to build and maintain the data engineering infrastructure that powers our analytics delivery. You'll work alongside data scientists and analytics teams to productionize manual solutions, build governed datapipelines, and deliver data to customers at scale. This is a high-impact, high-ownership role in a startup-like environment within Veeva — with immediate influence over how NitroAI delivers data across a growing multi-tenant platform. Own the Airflow codebase end-to-end (MWAA, ~100 DAGs currently): build reusable templates, scale patterns, enforce standards, improve testing infrastructure Serve as delivery teams' go-to on pipeline architecture and troubleshooting Own data onboarding when we connect to a new source including connection setup, schema discovery, initial pipeline design Operate and extend large-scale Spark pipelines on AWS Glue, including multi-TB joins and compaction jobs, and support migration of those workloads to Databricks as we move platforms Drive data model improvements around our commercial pharma data (patient claims, KOL and HCP data, CRM activity) with a focus on structure, lineage, and how it flows through the platform Contribute to and maintain our internal Python package used across the data team 5+ years building data models and pipelines 2+ years production Airflow experience 2+ years Spark at scale: hands-on experience with multi-TB datasets; AWS Glue experience preferred Strong python and SQL AWS fluency in S3, ECS/Fargate, Glue, IAM, Secrets Manager Clear communicator who can teach: onboarding analytics team to the codebase and developing team Airflow capability is a core part of this job, not a side responsibility Experience working with privacy-sensitive or governed data Databricks experience – we're actively migrating Glue workloads there Proficiency with Claude Code Exposure to data science workflows and ML pipeline tooling Background in life sciences or healthcare Experience with Open Meta data or similar data catalog and data dictionary tooling We value your time and believe in a transparent hi