About the Role
Apollo is building out a data lake using an already-established medallion architecture, and expanding it as we bring on new clients and migrate existing ones. We're looking for a Data Engineer to lead this effort — owning the data lake's architecture and health, leading external data imports as we onboard new clients, and leading all API-based data imports, which flow directly into the data lake rather than our transactional database.
This is a leadership role on the data side: you'll be the person setting direction on how data flows in, how it's structured, and how ETL pipelines are built and maintained, while also staying hands-on with the pipelines and queries themselves.
What You'll Do
- Lead the data lake build-out and expansion using our existing medallion architecture, onboarding both new and existing clients
- Own external data imports as part of client onboarding
- Lead all API-based data imports, ensuring API data is imported directly into the data lake rather than the transactional database
- Design, build, and maintain ETL pipelines using Python
- Write and optimize complex SQL queries against large datasets
- Build web scraping tools to source external data as needed
- Set technical direction and best practices for the team's data engineering work
What We're Looking For
- Experience with Azure Synapse (data lake)
- Experience with Azure Data Lake Studio
- Professional SQL experience in a leadership or senior capacity
- Experience building and maintaining ETL pipelines
- ETL experience using Python
- Web scraping experience
Nice to Have
- Experience with PySpark
- Experience with Kafka
- Experience with Flink
- Microsoft 365 experience
How We Evaluate
Our process includes a practical component — we're interested in how quickly you can get up to speed on an existing database/data lake architecture, and we'll ask you to work through some complex SQL queries.
Pay: $80,000.00 - $90,000.00 per year
Work Location: In person