Responsible for the design, development, and optimization of data pipelines using Azure Data Factory, Databricks, Spark, and PySpark, ensuring scalable ingestion and transformation of structured and unstructured data; implement and maintain Lakehouse architectures with Delta Lake, Unity Catalog, and Azure Data Lake Storage (ADLS), applying best practices for performance, governance, and security; build and automate CI/CD pipelines with Jenkins, Git, and Airflow, streamlining deployment and monitoring of production-grade data solutions; develop and coordinate data models including Star, and Snowflake schemas, applying partitioning, bucketing, and query optimization techniques for high-performance analytics; leverage cloud platforms including Azure and AWS (S3, EC2, EMR) to support enterprise-scale data processing, migration, and storage solutions; engineer ETL/ELT solutions with SQL, Spark SQL, and Python, enabling efficient integration across Snowflake, Teradata, Oracle, and Hadoop ecosystems; implement Data Governance, Data Quality, and Data Lineage using Ataccama, Manta, and Unity Catalog, ensuring compliance, transparency, and reliability; perform performance tuning of Spark, Hive, and Hadoop jobs to improve efficiency and reduce compute costs; collaborate with stakeholders and Agile teams, gathering requirements, delivering technical documentation, and supporting enterprise decision-making with robust data solutions; utilize and apply knowledge of Azure, Snowflake, Apache Spark, Hadoop Ecosystem, Python, SQL, AWS, Airflow, Jenkins CI/CD, Git, and data governance/quality tools to perform assigned tasks; and lead code reviews and resolve production issues in order to drive reliability and best practices across data engineering initiatives.
Location: Bentonville, AR and multiple undetermined worksites throughout the U.S.;
Salary: $107,744 per year.