Onsite 4 days in office 1 working from home
No fully remote options
Job Summary:
The Data Platform Lead is responsible for implementing, configuring, and governing the technical foundation of the company’s enterprise data Lakehouse on Azure Databricks. This role combines hands-on data platform ownership with data engineering leadership, ensuring that data from ERP systems, product SaaS PostgreSQL databases, APIs, and other enterprise sources is ingested, modeled, governed, secured, and prepared for analytics, reporting, automation, and AI use cases.
This position serves as the internal technical owner for Databricks platform standards, Bronze/Silver/Gold data architecture, Unity Catalog governance, pipeline design patterns, source-to-target mapping standards, data quality implementation, vendor technical review, and production readiness. The role is hands-on, with responsibility for platform configuration, environment setup, access controls, catalog/schema structure, compute standards, and operational readiness, while also providing technical directions to data engineers, contractors, vendors, and future internal team members through strong standards, practical architecture, and disciplined delivery.
Essential Functions, Roles and Responsibilities:
Essential duties and responsibilities include the following:
Data Platform & Lakehouse Architecture
- Own the technical architecture for the Azure Databricks Lakehouse, including workspace structure, catalogs, schemas, compute patterns, storage strategy, and environment separation.
- Define and maintain bronze, silver, and gold layer standards, including naming conventions, table ownership, audit columns, refresh patterns, and production readiness criteria.
- Implement and govern Unity Catalog standards for access control, lineage, data classification, catalog/schema organization, and least-privilege access.
- Partner with cybersecurity, infrastructure, and DevOps teams to align Databricks with enterprise identity, networking, secrets management, monitoring, and compliance expectations.
- Establish cost controls, cluster policies, job standards, and usage monitoring to ensure the platform is reliable, scalable, and cost-effective.
Data Engineering & Pipeline Delivery
- Design, build, and oversee production-grade data pipelines using Databricks, Spark, Python/PySpark, SQL, Delta Lake, and approved orchestration patterns.
- Lead ingestion from ERP systems, product PostgreSQL databases, SaaS platforms, APIs, files, and other enterprise data sources into the Lakehouse.
- Define engineering patterns for full loads, incremental loads, CDC where applicable, reprocessing, error handling, logging, reconciliation, and pipeline recovery.
- Ensure every production pipeline includes source-to-target mapping, ownership, data quality rules, monitoring, alerting, and operational handover documentation.
- Review vendor and contractor deliverables for technical quality, maintainability, security, performance, and production readiness.
Data Quality, Governance & Production Readiness
- Implement practical data quality controls for completeness, uniqueness, validity, freshness, referential integrity, and reconciliation to source systems.
- Support data governance by ensuring datasets have clear owners, stewards, classifications, lineage, refresh frequency, and required documentation before go-live.
- Work with business, ERP, and product teams to understand source system meaning, schema changes, business logic, and downstream impact.
- Enable certified silver and gold datasets that can support analytics, executive reporting, operational dashboards, automation, and AI/ML use cases.
- Own technical incident response for pipeline failures, data refresh issues, root cause analysis, and corrective actions.
Enterprise Data Architecture
- Define enterprise data architecture standards, reference architectures, and long-term roadmap.
- Establish data domain ownership and enterprise data governance models.
- Lead architecture decisions for analytics, AI/ML, master data, and enterprise reporting platforms.
- Define standards for semantic models, reusable data products, and self-service analytics.
Collaboration & Leadership
- Act as the internal technical authority for Databricks platform and data engineering decisions.
- Provide direction to vendors, contractors, and future internal data engineers to ensure delivery follows company standards.
- Partner with Cactus IT Leadership team on execution, prioritization, architecture decisions, production risk, and vendor acceptance.
- Collaborate with analytics, business applications, ERP, product engineering, cybersecurity, infrastructure, and business stakeholders.
- Promote engineering discipline, documentation quality, reusable patterns, and operational excellence across the data function.
Education, Training, Experience:
Experience
- 7+ years of experience in data engineering, data architecture, cloud data platforms, or enterprise data integration.
- Hands-on experience designing and supporting production data pipelines in cloud environments.
- Experience with modern Lakehouse architecture, preferably Azure Databricks and Delta Lake.
- Experience working with enterprise source systems such as ERP, CRM, SaaS applications, PostgreSQL, SQL Server, Oracle, or similar relational databases.
- Experience reviewing vendors or contractor technical deliverables and enforcing engineering standards.
Technical Skills
- Strong expertise with Azure Databricks, Apache Spark, Python/PySpark, advanced SQL, Delta Lake, and Lakehouse design patterns.
- Working knowledge of Unity Catalog, RBAC, lineage, data classification, metadata, and access governance.
- Experience with Azure Data Lake Storage, GitHub or Azure DevOps, CI/CD, secrets management, and cloud integration patterns.
- Strong understanding of data modeling, dimensional modeling, normalized models, medallion architecture, data quality, and reconciliation.
- Ability to design ingestion patterns for ERP data, application databases, APIs, files, and incremental source changes.
Preferred Qualifications
- Experience in manufacturing, oil & gas, field services, industrial operations, or ERP-heavy environments.
- Exposure to Power BI, Tableau, semantic models, reporting migration, or analytics product delivery
- Exposure to ML/AI pipelines, feature engineering, GenAI use cases, automation, or AI-ready data product development.
- Experience with infrastructure-as-code, automated testing, data observability, or enterprise data catalog tools.
Certifications, Licenses, Registrations:
- None required.
- Preferred: Databricks, Azure Data Engineer, Azure Solutions Architect, or related cloud/data certifications.
Job Knowledge, Skills, Abilities:
- Ability to operate as both a hands-on technical lead and a manager of delivery standards.
- Strong judgment to challenge designs that are not secure, scalable, documented, or production ready.
- Strong communication skills with technical teams, business stakeholders, vendors, and leadership.
- Ability to translate business data needs into scalable platform and engineering solutions.
- Strong ownership mindset, documentation discipline, and ability to work in a growing data organization with evolving standards.
Supervisory Responsibilities:
This role may directly or indirectly lead data engineers, contractors, and vendor delivery resources. The role is expected to provide technical direction, review deliverables, enforce standards, and support future team growth as the data platform matures.
Physical Demands:
The physical demands described here are representative of those that must be met by an employee to successfully perform the essential functions of this job. Reasonable accommodation may be made to enable individuals with disabilities to perform the essential functions.
- Regularly required to sit, stand, walk, talk, hear, and use hands to operate a computer and standard office equipment.
- Regularly required to view computer screens for extended periods and communicate through meetings, calls, and collaboration tools.
- Occasionally required to lift and/or move up to 10 pounds.
- Close vision, distance vision, color vision, peripheral vision, depth perception, and ability to adjust focus may be required.
Work Environment:
The work environment characteristics described here are representative of those an employee encounters while performing the essential functions of this job. Reasonable accommodation may be made to enable individuals with disabilities to perform the essential function.
Primarily working in a professional office or technology environment.
- May work with global teams, vendors, and stakeholders across different time zones.
- May occasionally visit operational, manufacturing, or field locations as business needs require.
- The noise level in the normal office work environment is usually low to moderate.
Disclaimer: This job description indicates the general nature and level of work expected of the incumbent. It is not designed to cover or contain a comprehensive list of activities, duties or responsibilities required of the incumbent. Incumbent may and probably will be asked to perform other duties as required. Each employee, regardless of classification, is required to maintain a safe, orderly and clean workplace, using safety precautions and always observing safety rules.
Equal Opportunity Employer
This employer is required to notify all applicants of their rights pursuant to federal employment laws. For further information, please review the Know Your Rights (https://www.eeoc.gov/poster) notice from the Department of Labor.