Job Description
Role: Senior Data Platform Engineer
Rate: $65.34/hour W2
Location: Minneapolis, MN - 3 days onsite required
About the Platform
The client's Data Analytics Platform (DAP) spans the full data lifecycle on Azure. Batch and file-based sources flow through a multi-zone Delta Lake medallion lakehouse on ADLS Gen2, processed with Azure Synapse Analytics. A real-time tier on Azure Data Explorer handles streaming and sensor telemetry. Azure SQL Database holds orchestration metadata and operational state, master data management provides enterprise golden records, containerized Python jobs on Azure Container Apps handle API extraction, and Logic Apps handle lightweight intake and notification workflows. Power BI delivers reporting tiered by data velocity: high-velocity real-time dashboards from Azure Data Explorer, near-real-time views from Azure SQL, and analytical reporting from curated serverless SQL marts in the lakehouse. Access is governed by a domain-based classification model aligned to Minnesota government data practices requirements, and all code deploys through Azure DevOps CI/CD.
Every ingestion pipeline runs through the client's metadata-driven orchestration framework: a central metadata database that drives source configuration, scheduling, watermarking, logging, and layer promotion through shared Synapse notebook toolboxes. New work is done inside this framework, not alongside it.
About the Role
This is a generalist, delivery-focused engineering role built for independent work inside an established platform. The role reports administratively to an IT manager. The client's Data Architect owns the platform's architecture, standards, and conventions and reviews and approves design and code; within those guardrails, the engineer owns how the work gets done. The work is broad: onboarding new data sources, extending and hardening existing pipelines, resolving defects across the estate, building curated data products, and improving the shared framework itself where it falls short.
We are not looking for a narrow specialist. We are looking for a specific working style: an engineer who can take an ambiguous request, investigate the systems involved, read the existing code and metadata before writing anything new, propose an approach, and carry the work to done. Direction comes from the platform's architecture and standards rather than step-by-step instruction: the Data Architect reviews and approves outcomes, and the engineer owns everything in between.
You will succeed in this role if you
-
Treat an unfamiliar system, API, or codebase as something to investigate, not a blocker. You read documentation, trace existing pipelines, test assumptions, and come back with findings and a proposed path, not just questions.
-
Work within an established framework and its conventions, and improve it through its own patterns when it falls short, rather than building one-off solutions around it.
-
Leave code better than you found it, with the discipline to keep changes surgical: fix what is broken, flag what is fragile, and avoid both unrequested rewrites and shortcuts that leave technical debt behind.
-
Communicate status honestly and early, including when something is harder than expected or when you find a problem nobody asked about.
Key Responsibilities
-
Source onboarding: Build ingestion for new data sources (REST APIs, SFTP drops, network file shares, vendor exports, sensor and telemetry feeds) through the metadata-driven orchestration framework, landing raw data and promoting it through the medallion zones with PySpark notebooks and Synapse pipelines.
-
Curated data products: Develop serverless SQL views and Delta Lake tables in the curated consumption layer that apply business rules and serve analytics, reporting, and downstream feeds, following the platform's schema-per-mart and access-control conventions.
-
Extraction jobs: Develop and maintain containerized Python jobs on Azure Container Apps for API extraction workloads, including managed identity authentication, structured logging to the orchestration database, and YAML-based CI/CD deployment.
-
Troubleshooting and defect resolution: Diagnose and resolve pipeline failures, data quality issues, and performance problems across the platform, tracing issues through orchestration metadata, Spark logs, and source systems to root cause.
-
Framework improvement: Extend and refactor shared toolbox notebooks and framework components when the backlog exposes gaps, with regression awareness across all sources that depend on them.
-
Metadata and documentation: Register new sources and datasets in platform metadata and document solutions for team knowledge sharing.
Required Qualifications
What we mean by experience: every item below requires sustained, hands-on production work. That is code the candidate personally designed, built, deployed, and supported in a live environment over time, not coursework, certifications alone, proof-of-concept exposure, or brief contact with a technology on someone else's project. Evaluation will include a detailed technical discussion of production work the candidate has personally delivered (design decisions, failure modes encountered, and how solutions held up under real data and real change) and a hands-on technical exercise using materials the client provides. Candidates will not be asked to show code or other proprietary artifacts belonging to prior employers or clients.
-
5+ years of data engineering experience with your own code running in production, including substantial hands-on work in Azure.
-
Production proficiency with Azure Synapse Analytics: Spark notebooks, pipeline development, triggers, and serverless SQL pools that you have built and operated in a live environment. Note that Databricks and Microsoft Fabric are not used in this environment; native Synapse experience matters.
-
PySpark and T-SQL at the depth that comes from regular production use: performance tuning, window functions, and incremental load patterns against Delta Lake tables on ADLS Gen2, with the ability to explain the choices behind your own code.
-
Production-grade general-purpose Python beyond Spark: API clients, authentication handling, parsing varied file formats, and code you have packaged and deployed for containerized execution.
-
Day-to-day working use of Azure DevOps (Git, YAML pipelines, pull-request workflow) and of managed identity and Entra ID authentication patterns for Azure services.
-
A demonstrated record of self-directed delivery: named, specific examples of taking on unfamiliar systems or vaguely specified problems and driving them to a working, reviewed solution in production.
-
Clear written and verbal communication, including the ability to document technical work for a small team.
Preferred Qualifications
Unlike the required items above, familiarity or lighter exposure is acceptable here.
-
Experience with Azure Data Explorer (KQL, ingestion mappings, continuous export) or other real-time analytics stores.
-
Experience working in a metadata-driven ingestion framework, where pipeline behavior is configured in database metadata rather than hard-coded per source.
-
Experience with Azure Container Apps or Docker-based job workloads, Logic Apps, and infrastructure-as-code (Bicep).
-
Familiarity with Power BI as a consumer of the platform, including DirectQuery and Government Community Cloud constraints.
-
Microsoft certifications such as Azure Data Engineer Associate or Azure Solutions Architect.
-
Prior work in government, transportation, aviation, utilities, or other regulated environments handling non-public data.