About the Role
JPG Infotech LLC is seeking a Hands-On Senior Databricks Engineer & Architect to lead data modernization and engineering initiatives on a mission-critical Federal contract. In this role, you will own the end-to-end Databricks ecosystem—from infrastructure deployment and security configuration to designing scalable Lakehouse architectures and executing complex data migrations from legacy environments.
You will bridge the gap between high-level systems architecture and hands-on PySpark/SQL development, working closely with federal stakeholders to transform complex data requirements into secure, high-performance analytical pipelines.
Key Responsibilities1. Databricks Installation, Configuration & Governance
- Platform Setup: Architect, install, and configure enterprise Databricks workspaces within a secure Federal Cloud environment (AWS GovCloud, Azure Government, or GCP).
- Cluster & Compute Management: Configure compute clusters, auto-scaling policies, and job orchestration environments to optimize performance and cloud consumption.
- Governance & Security: Implement Unity Catalog for centralized data governance, metadata management, fine-grained access control (RBAC/ABAC), and data lineage tracking.
- Compliance: Ensure all Databricks deployments align with federal security frameworks (e.g., FedRAMP, FISMA, NIST SP 800-53).
2. Data Architecture & Lakehouse Design
- Lakehouse Architecture: Design and deploy modern Medallion Architectures (Bronze, Silver, Gold layers) using Delta Lake.
- Data Modeling: Create conceptual, logical, and physical data models for transactional, dimensional, and analytical data stores.
- System Integration: Establish seamless API and pipeline integrations across multi-cloud services, enterprise data warehouses, and downstream reporting tools.
3. Data Migration & Pipeline Engineering
- Legacy Modernization: Lead the migration of on-premises data warehouses, relational databases, and legacy ETL workloads into Databricks Delta Lake.
- Pipeline Development: Build, test, and deploy robust ETL/ELT pipelines using PySpark, Python, and SQL for both batch and real-time streaming data.
- Orchestration: Configure automated scheduling, dependency management, and monitoring using Lakeflow / Databricks Workflows or Apache Airflow.
4. Data Analysis & Performance Tuning
- Query & Cluster Optimization: Diagnose and troubleshoot bottlenecks, optimize Spark query execution plans, and implement caching and partitioning strategies.
- Data Quality & Validation: Implement automated data-reconciliation, completeness, and schema-enforcement checks to ensure auditability and compliance.
- Stakeholder Collaboration: Partner with federal program managers, analysts, and engineering teams to translate operational requirements into actionable technical solutions.
Required QualificationsSecurity & Location
- Citizenship: Must be a U.S. Citizen (dual citizenship or non-citizen status cannot be accommodated due to federal contract mandate).
- Clearance: Ability to pass a federal background investigation and obtain/maintain a Public Trust or Secret clearance.
- Location: Preference for candidates residing in the Washington, D.C. Metropolitan Area (DC/MD/VA) with the flexibility to work on-site a few days per week as required by project milestones.
Experience & Technical Skills
- Experience: 5+ years of hands-on data engineering, data architecture, or big data platform experience.
- Databricks Expertise: Demonstrated hands-on experience installing, configuring, and maintaining production Databricks environments.
- Programming Languages: Deep proficiency writing production-grade code in Python (PySpark) and SQL (Scala is a plus).
- Core Technologies:
- Strong expertise with Apache Spark™ distributed processing and runtime internals.
- Extensive experience building Lakehouses with Delta Lake and Unity Catalog.
- Familiarity with major cloud infrastructure platforms (AWS, Azure, or GCP).
- Engineering Practices: Experience with Git-based CI/CD workflows, automated testing, and Infrastructure as Code (e.g., Terraform) for data platform deployment.
Preferred Qualifications
- Certifications: Databricks Certified Data Engineer Professional or Databricks Certified Data Engineer Associate.
- Public Sector Experience: Prior experience working on federal, state, or defense IT contracts.
- Advanced Databricks Tools: Experience with Delta Live Tables (DLT) / Lakeflow Declarative Pipelines, Databricks SQL Warehouses, or MLOps integrations.
- Data Integration: Familiarity with data contract standards, source-to-target mapping, and open-data government reporting requirements.
Why Join JPG Infotech LLC?
At JPG Infotech LLC, we deliver cutting-edge AI, cloud, and digital transformation solutions to federal and enterprise clients. You will be part of an agile, high-impact technical team working on meaningful government missions with access to the latest data and AI technologies.
Pay: $117,030.43 - $153,939.88 per year
Benefits:
- 401(k)
- 401(k) matching
- Health insurance
- Paid time off
- Retirement plan
Work Location: Hybrid remote in Arlington, VA 22203