Job Summary
We are seeking an experienced Senior Big Data Engineer with strong expertise in the Hadoop ecosystem, Apache Spark, Scala, Kafka, and large-scale data processing. The ideal candidate will have extensive hands-on experience designing, developing, and supporting scalable batch and real-time data pipelines across multiple data platforms.
This role requires strong technical capabilities in data ingestion, distributed data processing, data transformation, data modeling, and cloud-based big data technologies. Experience with Hadoop platform administration & Google Cloud Platform (GCP) Bigquery will be a plus.
Key Responsibilities
Big Data Engineering & Data Pipeline Development
Design, develop, implement, and maintain scalable, high-performance data ingestion and processing pipelines using Hadoop ecosystem technologies.
Develop and manage data pipelines supporting multiple data ingestion and processing modes, including:
* Batch data processing
* Real-time/event-driven data processing
* Streaming data processing
Build robust data ingestion solutions that:
* Read data from multiple structured, semi-structured, and unstructured data sources.
* Ingest batch data and real-time event streams, including Kafka events.
* Perform complex data validation, cleansing, enrichment, and transformation.
* Deliver processed data to target data stores, curated data layers, publishing zones, and downstream endpoints.
Develop and optimize Apache Spark applications using Scala for large-scale distributed data processing.
Design and implement Kafka-centric event processing and real-time data pipelines.
Develop streaming data transformation logic using Apache Spark Streaming and/or Spark Structured Streaming.
Build and maintain scalable batch processing solutions using Apache Spark.
Develop data processing and analytical solutions using HiveQL, Pig Latin, HBase, and custom MapReduce programs.
Develop data transformation and integration processes to move data from raw data zones to curated and published data warehouse layers.
Collaborate with data architects, application teams, business stakeholders, and platform teams to translate functional requirements into scalable technical solutions.
Hadoop Ecosystem & Platform Engineering
Work extensively with Hadoop ecosystem technologies, including:
* HDFS
* MapReduce
* Hive
* Pig
* Sqoop
* HBase
* ZooKeeper
* Oozie
* Apache Spark
* Scala
* Flume/Flume NG
* Kafka
* Hue
Apply strong knowledge of Hadoop architecture and core components, including:
* NameNode
* DataNode
* HDFS architecture
* JobTracker
* TaskTracker
* MapReduce programming and execution paradigm
Install, configure, integrate, and support Hadoop ecosystem components within Cloudera-based environments.
Work with distributed storage and processing frameworks to ensure scalability, reliability, fault tolerance, and high performance.
Monitor and optimize data pipeline performance, resource utilization, throughput, and processing efficiency.
Data Modeling & Data Integration
Design and implement data models, data transformation processes, and detailed technical designs.
Develop scalable data integration solutions to move data across raw, staging, curated, and publishing layers.
Support data warehouse and data lake architectures and ensure high-quality, reliable, and timely data delivery.
Implement data quality, validation, reconciliation, and error-handling mechanisms within data pipelines.
Ensure data solutions are scalable, maintainable, reusable, and aligned with enterprise data architecture standards.
Integrate on-premises Hadoop platforms with cloud-based data platforms where required.
Required Qualifications & Experience
Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, Data Science, or a related technical discipline.
7+ years of strong hands-on experience in Hadoop framework and the broader Hadoop ecosystem.
6+ years of hands-on experience developing data ingestion and integration solutions across multiple data platforms.
5+ years of strong hands-on experience in Apache Spark with Scala-based distributed data processing.
5+ years of experience in data modeling, data transformation, detailed technical design, and data integration.
Strong experience designing and developing large-scale batch and real-time data pipelines.
Strong experience with HiveQL, Pig Latin, HBase, and custom MapReduce programming.
Experience developing and managing Kafka-centric event-driven data pipelines.
Strong understanding of batch processing, stream processing, and event-driven architecture.
Hands-on experience with Spark Streaming and/or Spark Structured Streaming.
Experience installing and configuring Cloudera Hadoop ecosystem components, including Hive, HBase, ZooKeeper, Oozie, Spark, Sqoop, Flume, Pig, and Hue.
Strong understanding of Hadoop architecture, HDFS, distributed storage, and MapReduce concepts.
Strong analytical, problem-solving, debugging, and performance-tuning skills.
Excellent communication and collaboration skills.
Preferred / Nice-to-Have Skills:
Hadoop Platform Administration, GCP Bigquery
The following Hadoop administration and platform engineering skills are highly desirable:
Experience providing end-to-end Hadoop administration and production support.
Experience with Hadoop infrastructure setup, software installation, configuration, upgrades, patching, monitoring, troubleshooting, and maintenance.
Experience administering Hadoop distributions and platforms such as:
* Cloudera
* MapR
* Hortonworks
Experience installing, configuring, and managing Hadoop ecosystem components, including Hive, Pig, HBase, ZooKeeper, Oozie, Spark, and related services.
Experience managing and monitoring HDFS, distributed file systems, and Hadoop clusters.
Experience managing, monitoring, scheduling, and troubleshooting MapReduce and distributed processing jobs.
Experience with cluster capacity planning, resource management, health monitoring, and operational support.
Experience automating operational activities using scripting, including:
* Backup and restore processes
* Cluster monitoring
* Health checks
* Maintenance activities
* Operational reporting
Experience with version control, change management, release management, incident management, problem management, and root-cause analysis.
Key Competencies
Strong expertise in distributed data processing and big data architecture.
Deep understanding of batch, real-time, streaming, and event-driven data processing.
Strong hands-on programming skills in Scala and distributed data engineering frameworks.
Ability to design scalable, fault-tolerant, and high-performance data solutions.
Strong technical troubleshooting and root-cause analysis capabilities.
Ability to work independently while collaborating effectively with cross-functional teams.
Strong ownership, attention to detail, and commitment to data quality and operational excellence.
Pay: $100,696.26 - $110,268.61 per year
Work Location: In person