Event-Driven Systems, Observability and Data Lineage
Position Summary
We are seeking a hands-on Senior Python and Kafka Engineer to join an established engineering team supporting a correctness-critical, event-driven digital repository processing platform.
The platform performs downstream processing after digital objects have been deposited and archived. Its services create persistent identifiers, update operational metadata and search stores, generate viewable and streamable derivatives, distribute content to image, media and full-text services, and notify depositors. The pipeline is event-driven, with an initial archival event being processed by multiple independent services that publish and consume subsequent events.
The successful candidate will work directly within the existing team to improve structured logging, end-to-end event visibility, reconciliation, completion verification, and Kafka delivery reliability. This is an implementation-focused engineering position. The architecture is owned internally, and the selected engineer will build enhancements into a live system rather than redesigning or replacing it.
Key Responsibilities
Implement a consistent structured JSON logging standard across Python services, including standardized fields, logging levels, contextual information, and actionable failure details.
Improve Kafka producer and consumer reliability, with particular attention to:
o Offset commit behavior
o Consumer-group stability and rebalancing
o Retry processing
o Dead-letter topic handling
o Message-loss investigation
o Duplicate-processing and idempotency considerations
Instrument distributed services to provide greater visibility into how individual digital objects move through the processing pipeline.
Implement or contribute to an event-lineage capability that records:
o The processing stages an object passed through
o The order in which stages occurred
o The outcome of each stage
o Where processing stopped or failed
Develop reconciliation and completion-tracking capabilities that compare expected workflow stages with recorded events.
Help produce an actionable completion verdict for processed objects, including the ability to identify incomplete processing and support alerting or delivery holds.
Implement telemetry using OpenTelemetry and contribute to an OpenLineage-based lineage solution, with Marquez currently identified as the leading implementation option.
Develop and maintain Python services that interact with Kafka, PostgreSQL, MongoDB, object storage, and external APIs.
Work with shared internal Python packages that provide message-envelope models, Kafka producer and consumer wrappers, retry and dead-letter handling, shared-state storage, and operational-metadata access.
Write automated tests and support safe delivery through the existing Kubernetes, Kustomize, and ArgoCD environment.
Troubleshoot complex distributed-processing issues across services, messages, data stores, and external dependencies.
Pair closely with the internal engineer leading this work and contribute directly to implementation, testing, technical documentation, and operational readiness.
The principal areas of work are structured logging, event lineage and tracking, reconciliation and completion verification, and Kafka delivery semantics.
Required Qualifications
Candidates should have strong, production-level experience in all four of the following areas:
Apache Kafka
Significant hands-on experience operating or developing Kafka-based production applications.
Deep understanding of producer and consumer delivery semantics.
Experience with Kafka offsets, commit strategies, consumer groups, partition assignment, and rebalancing.
Demonstrated ability to diagnose missing, delayed, duplicated, or repeatedly retried messages.
Practical experience implementing retry and dead-letter processing patterns.
Python
Advanced Python software-engineering experience.
Strong understanding of concurrency, threading, and runtime behavior.
Experience building and supporting production services and shared Python libraries.
Ability to write maintainable, testable, and operationally supportable code.
Observability and Structured Logging
Experience defining and implementing structured logging standards, not only configuring logging products.
Ability to establish consistent log fields, severity levels, correlation information, error context, and diagnostic detail.
Understanding of observability practices for distributed and event-driven systems.
PostgreSQL
Strong working knowledge of PostgreSQL.
Experience using PostgreSQL JSON capabilities.
Ability to design and query operational, tracking, or event-related data structures.
These four capabilities are the primary qualification bar for the role. The source brief specifically notes that candidates are not expected to possess every technology listed in the environment.
Preferred Qualifications
Hands-on experience with OpenTelemetry.
Experience with OpenLineage, Marquez, or comparable data-lineage and event-lineage solutions.
Working knowledge of Kubernetes sufficient to deploy, troubleshoot, and operate applications in a Kustomize and ArgoCD environment.
Experience using Grafana and Grafana Loki, or equivalent metrics, dashboarding, and log-management platforms.
Strong understanding of event-driven architecture patterns, including:
o Idempotent processing
o Exactly-once delivery concerns
o At-least-once processing implications
o Transactional boundaries between a datastore and a message bus
o Failure recovery in distributed workflows
The environment values OpenTelemetry, lineage tooling, Kubernetes, Grafana/Loki, and practical knowledge of event-driven architecture, but does not require deep Kubernetes platform-engineering expertise.
Additional Relevant Experience
Experience with one or more of the following would be beneficial but is not required:
MongoDB
Apache Solr
AWS S3 or other S3-compatible object storage
Istio
AWS Secrets Manager
External Secrets Operator
AWS EFS and NFS-based shared storage
Image or video processing pipelines
Digital asset, library, archive, or repository platforms
Apache Airflow and PgBouncer
TypeScript or Node.js
MongoDB, Solr, S3-compatible storage, Istio, and image or video processing are identified as useful rather than required capabilities