DevOps Engineer – Observability & Platform Engineering
Location: Fully Remote – U.S.
Duration: 7 months with potential for extension
Schedule: Full-Time | Must Work East Coast Hours
Industry: Financial Services / Mortgage Technology
Candidates Considered: USC or GC Holder
Position Overview
Our client is seeking an experienced DevOps Engineer – Observability & Platform Engineering to join a Platform Foundation team supporting a large-scale technology transformation. This role will focus heavily on building and standardizing observability, tracing, logging, telemetry, CI/CD, and cloud infrastructure practices across complex applications and services.
This is not simply a pipeline-building DevOps role. Approximately 80% of the focus will be on observability, platform standards, and CI/CD enablement , with roughly 20% focused on directly building pipelines and infrastructure . The ideal candidate will bring strong AWS and Terraform expertise and have experience establishing reusable observability patterns that development teams can consistently adopt.
Key Responsibilities
- Help design and implement the organization's observability and telemetry architecture across applications, services, and cloud infrastructure.
- Establish standards and best practices for distributed tracing, logging, monitoring, metrics, and telemetry .
- Partner directly with development teams to define tracing standards and ensure requests can be followed end-to-end across multiple services .
- Create reusable observability patterns, libraries, templates, shared dashboards, and reporting standards.
- Provide guidance to engineering teams on implementing observability consistently across complex systems.
- Support root cause analysis (RCA) by helping teams identify where failures originated, why they occurred, and how similar issues can be prevented.
- Design and maintain secure, reliable, and scalable AWS infrastructure .
- Leverage Terraform for Infrastructure as Code and standardized infrastructure deployment.
- Support and enhance CI/CD pipelines , automation, and development workflows.
- Identify opportunities to use automation and AI-assisted development practices to improve engineering speed and operational efficiency.
- Work closely with architecture and platform teams while embedding with application development teams when needed.
- Help establish logging and tracing guidelines that can be adopted across the broader engineering organization.
- Ensure observability and telemetry solutions meet appropriate security, reliability, and regulatory requirements .
- Support observability for applications incorporating AI agents and AI-enabled solutions .
Required Qualifications
- Strong DevOps, Platform Engineering, SRE, or Cloud Engineering experience within complex enterprise environments.
- Deep hands-on experience with AWS infrastructure and services .
- Strong Terraform experience is mandatory , including Infrastructure as Code design and deployment.
- Demonstrated experience designing and deploying enterprise observability solutions .
- Strong knowledge of distributed tracing, logging, telemetry, metrics, monitoring, and alerting .
- Experience with observability platforms such as Datadog, AWS CloudWatch , or comparable enterprise tooling.
- Experience establishing observability standards across multiple applications, services, and development teams.
- Strong understanding of CI/CD pipelines, automation, and modern software delivery practices .
- Ability to troubleshoot complex distributed systems and contribute to detailed root cause analysis.
- Understanding of cloud security and secure infrastructure practices.
- Experience deploying solutions within complex, regulated enterprise environments .
- Ability to collaborate effectively with architecture, development, QA, security, and operations teams.
Preferred Qualifications
- Experience within a regulated industry , particularly financial services, mortgage/lending, healthcare, or similar environments.
- Experience supporting applications that leverage AI agents, LLMs, or AI-assisted development tools .
- Familiarity with LLM/AI observability , including monitoring AI-agent behavior and performance.
- Experience with modern observability technologies such as OpenTelemetry and related tracing frameworks.
- Experience developing shared observability libraries, templates, dashboards, and platform-level standards.
Ideal Candidate
The ideal candidate is a senior DevOps or Platform Engineer who understands observability from an engineering and architecture perspective , rather than simply administering a monitoring tool. They can work with development teams to understand how applications behave across distributed services, establish consistent tracing and logging practices, and create reusable patterns that make troubleshooting faster and more effective.
This person will be proactive, technically hands-on, and comfortable operating across platform architecture and application development teams while helping establish the foundation for secure, reliable, observable, and increasingly AI-enabled systems .