Senior observability Architect responsible for platform modernization, Grafana Cloud migration, operational excellence, production support readiness, and enterprise observability strategy across hybrid cloud environments.
Core Responsibilities:
Lead observability strategy, architecture, and operational excellence initiatives.
Provide servant leadership to Operations and Production Support teams.
Own end-to-end design across all layers (infra end user); define standards, SLOs, signal model
Drive continuous improvement programs focused on reliability, system health, and MTTR reduction.
Ensure operational readiness and supportability of all platform changes before production deployment.
Observability Platform Modernization:
Mature the enterprise observability platform and lead migration from New Relic to Grafana Cloud at scale.
SLO/SLI design, anomaly detection (Sift), predictive alerting
Design observability standards for AWS and on-premises hybrid environments.
Standardize telemetry collection using Grafana Alloy and OpenTelemetry.
Develop enterprise monitoring, logging, tracing, dashboarding, and alerting frameworks.
Grafana & OpenTelemetry Architecture:
Architect solutions leveraging Grafana Mimir, Loki, Tempo, IRM, and Grafana Cloud.
Define instrumentation patterns and validated templates for infrastructure, applications, and distributed services.
Design dashboards, alerts, service health views, SLIs, SLOs, and error-budget monitoring.
Ensure telemetry quality, scalability, and governance across the organization.
Production Support & Operational Readiness:
Own production support intake, triage, escalation, and stakeholder communications.
Validate monitoring, alerting, logging, and support coverage prior to go-live.
Coordinate incident response, root cause analysis, and service restoration activities.
Improve Mean Time to Detect (MTTD) and Mean Time to Resolution (MTTR).
Runbooks & Knowledge Management:
Create, maintain, and continuously improve Tier 2 and Tier 3 runbooks.
Develop onboarding documentation, templates, and operational standards.
Partner with support teams to ensure consistent observability adoption and execution.
Tools & Technologies:
Grafana Cloud, Grafana Alloy, OpenTelemetry, Mimir, Loki, Tempo, IRM
Splunk, New Relic, Power BI, SQL
AWS Cloud Services and Hybrid Infrastructure Monitoring
The pay range that the employer in good faith reasonably expects to pay for this position is $57.79/hour - $90.29/hour. Our benefits include medical, dental, vision and retirement benefits. Applications will be accepted on an ongoing basis.
Tundra Technical Solutions is among North America’s leading providers of Staffing and Consulting Services. Our success and our clients’ success are built on a foundation of service excellence. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable law, including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Unincorporated LA County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: client provided property, including hardware (both of which may include data) entrusted to you from theft, loss or damage; return all portable client computer hardware in your possession (including the data contained therein) upon completion of the assignment, and; maintain the confidentiality of client proprietary, confidential, or non-public information. In addition, job duties require access to secure and protected client information technology systems and related data security obligations.