Infrastructure Intelligence and Analytics (IIA) team builds and operates the data platform and AI agent infrastructure that powers proactive network monitoring and autonomous investigation for network operations. The Platform Engineer VI serves as the offshore lead for IIA platform engineering, responsible for production support during offshore hours, data pipeline health, cost optimization, building/leading the offshore platform engineering team, and developing the utilitarian application code, automation, and internal tooling that support these systems. This role requires the seniority to make independent architectural and operational decisions without escalating to onshore leadership during their shift.
Major and responsibilities:
Production Support and Operations:
Own production support during offshore hours. Serve as first responder for infrastructure incidents, pipeline failures, and agent downtime.
Monitor data pipelines feeding agents and anomaly detection models. Ensure upstream data is landing on time and in expected format.
Triage and resolve production issues independently. Escalate to onshore only for decisions requiring architectural changes or cross-team coordination.
Maintain and improve monitoring dashboards for pipeline health, agent availability, and data freshness.
Cost and Resource Optimization:
Build and maintain cost dashboards covering AWS infrastructure spend, AI compute/token costs, and third-party service costs.
Monitor token consumption across deployed agents. Identify cost anomalies and recommend optimization strategies.
Optimize resource utilization across compute, storage, and networking. Right-size infrastructure based on usage patterns.
Support pilot cost modeling: measure cost per user, project total cost at scale, and provide data for expansion decisions.
Data Availability and Pipeline Health:
Ensure agents are getting the data they need. Monitor data availability, freshness, and quality for all sources feeding the IIA Data Lake.
Handle upstream data availability issues: identify when sources stop landing, coordinate with upstream data owners for resolution.
Support data pipeline backfills and recovery when issues are detected.
Coordinate with data engineering team on pipeline health checks and SLA monitoring.
Application Development and Tooling:
Develop and maintain utilitarian application code: scripts, CLIs, small services, and automation utilities (primarily Python) that support data ingestion, deployment, environment provisioning, and operational workflows.
Write integration code and glue services that connect IIA systems with upstream data providers, downstream consumers, and external platforms.
Leadership:
Lead and grow the offshore platform engineering team. Mentor junior engineers and establish operational standards.
Define and enforce runbooks, escalation procedures, and operational best practices for the offshore team to follow during production support.
Coordinate across time zones with onshore platform engineers, data engineers, and data scientists to ensure smooth handoffs.
Participate in architectural decisions and provide input on infrastructure design based on operational experience.
Perform other duties as required.
Required Qualifications:
Skills/Abilities and Knowledge:
Ability to read, write, speak and understand English
Strong communication skills with ability to lead across time zones and coordinate with onshore teams
Expert-level experience with AWS services: EC2, S3, IAM, VPC, Glue, Athena, EMR, Secrets Manager, CloudWatch
Strong experience with infrastructure-as-code (Terraform preferred, CloudFormation acceptable)
Experience with production monitoring, alerting, and incident response at scale
Experience with cost monitoring and optimization in AWS environments
Experience with containerization (Docker) and CI/CD pipelines (GitLab CI preferred)
Proficiency with Linux-based operating systems and shell scripting
Experience managing and mentoring engineering teams
Demonstrated ability to make independent technical decisions and triage production issues without escalation
Understanding of data pipeline architectures (batch and streaming) and common failure modes
Proficiency in Python (or a comparable general-purpose language) for building automation, tooling, and applications
Solid software engineering fundamentals: Git-based workflows, code review, modular and reusable design, dependency management, and writing maintainable, documented code
Experience writing automated tests (unit/integration) for application and infrastructure code, and integrating those tests into CI/CD
Ability to write integration code against REST APIs and cloud SDKs (e.g., AWS SDK / boto3)
Preferred Qualifications:
Skills/Abilities and Knowledge:
Experience with Apache Spark (Scala preferred) for distributed data processing
Experience with Apache Kafka or similar streaming platforms
Experience with Airflow or similar workflow orchestration platforms
Experience with graph databases (AWS Neptune, Neo4j)
Familiarity with AI/ML infrastructure requirements (model serving, token cost monitoring, artifact management)
Experience in the telecommunications industry or other large-scale network operations environments
Experience with LLM/agent deployment infrastructure and observability
AWS certifications (Solutions Architect, DevOps Engineer, or similar)
Experience developing and operating small services or APIs (e.g., FastAPI/Flask) in a production environment
Education:
Bachelor’s degree in Computer Science, Information Technology, Systems Engineering, or related field, or relevant experience
Related Experience:
Bachelor's degree: 8+ years of platform/infrastructure engineering experience, including 2+ years in a lead or senior role
Master's degree: 6+ years of platform/infrastructure engineering experience, including 2+ years in a lead or senior role
The pay range that the employer in good faith reasonably expects to pay for this position is $39.30/hour - $61.40/hour. Our benefits include medical, dental, vision and retirement benefits. Applications will be accepted on an ongoing basis.
Tundra Technical Solutions is among North America’s leading providers of Staffing and Consulting Services. Our success and our clients’ success are built on a foundation of service excellence. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable law, including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Unincorporated LA County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: client provided property, including hardware (both of which may include data) entrusted to you from theft, loss or damage; return all portable client computer hardware in your possession (including the data contained therein) upon completion of the assignment, and; maintain the confidentiality of client proprietary, confidential, or non-public information. In addition, job duties require access to secure and protected client information technology systems and related data security obligations.