A track record leading platform or infrastructure teams that ran production systems for multiple product teams, with accountability for adoption, not just delivery
Hands-on credibility in modern AI infrastructure: LLM inference and serving, model gateways, vector search, guardrails, and evaluation systems
Frontier research fluency: you read post-training, reinforcement learning, and agentic-systems work as it lands and can turn it into engineering strategy; an engineer who reads research, not a researcher at engineering distance
Cloud platform depth (AWS, Azure, or GCP) with Kubernetes and infrastructure-as-code at production scale
Experience delivering in a regulated industry, ideally financial services, or demonstrable fluency in what model-risk and security review requires of a platform
A platform-as-product mindset: you can talk about golden paths, voluntary adoption, and developer research as naturally as architecture
Daily, hands-on use of AI coding assistants in your own work
What sets you apart
You have owned both an AI platform floor and shared business services (workflow, tenancy, billing/metering) in one charter
You have taken a model through post-training (RLHF, RLAIF, fine-tuning, or distillation to smaller models) into production
Published or open-source work in agent infrastructure or evaluation tooling
Cost management (FinOps) experience for LLM workloads
Financial services domain depth: you have shipped production systems for banks, insurers, or payment providers
The reference stack
The reference technology stack for this role is our supported paved road: self-hosted Lang Smith and Lang Graph Platform as the agent runtime and evaluation plane, model providers behind a swappable gateway seam, PostgreSQL with pg vector plus Click House and S3-compatible object storage as the data platform, Neo4j Enterprise as the semantic knowledge graph, an agent memory plane serving episodic and precedent memory over MCP, MCP-native connectors, Open Telemetry and Grafana for observability, all on CNCF-conformant Kubernetes with Helm and Argo CD, deployable to any hyper scaler or on-prem. A tool-for-tool match is not expected: analogous experience counts fully. If you have built and operated systems of this shape on comparable components (a different orchestration framework, graph engine, evaluation platform, or serving stack), you have what we are looking for.
How we work
Engineers write specs, harnesses, evals, and guardrails; AI agents execute the implementation loops. Review, not typing, is where engineering judgment goes.
Three human gates govern everything we ship: spec approval, merge, and release. Regulated code paths (money movement, authentication, cryptography, secrets) are always human-owned.
Small and senior by design. No separate QA function, no scrum masters; quality comes from evaluation gates and whole-team review rituals.
Domain experts (claims practitioners, payment scheme experts, clinicians) are full-time members of the product teams you will serve.
Success in year one
The first product line ships to its first enterprise client on platform services it chose to use, with platform cost attributed per line
Platform adoption is voluntary and measured; product teams' deployment frequency and change-failure rates improve after adoption
Bank or insurer model-risk teams accept the platform's evidence pack on first review
A written engineering strategy exists, is re-argued each quarter against frontier developments, and the first New Co-tuned model (LLM or SLM) serves production traffic behind evaluation gates