Are you an AI/GenAI engineer who loves shipping real systems? Join Stanford’s Enterprise Technology team to design, implement, and support AI solutions across university use cases. In this role, you’ll work hands-on to implement LLM/RAG services, build tool-using agents, and integrate with enterprise platforms (ServiceNow, Salesforce, Oracle Financials, etc.) using modern interoperability standards such as the Model Context Protocol (MCP), while following strong MLOps/SDLC practices. You’ll prototype, harden, and ship features, partnering closely with product, security, infrastructure, and application teams.
This is an applied engineering role (not research). You’ll learn rapidly, contribute code daily, write clear docs, and develop strong habits in quality, governance, and cost/latency optimization.
Agent Interoperability Protocols: Familiarity with agent-to-agent coordination standards such as A2A (Agent2Agent) for multi-agent workflows; awareness that production systems increasingly run MCP (agent-to-tool) and A2A (agent-to-agent) together.
Agentic Evaluation: Setting up evaluation frameworks for agents: LLM-as-a-Judge for reasoning/task-completion quality, plus tracking accuracy, latency, and cost across agent runs (rubrics, hallucination/bias checks, A/B tests, golden sets).
Deployment & Infrastructure: Docker and Kubernetes, CI/CD pipelines, and microservices architecture for serving modular, independently scalable agent components.
AI-Assisted Development: Productive use of AI coding assistants (e.g., Claude Code, Cursor, GitHub Copilot) in day-to-day engineering.
MLOps Tooling: MLflow, Kubeflow, Vertex Pipelines, SageMaker Pipelines; LangSmith/PromptLayer/Weights & Biases.
Open-Source Savvy: Experience working with, customizing, and improving open-source solutions; comfortable contributing fixes/features upstream.
Rapid Tech Adoption: Demonstrated ability to pick up a new technology/framework quickly and deliver production value with it.
GenAI Frameworks: LangChain, LlamaIndex, DSPy, Haystack, LangGraph, Agent Engine, Google ADK, AWS AgentCore, CrewAI/AutoGen.
Security & Governance: Implementing AI guardrails, red-teaming, and policy-enforcement frameworks.
Enterprise Integrations: ServiceNow, Salesforce, Oracle Financials, or others.
UI Development: React/Next.js/Tailwind for internal tools.
Prompt engineering at scale: Structured prompts (JSON/function-calling), templates, version control; automated/offline & online evals.
Parameter-efficient fine-tuning (LoRA/QLoRA/adapters), supervised instruction tuning; hosting open-weight models (Llama/Mistral/Qwen) with vLLM/TGI/Ollama.
Safety / guardrails frameworks (Guardrails.ai, NeMo Guardrails, Azure/AWS safety filters) and jailbreak/drift detection.
Hybrid search & reranking (BM25+dense, Cohere/Voyage/Jina rerankers), synthetic data generation, provenance/watermarking.
Telemetry & governance: prompt/model drift monitoring, policy-as-code, audit logging, red-teaming playbooks.