Own your opportunity to work alongside federal civilian agencies. Make an impact by providing services that help the government ensure the well being and support of U.S. citizens.
Job Description:
As an AI/ML Engineer Associate, the work you’ll do at GDIT will be impactful to the mission of the Diplomatic Security Bureau of the Department of State. You will play a crucial role as part of a team to develop, implement and maintain an AI powered solution leveraging existing Department of State data and reports that will deliver insights and assist in decision making for Diplomatic Security Leaders and Analysts.
Core responsibilities:
RAG Pipeline Development & Maintenance
- Implement and iterate document ingestion, chunking, and embedding pipelines (e.g., Nomic Embed v1.5)
- Tune retrieval parameters (chunk size, overlap, top-k, similarity thresholds) against evaluation sets
- Maintain and troubleshoot the vector store (PGVector on PostgreSQL) — indexing, query performance, schema updates
Model Serving & Inference Support
- Support day-to-day operation of the LLM serving layer
- Assist with model updates, version testing, and rollback procedures
- Monitor GPU utilization, memory usage, and inference latency
Application Integration
- Work within front-end integrations to wire up new features, prompt templates, or tool-calling workflows
- Build and maintain API integrations between the LLM layer and downstream applications (via PGBouncer/Postgres, Redis caching, etc.)
- Write and refine system prompts, few-shot examples, and prompt-engineering iterations for specific use cases
Evaluation & Quality
- Build/run evaluation harnesses to test retrieval accuracy and generation quality (hallucination checks, relevance scoring)
- Track regressions when models, embeddings, or chunking strategies change
- Document known failure modes and edge cases
- Infrastructure support (Junior level)
- Assist with environment setup, dependency management, and container/service configuration in development environments
- Support basic troubleshooting of Redis, PostgreSQL, PGAdmin as they relate to the RAG pipeline
- Escalate deeper infra/networking issues to senior engineers or platform team
Test Strategy & Planning
- Contribute to a test strategy for the RAG/LLM pipeline covering three distinct layers: retrieval quality (are the right chunks being pulled), generation quality (is the LLM producing accurate, grounded, non-hallucinated answers), and system/integration (does the pipeline work end-to-end under real conditions)
- Help define acceptance criteria for "good enough" retrieval and generation — e.g., minimum relevance score thresholds, acceptable hallucination rate, latency SLAs
- Participate in test planning for new features or model/embedding swaps — identify what could break (retrieval drift, prompt regressions, latency changes) before rollout
- Maintain a golden/reference dataset of representative queries and expected answers or expected retrieved sources, used as a stable benchmark across changes
Test Execution
- Execute manual exploratory testing for new features or edge cases automation doesn't yet cover — adversarial prompts, out-of-scope questions, ambiguous queries, multi-turn context handling
- Run pre-deployment validation checklists before pushing model, prompt, or pipeline changes to production
- Execute periodic regression passes on a schedule (not just at release time) to catch silent drift — since RAG/LLM systems can degrade without any code change (e.g., underlying model provider updates, data staleness)
- Validate fixes against the original defect/failure case plus the broader regression suite
Documentation & knowledge transfer
- Maintain technical documentation for pipelines, configs, and architecture decisions
- Document runbooks for common operational tasks (restarting services, common errors, model swap procedures)
Collaboration
- Collaborate with senior engineers on architecture decisions
- Participate in code review, both giving and receiving feedback
- Communicate technical constraints/tradeoffs to non-technical stakeholders as required
Technical Skills: Python, Machine Learning, Deep Learning, SQL, Data Science, PyTorch, Docker, TensorFlow, Artificial Intelligence, Natural Language Processing, Linux, JavaScript, MATLAB, Architecture, Data Analytics, Kubernetes, MSFT Azure Platform, Big Data, Hadoop, Visualization, Software Development, and Agile.