Overview
Progile Solutions is seeking a Data Scientist — AI Evaluation Analytics for an exclusive client engagement supporting a global consulting leader and a Fortune 100 enterprise AI Platform-as-a-Service provider.
Provide data science and analytical expertise to improve the measurement, reliability, and interpretation of AI system performance through data-driven evaluation approaches. This role will support AI capabilities used within enterprise PaaS, workflow, and ServiceNow-related applications.
Key Responsibilities
- Support the development and refinement of AI evaluation metrics, scoring approaches, and benchmark methodologies.
- Validate datasets, data-quality criteria, evaluation inputs, and analytical assumptions to improve the reliability and reproducibility of results.
- Analyze evaluation data, performance trends, product telemetry, and measurement signals to generate actionable insights.
- Apply statistical analysis and experimentation techniques to assess AI system effectiveness and compare models, prompts, tools, or configurations.
- Identify trends, patterns, performance gaps, and opportunities to improve AI quality, workflow efficiency, user adoption, and customer outcomes.
- Communicate analytical findings, limitations, and recommendations clearly to product, engineering, and business stakeholders.
Required Skills & Experience
- Minimum of five years of professional experience in data science, advanced analytics, applied statistics, or a related field.
- At least two years of experience focused on AI evaluation analytics, generative AI measurement, machine learning evaluation, or AI performance analysis.
- Experience designing metrics, analytical frameworks, experiments, benchmarks, or measurement approaches.
- Proficiency in Python, SQL, and commonly used data-analysis and statistical libraries.
- Experience working with large datasets, product telemetry, operational data, or customer-usage data.
- Familiarity with generative AI, large language models, AI agents, coding assistants, or enterprise AI applications.
- Familiarity with tools and platforms such as OpenAI models and APIs, Anthropic Claude, Claude Code, Claude Opus, GitHub Copilot, or comparable AI platforms.
- Understanding of statistical methods such as hypothesis testing, regression analysis, confidence intervals, sampling, variance analysis, and experiment design.
- Ability to determine whether evaluation results are valid, reliable, reproducible, and meaningful.
- Strong communication skills with the ability to explain analytical results clearly.
Preferred Experience
- Experience evaluating AI/ML systems, generative AI applications, or agentic AI workflows.
- Experience with experimentation, benchmarking, model comparison, or performance analysis.
- Experience analyzing enterprise SaaS or PaaS applications, workflow platforms, ServiceNow applications, or similar enterprise technologies.
- Experience with software engineering, developer productivity, service-management, customer-support, or quality metrics.
- Familiarity with AI evaluation or observability tools such as OpenAI Evals, LangSmith, MLflow, Ragas, DeepEval, or similar tools.
- Experience with dashboards and visualization platforms such as Tableau, Power BI, Looker, or Grafana.
- Familiarity with predictive analytics, forecasting, segmentation, anomaly detection, or causal inference when relevant to improving product and customer outcomes.
Pay: $50.00 - $75.00 per hour
Application Question(s):
- LinkedIn profile and work experience validation is mandatory. Please provide your LinkedIn Profile ID:
- This is a remote role; however, client prefers candidates who are located in West Coast. Work during the PDT hours is required. Can you work during PDT hours?
- Do you have hands on experience building, or evaluating AI-powered coding tools and software engineering agents, such as Claude Code, Devin, OpenCode, or comparable platforms?
Education:
Experience:
- designing metrics and analytical frameworks: 3 years (Required)
- large datasets and extracting meaningful insights: 3 years (Required)
- background in data science, analytics, statistics: 5 years (Required)
- evaluating AI/ML systems and agentic AI applications: 2 years (Required)
- experimentation, benchmarking, or performance analysis: 2 years (Required)
- OpenAI Evals, LangSmith, MLflow, Ragas, DeepEval: 3 years (Preferred)
Language:
Shift availability:
Work Location: Remote