About the Role:
We’re building a large-scale synthetic data generation engine to produce realistic observability datasets — metrics, logs, and traces — to support AI/ML training and benchmarking. You will design, implement, and scale pipelines that simulate complex production environments and emit controllable, parameterized telemetry data.
What You’ll Do • Design and implement generators for metrics (CPU, latency, throughput) and logs (structured/unstructured). • Build configurable pipelines to control data rate, shape, and anomaly injection. • Develop reproducible workload simulations and system behaviors (microservices, failures, recoveries). • Integrate synthetic data storage with Prometheus, ClickHouse, or Elasticsearch. • Collaborate with ML researchers to evaluate realism and coverage of generated datasets. • Optimize for scale and reproducibility using Docker containers.
✅ Who You Are • Strong programming skills in Python. • Familiarity with observability tools (Grafana, Prometheus, ELK, OpenTelemetry). • Solid understanding of distributed systems metrics and log structures. • Experience building data pipelines or synthetic data generators. • (Bonus) Knowledge of anomaly detection, time-series analysis, or generative ML models.
Pay $50 - 75/hr depending on experience Remote, flexible hours Project timeline: 5-6 weeks