Principal Reliability Engineer - EDS
with enterprise observability stacks (Prometheus, Grafana, Datadog, Splunk, Dynatrace, OpenTelemetry). Ability to design and enforce...
with enterprise observability stacks (Prometheus, Grafana, Datadog, Splunk, Dynatrace, OpenTelemetry). Ability to design and enforce...
/ components / tools in our stack: Python, Redpanda/Kafka, Databricks/Spark, AWS/S3, Terraform, Datadog, GitHub What You'll...
, testing, and release management. Experience with monitoring and alerting tools like CloudWatch, Datadog, or similar. Passion...
including Grafana, Dynatrace, Prometheus, Datadog, and Splunk;experience with SLO alerting, white/black box monitoring...
troubleshooting. Own the deployment and integration of Datadog monitoring for the AWS VDI stack, including metrics, logs, traces.../metrics/alarms, Route53 health checks/failover routing) as well as Datadog for serverless observability. AWS KMS Multi...
–5 internal .NET applications using Azure DevOps, SonarQube, Azure Key Vault, and Datadog. This is an implementation... pipelines, define Git and pull request standards, integrate SonarQube, set up Datadog observability, and document the approach...
incidents;participate in on-call rotation. Design and evolve the observability stack (metrics, logs, traces) using Datadog...
documentation, operational procedures, and production runbooks Monitor enterprise applications using Datadog, New Relic, Splunk... with ArgoCD deployment automation Experience using Datadog, New Relic, and Splunk for enterprise monitoring and observability...
fundamentals and secure configuration basics Bonus points (nice to have) Observability: Grafana stack, ELK, or Datadog...
across database engineering and platform operations. Monitor database and application performance using tools such as Datadog... Engineer –certifications are mandatory. Familiarity with monitoring and analytics tools like Datadog and Sentry. Proven...