Fullstack Jobs

Site Reliability Engineer - Telemetry

πŸ•’ +30 days ago

AMERπŸ‡§πŸ‡· BrazilπŸ‡ΊπŸ‡Ύ UruguayπŸ‡΅πŸ‡ͺ PeruπŸ‡¨πŸ‡΄ ColombiaπŸ‡¦πŸ‡· ArgentinaπŸ‡΅πŸ‡Ύ ParaguayπŸ‡¨πŸ‡¦ CanadaπŸ‡¨πŸ‡± Chile🏠 RemoteπŸ•› Full Timeβš™οΈ Mid-Level☁️ Devops Engineer
GrafanaGrafana LokiSplunkTempoTerraformKubernetes

πŸ“‹ Description

  • Work across metrics, logs, traces, alerting, dashboards, and profiling systems that make our platform observable at scale.
  • Help keep the systems that collect, store, query, and route operational data reliable and scalable.
  • Operate and improve the shared platform for metrics, logs, traces, alerting, dashboards, and profiling.
  • Maintain metrics collection, long-term storage, querying, dashboards, and alerting using Prometheus-compatible systems, VictoriaMetrics, Grafana, and modern alerting tools.
  • Operate log pipelines using Vector, Splunk, and Loki, including reliability, throughput, and troubleshooting.
  • Operate distributed tracing and profiling capabilities using Grafana Alloy, Tempo, OpenTelemetry, and Pyroscope.
  • Deploy and manage telemetry services using Terraform, Terragrunt, and container orchestration across multiple environments.
  • Troubleshoot missing data, slow queries, broken alerts, pipeline backpressure, and capacity issues.
  • Build reusable configuration and automation that helps teams manage dashboards, alerts, and telemetry integrations safely.
  • Participate in incident response and on-call, write runbooks, and improve the platform using what we learn from incidents.

🎯 Requirements

  • 3+ years of experience as a Site Reliability Engineer, Platform/Infrastructure Engineer, Observability Engineer, or similar production engineering role.
  • Comfortable managing production systems at scale that collect, process, store, and serve telemetry such as metrics, logs, traces, or profiles.
  • Experience with Prometheus or a Prometheus-compatible monitoring stack, including metrics collection, querying, and alerting.
  • Experience troubleshooting distributed production systems, including availability, latency, data flow, and capacity issues.
  • Experience with Infrastructure as Code, particularly Terraform, and CI/CD.
  • Experience operating containerised workloads with Nomad, Kubernetes, or similar platforms.
  • Solid scripting/programming ability and comfort using AI tools and agents (e.g., Claude) to accelerate delivery.
  • Strong incident response, documentation, and collaboration skills.
  • Experience with VictoriaMetrics, Grafana, Tempo, Loki, Vector, Splunk, Alertmanager, or OpenTelemetry.
  • Experience with PromQL or LogQL, dashboards or alerts as code.
  • Experience with maintaining operators and their related CRDs in Kubernetes.
  • Experience with Consul, Vault, AWS, and on-premises or datacentre infrastructure.
  • Experience operating high-volume logging, streaming, or data pipelines.
  • Experience making practical trade-offs between observability data volume, performance, and cost.
  • Background in highly regulated or financial services environments where change management and audit trails are critical.

🎁 Benefits

  • Applications are accepted on an ongoing basis.
  • Payward is an equal opportunity employer and celebrates diverse talents and backgrounds.
Apply