Site Reliability Engineer - Telemetry
π +30 days ago
AMERπ§π· BrazilπΊπΎ Uruguayπ΅πͺ Peruπ¨π΄ Colombiaπ¦π· Argentinaπ΅πΎ Paraguayπ¨π¦ Canadaπ¨π± Chileπ Remoteπ Full TimeβοΈ Mid-LevelβοΈ Devops Engineer
GrafanaGrafana LokiSplunkTempoTerraformKubernetes
π Description
- Work across metrics, logs, traces, alerting, dashboards, and profiling systems that make our platform observable at scale.
- Help keep the systems that collect, store, query, and route operational data reliable and scalable.
- Operate and improve the shared platform for metrics, logs, traces, alerting, dashboards, and profiling.
- Maintain metrics collection, long-term storage, querying, dashboards, and alerting using Prometheus-compatible systems, VictoriaMetrics, Grafana, and modern alerting tools.
- Operate log pipelines using Vector, Splunk, and Loki, including reliability, throughput, and troubleshooting.
- Operate distributed tracing and profiling capabilities using Grafana Alloy, Tempo, OpenTelemetry, and Pyroscope.
- Deploy and manage telemetry services using Terraform, Terragrunt, and container orchestration across multiple environments.
- Troubleshoot missing data, slow queries, broken alerts, pipeline backpressure, and capacity issues.
- Build reusable configuration and automation that helps teams manage dashboards, alerts, and telemetry integrations safely.
- Participate in incident response and on-call, write runbooks, and improve the platform using what we learn from incidents.
π― Requirements
- 3+ years of experience as a Site Reliability Engineer, Platform/Infrastructure Engineer, Observability Engineer, or similar production engineering role.
- Comfortable managing production systems at scale that collect, process, store, and serve telemetry such as metrics, logs, traces, or profiles.
- Experience with Prometheus or a Prometheus-compatible monitoring stack, including metrics collection, querying, and alerting.
- Experience troubleshooting distributed production systems, including availability, latency, data flow, and capacity issues.
- Experience with Infrastructure as Code, particularly Terraform, and CI/CD.
- Experience operating containerised workloads with Nomad, Kubernetes, or similar platforms.
- Solid scripting/programming ability and comfort using AI tools and agents (e.g., Claude) to accelerate delivery.
- Strong incident response, documentation, and collaboration skills.
- Experience with VictoriaMetrics, Grafana, Tempo, Loki, Vector, Splunk, Alertmanager, or OpenTelemetry.
- Experience with PromQL or LogQL, dashboards or alerts as code.
- Experience with maintaining operators and their related CRDs in Kubernetes.
- Experience with Consul, Vault, AWS, and on-premises or datacentre infrastructure.
- Experience operating high-volume logging, streaming, or data pipelines.
- Experience making practical trade-offs between observability data volume, performance, and cost.
- Background in highly regulated or financial services environments where change management and audit trails are critical.
π Benefits
- Applications are accepted on an ongoing basis.
- Payward is an equal opportunity employer and celebrates diverse talents and backgrounds.