استخدام Senior Site Reliability Engineer (SRE)
شرح موقعیت شغلی
Description:
Pegah is the technology group driving a wide range of digital products and businesses, including Cafe Bazaar, Tapsell, Metrix, Bazaar Pay, Metis AI, Gapify, Athena AI, Bebin TV, Beeptunes, Footballi, and more, serving over 50 million active users.
The Technology Team builds and operates the shared infrastructure, platforms, services, and core capabilities that enable engineering teams across the group to operate reliably at massive scale every single day.
We are looking for a Senior Site Reliability Engineer (SRE) to join our Technology team and help improve the reliability, scalability, and operational excellence of our production systems across the group.
Position Summary:
As a Senior Site Reliability Engineer (SRE), you will work closely with engineering teams to improve the reliability, availability, scalability, and performance of business-critical production services. You will help define and improve reliability practices, participate in incident response and on-call rotations, reduce operational toil through automation, and proactively identify risks before they impact users.
This role is a good fit for someone who enjoys understanding complex production systems, troubleshooting under pressure, collaborating across engineering domains, and continuously improving how services are operated and supported.
What You Will Do:
Pegah is the technology group driving a wide range of digital products and businesses, including Cafe Bazaar, Tapsell, Metrix, Bazaar Pay, Metis AI, Gapify, Athena AI, Bebin TV, Beeptunes, Footballi, and more, serving over 50 million active users.
The Technology Team builds and operates the shared infrastructure, platforms, services, and core capabilities that enable engineering teams across the group to operate reliably at massive scale every single day.
We are looking for a Senior Site Reliability Engineer (SRE) to join our Technology team and help improve the reliability, scalability, and operational excellence of our production systems across the group.
Position Summary:
As a Senior Site Reliability Engineer (SRE), you will work closely with engineering teams to improve the reliability, availability, scalability, and performance of business-critical production services. You will help define and improve reliability practices, participate in incident response and on-call rotations, reduce operational toil through automation, and proactively identify risks before they impact users.
This role is a good fit for someone who enjoys understanding complex production systems, troubleshooting under pressure, collaborating across engineering domains, and continuously improving how services are operated and supported.
What You Will Do:
- Maintain and improve a diverse range of production services, including in-house and open-source systems.
- Work closely with product and engineering teams to understand service behavior, dependencies, failure modes, and operational risks.
- Participate in on-call rotations, incident response, troubleshooting, and production support for critical systems.
- Troubleshoot complex production issues across applications and their underlying dependencies.
- Monitor system behavior and proactively identify reliability, capacity, availability, and performance risks.
- Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and other reliability targets.
- Lead or contribute to incident analysis, postmortems, and corrective actions to prevent recurring failures.
- Reduce operational toil through automation, better workflows, and elimination of repetitive manual work.
- Improve service observability through meaningful metrics, logs, traces, dashboards, and alerts.
- Review system architecture and production readiness, bringing reliability considerations into service design and evolution.
- Improve resilience through better failure handling, graceful degradation, recovery mechanisms, and capacity planning.
- Turn incident learnings into better runbooks, reliability practices, tooling, and operational standards.
- Collaborate across product, platform, infrastructure, delivery, and security to continuously improve production reliability and operational excellence.
What We Expect:
- 4+ years of experience in Site Reliability Engineering, Production Engineering, or a similar production-focused role.
- Strong problem-solving and troubleshooting skills in complex production environments.
- Solid knowledge of Linux systems and underlying concepts.
- Good understanding of networking fundamentals: TCP/IP, DNS, load balancing, and proxies.
- Good understanding of distributed systems, failure modes, high availability, fault tolerance, and recovery patterns.
- Solid understanding of core SRE concepts such as SLIs, SLOs, error budgets, toil, and service reliability.
- Proven experience owning production systems with a high degree of autonomy, driving incidents through to resolution, and collaborating effectively across engineering teams.
- Hands-on experience with Kubernetes and containerized workloads in production.
- Experience with observability systems such as Prometheus, Grafana, VictoriaMetrics, or similar technologies.
- Experience with centralized logging systems such as Loki or Elasticsearch/ELK, along with distributed tracing and instrumentation using OpenTelemetry or similar technologies.
- Good understanding of service mesh and traffic management concepts, including the roles of service meshes such as Istio or Linkerd and proxies such as Envoy.
- Proficiency in Go or Python plus solid shell scripting skills.
- Strong mindset for automation and operational efficiency.
- Willingness to participate in scheduled on-call rotations and incident response.
Nice to Have:
- Experience defining and managing SLIs/SLOs and error budgets in practice.
- Hands-on experience operating service meshes such as Istio or Linkerd, or proxies such as Envoy, in production.
- Familiarity with Infrastructure as Code and GitOps tooling such as Terraform, Ansible, Helm, or Argo CD.
- Experience with software delivery tooling and workflows, including GitLab CI/CD, runners, and artifact repositories such as JFrog Artifactory or Nexus.
- Experience operating or supporting highly available databases, messaging systems, stateful services, or distributed data platforms such as PostgreSQL, MySQL, or Kafka.
- Practical experience using AI-assisted engineering tools such as Claude, Codex, Gemini, Cursor, etc. to accelerate troubleshooting, automation, documentation, and day-to-day technical workflows.
مهارتهای مورد نیاز
- DevOps
- عیب یابی
- Linux
- Grafana
- SRE
حداقل سابقه کار
- سه تا شش سال
جنسیت
- مهم نیست
وضعیت نظام وظیفه
- معافیت تحصیلی معافیت دائم پایان خدمت