Description: Pegah is the technology group driving a wide range of digital products and businesses, including Cafe Bazaar, Tapsell, Metrix, Bazaar Pay, Metis AI, Gapify, Athena AI, Bebin TV, Beeptunes, Footballi, and more, serving over 50 million active users. The Technology Team builds and operates the shared infrastructure, platforms, services, and core capabilities that enable engineering teams across the group to operate reliably at massive scale every single day. We are looking for a Senior Site Reliability Engineer (SRE) to join our Technology team and help improve the reliability, scalability, and operational excellence of our production systems across the group.
Position Summary: As a Senior Site Reliability Engineer (SRE), you will work closely with engineering teams to improve the reliability, availability, scalability, and performance of business-critical production services. You will help define and improve reliability practices, participate in incident response and on-call rotations, reduce operational toil through automation, and proactively identify risks before they impact users. This role is a good fit for someone who enjoys understanding complex production systems, troubleshooting under pressure, collaborating across engineering domains, and continuously improving how services are operated and supported.
What You Will Do:
Maintain and improve a diverse range of production services, including in-house and open-source systems.
Work closely with product and engineering teams to understand service behavior, dependencies, failure modes, and operational risks.
Participate in on-call rotations, incident response, troubleshooting, and production support for critical systems.
Troubleshoot complex production issues across applications and their underlying dependencies.
Monitor system behavior and proactively identify reliability, capacity, availability, and performance risks.
Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and other reliability targets.
Lead or contribute to incident analysis, postmortems, and corrective actions to prevent recurring failures.
Reduce operational toil through automation, better workflows, and elimination of repetitive manual work.
Improve service observability through meaningful metrics, logs, traces, dashboards, and alerts.
Review system architecture and production readiness, bringing reliability considerations into service design and evolution.
Improve resilience through better failure handling, graceful degradation, recovery mechanisms, and capacity planning.
Turn incident learnings into better runbooks, reliability practices, tooling, and operational standards.
Collaborate across product, platform, infrastructure, delivery, and security to continuously improve production reliability and operational excellence.
What We Expect:
4+ years of experience in Site Reliability Engineering, Production Engineering, or a similar production-focused role.
Strong problem-solving and troubleshooting skills in complex production environments.
Solid knowledge of Linux systems and underlying concepts.
Good understanding of networking fundamentals: TCP/IP, DNS, load balancing, and proxies.
Good understanding of distributed systems, failure modes, high availability, fault tolerance, and recovery patterns.
Solid understanding of core SRE concepts such as SLIs, SLOs, error budgets, toil, and service reliability.
Proven experience owning production systems with a high degree of autonomy, driving incidents through to resolution, and collaborating effectively across engineering teams.
Hands-on experience with Kubernetes and containerized workloads in production.
Experience with observability systems such as Prometheus, Grafana, VictoriaMetrics, or similar technologies.
Experience with centralized logging systems such as Loki or Elasticsearch/ELK, along with distributed tracing and instrumentation using OpenTelemetry or similar technologies.
Good understanding of service mesh and traffic management concepts, including the roles of service meshes such as Istio or Linkerd and proxies such as Envoy.
Proficiency in Go or Python plus solid shell scripting skills.
Strong mindset for automation and operational efficiency.
Willingness to participate in scheduled on-call rotations and incident response.
Nice to Have:
Experience defining and managing SLIs/SLOs and error budgets in practice.
Hands-on experience operating service meshes such as Istio or Linkerd, or proxies such as Envoy, in production.
Familiarity with Infrastructure as Code and GitOps tooling such as Terraform, Ansible, Helm, or Argo CD.
Experience with software delivery tooling and workflows, including GitLab CI/CD, runners, and artifact repositories such as JFrog Artifactory or Nexus.
Experience operating or supporting highly available databases, messaging systems, stateful services, or distributed data platforms such as PostgreSQL, MySQL, or Kafka.
Practical experience using AI-assisted engineering tools such as Claude, Codex, Gemini, Cursor, etc. to accelerate troubleshooting, automation, documentation, and day-to-day technical workflows.
ما در هلدینگ پگاه داده کاوان شریف، در کنار یکدیگر برای دستیابی به رویاهایمان تلاش میکنیم و آنچه را که در ذهن داشتیم به واقعیت بدل میکنیم و از مشاهده موفقیتهای خود و سازمان مان لذت میبریم. روحیه کار تیمی و فضای سرشار از همراهی و همیاری یکدیگر باعث شده هیچ مانعی نتواند بر سر موفقیت های ما قرار بگیرد. ما از سال ۱۳۹۴ فعالیت خودمان را در حوزه ی تبلیغات آنلاین ( Ads Network ) آغاز کردیم. برند تپسل سرآغاز مسیر ما بود، در حال حاضر تپسل یکی از اصلی ترین کسب و کار های ارائه دهنده تبلیغات آنلاین در ایران است که بیش از ۴۰ میلیون کاربر منحصر به فرد آنلاین دارد. پس از ۴ سال در سال ۱۳۹۸ قدمی نو برداشته و کسب و کار فانتوری را راه اندازی کردیم، فانتوری فعالیت خود را با تولید و نشر بازی های موبایلی آغاز کرده و رفته رفته با نگاهی بینالمللی در حال حاضر برای میلیون ها کاربر در سراسر دنیا بازی میسازد. در سال ۱۳۹۹ بار دیگر حوزه فعالیتمان را گسترش داده و کسب و کار متریکس را در حوزه مارکتینگ تکنولوژی راهاندازی کردیم. پس از موفقیت های بزرگ در حوزه های مختلف و استقبال مخاطبان، پا به حوزه فرهنگی و خلاق گذاشته و در سال ۱۴۰۰ کسب و کار مدیاهاوس با دو محصول اصلی سلام سینما و بیپتونز به مجموعه پگاه پیوست. ما در هلدینگ پگاه همچنان به رشد و توسعه میاندیشیم و با نبوغ و خلاقیت اعضای تیممان در آینده کسب و کارهای نوآور و دانش بنیان دیگری را نیز به مجموعه خود خواهیم افزود.