تپسل | Tapsell

تاسیس در ۱۳۹۴ کامپیوتر، فناوری اطلاعات و اینترنت ۲۰۱ تا ۵۰۰ نفر www.tapsell.ir

استخدام Senior Site Reliability Engineer (SRE)

  • دسته‌بندی شغلی

    IT / DevOps / Server
  • موقعیت مکانی

    تهران ، تهران
  • نوع همکاری

    تمام وقت
  • حداقل سابقه کار

    سه تا شش سال
  • حقوق

    توافقی

شرح موقعیت شغلی

Description:
Pegah is the technology group driving a wide range of digital products and businesses, including Cafe Bazaar, Tapsell, Metrix, Bazaar Pay, Metis AI, Gapify, Athena AI, Bebin TV, Beeptunes, Footballi, and more, serving over 50 million active users.
The Technology Team builds and operates the shared infrastructure, platforms, services, and core capabilities that enable engineering teams across the group to operate reliably at massive scale every single day.
We are looking for a Senior Site Reliability Engineer (SRE) to join our Technology team and help improve the reliability, scalability, and operational excellence of our production systems across the group.

Position Summary:
As a Senior Site Reliability Engineer (SRE), you will work closely with engineering teams to improve the reliability, availability, scalability, and performance of business-critical production services. You will help define and improve reliability practices, participate in incident response and on-call rotations, reduce operational toil through automation, and proactively identify risks before they impact users.
This role is a good fit for someone who enjoys understanding complex production systems, troubleshooting under pressure, collaborating across engineering domains, and continuously improving how services are operated and supported.

What You Will Do:

  • Maintain and improve a diverse range of production services, including in-house and open-source systems.
  • Work closely with product and engineering teams to understand service behavior, dependencies, failure modes, and operational risks.
  • Participate in on-call rotations, incident response, troubleshooting, and production support for critical systems.
  • Troubleshoot complex production issues across applications and their underlying dependencies.
  • Monitor system behavior and proactively identify reliability, capacity, availability, and performance risks.
  • Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and other reliability targets.
  • Lead or contribute to incident analysis, postmortems, and corrective actions to prevent recurring failures.
  • Reduce operational toil through automation, better workflows, and elimination of repetitive manual work.
  • Improve service observability through meaningful metrics, logs, traces, dashboards, and alerts.
  • Review system architecture and production readiness, bringing reliability considerations into service design and evolution.
  • Improve resilience through better failure handling, graceful degradation, recovery mechanisms, and capacity planning.
  • Turn incident learnings into better runbooks, reliability practices, tooling, and operational standards.
  • Collaborate across product, platform, infrastructure, delivery, and security to continuously improve production reliability and operational excellence.

What We Expect:

  • 4+ years of experience in Site Reliability Engineering, Production Engineering, or a similar production-focused role.
  • Strong problem-solving and troubleshooting skills in complex production environments.
  • Solid knowledge of Linux systems and underlying concepts.
  • Good understanding of networking fundamentals: TCP/IP, DNS, load balancing, and proxies.
  • Good understanding of distributed systems, failure modes, high availability, fault tolerance, and recovery patterns.
  • Solid understanding of core SRE concepts such as SLIs, SLOs, error budgets, toil, and service reliability.
  • Proven experience owning production systems with a high degree of autonomy, driving incidents through to resolution, and collaborating effectively across engineering teams.
  • Hands-on experience with Kubernetes and containerized workloads in production.
  • Experience with observability systems such as Prometheus, Grafana, VictoriaMetrics, or similar technologies.
  • Experience with centralized logging systems such as Loki or Elasticsearch/ELK, along with distributed tracing and instrumentation using OpenTelemetry or similar technologies.
  • Good understanding of service mesh and traffic management concepts, including the roles of service meshes such as Istio or Linkerd and proxies such as Envoy.
  • Proficiency in Go or Python plus solid shell scripting skills.
  • Strong mindset for automation and operational efficiency.
  • Willingness to participate in scheduled on-call rotations and incident response.

Nice to Have:

  • Experience defining and managing SLIs/SLOs and error budgets in practice.
  • Hands-on experience operating service meshes such as Istio or Linkerd, or proxies such as Envoy, in production.
  • Familiarity with Infrastructure as Code and GitOps tooling such as Terraform, Ansible, Helm, or Argo CD.
  • Experience with software delivery tooling and workflows, including GitLab CI/CD, runners, and artifact repositories such as JFrog Artifactory or Nexus.
  • Experience operating or supporting highly available databases, messaging systems, stateful services, or distributed data platforms such as PostgreSQL, MySQL, or Kafka.
  • Practical experience using AI-assisted engineering tools such as Claude, Codex, Gemini, Cursor, etc. to accelerate troubleshooting, automation, documentation, and day-to-day technical workflows.

معرفی شرکت

ما در هلدینگ پگاه داده کاوان شریف، در کنار یکدیگر برای دستیابی به رویاهای‌مان تلاش می‌کنیم و آنچه را که در ذهن داشتیم به واقعیت بدل می‌کنیم و از مشاهده موفقیت‌های خود و سازمان مان لذت می‌بریم. روحیه کار تیمی و فضای سرشار از همراهی و همیاری یکدیگر باعث شده هیچ مانعی نتواند بر سر موفقیت های ما قرار بگیرد. ما از سال ۱۳۹۴ فعالیت خودمان را در حوزه ی تبلیغات آنلاین ( Ads Network ) آغاز کردیم. برند تپسل سرآغاز مسیر ما بود، در حال حاضر تپسل یکی از اصلی ترین کسب و کار های ارائه‌ دهنده تبلیغات آنلاین در ایران است که بیش از ۴۰ میلیون کاربر منحصر به فرد آنلاین دارد. پس از ۴ سال در سال ۱۳۹۸ قدمی نو برداشته و کسب و کار فانتوری را راه اندازی کردیم، فانتوری فعالیت خود را با تولید و نشر بازی های موبایلی آغاز کرده و رفته رفته با نگاهی بین‌المللی در حال حاضر برای میلیون ها کاربر در سراسر دنیا بازی می‌سازد. در سال ۱۳۹۹ بار دیگر حوزه فعالیت‌مان را گسترش داده و کسب و کار متریکس را در حوزه مارکتینگ‌ تکنولوژی راه‌اندازی کردیم. پس از موفقیت های بزرگ در حوزه های مختلف و استقبال مخاطبان، پا به حوزه فرهنگی و خلاق گذاشته و در سال ۱۴۰۰ کسب و کار مدیاهاوس با دو محصول اصلی سلام سینما و بیپ‌تونز به مجموعه پگاه پیوست. ما در هلدینگ پگاه همچنان به رشد و توسعه می‌اندیشیم و با نبوغ و خلاقیت اعضای تیم‌مان در آینده کسب و کارهای نوآور و دانش بنیان دیگری را نیز به مجموعه خود خواهیم افزود.
  • مهارت‌های مورد نیاز

    DevOps عیب یابی Linux Grafana SRE
  • جنسیت

    مهم نیست
  • وضعیت نظام وظیفه

    معافیت تحصیلی معافیت دائم پایان خدمت
  • حداقل مدرک تحصیلی

    کارشناسی

مشاغل مشابه

چه موردی را می‌خواهید گزارش کنید؟

از اینجا شروع کنید

در شغل بهتری استخدام شوید! رایگان!

  • جستجو و ارسال رزومه به آگهی‌های استخدام بیش از ۱۰۰,۰۰۰ شرکت ایرانی
  • رزومه‌ساز رایگان
  • دریافت فرصت‌های شغلی جدید مرتبط از طریق ایمیل (Job Alert)
  • شناخت محیط کار و فرهنگ سازمانی شرکت‌های در حال استخدام
image/svg+xml