استخدام (Site Reliability Engineer(SRE
شرح موقعیت شغلی
About Us
We are building a scalable SaaS platform where reliability, availability, and operational excellence are at the core of our engineering culture.
We are building a scalable SaaS platform where reliability, availability, and operational excellence are at the core of our engineering culture.
Role Overview
We are looking for a Senior Site Reliability Engineer to improve the reliability, availability, and scalability of our production platform. You will drive automation, observability, incident management, and reliability engineering practices while collaborating closely with software and infrastructure teams.
Key Responsibilities
- Monitor and improve production service reliability and availability.
- Define and maintain SLIs, SLOs, and SLAs.
- Build and enhance monitoring, logging, and observability solutions.
- Lead production incident response, RCA, and postmortems.
- Automate operational processes to improve efficiency.
- Improve High Availability (HA), Disaster Recovery (DR), and system resilience.
- Support capacity planning and infrastructure performance optimization.
- Collaborate with Engineering, DevOps, Security, and Product teams.
- Mentor engineers and promote SRE best practices.
Required Skills & Experience
- 5+ years of experience in SRE, Infrastructure, or Platform Engineering.
- Strong Linux and networking knowledge.
- Hands-on experience with Kubernetes and Docker.
- Experience with Prometheus, Grafana, Loki, and ELK.
- Experience with Terraform and Ansible.
- Proficiency in Python, Bash, or Go.
- Strong understanding of CI/CD and distributed systems.
- Strong analytical, incident management, and communication skills.
Success Metrics
- Achieve SLA/SLO targets.
- Reduce MTTD and MTTR.
- Minimize recurring production incidents.
- Improve automation, observability, and operational efficiency.
What We Value
- Ownership and accountability
- Systems thinking
- Continuous improvement
- Collaboration and knowledge sharing
Why Join Us?
- Work on a modern SaaS platform.
- Build reliable systems at scale.
- Solve complex production challenges.
- Grow with a high-performing engineering team
- Flexible remote work policy
- Performance-based bonus
مهارتهای مورد نیاز
- نگهداری سرور
- SRE
- Docker
حداقل سابقه کار
- سه تا شش سال
جنسیت
- مهم نیست
وضعیت نظام وظیفه
- معافیت تحصیلی معافیت دائم پایان خدمت