استخدام SRE) Site Reliability Engineer-امریه سربازی-غائبین)
دستهبندی شغلی
IT / DevOps / Server
موقعیت مکانی
تهران
، تهران
نوع همکاری
تمام وقت
حداقل سابقه کار
سه تا شش سال
حقوق
توافقی
شرح موقعیت شغلی
About the Role
At Hami System Sharif, we build and operate large-scale national systems in payments, identity verification, and capital markets, systems that serve hundreds of thousands of users, where every minute of downtime has a cost. We are looking for a strong Site Reliability Engineer to guarantee their stability, performance, and availability. In this role, we expect you to own the reliability of our production services, design and implement observability and monitoring, bring structure to incident management, and work with the backend and DevOps teams to reduce toil and increase automation. We are looking for someone who understands both engineering and operations, stays calm and methodical in a crisis, is driven by measurement and data, and fixes problems at the root rather than patching them.
Responsibilities
• Define and track SLIs, SLOs, and error budgets for core services and follow up on SLA commitments to clients • Design and implement observability (metrics, logging, and tracing) with tools such as Prometheus, Grafana, and ELK • Design and manage intelligent alerting and reduce alert fatigue • Manage on-call rotations, respond to incidents, and lead the incident resolution process • Run blameless postmortems and follow corrective actions through to completion • Perform capacity planning and analyze service performance under real and test load • Design and run load tests and chaos engineering experiments to find weak points before they cause outages • Automate repetitive operational work (toil) with scripting and infrastructure-as-code tools • Work with the DevOps team to improve CI/CD, deployment strategies (blue-green, canary), and safe rollbacks • Review service architecture from the perspective of high availability, fault tolerance, and disaster recovery • Manage and optimize Kubernetes, containers, and related infrastructure services • Improve the stability and performance of databases, caches, and message queues together with the backend team • Write runbooks and operational documentation for services and failure scenarios • Monitor security requirements and infrastructure compliance with supervisory-body standards together with the network team
Requirements
• At least 3 years of relevant experience in SRE, DevOps, backend, or production operations roles • Bachelor's degree or higher in Computer Engineering, Software Engineering, IT, or a related field • Strong command of Linux and networking fundamentals (TCP/IP, DNS, load balancing, TLS) • Hands-on experience with Docker and Kubernetes in production • Experience with monitoring and observability tools such as Prometheus, Grafana, ELK, or equivalent • Programming and scripting ability in at least one of Python, Go, or Bash • Familiarity with infrastructure-as-code and configuration management tools such as Terraform or Ansible • Solid understanding of distributed systems, microservices, and high-availability concepts • Experience with incident management and root cause analysis on real-world services • Proficiency in Office and Google Workspace tools • Working knowledge of English for reading technical resources and documentation
Nice to Have
• Experience with fintech, payment, or banking systems with high availability requirements • Experience introducing SLOs / SLIs and an error-budget culture in an organization • Familiarity with PostgreSQL / MySQL, Redis, and Kafka from an operations and tuning perspective • Performance testing experience with tools such as k6, JMeter, or Gatling • Certifications such as CKA, CKAD, RHCE, or equivalent • Experience working in growing organizations or fast-paced environments
What We Value
• Ownership: You treat production stability as your responsibility, not just closing tickets. • Blameless Mindset: When investigating incidents, you look for systemic causes, not culprits. • Automation First: You see every repetitive task as an opportunity to automate. • Calm under Pressure: During an outage, you make decisions methodically and communicate clearly. • Continuous Improvement: You are always looking for ways to improve reliability, performance, and operational processes.
شرایط کاری و مزایا :
• فرصت تجربهاندوزی و بستری مناسب برای ارتقای مهارت و دانش شغلی • بیمه تامین اجتماعی و تکمیلی از روز اول شروع به کار • محیط کاری پویا، حرفهای و دوستانه • امکان ارتقا و پیشرفت کاری • موقعیت جغرافیایی بهینه و در دسترس با وسایل حمل و نقل عمومی به تمامی نقاط شهر • محل کار: تهران - طرشت (نزدیک ایستگاه مترو شریف) • فرصت عالی برای امریه سربازی آقایان (غائبین)
فعالیت «حامی سیستم شریف» بر اجرای پروژههای نرمافزاری بزرگمقیاس متمرکز است.
طی سالهای گذشته «همانا»، پلتفرم احراز هویت غیرحضوری و «پیوب» پلتفرم پرداختیاری توسعه یافته در حامی سیستم شریف، میزبان صدها هزار کاربر نهایی بوده است.
توسعه اپلیکیشن و درگاههای «دولت همراه»
پیادهسازی سامانههای بورسی در «شرکت سپردهگذاری مرکزی (سمات)»
پیادهسازی پتلفرم ارائه «سهام عدالت»
و همچنین بیش از ده سال پیمانکاری توسعه نرمافزارهای «همراه اول» از جمله ارزشآفرینیهای حامی سیستم شریف به عنوان شریک راهبردی نهادها، سازمانها و کسبوکارهای مطرح کشور است.