Part AI Research Center is looking for a Senior DevOps Engineer to help us build, evolve, and operate the infrastructure behind our growing AI products.
We’re looking for someone who is not only comfortable working with modern infrastructure and DevOps tooling, but who can also bring their own experience, ideas, and best practices to the team.
You’ll have the opportunity to work closely with backend, AI/ML, and QA teams, help shape our infrastructure, improve reliability and delivery, and solve interesting problems as our systems scale.
WHAT YOU’LL WORK ON :
Build, maintain, and improve infrastructure using Kubernetes, Docker, and Ansible
Design, maintain, and improve CI/CD pipelines using GitLab CI and Argo CD
Manage object storage solutions such as MinIO or Ceph, including data lifecycle and retention policies
Manage secrets, credentials, and access control using HashiCorp Vault or similar solutions
Build and improve monitoring, observability, and logging using tools such as Prometheus, Grafana, ELK, and OpenSearch
Write automation and operational tooling using Python and Bash
Work closely with developers, AI/ML engineers, and QA to make deployments and services reliable and easy to operate
Help improve infrastructure reliability, scalability, security, and automation
Contribute to infrastructure architecture and technical decisions as the platform evolves
AI & MODERN INFRASTRUCTURE
Experience or familiarity with AI infrastructure and AI/ML systems is a strong plus.
You don't necessarily need to be an AI engineer, but understanding how AI workloads are deployed and operated — and being comfortable working alongside ML engineers — would be very valuable.
Experience with systems such as OpenSearch, Grafana, Vault, Prometheus, Kubernetes, and similar infrastructure and observability platforms is highly useful.
We’re also very interested in people who have experience solving infrastructure problems at scale and can share what has worked for them in previous environments.
DISASTER RECOVERY & RELIABILITY
Experience in Disaster Recovery Planning (DRP), and especially hands-on experience designing and implementing disaster recovery strategies, is highly valued.
If you’ve previously designed DR plans, backup and restore strategies, failover mechanisms, high-availability architectures, or have experience recovering production systems from real incidents, we’d genuinely like to learn from that experience.
BONUS SKILLS
Experience with some of the following would be a strong advantage:
PostgreSQL and MongoDB — performance tuning, replication, backup, and recovery
RabbitMQ or other messaging systems
Redis and caching architectures
•Networking and load balancing with HAProxy, Nginx, Traefik, or similar tools
Linux system administration and troubleshooting
High-availability and fault-tolerant architectures
Infrastructure security and access management
Performance monitoring and capacity planning
ON-CALL & PRODUCTION OPERATIONS
Our infrastructure supports production services, so on-call availability is an important part of this role.
We need someone who is comfortable with the reality of operating production infrastructure: occasionally investigating an incident outside normal working hours, responding to infrastructure issues, and helping keep critical services running.
This doesn't mean being constantly available or working around the clock. We aim to have a structured and reasonable on-call process, but being willing and prepared to participate in on-call rotations is a requirement for this position.
WHAT WE VALUE
More than knowing a particular list of tools, we value:
Strong production experience and solid troubleshooting skills
The ability to understand systems end-to-end
A practical approach to reliability and automation
Curiosity and willingness to learn new technologies
The ability to explain technical decisions and share knowledge with the team
Experience that can help us avoid reinventing the wheel
Someone who can say: “I’ve seen this problem before, and here’s how we solved it.
Our stack will continue to evolve, so we’re not looking for someone who only follows a predefined playbook. We’re interested in someone who can help us improve that playbook and bring their own ideas and experience to the table.
If you have strong experience in a particular area that isn't explicitly listed here,we'd still love to hear about it .
«شرکت دانشبنیان پردازش اطلاعات مالی پارت» با توسعه راهکارهای فناورانه و خدمترسانی در سه حوزه «مالی و سرمایهگذاری»، «هوش مصنوعی» و «فناوری و زیرساخت»، موفق شده است که میزبان بیش از 10 میلیون کاربر و سازمان گوناگون باشد.
«مرکز تحقیقات هوش مصنوعی پارت» به عنوان بازوی هوش مصنوعی این شرکت، فعالیت خود را در سال 1396 باهدف تبدیلشدن به بزرگترین و پیشروترین بازیگر هوش مصنوعی ایران و خاورمیانه آغاز کرد. در طی این سالها، پارت موفق شد با توسعه زیرساختهای پردازشی و نرمافزاری قدرتمند و تشکیل عظیمترین مجموعه هوش مصنوعی کشور از لحاظ نیروی انسانی از طریق گردآوری بهترین متخصصان این حوزه، دسترسی طیف وسیعی از مشتریان خود به بیش از 100 سرویس توسعه یافته هوشمند را تسهیل کند و به شعارش که همان «هوشمندسازی فرایندهای زندگی» است رنگ واقعیت ببخشد.
گروههای چهارگانه بینایی ماشین، پردازش زبان طبیعی، پردازش گفتار و دادهکاوی مجموعه پارت، سبد گستردهای از محصولات هوشمند را توسعه دادهاند. از جمله:
فراشناسا (راهکار جامع احراز هویت الکترونیکی)؛
دانابات (پاسخگوی هوشمند کسبوکارها)
آواشو (خوانشگر هوشمند متون)
آوانگار (نگارشگر هوشمند گفتار)
نویسهنگار (تبدیگر هوشمند نگاره به متن)
کالج تخصصی هوش مصنوعی؛
دیدهبان هوش مصنوعی؛
و ...
پارت میکوشد تا با فراهمآوردن فضایی پویا و زمینهسازی رشد و توسعه فردی شما، رنگ تازهای به تجربه شغلیتان ببخشد.