آگهی‌های استخدامی

استخدام Senior Data Engineer

بیت‌ پین | Bitpin
تهران، تهران

شرح موقعیت شغلی

As a Data Engineer on Bitpin's AI & Data team, you will help architect and operate the data platform behind our cryptocurrency exchange. You will own the data's journey from source to value  building pipelines from our operational PostgreSQL databases into our ClickHouse analytical warehouse, serving governed data marts to BI consumers, and managing the data ecosystem for our Qdrant-powered AI applications.

Key Responsibilities:

  • ETL/ELT Pipeline Development: Design, build, and maintain robust, scalable data pipelines using Apache Airflow and PySpark to transform data from PostgreSQL sources into our ClickHouse analytical warehouse.
  • Real-Time & Streaming Ingestion: Build and operate streaming ingestion paths that move data from operational systems into the warehouse with low latency, ensuring correct ordering, deduplication, and delivery semantics.
  • ClickHouse Modeling & Optimization: Design performant ClickHouse data models and write highly optimized SQL for large-scale analytical workloads.
  • Performance Tuning: Continuously profile and optimize query performance, storage layout, and resource utilization across the analytical warehouse as data volumes grow.
  • Data Quality & Reliability: Implement validation, reconciliation, and completeness checks across pipelines to guarantee that downstream consumers can trust every number — no silent data loss, no unnoticed drift.
  • Monitoring & Observability: Instrument pipelines and data services with metrics, alerting, and freshness monitoring so that failures and anomalies are detected before stakeholders notice them.
  • Application-Level Service Administration: Take full ownership of the configuration and administration of our data services. This includes creating Kafka topics and managing ACLs, designing database schemas, managing users and permissions, and setting up S3-compatible object storage buckets.
  • Workflow Orchestration: Develop, schedule, and monitor complex data workflows (DAGs) in Apache Airflow, ensuring they are efficient, reliable, and well-documented.
  • Serving Analytics & Governance: Build and maintain governed data marts consumed by our BI layer (Metabase) and data catalog (OpenMetadata), including supporting access-control enforcement jobs.
  • AI Data Provisioning: Develop processes to prepare, embed, and load data into our Qdrant vector database, often orchestrating these jobs within Airflow.
  • Collaboration & Ownership: Work closely with software developers, analysts, and AI specialists to define their data needs and build the solutions to meet them.
What We're Looking For:


  • Proven experience as a Data Engineer.
  • Hands-on experience with ClickHouse and its feature set.
  • Experience with CDC pipelines (Debezium, Kafka Connect) and streaming data processing paradigms.
  • Expert-level proficiency in Python for data processing, with experience in PySpark for transformation workloads.
  • Deep knowledge of SQL and experience with relational databases (PostgreSQL is a must).
  • Hands-on experience developing, scheduling, and monitoring complex data pipelines using Apache Airflow.
  • Experience with ETL/ELT architecture, data modeling, and data warehousing concepts.
  • Familiarity with containerizing applications using Docker.
Bonus Points:

  • Familiarity with BI tools (Metabase) and data catalogs (OpenMetadata).
  • Experience with vector databases (Qdrant, Weaviate, etc.).

مهارت‌های مورد نیاز

  • Python
  • SQL
  • Data engineer

حداقل سابقه کار

  • سه تا شش سال

جنسیت

  • مهم نیست

وضعیت نظام وظیفه

  • مهم‌ نیست

نوع همکاری:

تمام وقت

تاریخ انتشار آگهی:

۱۴۰۵/۰۵/۰۴
ارسال رزومه