آگهی‌های استخدامی

استخدام Data Engineer (مشهد)

گروه حقوق آسیا | Asia Law Group
خراسان رضوی، مشهد

شرح موقعیت شغلی

About the Role

We are looking for a data engineer to build reliable pipelines that transform diverse Persian-language documents into accurate, traceable, and up-to-date knowledge-base content.

Project details will be shared during later recruitment stages under a confidentiality agreement. We prefer on-site collaboration in Mashhad.

Responsibilities

• Build ingestion pipelines for files, APIs, and authorized web sources, including PDF, scanned documents, Word, and HTML.
• Implement and improve Persian OCR, flagging low-quality outputs for review.
• Clean and normalize Persian text while preserving headings, tables, footnotes, page references, and relationships between sections.
• Design storage for original files, processed text, and metadata, including identifiers, topics, versions, permissions, publication dates, validity periods, and ingestion timestamps.
• Detect duplicates, preserve meaningful version differences, and maintain historical records and data lineage.
• Structure narrative records and case studies into problems, actions, outcomes, and linked evidence; remove identifying information before shared use.
• Prepare, chunk, and index data according to designs developed with the AI/RAG engineer.
• Monitor sources, detect changes, perform incremental updates, and propagate modifications and deletions to processed text, chunks, search indexes, and vector data.
• Build scheduled and batch workflows with retries, failure recovery, reprocessing, and duplicate-safe execution.
• Implement automated quality checks, exception review workflows, logging, and alerts for processing failures or delayed updates.
• Optimize processing speed, cost, and resource usage for large document collections.
• Enforce access controls, user data isolation, and retention and deletion policies.
• Document data structures, workflows, and backup and recovery procedures.

Required Skills

• Strong Python and SQL skills and practical experience with ETL/ELT pipelines.
• Experience with relational databases such as PostgreSQL and file or object storage.
• Experience processing documents, using OCR, and handling Persian text challenges.
• Familiarity with API integration and authorized web data extraction.
• Ability to design data schemas, manage metadata, deduplicate records, and maintain versions.
• Experience with workflow scheduling, error handling, and data quality controls.
• Proficiency with Git and familiarity with Linux, Docker, and maintainable code development.
• Ability to read English technical documentation and collaborate with AI, backend, and content teams.
• Careful handling of confidential information and adherence to data protection requirements.

Preferred Qualifications

• Understanding of LLMs and the differences between model training, fine-tuning, and RAG.
• Experience building document ingestion pipelines for AI knowledge bases while preserving structure, provenance, and versions.
• Familiarity with chunking, embeddings, vector indexing, keyword and semantic search, and hybrid retrieval.
• Experience maintaining knowledge-base updates and deletions across dependent components.
• Familiarity with preparing training, fine-tuning, and evaluation datasets, including formatting, deduplication, and preventing train–test leakage.
• Ability to troubleshoot data, chunking, and metadata issues with the AI/RAG engineer.
• Experience with Qdrant, pgvector, Elasticsearch, or OpenSearch.
• Experience with Airflow, Prefect, Dagster, or similar orchestration tools.
• Experience extracting complex Persian document layouts, tables, and footnotes.
• Experience with parallel processing, task queues, large document collections, change detection, and data lineage.
• Experience detecting personal information and preparing confidential data.

A demonstrable RAG knowledge-base ingestion project is a strong advantage.

Expected Deliverables and Role Scope

Reliable, monitored ingestion and update pipelines; structured, traceable data; quality reports; and maintenance documentation.

This role owns data infrastructure and preparation and collaborates with the AI/RAG engineer on retrieval integration. Subject-matter specialists are responsible for validating source content.

Working Arrangement and Application

On-site collaboration in Mashhad is preferred. Please include your current location and on-site availability.

Where possible, submit a relevant project example describing your contribution, data scale, error handling, and quality controls. For confidential projects, a high-level description without disclosing protected information is sufficient.

مهارت‌های مورد نیاز

  • SQL
  • Python
  • Data engineer
  • Big Data

حداقل سابقه کار

  • سه تا شش سال

جنسیت

  • مهم نیست

وضعیت نظام وظیفه

  • مهم‌ نیست

نوع همکاری:

تمام وقت

تاریخ انتشار آگهی:

۱۴۰۵/۰۶/۳۰
ارسال رزومه