Our platform is built on large-scale, real-world B2B data — people, companies, and unstructured professional content. Infrastructure, data pipelines, and our Elasticsearch stack are owned by a capable engineering team; your focus is the intelligence layer: making messy data clean and trustworthy, turning unstructured text into structured attributes, resolving entities across sources, and building the similarity and relationship models that power matching across the platform.
What You'll Do
Design the cleaning, normalization, and validation logic for messy real-world data (malformed dates, inconsistent fields, free text) — combining rules and models, with measurable quality metrics.
Build extraction and inference models that turn unstructured and semi-structured text into reliable structured attributes, and maintain the resulting taxonomies.
Own entity resolution: detect and merge duplicate or matching records (people, companies) across noisy, inconsistent sources.
Build and continuously improve similarity and matching models between entities — embeddings, semantic matching, and ranking — with rigorous relevance evaluation.
Model relationships between entities, including multi-hop connections, and design meaningful scoring for them.
Apply LLMs pragmatically for extraction and classification where they outperform classical methods — with real evaluation sets and cost budgets.
Define precision/recall targets, build golden datasets, and set up human-in-the-loop labeling where needed.
Work closely with engineering to get your models into production — you own the logic and quality; they own pipelines and serving.
What We're Looking For
4+ years as a Data Scientist / ML Engineer working on messy real-world data, with a focus on NLP / information extraction or data mining.
Hands-on experience with entity resolution / record linkage and large-scale data cleaning.
Strong NLP fundamentals: text classification, NER, embeddings — plus practical, well-evaluated use of LLMs for extraction.
Experience building or substantially improving similarity / semantic matching systems.
Strong Python and excellent SQL.
Evaluation discipline: you measure quality with golden sets and precision/recall, not anecdotes.
Ability to explain data decisions clearly to technical and non-technical stakeholders.
Nice to Have
Graph data modeling and multi-hop traversal (Neo4j, NetworkX, link prediction) — a strong plus; this role grows into it.
Familiarity with Elasticsearch and vector/kNN search.
PyTorch for custom models; MLflow or similar tooling.
Familiarity with B2B data (firmographics, professional profiles).
روند AI یک استارتاپ نوپا در حوزه دیجیتال مارکتینگ و هوش مصنوعی است که با هدف کمک به کسبوکارها برای رشد هوشمندتر و استفاده موثرتر از داده، محتوا و تکنولوژی شکل گرفته است.
ما با ترکیب استراتژی بازاریابی، تولید محتوا، اتوماسیون و ابزارهای مبتنی بر هوش مصنوعی، به برندها کمک میکنیم فعالیتهای بازاریابی خود را هدفمندتر، سریعتر و قابلاندازهگیریتر پیش ببرند.
تمرکز ما روی استفاده کاربردی از AI در فرآیندهای واقعی بازاریابی است؛ از تحلیل داده و شناخت مخاطب گرفته تا تولید محتوا، بهینهسازی کمپینها و خودکارسازی فعالیتهای تکرارشونده.
بهعنوان یک تیم نوپا، محیط کاری ما بر یادگیری، تجربهگرایی، سرعت و همکاری نزدیک میان اعضای تیم بنا شده است.
هدف ما ساخت راهکارهای پیچیده نیست؛ هدف ما استفاده هوشمندانه از تکنولوژی برای حل مسائل واقعی بازاریابی و ایجاد رشد قابلاندازهگیری برای کسبوکارهاست.