This is the most foundational engineering role on the platform. You will own and lead the implementation of two core backbone services: a decision orchestration engine that coordinates multi-step, stateful workflows, and a policy / authorization engine that governs what those workflows are allowed to do. Together these form the deterministic, fail-closed, auditable core that the rest of the platform depends on.
You will set the technical direction for these services — the distributed-systems design, the state and consistency model, the API contracts, and the safety and audit guarantees — and lead their build to a production standard. You will work closely with our AI engineering team, who own the AI infrastructure, model gateway, and inference layers, while you provide the governed execution backbone beneath them.
What you will do
Architect and implement a workflow orchestration engine for multi-step, long-running, stateful processes — including state-machine modelling, event-driven execution, and durable workflow state.
Design and build a policy / authorization engine supporting fine-grained access control (RBAC / ABAC or OPA-style policy-as-code), evaluated deterministically and fail-closed.
Develop a rule / decision DSL and the evaluation engine behind it, so that governance logic is expressed declaratively, versioned, and testable.
Guarantee correctness under failure through idempotency, distributed locking, exactly-once / at-least-once semantics where appropriate, retries, and fault recovery.
Own the API contracts between the backbone and every consuming service — clear, versioned, backward-compatible interface design.
Build for multi-tenancy and isolation with security, data separation, and tenant-aware policy enforcement as first-class concerns.
Make the platform auditable with comprehensive, tamper-evident audit logging and decision provenance suitable for regulated, compliance-driven environments.
Set engineering standards for the backbone — design reviews, testing strategy, observability, and the deterministic / fail-closed patterns the rest of the team builds on.
What we are looking for (required)
Distributed systems depth
Workflow orchestration & state machines
Event-driven architecture — strong command of asynchronous, message- or event-driven design.
Correctness primitives — hands-on experience with idempotency, distributed locking, consistency models, and fault / failure recovery.
API contract design —
Policy / authorization engines
DSL / rule-engine development — you have designed or implemented a domain-specific language or rule/decision engine.
Security, multi-tenancy & audit
Technical skills & stack
Languages — strong command of a systems / backend language such as Go, Rust, or Java / Kotlin, with the maturity to choose the right tool for the job. Python familiarity for working alongside the AI engineering team.
Workflow orchestration — durable-execution / workflow engines such as Temporal (or Cadence, Camunda / Zeebe, AWS Step Functions); building long-running, stateful, recoverable workflows.
Policy & authorization — policy-as-code and fine-grained authorization in production: OPA / Rego, OpenFGA or Zanzibar-style systems (e.g. SpiceDB), Cerbos, Oso, or Casbin.
Rule engines & DSLs — expression / rule evaluation (e.g. CEL, Drools) and parser / interpreter work (ANTLR or hand-built); having designed and shipped a domain-specific rule language.
Messaging & eventing — event-driven, asynchronous architectures on Kafka, NATS, RabbitMQ, or Pulsar.
Data & state — PostgreSQL with strong transactional and consistency modelling; event-sourcing / append-only stores; Redis. Distributed SQL (CockroachDB / Spanner-style) a plus.
Coordination & locking — distributed locks, leader election, and coordination using Redis, etcd, ZooKeeper, or Consul.
Infrastructure & deployment — Docker, Kubernetes, and Terraform; operating services on a major cloud (AWS, GCP, or Azure).
Observability & operations — OpenTelemetry, Prometheus, Grafana, and distributed tracing; you instrument and run what you build.
Security & audit — secrets management (e.g. Vault), encryption in transit and at rest, tamper-evident / append-only audit logging, and compliance-aware design.
Testing & correctness — contract testing, property-based testing, and deterministic / simulation testing for distributed systems.
Strongly preferred
Model-governance patterns — familiarity with governing, gating, or controlling AI/ML model behaviour in production.
Deterministic execution — experience building systems where reproducibility and deterministic outcomes are a hard requirement.
Fail-closed safety mechanisms — you default to safe failure: when in doubt, the system denies rather than guesses.
Regulated-domain experience — prior work in healthcare, fintech, aerospace, defense, or another compliance-heavy industry.
eSora Labs is for builders, thinkers, designers, engineers, and domain experts who want to work across AI, software, design, healthcare, semiconductor, aerospace, and regulated technology.