Role Overview
At NOVA, our agent execution runtime is the beating heart of everything. It is the layer that schedules, isolates, checkpoints, and resumes hundreds of thousands of concurrent AI workloads with deterministic precision. As a Distributed Systems & Runtime Engineer, you will own this critical infrastructure — designing the primitives, protocols, and scheduling algorithms that allow enterprise customers to trust NOVA with mission-critical autonomous operations.
Responsibilities
Design and implement the core ephemeral agent execution runtime capable of sustaining 100,000+ concurrent state machines with sub-50ms p99 latency.
Build distributed task scheduling systems using consistent hashing, work-stealing, and priority queuing with automatic back-pressure.
Architect fault-tolerant checkpoint and resume protocols so agent workloads survive hardware failures and network partitions without data loss.
Own the agent lifecycle management layer: spawn, suspend, resume, snapshot, migrate, and terminate with zero-downtime semantics.
Collaborate with the AI Safety team to embed sandboxed tool execution environments within the runtime using Linux namespaces and seccomp profiles.
Requirements
7+ years of systems engineering experience, ideally in distributed systems, streaming infrastructure, or high-throughput compute platforms.
Deep expertise in one of: Go, Rust, or C++; strong familiarity with at least one additional systems language.
Strong understanding of consensus protocols (Raft, Paxos) and distributed coordination primitives (etcd, ZooKeeper).
Proven experience designing and operating systems at 10,000+ RPS in production environments.
Compensation & Benefits
Competitive base salary ($220K–$290K) plus generous early-stage equity
Full medical, dental, and vision coverage for you and dependents
10-year equity exercise window
Minimum recommended vacation with mandatory year-end company break
