A distributed AI inference platform — multi-tenant, event-driven, and built to be operated: queued inference requests dispatched across a worker pool, with the failure modes, backpressure, and observability that a real serving platform needs.
Polyglot on purpose: a TypeScript edge, a Go control plane, Go workers.
Status: in development. The architecture below is the target design. This README is honest about what exists — see Build status for what is actually running today. Nothing here is claimed as complete before it is.
Most inference examples are a single process wrapping a model behind an HTTP handler. That hides everything that is actually hard about serving inference at scale:
- Requests are long-running and expensive, so the queue — not the model — is where the system succeeds or fails.
- Load is bursty and unfair; one tenant can starve every other one.
- Workers die mid-inference, and a retry that re-runs a completed job costs real money.
- Backpressure has to be explicit. An inference queue with no admission control degrades into unbounded latency rather than clean rejection.
MiniInfer is built around those problems rather than around the model call.
┌──────────────┐
client ────────► │ gateway-ts │ auth (OIDC), rate limit, admission control,
│ (Node/TS) │ correlation IDs, streaming responses
└──────┬───────┘
│ gRPC
▼
┌──────────────┐ ┌────────────┐
│dispatcher-go │◄──────►│ Postgres │ task state, outbox
│ (Go) │ └────────────┘
│ task state, │ ┌────────────┐
│ scheduling, │◄──────►│ Redis │ leases, rate limits
│ outbox │ └────────────┘
└──────┬───────┘
│ Kafka (task.submitted / task.completed)
▼
┌──────────────┐
│ worker-go │ pulls work, runs inference, reports
│ (Go, ×N) │ progress, honours cancellation
└──────────────┘
cross-cutting: Keycloak (identity) · Prometheus + Grafana (metrics)
OpenTelemetry traces · Kubernetes (multi-node, from the K8s phase)
Boundary rule: no shared database and no shared code between services. They talk over gRPC and Kafka only. Each Go service is its own module; each Node service its own package.
| Service | Language | Responsibility | Owns |
|---|---|---|---|
gateway-ts |
TypeScript / Node | Client-facing API, authn/authz, rate limiting, admission control, correlation IDs, response streaming | no persistent state |
dispatcher-go |
Go | Task lifecycle and state machine, scheduling and fairness, transactional outbox, retry and cancellation | Postgres (task state), Redis (leases) |
worker-go |
Go | Consumes tasks, executes inference, reports progress, honours cancellation and lease expiry | no persistent state |
The tradeoffs — and, more importantly, the alternatives that were rejected and what they would have cost — are recorded as the system is built. Each entry names the decision, the alternative, and the measured or expected cost.
Decisions settled so far:
| Decision | Chosen | Rejected | Why |
|---|---|---|---|
| Service boundaries | 3 services split by capability | Entity-per-service (UserService, TaskService, …) |
Boundaries follow business capability and data ownership; entity-splitting produces a distributed monolith where one feature touches every service |
| Inter-service state | No shared DB or shared code | Shared schema for convenience | Two services writing one table are one service in disguise; the constraint is what makes independent deployability real |
Further decisions — tenancy model, delivery semantics, dispatch fairness, storage layout, cluster topology — are added here as they are made and measured.
| Component | State |
|---|---|
| Repository scaffold | ✅ |
gateway-ts |
⬜ not started |
dispatcher-go |
⬜ not started |
worker-go |
⬜ not started |
infra/compose |
⬜ not started |
infra/k8s (multi-node) |
⬜ not started |
Nothing is runnable yet. Setup instructions land with the first service, and are kept accurate — if a command is in this README, it works.
services/
gateway-ts/ TypeScript — client API, auth, streaming
dispatcher-go/ Go — task state, scheduling, outbox
worker-go/ Go — inference execution
infra/
compose/ docker-compose, split per phase (bring up only what you need)
k8s/ Kubernetes manifests — multi-node via kind, 3 workers
prometheus/ scrape config and alert rules
grafana/ dashboards and provisioning