Technical overview
How Holio Cloud runs your service
What happens between pushing an image and serving traffic, what each plan is allowed to use, and what is not built yet. Written for the engineer who has to say yes.
- Runtime
- Kubernetes, gVisor per tenant
- Images
- OCI, pinned by digest
- Location
- Our own servers in Germany

Architecture
Six parts, each with one job
Control plane
An API that stores desired state, and a separate reconciler that acts on it. The API never holds the credentials that change the runtime.
Sandbox per tenant
Each company gets its own namespace: gVisor, restricted pod security, non-root containers, a quota from the plan and deny-by-default networking.
Immutable revisions
A revision pins the image digest, port, concurrency, timeout, instances, environment and probes. Nothing about it changes after it exists.
Edge
HTTPS only. Custom domains are verified by DNS, certificates are issued and renewed for you, and traffic can be split between revisions by weight.
Secrets
Write-only. Encrypted per tenant with AES-256-GCM, decrypted only at deploy time, and delivered to one revision.
Logs and usage
Logs per revision with secrets redacted. CPU, memory, requests and egress metered per service, readable by your company only.
The life of a deploy
A rollout never replaces a healthy revision with one that has not proved itself.
Admitted
The image must come from an approved registry, be pinned by digest and pass the signature and vulnerability policy. A tag is refused.
Queued
The API stores the revision and answers 202. A reconciler picks it up from a durable queue, holding a lease so no other worker starts the same rollout.
Staged and probed
The candidate starts next to the live revision. A startup probe gates readiness and liveness, so a slow boot is not mistaken for a hang.
Switched, or retired
Traffic moves only when the candidate is ready. A candidate that times out is retired and the last healthy revision keeps serving. Rollback points traffic at an earlier revision.
Plan limits
Each plan is a hard ceiling, enforced by the container's cgroup and the tenant's namespace quota.
| Starter | Standard | Scale | |
|---|---|---|---|
| CPU per instance | 0.25 vCPU | 1 vCPU | 1 vCPU |
| Memory per instance | 512 MiB | 1 GiB | 1 GiB |
| Instances (max) | 1 | 2 | 8 |
| Ephemeral storage | 1 GiB | 4 GiB | 8 GiB |
| Processes (PIDs) | 128 | 256 | 512 |
| Egress bandwidth | 10 Mbit/s | 50 Mbit/s | 100 Mbit/s |
| Deploys per hour | 6 | 30 | 60 |
| Secrets | 20 | 50 | 100 |
| Requests per month | 5 million | 25 million | 50 million |
| Egress per month | 25 GiB | 100 GiB | 200 GiB |
Crossing a monthly allowance records an audit event and the service keeps serving. The ceilings above are never crossed.
Revision settings
A revision snapshots everything it runs with. Changing any of it means a new revision.
| Setting | Range | Default |
|---|---|---|
image | OCI reference pinned by @sha256: digest | required |
port | 1–65535 | 8080 |
containerConcurrency | 1–1000 requests per instance | 80 |
timeoutSeconds | 1–3600, enforced at the edge | 300 |
minInstances / maxInstances | 0 … plan maximum | 1 / plan |
env | up to 64 variables, 32 KiB in total | — |
command / args | argv lists, up to 64 entries each | the image's own |
| Probes | startup, readiness and liveness, TCP or HTTP | readiness + liveness |
| Private images | a pull secret named from the service's secrets | public |
API
| What | Limit |
|---|---|
| Rate limit, reads | 20/s per company, burst 100 |
| Rate limit, writes | 5/s per company, burst 30 |
| Over the limit | 429 with Retry-After |
| Writes | require an Idempotency-Key header, so a retry never deploys twice |
| Deploy | answers 202 queued; the revision's state tells you what happened |
A deploy, as the API sees it
One request: POST /v1/services/{serviceId}/revisions with an Idempotency-Key header. Only image is required; everything else falls back to the defaults in the table above.
| Field | Example |
|---|---|
image | ghcr.io/acme/shop-api@sha256:9c1e…a87b |
port | 8080 |
containerConcurrency | 80 |
timeoutSeconds | 300 |
minInstances / maxInstances | 1 / 2 |
env | LOG_LEVEL=info |
The answer is 202 queued. The same route is behind the Cloud tab in Holio and the holio CLI, and is authorised by the roles your company already has in Holio. A mutable tag such as :latest is refused.
Not built yet — early access, honestly
Can a service scale to zero?
The contract is there: a revision may set minInstances to 0. Request-driven scaling is switched off during early access, so a revision rests at its floor.
Can Holio Cloud build from source?
Not yet offered. You deploy an image you have built. Source builds pinned to a commit are on the roadmap.
Is there a WAF?
Every service has an edge policy: a per-client-IP rate limit with burst, an in-flight ceiling, an optional IP allow list and a body cap. Managed WAF rule sets and IP reputation are not built.
How many regions?
One: our own servers in Germany. There is no multi-region failover or anycast edge yet.
Does a timeout cut a streaming response?
No. The timeout bounds how long a response may take to START. A response that is already streaming is not cut.
Want to see it with your own image?
Tell us what you run. We answer within one working day.