Cold-start
Stop paying the cold-start tax on every tenant.
isovert lets tenants safely share one warm model instead of loading per tenant. On a single node, across three transformer families, we measured 90×–1000× cold-start reduction. We're looking for one serving partner to prove it at fleet scale — and capture the saving first.
The cost we remove
The per-tenant load penalty, at scale
isovert changes the shape of the problem: a shared warm model, provable isolation between tenants, and the per-tenant load penalty largely gone.
Serving cost
Every tenant that needs its own model instance burns GPU-hours idling warm — or pays to reload on demand. The bill scales with tenants, not usage.
First-token latency
A cold tenant waits for a full model load before the first token. The usual fix — keeping everything warm — just moves the cost back onto the GPU bill.
The trade-off
Keep everything warm and it is expensive; accept cold-start and it is slow. Either way you are paying the per-tenant load penalty at scale.
What we measured — and what we didn't
90×–1000×, and the caveats that make it credible
One capture per model, on a single node. The spread is large and model-dependent — so here is the full range, each figure with its note, not a cherry-picked headline.
Scope. Measured on a single hardware configuration, transformer-only — it does not port to state-space models. One capture per model; variance is large and model-dependent. Cross-hardware and fleet-scale are unproven — that is the design-partnership deliverable, not a claim.
The joint-validation offer
A bounded benchmark on your stack
- →
A bounded, roughly two-week benchmark run on your serving stack.
- →
You set the success bar and the kill criterion up front.
- →
A fast, honest no if it does not port to your hardware — the caveats are the point.
- →
If it holds, you capture the serving-cost saving before anyone else.
Why safe sharing, not just caching
The saving comes from provable isolation
Caching tricks can warm a model, but they can't let different tenants safely share one — shared model state leaks across tenants. The saving here comes from provable multi-tenant sharing, the isovert isolation core, which is what makes one warm model safe to share in the first place.
That means the cost win composes with a security property your regulated customers already want: per-tenant isolation you can prove, on commodity hardware, with no foreign hardware root of trust.