The slow way, and why it is the obvious way
The intuitive design is to do what you would do by hand: create a volume, attach an installer ISO, boot it, answer the questions from a preseed file, install the packages, reboot, configure networking, add the user's SSH key.
It works, it is easy to reason about, and it takes three to five minutes — most of which is the installer unpacking and writing packages that will be identical on every single machine you ever build.
That is the tell. Work repeated identically every time is work that should have been done once.
Cloning instead
So it is done once. A golden image per operating system is built ahead of time: installed, updated, hardened, cloud-init present, machine-specific identity stripped out. Provisioning then copies that image instead of creating one.
Copying 2 GB still takes time, so the copy is avoided too. A copy-on-write layer over a shared read-only base means the new instance starts with an empty overlay and only writes blocks that actually change:
# the base, built once, never written to again
/images/ubuntu-24.04-base.qcow2
# a new instance: a few hundred KB, not 2 GB
qemu-img create -f qcow2 \
-b /images/ubuntu-24.04-base.qcow2 -F qcow2 \
/instances/vm-1234.qcow2 80G
That command returns in milliseconds. The disk is logically 80 GB and physically almost nothing. Three minutes of installer just disappeared.
Where the 38 seconds actually go
| Step | Roughly |
|---|---|
| Validate order, pick a host with headroom | < 1 s |
| Allocate IPv4 and IPv6, write DNS | ~2 s |
| Create the copy-on-write disk | < 1 s |
| Write the cloud-init seed (keys, hostname, network) | ~1 s |
| Define and start the domain | ~2 s |
| Guest boot to multi-user | ~18 s |
| cloud-init: identity, keys, first-boot work | ~9 s |
| Network converges, health check passes | ~4 s |
Over two thirds of it is the guest booting and configuring itself. Nothing in the control plane is the bottleneck any more, which is the correct place to end up.
The last ten seconds are the hard ones
Cutting minutes was easy — it was one architectural decision. Cutting the remaining seconds means arguing with a Linux boot, and the wins are small and specific:
- Disable services the image does not need.
systemd-analyze blameis unglamorous and effective. On a stock cloud image there is usually four or five seconds of things nobody wants. - Do not wait for network-online.target unless something genuinely requires it. It is a common and expensive default.
- Pre-seed the machine ID and SSH host keys per instance rather than generating them at first boot. Host key generation alone can cost a second or two of entropy-gathering.
- Trim cloud-init modules to the ones that do something. The default set runs a great deal that is irrelevant on a first boot you control completely.
Each is worth a second or two. Together they are the difference between 38 seconds and just under a minute, which is the difference between "that was fast" and "is it broken?".
The optimisation nobody should make
Instances could be pre-booted and held in a warm pool, handed out on demand in about two seconds. Almost nobody does, for two reasons: a pre-booted instance consumes RAM on a host while nobody owns it, and a pool sized for a spike is idle the rest of the time. The cost of that idle capacity ends up in the price of every server, including everyone's who did not need it in two seconds.
Thirty-eight seconds is fast enough that nobody is waiting on it, and the alternative is a permanent tax on people who never noticed the difference.
Windows is its own thing
Everything above describes Linux. Windows images are larger, sysprep is slow, and first boot does a lot of work that cannot be skipped. A Windows VPS takes five to ten minutes and we say so up front rather than quoting the Linux number and hoping.
You can watch the Linux number for yourself on any KVM VPS. If one ever takes materially longer, that is a fault and we would like to hear about it on +91 75994 50220.