The class of failure
Most capacity problems announce themselves gradually. CPU gets slower, memory pressure shows up in swap, disks fill at a rate you can extrapolate. You get warning, and the warning is proportional to the problem.
IPv4 addresses do not behave like that. A pool with two addresses left behaves exactly like a pool with two hundred, right up until the moment it behaves like a pool with none. There is no degradation, no slow-down, no partial service. It works, and then provisioning fails outright.
This is worth understanding whether you run a hosting platform or a single VPS, because the same shape of failure appears anywhere a resource is allocated rather than shared.
Why it is different from CPU and RAM
CPU and memory can be overcommitted. Hosts routinely allocate more vCPU than there are physical cores, because not every guest is busy at once. It is a bet, it usually pays, and when it does not the symptom is contention rather than failure.
An IP address cannot be overcommitted. Two machines cannot share one routable address. The allocation is binary and exclusive, so the pool is a hard count and the failure is a cliff.
Anything with that property — floating IPs, licence seats, VLAN tags, port numbers on a NAT gateway — needs a different kind of monitoring from the things that degrade.
The check that reports green while the pool is empty
The naive health check asks: can I allocate an address? It allocates one, confirms it worked, releases it, reports healthy.
That check is green when one address remains. It is green when one address remains and forty servers are being ordered. It only goes red after the failure it was supposed to predict.
# the check that tells you nothing
def healthy(pool):
addr = pool.allocate() # works with 1 left
if addr is None:
return False # by now it is too late
pool.release(addr)
return True
A check on an exclusive resource has to report the count, not the possibility. And the threshold has to account for how fast the count can move, not just where it is:
# what it should ask
def status(pool):
free = pool.free_count()
burn = pool.allocations_per_hour(window='24h')
hours_left = free / max(burn, 0.1)
if hours_left < 24: return CRITICAL
if hours_left < 168: return WARNING
return OK
Now the alert fires with a week of notice, and it fires earlier when demand is climbing — which is exactly when you need it to, and exactly when a static "alert below 10" threshold does not.
The second trap: leaks
Free count is only meaningful if addresses actually come back. They frequently do not, for three reasons:
- Teardown failed halfway. The instance is gone; the address is still marked allocated to it.
- A reservation was never released. Provisioning reserves an address before the instance exists. If provisioning fails between those two steps, the reservation can outlive everything.
- Someone allocated one by hand for a test in 2024 and nobody wrote it down.
The fix is reconciliation rather than trust: walk the pool, walk the live instances, and report every address allocated to something that does not exist. Run it hourly, alert on anything older than a few minutes, and hold reservations on a TTL so a crashed provisioner releases them by expiry rather than by cleanup.
What we do about it
Pools are monitored by time-to-exhaustion rather than by count, reconciled hourly against live instances, and reservations expire on a TTL. Headroom is held per region rather than globally, because an address in Delhi does not help an order in Mumbai — aggregate figures hide exactly the shortage you care about.
And capacity alerts page a human rather than filing a ticket. A week of notice is only useful if somebody sees it in that week.
The general lesson
Ask of any resource: does running out degrade it, or stop it? Degrading resources can be monitored by utilisation. Cliff-edged ones need time-to-exhaustion, and their health checks must never be able to report green while the resource is empty.
If you run your own VPS, the same reasoning applies to disk inodes, conntrack table entries, file descriptors and database connections. All four are cliffs, all four have a count you can read, and all four will take your service down without warning if you monitor them by asking "does it still work?".