Talk to a real engineer in Agra, 24×7 — +91 75994 50220 support@bigdomainhost.com
Engineering

What happens when an IP pool runs dry

A capacity check that reports headroom while provisioning fails is worse than no check at all. IPv4 allocation is the classic case, and the reason is that it is the one resource you cannot overcommit.

The class of failure

Most capacity problems announce themselves gradually. CPU gets slower, memory pressure shows up in swap, disks fill at a rate you can extrapolate. You get warning, and the warning is proportional to the problem.

IPv4 addresses do not behave like that. A pool with two addresses left behaves exactly like a pool with two hundred, right up until the moment it behaves like a pool with none. There is no degradation, no slow-down, no partial service. It works, and then provisioning fails outright.

This is worth understanding whether you run a hosting platform or a single VPS, because the same shape of failure appears anywhere a resource is allocated rather than shared.

Why it is different from CPU and RAM

CPU and memory can be overcommitted. Hosts routinely allocate more vCPU than there are physical cores, because not every guest is busy at once. It is a bet, it usually pays, and when it does not the symptom is contention rather than failure.

An IP address cannot be overcommitted. Two machines cannot share one routable address. The allocation is binary and exclusive, so the pool is a hard count and the failure is a cliff.

Anything with that property — floating IPs, licence seats, VLAN tags, port numbers on a NAT gateway — needs a different kind of monitoring from the things that degrade.

The check that reports green while the pool is empty

The naive health check asks: can I allocate an address? It allocates one, confirms it worked, releases it, reports healthy.

That check is green when one address remains. It is green when one address remains and forty servers are being ordered. It only goes red after the failure it was supposed to predict.

# the check that tells you nothing
def healthy(pool):
    addr = pool.allocate()      # works with 1 left
    if addr is None:
        return False          # by now it is too late
    pool.release(addr)
    return True

A check on an exclusive resource has to report the count, not the possibility. And the threshold has to account for how fast the count can move, not just where it is:

# what it should ask
def status(pool):
    free = pool.free_count()
    burn = pool.allocations_per_hour(window='24h')
    hours_left = free / max(burn, 0.1)

    if hours_left < 24:  return CRITICAL
    if hours_left < 168: return WARNING
    return OK

Now the alert fires with a week of notice, and it fires earlier when demand is climbing — which is exactly when you need it to, and exactly when a static "alert below 10" threshold does not.

The second trap: leaks

Free count is only meaningful if addresses actually come back. They frequently do not, for three reasons:

  • Teardown failed halfway. The instance is gone; the address is still marked allocated to it.
  • A reservation was never released. Provisioning reserves an address before the instance exists. If provisioning fails between those two steps, the reservation can outlive everything.
  • Someone allocated one by hand for a test in 2024 and nobody wrote it down.

The fix is reconciliation rather than trust: walk the pool, walk the live instances, and report every address allocated to something that does not exist. Run it hourly, alert on anything older than a few minutes, and hold reservations on a TTL so a crashed provisioner releases them by expiry rather than by cleanup.

What we do about it

Pools are monitored by time-to-exhaustion rather than by count, reconciled hourly against live instances, and reservations expire on a TTL. Headroom is held per region rather than globally, because an address in Delhi does not help an order in Mumbai — aggregate figures hide exactly the shortage you care about.

And capacity alerts page a human rather than filing a ticket. A week of notice is only useful if somebody sees it in that week.

The general lesson

Ask of any resource: does running out degrade it, or stop it? Degrading resources can be monitored by utilisation. Cliff-edged ones need time-to-exhaustion, and their health checks must never be able to report green while the resource is empty.

If you run your own VPS, the same reasoning applies to disk inodes, conntrack table entries, file descriptors and database connections. All four are cliffs, all four have a count you can read, and all four will take your service down without warning if you monitor them by asking "does it still work?".

Support that picks up the phone.

24×7, from our office in Agra, in IST — Hindi or English. Sales, migration and emergencies all reach the same engineers. No offshore queue, no 48-hour first reply.

Questions about any of this?

Call +91 75994 50220. The people who wrote this are the people who answer.