← Liana Grigory
Insights

Rate limiting inside a stateless function

An in-memory counter in a serverless function is not a rate limit, it is a rate suggestion. That turns out to be worth having anyway, as long as you are honest about which attack it stops.

By Liana Grigory · 8 September 2026 · 9 min read

Every endpoint in my portfolio has a rate limiter on it. That is a rule I wrote for myself rather than one anybody imposed, and the reasoning was simple: an endpoint without a ceiling is an endpoint whose worst case is set by whoever finds it.

What I want to write about is the uncomfortable part. Most of those limiters are in-process counters living in the memory of a serverless function, and an in-process counter in a horizontally scaled runtime is a much weaker control than it appears. I kept them anyway. The interesting question is why that is a defensible decision rather than a lazy one.

What the naive version actually does

The shape is familiar. A module-scope map, keyed by caller identity, holding a count and a window start. On each request you look up the key, reset the window if it has expired, increment, and reject over the ceiling.

In a long-lived server this is a real rate limit. In a serverless function it is three things at once, and only one of them is what you wanted:

  • The counter is per instance. The platform runs as many concurrent instances as it feels like. Ten instances with a ceiling of a hundred is a ceiling of a thousand, and you do not control the multiplier. It moves with load, which means it is loosest exactly when you are under pressure.
  • The counter dies on cold start. Instances are recycled constantly. An attacker does not need to defeat the limiter; they need only to be unlucky enough to be spread across fresh instances, which happens by itself.
  • The map is a memory leak with good manners. Every distinct key allocates an entry. Unless something evicts them, a long-lived instance accumulates one entry per caller, forever, and the eviction is the part people forget.

So the honest description is not one hundred requests per hour per caller. It is roughly one hundred requests per hour per caller per instance, when the instance happens to be warm. If you write the first sentence in a comment above the second implementation, you have written a lie that will be believed for years.

The distinction that makes it worth keeping

The reason I did not rip these out is that abuse is not one thing. It is at least two, and they have different shapes.

Sustained, distributed abuse is what an in-process limiter fails against. A patient attacker with a pool of addresses, pacing themselves, is not meaningfully slowed by a counter that resets whenever the platform feels like it.

An accidental loop, a retry storm, or a naive script is what it stops very well. And in my experience that is overwhelmingly the traffic that actually shows up. A misconfigured cron calling an endpoint every second, a client retry with no backoff, somebody enumerating an endpoint from one machine in a tight loop: all of these hammer whatever instance they land on, which is precisely where the counter is.

The naive limiter is therefore a decent control against the failure mode that is common and a poor control against the one that is rare and targeted. That is a reasonable place to be, provided you do not confuse it with security. It buys availability and cost protection. It does not buy protection against a determined adversary, and it must never be the only thing standing between the internet and something expensive.

Where the real ceiling has to live

For anything where being wrong is expensive, the state has to leave the process. The options, in rough order of how much machinery they drag in:

  • The edge, before your code runs. A platform-level rate limiting rule is the cheapest correct answer, because the request is rejected before it becomes an invocation you pay for. Anything that can be expressed as a ceiling per address per path belongs here, not in your handler.
  • A shared store with atomic increment. A key-value store with a real increment-and-expire gives a genuine global count. The cost is a network round trip on every request, including the ones you were going to allow.
  • A single-threaded coordination object. Where the platform offers a durable, single-instance object addressable by key, you get exact counting with no race, because there is precisely one of them per key. This is the correct primitive and it is the one I would reach for if the ceiling genuinely had to hold.
  • The record itself. Frequently overlooked and often the best answer: if the thing being limited creates a document, the document is the counter. A uniqueness constraint or a conditional write enforces at most one, exactly, forever, with no window arithmetic at all.

Fail closed, and only where it matters

The direction a limiter fails in is a product decision disguised as an error handler.

If the backing store is unreachable and you cannot read the count, you either let the request through or you reject it. Letting it through means an attacker who can degrade your store has removed your ceiling. Rejecting means a store outage is a full outage of the endpoint.

I do not think there is a universal answer, but there is a rule: fail closed wherever exceeding the ceiling costs money that does not come back. A metered third-party call, a generation endpoint, anything with a per-unit price. For those, an unavailable limiter must mean a refused request, and the hard monthly cap behind it must be enforced by something more durable than a counter in memory. For a read-only endpoint where the worst case is some wasted compute, failing open is usually the better trade.

Two identities, not one

The other thing worth doing properly is keying. An address-only limiter punishes everyone behind a shared network address and does nothing about an authenticated abuser with a home connection. An account-only limiter does nothing at all before sign-in, which is exactly where credential stuffing lives.

So both, on the endpoints that have both: a per-address ceiling that applies to everyone including anonymous callers, and a per-account ceiling on top for authenticated routes. They are different controls answering different questions, and the second one is the only one that means anything once an attacker has a valid session.

Say 429, and say when

A rejected request should return the status that means rejected-for-rate, and it should carry a Retry-After. This is not politeness. A client that receives a bare error has no basis for choosing a backoff and will typically retry immediately, which converts your rate limiter into an amplifier. Telling the caller when to come back is how you get the traffic to actually go away.

The corollary is that your own automation must read it. A job that ignores Retry-After and retries in a tight loop is indistinguishable from the abuse the limiter exists to stop.

What I would tell myself earlier

Put a limiter on everything, including the endpoints that feel too boring to attack, because the cost of adding one is minutes and the cost of not having one is unbounded. Then be exact in the comment about what it does: name the fact that the count is per instance, so that the next person to read it does not build something expensive on top of a guarantee that was never there.

The failure I care about is not the weak limiter. It is the weak limiter with a confident comment above it. A control you have overestimated is worse than no control, because you stop looking at the thing it was supposed to be protecting.

Written by Liana Grigory, Entrepreneur and Software Engineer, from work on The Care Royal, Tegula Stone and Unified Savers. Everything above describes decisions actually made on those systems, including the ones that turned out to be wrong.

Home · Terms of Use · Privacy Notice