← Liana Grigory
Insights

Keeping a serverless portfolio inside the free tier

The only metered line allowed to grow is the number of accounts. Notes on the specific Firebase and Cloudflare decisions that keep several live products flat, including the two I got wrong.

By Liana Grigory · 28 August 2026 · 10 min read

The invoice that ends most small serverless products does not arrive because something broke. It arrives because something worked.

A page gets shared. A crawler finds the site. A directory that had forty rows has four hundred. None of that is a failure, and on a badly shaped architecture every one of those is a bill. The defect is not in the traffic. It is in a system where the cost of being looked at scales with how often you are looked at.

I run several products on one Firebase project and a handful of Cloudflare Workers and Pages sites, and the constraint I hold them to is deliberately blunt: the only metered line that is allowed to grow is the number of user accounts. Reads, writes, storage and invocations stay flat, by design, whatever happens to traffic. Not "we will optimise it later" flat. Flat as an architectural property, checked before a feature ships.

This is a note on the specific decisions that produce that, and the two that I got wrong first.

The test that settles most design arguments

Before adding anything that reads data, I ask one question:

If this feature went viral tomorrow, what would the bill be?

If the answer scales with views, the design is wrong and no amount of caching bolted on afterwards fixes it properly. If it scales with signups, it is fine, because a signup is a human being who chose to be there and is the thing the business is actually for.

That single question kills a surprising number of otherwise attractive ideas, and it kills them at design time, which is the only time killing them is cheap.

Public data goes in one document, not a query over many

This is the highest-leverage decision on the list, so it goes first.

A public board — a directory, a listing page, a roster, anything anonymous visitors can see — built as a query over a collection bills one document read per row returned. Per visit. For ever. Forty rows and a hundred visitors a day is four thousand reads a day, and it grows on both axes at once: more rows, more visitors, multiplied.

The same data held in a single aggregate document costs one read. One. Whether the directory has forty entries or four hundred, and regardless of how many people load the page, because the whole payload is one document.

The objection is obvious: a shared document that everyone writes to sounds like a security problem. It is not, because rules can constrain a write to specific keys:

match /directory/{shard} {
  allow read: if true;
  allow update: if request.auth != null
    && request.resource.data.diff(resource.data)
         .affectedKeys().hasOnly([request.auth.uid]);
}

Each participant owns exactly one key, named for their uid. Nobody can edit anyone else's entry and nobody can clear the document. One read for readers, one constrained write for writers.

Two things belong in the file alongside it. A document has a hard size ceiling, so write the split plan — by locality, by ZIP prefix, by alphabet — and write the size at which the split becomes due, in a comment, next to the rule. And do not do the split before that number. An optimisation for a table with nine rows in it is its own kind of waste.

Never open a live listener

onSnapshot is the single most expensive habit available in this stack, and it is the default in most tutorials.

A listener bills a read per changed document per connected client. Ten people with a dashboard open and one document changing is ten reads, and it is ten reads again on the next change, and the tab left open on somebody's second monitor is billing all afternoon. The costs are invisible in development, where there is one client and nothing changes, and they are not invisible at the end of the month.

On-demand reads only. Fetch when the user acts. A refresh button that costs one read and that the user pressed deliberately is better, cheaper and more honest than live data nobody is watching.

There is a real trade here and it is worth being straight about it: genuinely collaborative features — a shared cursor, a live chat — want a listener, and for those the right answer is usually a Durable Object or a WebSocket rather than a Firestore listener. But "the dashboard updates by itself" is not a collaborative feature. It is a preference, and it is one that scales its price with how many people leave a tab open.

Counters, not counts

Wanting to display "48 listings" and getting it with getDocs() and reading .length costs 48 reads to render one number. On every page load.

Keep a counter and increment it on write. One write when something changes, zero reads when something is displayed. If the count only needs to be approximately right, and most displayed counts do, an occasional reconciliation is cheaper than continuous accuracy nobody is checking.

Where an exact count is genuinely needed, Firestore's aggregation query bills far less than reading the documents, so use it — but the reflex to reach for first is the counter, because the cheapest read is the one that does not happen.

Anything a crawler can reach is generated at build time

This is where I got it wrong, and it is the mistake I would most want someone else to avoid.

A sitemap generated at request time is a read amplifier of the worst kind, because the thing requesting it is a bot, bots do not have a natural rate limit, and a single crawl can hit it repeatedly. It is a URL that costs money and produces no user. So can a category page, a public profile, or any listing rendered per request.

Generate them at build time into static files. Cloudflare Pages serves a static asset for effectively nothing and it is faster than anything you can compute. If a page must be dynamic, cache it at the edge with a real TTL and make sure the cache actually engages — a Cache-Control header on a route that also sets a cookie, or varies, may not be cached at all, and the way to know is to look at the response headers on a live request rather than to assume.

The rule I use now: if an anonymous visitor or a bot can reach it, it is static or it is cached. No exceptions, and prove it with a live request rather than by reading the config.

One query per screen, with a limit, and filter in the browser

Fetch a bounded page of data once, then let filtering, sorting and searching happen client-side over what is already in memory. Typing in a search box must cost nothing, and if each keystroke fires a query it costs on every keystroke, on every session, for every user, to produce a result the user is going to narrow again in half a second anyway.

Every list query carries a limit(). Not as a performance nicety — as a hard rule, because an unbounded query is a bill with no ceiling, and the day it becomes expensive is the day the product succeeds.

Denormalise for locality so this stays true as the data grows. If a search is naturally geographic, store a ZIP prefix on the document and query the neighbourhood rather than the collection. The point is that cost stays flat as rows are added, rather than growing with the size of a table nobody is reading all of.

No server that has to be running

The second thing I got wrong, and it cost real money before it cost anything else.

A scheduled job that scans a collection to find work to do is billing every time it runs, whether or not there is anything to do. A cron every five minutes over a few hundred documents is a five-figure read count per month to discover, almost always, that nothing has changed.

Worse, the deploy of the function itself is metered on some platforms — a container build per deploy, image layers retained and billed. I have shipped monitoring that cost more than the thing it was monitoring, which is a specific and embarrassing category of failure. A tool built to watch cost must not itself cost.

The alternative that has held up: if a client action can do the work with the acting user's own permissions, that is the design. Work happens on the user's device, with their token, at the moment they act. No server to keep warm, no cron to scan, and — because it runs under that user's credentials and their row-level rules — no ambient privilege sitting around waiting to be misused.

Where something genuinely must happen on a schedule, make it event-driven rather than poll-driven, and make the schedule remove itself when the work is finished. A job that has to be manually turned off is a job that will not be.

Uploads are the line that grows quietly

Storage is the one metered resource that only ever goes up, because nothing deletes itself.

Resize images in the browser before upload rather than storing a twelve-megapixel photo to display it at 200 pixels. Link out to a video host rather than serving video from your own bucket, which is bandwidth priced like a service you did not mean to run. And give every stored object a retention policy and a deletion path on day one, because a deletion path added later has to reconcile with everything already accumulated, and it never quite does.

The deletion path is also a legal requirement, not just a cost control — an account deletion that leaves the user's uploads sitting in a bucket is not a deletion. Building it once, at the start, satisfies both.

Write the cost down when you add the feature

The practice that has done the most to keep this honest is unglamorous: state the reads and writes per user action in the commit message.

"Adds the roster screen. One read per load (aggregate doc), one write per edit. No listener, no scheduled job." It takes ten seconds. It makes the cost of a feature a thing that was considered rather than a thing that was discovered, and it gives you something to grep when the bill moves and you need to know what changed.

A feature whose cost nobody wrote down is a feature nobody can defend.

What the discipline is actually for

None of this is austerity for its own sake. Staying inside the free tier on a portfolio of live products is not a badge; it is what makes it possible to run several of them at once, keep them running while they are small, and not have the survival of a product depend on it monetising before the invoice does.

And the same shape produces the security posture. An application that holds no bulk data at rest, reads only the slice a signed-in user is entitled to, opens no listeners and runs no ambient server has very little to steal and very little surface to attack. Cheap to run and hard to breach are not two goals that have to be balanced against each other. They are the same property, reached by the same decisions, and the second one is the reason to keep making them after the bill stops being scary.

Written by Liana Grigory, Entrepreneur and Software Engineer, from work on The Care Royal, Tegula Stone and Unified Savers. Everything above describes decisions actually made on those systems, including the ones that turned out to be wrong.

Home · Terms of Use · Privacy Notice