← Liana Grigory
Insights

Pages that live in a database

A page generated at build time costs nothing to look at. A page read from a database costs a read every time somebody looks at it. That difference decides more architecture than it gets credit for.

By Liana Grigory · 8 September 2026 · 9 min read

I run one system where published pages live as documents in a database and are served by a Worker that fetches the document and returns its stored HTML. I run two others where every page is written to disk at build time and served as a static asset. The three sites look identical to a visitor. Their cost curves are not remotely the same.

This is a small architectural decision that quietly determines whether traffic is something you want or something you fear.

The appeal of the database-backed page

The database version is genuinely nicer to operate, and I want to be fair to it before arguing against it.

Publishing becomes an API call rather than a deploy. There is no build, no CI run, no propagation wait. A page can be created, corrected or unpublished in the time it takes one request to complete. Nothing has to be checked into a repository, which means the publishing system does not need commit rights to anything. And a correction to a live page is a single write, not a rebuild of a site that may contain thousands of unrelated pages.

For a system that generates pages programmatically, that is a real reduction in moving parts. The alternative, where a machine opens pull requests against a content repository and waits for a build, is more machinery to own and more places for it to jam.

The bill you signed up for

The cost is that every view is a read.

That sentence sounds mild and is not. A static page has a marginal serving cost that rounds to nothing: the asset sits on a CDN and a million requests cost about what a thousand do. A database-backed page has a marginal cost per view, denominated in database operations, which are metered.

Run the thought experiment that matters: if this page went viral tomorrow, what would it cost? For the static version the answer is nothing worth discussing. For the database version the answer scales linearly with attention, and attention is the thing you built the page to attract. You have constructed a system where success and cost are the same variable.

Worse, it is not only humans. Crawlers hit published pages relentlessly, and they are not distributed like human traffic. A single crawler working through a few thousand generated pages will produce a few thousand reads in a short window, on a schedule you do not control, for pages nobody looked at.

Caching is the fix, and it has to be deliberate

The obvious answer is to put a cache in front, and it works, but only if you mean it.

The rule I use is that the database read must be the exception, not the default path. That means the response carries real cache headers with a long lifetime, the edge is allowed to serve a stale copy while revalidating, and publishing explicitly purges the affected key rather than relying on expiry. If the cache is a nice-to-have that some requests skip, you have not changed the cost model, you have made it harder to reason about.

The trap is a cache that looks like it is working. Hit rates on a long tail of rarely-visited pages are poor almost by definition: the pages that get one visit a week are the ones whose cache entry has always expired. Those pages are the majority of a generated corpus, and they are the ones a crawler visits.

What I would keep in the database

Having argued against it, here is where I think it is right.

  • When the page must change without a deploy. If correcting a live page has to happen in seconds, a build step is a liability.
  • When the writer is a machine that should not have commit access. Giving an automated process the ability to write to a content store is a much smaller grant than giving it the ability to push to a repository that also contains application code.
  • When the corpus is large and mostly cold. Rebuilding thousands of pages to change one is its own kind of waste.
  • When the content is genuinely per-user. This is not the same case at all, and it is the one where a live read is correct: nobody else can be served that page, so there is nothing to cache and nothing to pre-render.

What belongs on disk

Everything else, and specifically everything a crawler can reach.

Sitemaps are the sharpest example. A sitemap that queries the database at request time is a read amplifier pointed directly at the machines most likely to fetch it. One uncached request can produce a read per listed URL, and the whole purpose of the file is to invite exactly that. It should be a static file, generated when the content changes, with zero runtime reads. The same goes for index and category pages, which are crawled far more often than they are visited.

The pattern that resolves it

What I have settled on is a split by who can see it rather than by how the page is produced.

Public and shared content is written to disk, or written to the database once and then rendered through a cache that is treated as part of the architecture rather than an optimisation. Private, per-user content is read live, keyed on the signed-in identity, with a limit on the query.

That split has a pleasant property: it is also the security boundary. The pages that are cheap to serve are the ones with nothing sensitive in them, and the pages that cost a read are the ones scoped to a single authenticated user, which means there is no bulk surface to scrape even if somebody wants one. Cheap to run and hard to breach turn out to be the same design, reached the same way.

Write the number down

The habit I would recommend regardless of which side you land on: when you add a feature, state its cost in reads and writes per user action, in the commit message. Not an estimate of traffic, just the per-action arithmetic.

It takes one line and it makes the expensive design visible at the moment it is cheapest to change. A feature whose cost nobody wrote down is a feature nobody can defend later, and by the time it shows up on a bill it has usually acquired dependencies.

Written by Liana Grigory, Entrepreneur and Software Engineer, from work on The Care Royal, Tegula Stone and Unified Savers. Everything above describes decisions actually made on those systems, including the ones that turned out to be wrong.

Home · Terms of Use · Privacy Notice