01Work 02Impact 03About 04Experience 05Stack 06Contact 07Lab
CRSR 0.000 · 0.000
SCRL 0.00 %
← lab / experiments
07 / Lab EXPERIMENT · LIVE

Multi-tenant SaaS.

One codebase serving many customers, with zero data bleeding between them. Every interesting decision lives inside that word "zero" — so let's poke at it. Flip the switch below and watch a naive query leak.

01 · The data-leak demoROW-LEVEL SECURITY

You're logged into a SaaS billing app. The app runs SELECT * FROM invoices and forgets the tenant filter — a one-line mistake that ships constantly. Postgres Row-Level Security is the seatbelt: flip it off and the same query leaks every other customer's invoices.

Logged in as:
Row-Level Security: ON
InvoiceCustomerAmountStatusTenant
02 · Three ways to isolate a tenantPOOL · BRIDGE · SILO

RLS handles the shared case, but "where does a tenant's data physically live?" has three canonical answers. Cheaper and denser on the left; more isolated and expensive on the right. Pick your poison per the compliance and cost you can stomach.

03 · How a request becomes a tenantTENANT ROUTING

Before any of that isolation matters, the app has to know which tenant is asking. A subdomain (or a JWT claim) gets resolved to a tenant id, which is pinned to the DB session — and from there RLS does the rest. Switch domains and watch the context change.

04 · The noisy-neighbour problemSHARED VS ISOLATED

Pooling is efficient until one tenant runs a report that eats every connection. In a shared pool, everyone's latency climbs together. Silo them and the blast stays contained. Spike a tenant and compare.

Spike a tenant:

05 · What bites you in prodHARD-WON
01

The forgotten WHERE tenant_id

It will get forgotten — by you, by an ORM, by a junior at 5pm on a Friday. RLS is the backstop that makes that mistake a non-event instead of a breach. Belt and braces: filter in the app and enforce in the database.

02

Migrations multiply

In pool it's one ALTER TABLE. In silo it's the same migration across N databases, and if #47 fails halfway you now have two schema versions in prod. You need a runner that's idempotent, ordered, and tells you exactly which tenants are behind.

03

The session variable must actually be set

RLS keys off current_setting('app.current_tenant'). Set it per transaction, from a trusted middleware — never from user input. Miss it and, depending on your policy, you either see nothing or (worse) everything.

04

Admin queries need an escape hatch

Support tooling and cross-tenant analytics have to bypass RLS deliberately — a separate role with BYPASSRLS, audited and boring. Don't reuse the app role "just this once."

05

Connection-pool math ambushes the pool model

One shared Postgres has a hard connection ceiling. Thousands of pooled tenants behind PgBouncer is fine — until a few of them each open long transactions. Capacity-plan the pool, not just the CPU.

06

Backups and deletes are per-tenant problems

"Restore just Globex to last Tuesday" is trivial in silo and genuinely hard in pool. Same with GDPR deletes. Your isolation model quietly decides how painful your worst-day operations are.