Building a multi-tenant retail POS backend to scale on the edge
How I structured a retail point-of-sale backend on Cloudflare Workers to handle many tenants and spiky traffic while staying maintainable, using an unopinionated framework, a modular monolith, and two databases split by responsibility.
I built the backend for a multi-tenant retail point-of-sale platform from an empty repository to production. Retail is a demanding shape of “scalable.” Many businesses share one backend, traffic is spiky and tied to store hours, every sale touches money so correctness is not negotiable, and the product keeps growing new modules (returns, purchasing, quotations, promotions) long after launch. So “scale” here had to mean two things at once: hold up under load and many tenants, and stay changeable as the surface area grows.
This is how I approached it, and the decisions I would defend again.
Start on the edge, then solve what the edge does not
The backend runs on Cloudflare Workers with Hono. Choosing the edge decided the easy half of scaling for me: horizontal scale is the platform’s job. There is no server to size, no autoscaling group to tune, and requests run close to the user. That is a real head start.
But it also moves the hard part somewhere else. When compute is effectively free to scale, your bottlenecks become structure and data: how you keep a growing codebase from collapsing into a mess, and how you read and write state without contention. The rest of this post is about those two things.
Why an unopinionated framework mattered
Hono is deliberately unopinionated. It gives you a fast router, a typed request context, and composable middleware, and then it gets out of your way. There is no folder layout it expects, no controllers directory, no convention that a route must live in a certain place the way Rails, NestJS, or Django would insist. At first that feels like less help. In practice it was the thing that let me organize the codebase around the business instead of around the framework.
I keep a tiny factory that returns a Hono app typed with the Cloudflare bindings, and the root app composes one router per bounded context:
export function createHonoApp() {
return new Hono<{ Bindings: Env }>();
}
const app = new OpenAPIHono();
// An explicit, ordered middleware chain I control, scoped by path.
app.use("/v1/*", createApiSecurityHeaders());
app.use("/v1/*", bodyLimit({ maxSize: MAX_REQUEST_BODY_SIZE }));
app.use("/v1/*", createGlobalRateLimit());
app.use("/v1/*", requestScopeMiddleware(rootContainer)); // per-request DI scope
app.use("/v1/*", createIdempotencyMiddleware(rootContainer));
app.use("/v1/*", createRequestLogger());
// One mount per bounded context.
app.route("/v1/products", productRoutes);
app.route("/v1/orders", salesOrdersRoutes);
app.route("/v1/inventory", inventoryRoutes);
// ...one line per module
Two things fall out of this. First, app.route() maps one-to-one onto module boundaries: a module is a folder that exposes a router, and mounting it is a single line, so adding the nineteenth module meant one import and one route call, not a framework config change. Second, cross-cutting concerns are an explicit, ordered chain I can reason about and scope by path (/v1/* versus /webhooks/*). The order genuinely matters here (request scope must run before idempotency so the cache lookup can authenticate the caller), and Hono makes that order visible instead of magic.
The trade-off is honest: an unopinionated framework makes you bring the structure. Hono will happily let you dump two hundred routes in one file. The discipline (the modules, the boundaries, the layering) is mine, not the framework’s. But that is exactly why it fit. I wanted to impose a specific architecture, and Hono neither fought me nor made me fight it.
A modular monolith, not microservices
On top of that routing freedom I built a single modular monolith, organized by business domain with domain-driven design. There are around nineteen bounded contexts under src/modules/ (products, inventory, sales orders, returns, cash operations, invoicing, payments, and so on), and each follows the same hexagonal layout: api/ for HTTP, application/ for use cases, domain/ for entities and rules, persistence/ for repositories, and infrastructure/ for adapters.
The scaling this buys is not runtime scaling, it is change scaling. A monolith with clean internal boundaries lets me ship a cross-cutting feature in one deploy and reason about the whole system without chasing a request across five services and their failure modes. The rule that keeps it from rotting is strict: a module may only import another module through its public index.ts facade, never its internals. When one module needs a fact another owns (“is this variant active?”, “reserve this stock”), it depends on a port interface declared in shared/ports/, and the owning module provides the adapter.
// A module depends on a small port, never on another module's internals.
interface StockPort {
reserve(variantId: string, qty: Qty): Promise<void>;
release(variantId: string, qty: Qty): Promise<void>;
}
Everything is wired with Awilix, a dependency-injection container, and the container is request-scoped: each request gets its own scope carrying the authenticated principal, the tenant context, and a logger with a correlation id. Use cases receive their dependencies by constructor injection, so nothing reaches for a global. When I did extract a service (a dedicated PDF Worker, a durable import workflow), I did it because it had genuinely different resource needs, not on principle. Extract for a reason, not by reflex.
Two databases, split by responsibility
The most important data decision was using two databases for two different jobs, rather than forcing one to do both.
- Cloudflare D1 (SQLite at the edge) holds platform and access-control data: tenants, users, roles, permissions, and auth sessions. It is global, lives at the edge, and every request touches it to authenticate.
- A managed PostgreSQL holds the business data: products, orders, inventory, customers, invoicing. It is tenant-isolated with row-level security and gives me real ACID transactions and relational queries for reporting.
That split immediately raises a problem: business data in Postgres constantly needs to reference platform entities that live in SQLite (“who created this order,” “which tenant owns this row,” permission checks in reporting joins). Doing a cross-database lookup on every one of those would be miserable. So a narrow, one-way mirror solves it: a small set of access-control tables (users, tenants, permissions) is mirrored from D1 into Postgres through retryable event handlers. Postgres gets local, joinable reference copies, while D1 stays the single source of truth for auth.
flowchart TB C[POS / dashboard] --> W[Worker API - Hono] W -->|platform and auth| D[(D1 - SQLite, edge)] W -->|business data| N[(Managed Postgres, RLS per tenant)] D -->|one-way sync via retryable events| M[auth mirror tables in Postgres] N -.->|local joins reference| M
The important discipline is that the mirror is reference data, not a second source of truth. Auth is authoritative in D1, business state is authoritative in Postgres, and the sync only ever flows one way. Nothing reads auth decisions out of the mirror.
Do less on the request path
The fastest way to make requests scale is to make them do less. Much of what a POS backend does after a sale does not need to happen before the customer gets their response: writing an audit record, keeping the auth mirror in sync, generating a receipt PDF, sending an email, running downstream inventory and invoicing effects.
All of that hangs off a typed event bus. In production the bus is backed by Cloudflare Queues (in local dev it is a simple in-memory bus, so events still fire), and handlers subscribe to domain events like “order executed” or “invoice enrolled.” Audit logging, the D1-to-Postgres mirror sync, and module side effects all run as asynchronous handlers, so the synchronous request returns as soon as the sale is durably recorded. Periodic work runs on Cloudflare cron triggers rather than on a user’s request: expiring layaway orders, sweeping for payments that need reconciliation, emitting stock and batch-expiry alerts. The user-facing path stays lean, and the expensive work happens where nobody is waiting on it. Durable, resumable jobs like bulk product imports run on a Cloudflare Workflow, and generated documents land in R2.
Keep the hot reads cheap
Every authenticated request has to resolve the caller’s roles and permissions. Doing that against the database on every call is a self-inflicted bottleneck, so authorization resolution is cached in Workers KV, with an even shorter-lived per-request memo in front of it and explicit invalidation on change. A permission check becomes a cheap read instead of a query.
For list endpoints, pagination happens in the database with proper LIMIT/OFFSET and window functions, never by fetching rows and slicing in memory, which is the classic way a listing that was fine at launch falls over once a tenant has real data. That went hand in hand with adding composite and partial indexes for the actual query shapes, eliminating N+1 queries, batching bulk inserts, and parallelizing independent I/O with Promise.all. A slow-query monitor on the Postgres side keeps me honest about which queries are drifting. None of these are clever individually; together they are the difference between a backend that degrades gracefully and one that falls off a cliff at a certain tenant size.
Safe under load and retries
At scale, requests get retried, clients double-submit, and two schedulers occasionally do the same work at the same moment. If your writes are not safe to repeat, load turns into duplicate orders and double charges.
So mutating endpoints run behind idempotency middleware keyed on a client-supplied X-Idempotency-Key (validated as a CUID2 and cached in KV), uniqueness constraints guard the tables that must not have duplicates, and outbound calls use timeouts and retry-with-backoff. The goal is that “this happened twice” is a non-event, not an incident. (I wrote about a payment race this exact discipline prevents in an earlier post.)
Multi-tenancy and correctness as first-class concerns
Everything above assumes tenant isolation is airtight. The tenant is carried in the auth token, resolved into the request scope, and enforced with row-level security in Postgres, and every mutation includes the tenant id in its WHERE clause even when the row id is already known, because matching on id alone is how cross-tenant regressions sneak in. Business rules that legitimately differ per tenant (stock policy, discount limits, layaway rules) are expressed as per-tenant policy ports, so the same code serves every tenant while honoring each one’s configuration.
Two more correctness choices earned their keep. Every handler validates its response against a Zod schema at runtime, across roughly a hundred endpoints, so a shape drift is caught at the boundary instead of downstream. And all money is a BigInt-based value type, never a float, because a retail backend that rounds wrong is a retail backend nobody should trust.
What “scalable” actually meant
Looking back, almost none of the scaling work was about raw throughput. The edge handled that. The real work was:
- A framework that stayed out of the way, so the structure could follow the business.
- Structure that scales change, so adding the tenth or fifteenth module did not make the system harder to reason about.
- Two databases matched to their jobs, with a narrow one-way mirror instead of forcing one store to do everything.
- Moving work off the request path, so the synchronous path stayed short.
- Cheap hot reads and database-level pagination, so per-tenant data growth did not quietly degrade every list.
- Idempotency everywhere it matters, so retries and concurrency were boring instead of dangerous.
Scalability, in practice, was less about handling a big number and more about making sure nothing on the critical path got more expensive as the system and its tenants grew.