All posts
28 Sep 2026

How Onescle scales to 10k shops without slowing down

I loaded Onescle with 10,000 shops and 554,000 products and measured it. One missing index was making every shop pay for every other shop. Here is what I found and how Onescle grows.

How Onescle scales to 10k shops without slowing down

When people hear "SaaS for thousands of shops", the first question is usually: what happens when it gets big? Will it slow down? Will one busy shop make every other shop slow?

For Onescle I didn't want to guess. So before calling it ready, I loaded it with 10,000 shops and 554,000 products and measured what happened. This post is about what I found, the one bug that almost ruined everything, and how Onescle grows without a rewrite.

How Onescle is put together

Every request, from a shopper on a merchant's own domain or a merchant in the dashboard, enters through Cloudflare first. Cloudflare handles DNS, SSL for custom domains, the firewall and caching for images. Behind it are the apps, one API, and a background worker.

Onescle architecture
Follow a request from top to bottom: users, Cloudflare, apps, API, data, integrations

A few decisions matter most for scaling:

  • The API is stateless. It keeps nothing in memory between requests. That means you can run two, five or ten copies of it side by side, and any copy can answer any request.
  • Slow work goes to a separate worker. Sending SMS, talking to couriers, sending events to Meta: all of that is put on a RabbitMQ queue and done by a background worker, not while the shopper waits. Jobs are retried, and failed ones go to a dead-letter queue instead of disappearing.
  • Redis caches the hot data, like which shop a domain belongs to, so the database is not asked the same question thousands of times.
  • PgBouncer sits in front of PostgreSQL. Every API copy opens database connections. Without a pooler, adding API copies eventually runs PostgreSQL out of connections. With PgBouncer, the number of API copies no longer depends on the database's connection limit.

The test

I created a separate database and filled it the way real shops look: most stores with a dozen products, some with a few hundred, a few big ones with 2,000. In total, 10,000 shops and 554,000 products.

Then I sent realistic traffic at it. Not just the homepage: product lists, product pages, categories, search, filters, carts, and full checkouts that actually place an order.

The bug: every shop was paying for every other shop

The first results were bad. At 5,000 shops the API handled only 89 requests per second, and PostgreSQL was using up to ten CPU cores.

The cause was one missing database index. When a shopper opened a product list, Onescle checked stock for each product. The query looked up stock by product variant, but there was no index for that column alone. So PostgreSQL read the entire stock table, for every shop on the platform, on every product list.

That is the worst possible bug in a multi-shop system. A small shop with 12 products paid the same price as a big one with 2,000. And every time any merchant added stock, every other shop got a little slower. During one test run PostgreSQL did 22,611 full table scans and read 1.23 billion rows.

The fix was one line: an index on the variant column, added without locking the table. The same query went from 50 ms to 0.2 ms.

Before and after the index
Throughput at 5,000 shops, before and after one index

After the fix, throughput at 5,000 shops went from 89 to 223 requests per second, the typical product list got 2.5 times faster, and PostgreSQL dropped from up to ten cores to about one.

More shops, almost the same speed

With the index in place, I ran the same test four times, spreading the traffic across 100, 1,000, 5,000 and 10,000 shops, always with the full database.

Throughput from 100 to 10,000 shops
A 100 times increase in shops costs about 19% of throughput, with zero errors

Going from 100 shops to 10,000 shops cost about 19% of throughput and 21 milliseconds on a typical request. Zero errors in any run. And this is the pessimistic case: in the test every request went to a random shop, so the cache almost never helped. Real traffic is not like that. A small number of busy shops get most of the visits, and those stay in the cache.

I also counted database queries per request. A product list of 24 items runs 13 queries, whether it shows 5 products or 24. No query count grows with the size of the page.

Where it grows from here

After the fix, the bottleneck moved from the database to the Node.js API process. That is exactly where you want it. A busy database is hard to scale. A busy stateless API is easy: you add another copy behind the load balancer.

So when Onescle grows, the plan is simple:

  1. 1Add more API copies. They are stateless and PgBouncer keeps the database connections under control.
  2. 2Add a PostgreSQL read replica for heavy reading, like reports and storefront listings.
  3. 3Split the busiest modules into their own services. Onescle is a modular monolith, with 53 modules that already talk through clear boundaries and events, so this is a move, not a rewrite.

Being honest about the numbers

These tests ran on one 8-core development machine that also ran PostgreSQL, Redis and the load generator. So the absolute numbers describe that machine, not production. What they do prove is the *shape*: the cost of a request no longer grows with the number of shops on the platform. That was the goal.

If you are building a multi-tenant product, my one piece of advice is this: test with many tenants, not one big tenant. The bugs that hurt you are the ones where one customer quietly pays for all the others.