Scaling from 10K to 5M Monthly Active Users with a Lean Backend Team
← Back
5.5.26

Scaling from 10K to 5M Monthly Active Users with a Lean Backend Team

Scaling to 5M MAU: Maximizing ROI through MySQL domain splitting and 'phantom services'

Gal Lerman
ByGal Lerman
Tech Lead

This post is based on a podcast conversation. Watch on Youtube, Listen on Spotify



We scaled a real-money gaming platform from 10K to 5M MAU —starting with a tiny backend team, no DevOps, and no budget for overengineering. We outgrew Elastic Beanstalk and built a Kubernetes cluster fromscratch to get shared infrastructure, fast deploys, and independent scaling. Weevolved our CI/CD when local deploy scripts stopped scaling, split MySQL into9 domain-specific clusters, and invented "phantom services" to getmicroservice benefits without microservice complexity. The thread connectingall of it: maximize ROI, exhaust the simple options first, and validate everythingin production.

By the Numbers

10K → 5M MAU 1 service/machine → shared infrastructure 1 → 9 MySQL clusters 1 monolith → 7 phantom services


Papaya Gaming is a competitive real-money, skill-based gaming platform - that offers games like Solitaire and Bubble Shooter. In every tournament, each player gets the same game state, plays under the same conditions, and winning is determined by who performs best. Because real money is on the line, the backend must be stable, fast, and fair - always. 

I joined Papaya as the first backend engineer when the team was about ten people and we had 10K monthly active users. Today, the team has grown to over 400, the platform serves up to 5M MAU, and we handle roughly one million requests per minute at peak. 

In the early years, the backend team was around five engineers -responsible for everything from features to deployments to database management. No DevOps, no infrastructure specialists. As we grew, so did the team. But the lean culture that carried us through never changed. 

The Guiding Principle 

Every infrastructure investment had to justify itself against the alternative of shipping another feature. Before adopting anything new, we asked: 

  1. Have we exhausted the low-hanging fruit? Query optimizations, caching, config changes?
  2. What's the cost and risk? Engineering time, migration risk, knowledge we'd need to build?
  3. How long will it last? Six months isn't worth it. Years is.

This kept us from over-engineering early and under-investing late. 


One Service Per Machine Wasn't Going to Scale 

The platform started on AWS Elastic Beanstalk - simple, managed, reasonable for a small team. But the model had hard limits we were hitting from every direction: 

  • Resource waste - each service was locked to a dedicated EC2 instance, regardless of how much capacity it actually used 
  • Slow deploys - deploying meant replacing the entire machine, around 40 minutes per deploy. For hotfixes during incidents, that was unacceptable 
  • No independent scaling - couldn't scale one service without scaling the whole instance 
  • No self-healing - if a process crashed, there was no automatic recovery 

Kubernetes addressed all of these at once: shared infrastructure, fast rolling deploys, per service scaling, health checks, and a declarative operational model. It also became the foundation for nearly every infrastructure decision that followed - including the phantom services pattern we'd adopt later. 

The problem: no DevOps team, no one who'd ever built a cluster, and no AI to help - this was 2020. I had some Kubernetes experience, but only as a consumer - deploying onto clusters someone else maintained. Building one from scratch was entirely different. 

Weeks of reading docs, failed attempts, and iteration later - we had a working cluster on AWS. 

For the migration itself, we scheduled a one-hour meeting with the entire team. We used Route 53 to shift traffic at the DNS level - starting with a small percentage, then quickly ramping up. The migration was complete in 40 minutes. We cut the meeting short. I'd expected a week-long gradual process, but our CTO Andrey and VP R&D Alex pushed for speed. True startup spirit - move fast when you've done the preparation. 

How Our Deploy Scripts Stopped Scaling 

For two years, our deployment pipeline was a Node.js script that built Docker images locally and pushed them to AWS. It worked - until it didn't. 

The script depended on the developer's local machine and internet connection. In the office, that was fine. But as the team grew and people worked from different locations with varying internet quality, the multi-hundred-megabyte Docker image uploads became unreliable. Builds took 10 minutes locally, uploads failed after 7 minutes of retries, and the whole cycle could take 20 minutes with roughly a one-in-ten success rate. 

We migrated to Jenkins - builds happened on a server with a stable connection. Developers triggered builds remotely; the heavy lifting happened on infrastructure we controlled. The problem vanished overnight. 

Lesson: keep it simple until you can't anymore. The Node.js script was the right tool until it wasn't. When we outgrew it, we replaced it - no sooner, no later. 

Taming MySQL: Purging and Domain Splitting 

MySQL has been our primary database from the start, paired with Redis for caching. As we scaled, it became the most frequent bottleneck. Two strategies saved us.

Aggressive purging. Tournament data has a natural lifecycle - created, played, judged, prizes distributed, done. After that, the production system doesn't need the operational data. We built purging jobs that archive expired data to analytics and delete it from production. Tables that would grow to billions of rows stay small and fast. 

Key detail:

purging runs during off-peak hours only. Large DELETEs acquire locks and consume I/O — running them during peak traffic makes things worse, not better. 

Domain splitting. We eventually hit MySQL's write ceiling — single-writer architecture means one node accepts all writes. We considered alternatives, but adopting a new database technology at production scale is non-trivial effort — not because we're locked to a specific stack, but because building equivalent operational expertise takes years, not weeks. With real money transactions on the line, we weren't ready to take that on. 

Instead, we split by domain. Tournaments got their own MySQL instance. Live operations got another. Metrics a third. Each domain gets its own writer — write throughput multiplied. 

Today: 9 MySQL clusters, each serving a specific domain. The tradeoff is no cross-domain JOINs, but we split along natural boundaries, so most queries and transactions are domain internal anyway. 

Phantom Services: The Monolith Trick 

Since our client is a complex game application, our monolith grew to over 190 API endpoints. When memory spiked or CPU maxed out, pinpointing the cause was nearly impossible. The standard answer: decompose into microservices. The standard cost: months of extraction work. 

We found a shortcut. 

Kubernetes lets you run the same container image in multiple deployments with different configurations. We created additional deployments of our monolith — identical code, identical image — and used ingress rules to route specific endpoints to specific deployments. Tournament endpoints go to one deployment, app startup to another, general traffic to the original.

Zero code changes.

No shared-state extraction, no new protocols, no changes to the pipeline. Just new Kubernetes manifests. Each "phantom service" scales independently, has its own resource limits, and its own monitoring — but they all share one codebase.

Today we run seven phantom services. The pattern proved durable enough that we still use it alongside a number of conventional services. 

What We Learned 

Start with the simplest thing that works. Elastic Beanstalk, Node.js scripts, a single MySQL instance — none were the final architecture. Each was right for its moment. 

Exhaust optimizations before replacing technology. MySQL domain splitting extended MySQL's useful life by years. We haven't needed alternatives yet. 

Validate in production. No staging environment faithfully reproduces real traffic. Start small, observe, increase. Repeat for every major change. 

Stay with technology you know deeply. Expertise compounds. Deep MySQL knowledge let us confidently operate 9 clusters. Adopting a new database isn't about being locked in — it's about the non-trivial effort of building operational expertise while running real-money transactions. 

The lean approach doesn't expire. It works at 10K MAU with five engineers, and at 5M MAU with a larger team. The biggest mistake is over-engineering for scale you don't have. The second biggest is refusing to invest once scale arrives. The art is timing — driven by data, not by what other companies are doing. 

Gal Lerman is a backend engineering lead at Papaya Gaming, where he has been building and scaling the platform's backend infrastructure since the company's earliest days.