How Booking.com cut Node.js costs by 38% with Watt
Breaking my own rule to share a win: Booking.com just published a case study showing how moving from pm2 to Watt cut their Node.js compute costs by 38%: fewer pods, 20% less memory, lower latency, and not a single line of service code changed.
Hi Everyone!
I usually don't mention anything that my startup has published in my newsletter; however, as I've been looking forward to this one for some time, I'm making an exception, as this is such an important milestone for us!
Booking.com has very recently published a case study about the operation of Node.js on a large scale using Watt: Rebalancing Node.js at scale, along with the claim that it reduced our costs by 38%.
The numbers are excellent since, when they reconfigured their main Node.js rendering service, they achieved a 38% reduction in their computing costs by running 30% fewer pods, reducing memory usage by 20% for each of the remaining pods, and lowering latency by as much as 10% at the p75, p99, and p99.9 levels; no new hardware was employed, and there were no modifications made to the service's code.
I used to say, "friends don't let friends use pm2." They moved from pm2's cluster module to Watt's worker threads, and the details matter: each worker accepts connections directly from the kernel via SO_REUSEPORT instead of having every request pass through a supervisor via IPC. This means one less hop in the request path, and the benefit is evident under load: in mixed-workload tests, they can handle 40 to 50% more requests while still meeting their latency SLO.
There's also a quite ingenious approach to load balancing. When SO_REUSEPORT is used, the system determines which worker to assign a connection to by hashing the addresses and ports of that connection, and since tcp_tw_reuse allows loopback connections to be reused, the hash always ends up on the same worker, causing the work to be distributed in an uneven way, with the busiest threads being as much as 40% more loaded than the least busy ones (this is very much dependent on the workflows). Their solution was neat in that instead of going against the kernel, each worker listened on its own port and left it to the nginx that was already in use to carry out the load balancing in front of them; round robin worked well, and for one service they even assigned the render-intensive and the lightweight endpoints to separate workers (4+1) using least_conn in order to pick up a few percent from the tail. This mode of operation is fully supported by Watt thanks to the portAssignment configuration perWorkerIncrement.
It is also worth mentioning in the article that Watt smoothly and automatically monitors threads and restarts any that cause the event loop to block. This was the behavior both that I developed and promoted; when a worker stops responding, it should be restarted rather than being monitored, and this restart should be visible without the need to look at the dashboard. It works like that.
Booking ran the test as a proper A/B production trial: two complete installations ran side by side, both with the same number of pods and the same resources; traffic was split at the routing level; and the extra headroom was used to tune the V8 GC. This is a thorough, honest way to verify a migration, and it is exactly the kind of engineering care that makes a case study worthwhile.
If you are curious and wants to know more, just hit "reply", or contact me at matteo@platformatic.dev.
Thanks for reading!