A routine load test

I was load testing a messaging system. The goal was narrow: push a large volume of WhatsApp messages through it and watch how the RabbitMQ consumers behaved under pressure. Did they keep up? Did the queues back up? Did anything fall over?

I wrote the load script in Artillery. For recipients, I generated phone numbers in bulk. They had to be valid, or the system would reject them before they ever reached the queue.

The message itself was simple: Test RMQ {timestamp}

I queued millions of requests. It was a dev environment. What could go wrong?

The errors that told us

Then Kibana started filling up with the same error, over and over.

Meta was rejecting the messages. The recipient had no active session window, and the message was free-form text rather than an approved template, so WhatsApp refused to deliver it.

It took a moment for that to land. Meta rejecting our messages meant our messages were reaching Meta. From a dev environment, that should have been impossible.

The environment was connected to a live WhatsApp configuration. Our test traffic was going out to real phones.

It escalated fast. I was scared. Properly, stomach-dropping scared. Millions of messages, a live configuration, and no idea yet how many had gone out.

The rule that saved us

WhatsApp has a rule I have never been more grateful for.

A business can send free-form messages to someone only within 24 hours of that person's last message to the business. Outside that window, it can send only pre-approved template messages. Everything else is rejected.

Our load test sent free-form text. So for almost every number in the range, Meta said no. That was the wall of errors in Kibana.

The messages that did get through went to numbers that happened to have an open 24-hour window with the sending number. That came to tens of thousands.

Tens of thousands of people received a WhatsApp message that said Test RMQ, followed by a timestamp.

That is still a lot of people, and not something I ever want to repeat. But it was tens of thousands, not millions.

How it actually happened

The WhatsApp configuration for that environment lived in a database. It was a shared environment, so I had a habit: I always kept that configuration deliberately invalid. If anything ever tried to send a real message, it would fail.

That habit was my safety net. What I didn't know was that it had been undone. Another test on the same shared environment needed a working configuration, so it had been switched to a live one, and it stayed that way.

Nobody did anything unreasonable. One test needed a working configuration; mine needed a broken one. The environment allowed both, and told no one when it changed.

Look at what had to line up:

  • A shared environment where anyone could change a live configuration.
  • A change nobody announced, because nothing prompted anyone to.
  • Test data that was real. Every number I generated was a valid, dialable mobile number, because "valid" was all the validator asked for.

Take away any one of those and nothing happens. The invalid config fails every send. An announcement stops me before I start. Numbers nobody owns go nowhere.

All three were in place.

The only safeguard that held was one we didn't own: WhatsApp's 24-hour rule.

What I do differently now

Check where the traffic will go, not just what the script does. Before a high-volume run, I read the configuration the environment is using right now. My test was correct. The environment around it had changed.

Use test data that can't reach anyone. Passing validation is not the same as being safe. Use numbers your team owns, or the provider's test numbers, and never a range of real ones.

Start with ten, not millions. Send a handful of messages first and confirm exactly where they land. A small first step turns a disaster into a non-event.

Reserve the environment and tell people. On a shared environment, announce load tests, reserve the environment and its configuration for the duration, and put things back when you're done.

Make the environment enforce it. A person remembering to keep a config invalid is a habit, not a guardrail. Non-production environments shouldn't be able to hold live credentials at all, or should deliver only to an allowlist of numbers the team controls. Better still, point them at stubs of the outside world, as I wrote about in One Script, One Laptop, the Whole Stack.

The lesson I keep coming back to

None of this was a testing failure in the usual sense. The script did what it was written to do. The queues behaved. The failure was in the setup, and it was sitting there before I pressed run.

Quality is built long before the first test runs. That includes the environment you run it in.

Every test depends on the environment around it. Check your setup. Know where your traffic is going.

And be careful with the Send button.