Three Minutes of Trouble, 20 million events

This is a reflection of a post written by Sean aka webframp on a serious load test performed against the club. An insight into how it looked from the other side.

I spent this last Saturday hacking on Swamp Club building a combat log and a couple handfuls of other non-user facing performance and architecture changes. Quick aside: The combat log is awesome. It exists to answer how users at the top of the leaderboard get there. Ultimate transparency. What are they doing, exactly?

Anyway, to further set the stage, I enjoy working over the weekend while chatting on Discord. Get a little Claude going, talk a little shit, dream a little, maybe watch some world cup. It’s cozy, it’s nice, most of all, it’s fun.

Fast forward to the evening, I’m almost finished with the campaign in the new Torchlight Infinite season, and I see these 2 messages from Mat Greten in the span of 30 seconds.

🚨 swamp club site may be down

It’s back for me now

Check the site myself. Seems slow. Pop open Grafana, look at the nodes. Looks bad. Ask homie claude to look at the logs…

CPU, memory, network, and disk during the load test

Traefik keeps getting oom-killed. Telemetry service is struggling. Something is happening. Fun is happening.

In the end, all that needed to change was 2 things. Scale Traefik up from 2 instances to 6. Scale our ingest servers up from 2 to 4. That was enough to sustain the attack 600 event/second burst of usage and drain the backlog at 2,400 events/second.

3 minutes of instability. ~45 minutes to drain. No data lost.

Growth

On Saturday morning, Swamp Club had processed ~12.5 million automation events. That’s anytime someone runs a swamp command with telemetry enabled. That’s you, that’s me, that’s automation, that’s CI. Swamp delivering value to someone, somewhere.

By midnight on Saturday evening, Swamp Club was at 32 million events. 9 am Sunday morning: 36.2 million. As I write this: 41.3 million. That’s 20 million events in an evening.

To put this in perspective: it took ~60 days to go from 0 to 1 million. ~8 days to go from 1 to 2 million. On July 6th we celebrated 10 million.

And we’re still just getting started.

Comms and the aftermath

I had a hunch, based on a number of data points, about who the culprit was. I sent webframp a discord DM

are you running something insane rn?

Turns out yes, yes he was. He gave me the details he shared in his post. Apologized for bringing stuff down. My response? “yeah, don’t stop. you’re completely fine this is helping”

You see - it’s one thing to design for scale. It’s another to live it. When someone gives you the opportunity, you don’t shy away. You embrace it. Own it.

I called the shot after a little back and forth.

this is exactly the type of demand we need to meet, and we should design for, and don’t apologize! this is how stories are made

Yep. That’s it. The entire post can be boiled down to this sentence.

As Sean points out, there’s a lot of ways to respond to a situation like this, downtime on the line, when the perceived stakes are low. I could have told him to turn it off and knock it out. It would have been valid, even.

But honestly; that’s not fun. I’d rather know when the stakes are low. He’d rather know, too. How does this small team react when things go bad? Now he knows.

Fun

It’s hard in retrospect to convey just how much fun I was having: to see it all playing out exactly as I’d hoped. To dare the test to continue. To harden the design afterwards.

And I think that’s what it all boils down to in the end. I’m having a blast building Swamp Club for you all, interacting with our community of builders every day. Software may be dead, but whatever this is? It’s alive.

And that’s all that matters.