Some of you got woken up over the past few weeks for sites that were fine. You'd open the laptop at 2am, load the page, and there it was, up. After that happens three or four times you stop trusting the alerts, and an alert you don't trust is worth nothing.
What we assumed, and what the data said
Moonitor checks your site from Helsinki. Before it opens an incident it asks a second machine in Los Angeles to look too. The idea is that if our Helsinki box is having a bad day, Los Angeles disagrees and we keep quiet.
So we assumed Los Angeles had gone flaky and was calling things down that weren't.
It hadn't. We pulled the records, and every incident from the previous week had been independently confirmed by Los Angeles. The box was healthy and barely working. It wasn't inventing anything.
The problem was ours, and it was subtler than a broken box. Both machines were checking at the same instant. So when a site hiccuped for a second, both of them genuinely saw the hiccup, agreed with each other, and we treated that agreement as proof of an outage.
On top of that, we opened an incident on the very first failed check. One unlucky sample was enough to wake you up.
One customer's site made this painfully obvious. It sits behind a web server that returns an error on roughly one request in five, at random. We tested from three continents and got the same thing everywhere.
What changed
A fault now has to clear two bars before we'll alert you.
The first is agreement. A majority of the checking locations that answer have to see the same thing. One machine's opinion isn't enough.
The second is the one that actually matters: it has to still be broken a few seconds later. When a check fails we wait, then check again from every location. Site's back? It was a blip. We log it and leave you alone. Still down? You get the alert.
The same rule runs in reverse now, which fixes a separate complaint we've had: a site that flickers back up for one check no longer gets marked recovered, then down again ten minutes later. No more alert storms out of one wobbly server.
A third continent
We've added a checking location in Johannesburg, alongside Helsinki and Los Angeles. Different provider, different network.
Part of that is tie-breaking. With two locations they could only agree or deadlock, and a deadlock meant we froze and did nothing at all. On one monitor that happened 23 times in a month, and we were blind to it the whole time. A third location settles the argument.
A fourth continent is being added at the moment.
Your uptime number stays honest
When the locations disagree, or when a fault vanishes on the second look, we record the check and show it in your history. It does not count against your uptime percentage. If we can't demonstrate a fault happened, we're not putting it on your report card.