elvis/portfolio
back to all posts
BackendApril 28, 2026·5 min read

Cron jobs, Redis, and the illusion of simplicity

Pinging an API every minute sounds trivial until failures start compounding.

The pitch for the API monitoring platform I built took about one sentence to explain to a friend: register a URL, ping it every minute, tell someone if it stops responding. He nodded like I'd told him water is wet. I nodded back, agreeing it was simple, and then went and built the naive version in about forty minutes, and then spent the next several evenings finding out exactly how not-simple "ping it every minute" actually is once you stop imagining a URL that always responds politely on time.

Version one, which worked right up until it didn't

The first version was embarrassingly straightforward, and I mean that as a compliment to how little code it took, not as a compliment to how well it worked:

cron.schedule("* * * * *", async () => {
  const endpoints = await getAllEndpoints();
  for (const endpoint of endpoints) {
    const result = await pingEndpoint(endpoint.url);
    await saveResult(endpoint.id, result);
  }
});

Every minute, loop through every registered endpoint, ping it, save the result. It's the loop you'd sketch on a whiteboard in about ten seconds, and for the first few dozen endpoints, sitting comfortably on fast, cooperative test URLs, it worked exactly as advertised.

It fell over the moment I registered an endpoint that was slow, not down. A staging API I had lying around with an artificial two-second delay on one route. Nothing dramatic, no crash, no error. Just slow. And that one slow endpoint was enough to reveal that the entire design had an assumption baked into it that I hadn't noticed I was making: this loop assumes every ping finishes well within the interval before the next one starts.

It doesn't take a very slow endpoint to break that assumption. It just takes one endpoint whose response time creeps close to, or past, the gap between runs.

What actually happens when a ping runs long

Here's the failure mode, and it's a quieter, more insidious one than a crash: the loop is still await-ing that slow endpoint when the next minute's cron tick fires. Depending on how the scheduler and the process are wired, you get one of two bad outcomes. Either the next tick queues up and fires late, so your "every minute" monitor quietly drifts to being an "every ninety seconds, sometimes, depending on load" monitor. Or worse, if nothing's actually blocking a second invocation from starting, you now have two overlapping runs both hammering the same slow endpoint, both writing results to the same row, in whatever order they happen to finish.

Neither of those failure modes announces itself. There's no error message that says "your monitoring interval has degraded." The dashboard just quietly starts being less accurate than it claims to be, and you find out weeks later when someone asks why an outage that lasted six real minutes only shows four minutes of downtime in the history.

That's the part that actually got to me. I hadn't built something that crashes under load. I'd built something that lies gently under load, and a monitor that lies is arguably worse than one that's honestly offline, because at least an offline monitor tells you to go check on it.

Why a queue actually changes the shape of the problem

The fix isn't "make the ping faster" — some endpoints are just going to be slow, that's real information the monitor should capture, not something to engineer away. The fix is decoupling scheduling a ping from executing it, so a slow execution can never push the next scheduling tick off course. That's what a queue is actually for here, and it's the first time Redis clicked for me as something other than "a cache you put in front of a database":

cron.schedule("* * * * *", async () => {
  const endpoints = await getAllEndpoints();
  for (const endpoint of endpoints) {
    await pingQueue.add("ping", { endpointId: endpoint.id });
  }
});

The scheduler's entire job now is enqueueing lightweight jobs, which takes milliseconds regardless of how slow any endpoint is, because it's not waiting on the endpoint at all. A separate pool of workers pulls jobs off that queue and does the actual pinging, each on its own clock:

pingWorker.process(async (job) => {
  const endpoint = await getEndpoint(job.data.endpointId);
  const result = await pingEndpoint(endpoint.url);
  await saveResult(endpoint.id, result);
});

Now a slow endpoint only ever slows down the one worker handling it. The scheduler keeps its minute-by-minute rhythm regardless of what any individual ping is doing, because scheduling and executing are no longer the same synchronous unit of work fighting over the same clock tick.

The failure I didn't expect: retries turning a blip into a false alarm

Decoupling the queue from the schedule fixed the drift problem, but it exposed a second one almost immediately, which is the kind of thing you only find by actually running the thing, not by reasoning about it on paper. A single dropped packet, one genuinely transient network hiccup on an endpoint that's otherwise perfectly healthy, was enough to flip that endpoint to "down" and fire an alert. Technically accurate in the sense that the ping did fail. Practically useless, because now I'm getting paged for noise, and noise is exactly the thing that trains you to stop trusting a monitor.

The fix was giving failures room to be wrong once before believing them:

pingWorker.process(async (job) => {
  const endpoint = await getEndpoint(job.data.endpointId);
  try {
    const result = await pingEndpoint(endpoint.url);
    await saveResult(endpoint.id, { ...result, status: "up" });
  } catch (err) {
    if (job.attemptsMade < 2) {
      throw err; // let the queue's built-in backoff retry this job
    }
    await saveResult(endpoint.id, { status: "down", error: String(err) });
    await sendAlert(endpoint);
  }
});

Two retries with backoff before an endpoint is actually marked down and an alert goes out. That's not a complicated idea, but it's the exact kind of nuance that only shows up once a naive version is already running against something real. On paper, "ping it and record the result" sounds complete. In practice, the difference between a monitor you trust and one you learn to ignore lives entirely in how it handles the boring, unglamorous case of it failed once, was that real.

What I actually learned

I don't think the lesson here is "always use a job queue for scheduled work" — plenty of cron jobs really are simple, fast, and fine running exactly the way they look on the whiteboard. The lesson is narrower: a scheduler and the work it schedules are two different responsibilities with two different failure modes, and the moment you let one endpoint's slowness become the scheduler's problem, your "every minute" stops meaning every minute. A queue isn't there to make things faster. It's there so that one bad citizen can't quietly corrupt the timing guarantee everything else depends on.

That's not a lesson I'd have believed from a blog post before I built this. I'd have nodded along the same way my friend nodded when I explained the pitch — sure, obviously, decouple your concerns, everyone knows that. It took watching my own dashboard quietly under-report a real outage, with no error anywhere in the logs to point at, before "decouple scheduling from execution" stopped being advice I agreed with and started being a rule I actually reach for by instinct now.