elvis/portfolio
back to all posts
BackendJanuary 5, 2026·6 min read

Background workers: the part of backend nobody shows you

Most tutorials stop at the request/response cycle. Real systems live in the queue.

Every backend tutorial I ever learned from taught the same shape, and taught it well: a request comes in, a handler does some work, a response goes out. req in, res out, and everything in between happens synchronously enough that you can trace it top to bottom in one linear read. That shape is genuinely correct for an enormous amount of real backend work, and I don't want to undersell how far it gets you.

It's also a shape that quietly stops applying the moment "the work" isn't something that fits inside a request's lifetime, and building an API monitoring platform was the first time I ran face-first into that limit, because the entire point of the product is doing work that has no request attached to it at all.

The work that has no request to belong to

Nobody sends an HTTP request that means "please ping every registered endpoint right now." That work has to happen on its own, on a schedule, forever, with no client waiting on the other end and no response to send back when it's done. The first time I actually sat with that, it was a little disorienting — my entire mental model of "backend" had an implicit request sitting at the root of every call stack, and this had none.

The honest answer to "how does this work run" is: a process, started separately from the web server, that just runs continuously, entirely decoupled from anyone asking it to:

// worker.ts — not attached to any route
async function startPingLoop() {
  while (true) {
    const dueEndpoints = await getEndpointsDueForCheck();
    for (const endpoint of dueEndpoints) {
      await pingQueue.add("ping", { endpointId: endpoint.id });
    }
    await sleep(5000);
  }
}

startPingLoop();

That's not a route handler. It's not triggered by anything external at all. It's a process whose entire job is existing, indefinitely, and doing its work on its own clock — closer in spirit to a daemon than to anything a request/response tutorial ever prepared me to write.

Failure without anyone watching

Here's the part that request/response thinking genuinely didn't prepare me for: in a normal API handler, if something throws, the framework catches it, sends a 500, and a client somewhere sees an error and (hopefully) does something about it. There's always, structurally, someone downstream who finds out something went wrong.

A background worker has no one downstream. If the ping loop throws an unhandled exception and the process dies, there is no client anywhere in the world who receives an error. Nothing at all happens, from the outside, except that endpoint monitoring silently stops. The very first version of this worker did exactly that during testing — one malformed endpoint URL threw deep inside the ping logic, the loop's try block didn't actually wrap the right section, and the entire monitoring process quietly exited. I only found out because I happened to check the dashboard an hour later and noticed every single endpoint, including ones that were completely fine, had simply stopped reporting anything.

That's a genuinely different category of bug from anything a request/response handler produces, because there's no natural signal that it happened at all. The fix isn't clever, but it has to be deliberate in a way request handlers don't require, because nothing forces you to think about it otherwise:

async function startPingLoop() {
  while (true) {
    try {
      const dueEndpoints = await getEndpointsDueForCheck();
      for (const endpoint of dueEndpoints) {
        await pingQueue.add("ping", { endpointId: endpoint.id });
      }
    } catch (err) {
      logger.error("ping loop iteration failed", err);
      await alertOncall("Monitoring loop error", err);
      // deliberately NOT re-thrown — one bad iteration shouldn't kill
      // the whole loop, unlike a route handler where propagating an
      // error to the framework is exactly the right move
    }
    await sleep(5000);
  }
}

That inline comment is doing real work, because it's the exact opposite instinct from writing a route handler. In a request handler, letting an error propagate up to the framework is almost always correct — the framework knows what to do with it, turn it into a 500, log it, move on. In a loop with no framework catching anything on your behalf, letting an error propagate means the entire process dies, silently, with nobody watching. The right move flips completely: catch broadly, log loudly, alert explicitly, and keep the loop alive no matter what any single iteration does.

Idempotency, or: what happens when the same job runs twice

The other assumption request/response thinking quietly bakes in is that each unit of work happens exactly once — one request, one handler execution, one response. Queued background jobs don't get that guarantee for free. A worker can crash mid-job and get restarted, a job can be retried after a timeout that turns out to have been a false alarm, a deploy can restart the worker process while a job is mid-flight. Any of those can mean the same ping job executes twice.

For a read-only operation like pinging an endpoint, running twice is mostly harmless — you'd just get one duplicate result row. But the moment a background job triggers something with a side effect, like sending an alert email, "maybe runs twice" stops being harmless and starts being a real bug: nobody wants two identical "your API is down" emails four seconds apart. I ended up wrapping alert-sending in an explicit check against a short-lived flag, keyed to the specific incident, before actually sending anything:

async function sendAlertIfNotAlreadySent(endpoint: Endpoint) {
  const alreadySent = await redis.get(`alert-sent:${endpoint.id}`);
  if (alreadySent) return;
  await redis.set(`alert-sent:${endpoint.id}`, "1", "EX", 300);
  await sendAlert(endpoint);
}

That five-minute window is a deliberate, somewhat arbitrary tradeoff, not a universal constant — long enough to absorb a duplicate job firing seconds apart, short enough that a genuinely new outage starting a few minutes later still triggers its own fresh alert instead of getting silently swallowed by the same flag.

What I actually learned

I don't think the lesson is "background workers are just harder than request handlers, full stop" — they're not harder so much as they're a genuinely different shape of problem, and the request/response tutorials I learned from were never wrong, they just only ever taught one shape. The lesson is narrower: the moment work exists without a request attached to it, every assumption that shape quietly gave you for free — someone downstream to receive an error, one execution per unit of work — has to be rebuilt on purpose, because nothing hands it to you anymore. Error handling stops being "let the framework deal with it" and starts being something you have to actively design. Idempotency stops being invisible and starts being a decision.

I don't think I'd have internalized any of that from a tutorial, honestly, no matter how well written. It took a monitoring dashboard silently going dark for an hour, with nothing in any log screaming about it, to make the difference between those two shapes feel real instead of theoretical.