elvis/portfolio
back to all posts
System DesignFebruary 2, 2026·7 min read

I built a URL shortener I'll never ship, and it was still worth a weekend

It's the classic interview question for a reason — caching, rate limits, and collisions all show up fast, even in a toy you build for yourself.

I want to be upfront about what this actually was, because I think the honesty matters more than the polish: I didn't build a URL shortener because the world needs another one, or because I planned to launch it. I built it the way you'd run drills before a game — a deliberate practice exercise, the classic system design interview question, worked all the way through on my own machine instead of talked through out loud in front of someone with a whiteboard marker. No users, no deployment, no README promising it solves a real problem. Just me, seeing how far "shorten a URL" actually goes once you stop treating it as a toy.

It goes surprisingly far. Further than the one-sentence pitch has any right to suggest.

The boring version, on purpose, first

The instinct with any system design exercise is to reach straight for the interesting parts — the caching layer, the rate limiter, the distributed ID generation. I made myself resist that and build the boring version first, because I think skipping straight to the interesting parts is exactly how you end up over-engineering solutions to problems you haven't actually confirmed you have yet:

app.post("/shorten", async (req, res) => {
  const code = generateRandomCode(7);
  await db.urls.insert({ code, longUrl: req.body.url });
  res.json({ shortUrl: `https://short.ly/${code}` });
});

app.get("/:code", async (req, res) => {
  const entry = await db.urls.findOne({ code: req.params.code });
  if (!entry) return res.status(404).send("Not found");
  res.redirect(entry.longUrl);
});

That's genuinely the whole thing. Two routes, one table. It handles the actual job correctly, and I think it's worth sitting with how small that is, because everything that follows exists only because a real system has to survive conditions this version quietly assumes away.

Collisions: the thing that seems unlikely until you do the math

generateRandomCode(7) picks seven random characters from a 62-character alphabet. That's around three and a half trillion possible codes, which feels safely, comfortably unlimited — right up until you actually calculate when two random picks are likely to collide, and the birthday paradox turns "trillions of options" into "meaningfully likely long before you'd expect," once you're generating millions of codes rather than a handful.

The fix isn't a bigger alphabet or more characters, which just delays the same problem. It's checking for collisions honestly, at write time:

async function generateUniqueCode(): Promise<string> {
  let code: string;
  let exists: boolean;
  do {
    code = generateRandomCode(7);
    exists = await db.urls.exists({ code });
  } while (exists);
  return code;
}

That loop is correct and also, on paper, a little alarming — an unbounded retry inside a request handler is the kind of thing that looks fine until load makes it not fine. In practice, at the collision rates involved, the loop runs zero or one extra times in the overwhelming majority of calls, so it stays cheap. But it's exactly the kind of code where "it'll almost always be fine" needs to actually be measured, not just assumed, before it goes anywhere near production traffic — and this exercise was the first time I'd sat with that specific tension long enough to actually feel uncomfortable asserting it without numbers.

The read path is where the real system design lives

Here's the part of the exercise that actually earned its reputation as a real system design question, not a toy one: a URL shortener's redirect endpoint is read-heavy in a way that's almost extreme. Every single click hits GET /:code, and any popular link gets hit constantly, while writes — someone actually creating a new short link — happen comparatively rarely. That asymmetry is the whole game, and it's exactly the shape Redis exists for.

app.get("/:code", async (req, res) => {
  const cached = await redis.get(`url:${req.params.code}`);
  if (cached) return res.redirect(cached);

  const entry = await db.urls.findOne({ code: req.params.code });
  if (!entry) return res.status(404).send("Not found");

  await redis.set(`url:${req.params.code}`, entry.longUrl, "EX", 3600);
  res.redirect(entry.longUrl);
});

Popular links, the ones actually generating meaningful load, end up served almost entirely out of Redis after their first hit, and Postgres only sees the cache misses. That's not a premature optimization bolted onto a working system — it's a direct, structural response to the actual read/write ratio this specific kind of system has. I think that distinction, "optimize because you've identified the real shape of the load" versus "optimize because caching sounds like a thing serious systems do," is the single most useful thing this exercise actually taught me, more than any individual technique.

Rate limiting: protecting the system from its own convenience

The last piece I added was rate limiting on link creation, and it's worth naming honestly why it matters: a URL shortener that lets anyone create unlimited links, instantly, for free, is a trivially good tool for spam and phishing — you generate a short, trustworthy-looking link pointing at anywhere you want. Rate limiting isn't a performance feature here. It's the thing standing between "useful tool" and "tool that gets abused within a day of being reachable."

const limiter = rateLimit({
  windowMs: 60 * 1000,
  max: 10,
  keyGenerator: (req) => req.ip,
});

app.post("/shorten", limiter, async (req, res) => {
  // ...
});

Ten links a minute per IP, which is generous for a real person and genuinely annoying for a script trying to churn out thousands of throwaway redirects. It's a small addition, but it's the one that made this feel less like a toy and more like a system that had actually reckoned with how it would get misused if it existed in the real world.

What I actually learned

I don't think the lesson is "you should build every classic interview question as a real project" — plenty of them are genuinely fine as pure whiteboard conversations, and this one probably would have been too. The lesson is narrower: talking through caching, collision handling, and rate limiting in the abstract and actually implementing them against a database that returns real, sometimes annoying, results are different exercises that happen to share vocabulary. The birthday paradox math didn't feel real to me until I was staring at an actual collision in an actual table. The read/write asymmetry didn't feel real until I'd written the boring version first and watched exactly which endpoint would obviously buckle under load.

I still have no plans to deploy this thing anywhere. That was never the point. The point was finding out that a question I could already answer correctly in an interview room felt meaningfully different once I'd actually built the thing the interview room was only ever asking me to describe.