How to Design Distributed Rate Limiting for a Multi-Tenant API
13 min read

How to Design Distributed Rate Limiting for a Multi-Tenant API

By Govind Sharma

How to Think About Rate Limiting in a Multi-Tenant API

A rate limiter on a multi-tenant API has one job: stop one tenant from spending capacity that belongs to the others. Everything else people ask of it (abuse prevention, cost control, plan enforcement) is the same mechanism with a different key and a different number.

That framing decides most of the design. The limiter sits on the hot path of every request, so it has to be cheap. It runs on every API instance, so its state has to be shared. And it is a dependency that can fail, so each route needs a decided answer to the question of what happens when the limiter itself is unavailable.

This post walks through one design that holds up in production: GCRA in a single Redis Lua script, layered tenant-scoped keys, standard response headers, an explicit failure mode per route, and the same primitive reused for outbound calls to third-party providers.

Requirements worth writing down first

  • Atomic: two instances deciding at the same instant must not both admit the last request.
  • One round trip: the decision and the state update happen in a single Redis call.
  • Bounded memory: state per key is constant, and idle keys expire on their own.
  • Honest feedback: a rejected client is told exactly how long to wait.
  • Explicit failure mode: every policy declares fail-open or fail-closed.

Choosing the Algorithm

The four common choices differ in what they store per key and how they behave at the edges of a window. The differences matter more than they look on a whiteboard.

  • Fixed window: one counter per key per window. Cheapest, but a client can send a full quota at the end of one window and another at the start of the next, which is twice the limit in a short span.
  • Sliding window log: a sorted set of request timestamps. Exact, but memory grows with the limit, and a tenant allowed 10,000 requests per minute costs 10,000 entries.
  • Sliding window counter: two fixed windows blended by overlap. Constant memory and a good approximation, but it assumes traffic was evenly spread across the previous window.
  • Token bucket: a token count and a last-refill timestamp. Models a sustained rate plus a burst allowance, which is how most API quotas are actually described.
  • GCRA: the same behaviour as a token bucket, stored as a single timestamp.

Fixed window also has an implementation trap that shows up in a lot of first attempts. The counter and its expiry are set in two separate commands:

fixed-window-naive.tsts
// Do not ship this.
const windowKey = "rl:" + tenantId + ":" + Math.floor(Date.now() / 60_000);

const count = await redis.incr(windowKey);
if (count === 1) {
  // If the process dies between INCR and EXPIRE, this key never expires.
  await redis.expire(windowKey, 60);
}

return count <= limit;

The fix is to make the read, the decision and the write one atomic unit. Once the logic is in a script anyway, there is little reason to stay with fixed windows.

Why GCRA

The generic cell rate algorithm keeps one number per key: the theoretical arrival time, or TAT. It is the moment at which the bucket would be full again if no more requests arrived. Each admitted request pushes the TAT forward by one emission interval (the period divided by the limit). A request is admitted if the TAT, after that push, is no further ahead of now than the burst allowance.

  • State is one value with a TTL, so memory per key is constant and idle keys disappear.
  • There is no refill job and no window boundary to straddle.
  • The time until the next request would be admitted falls out of the arithmetic, so Retry-After is exact.
  • Weighted requests are a multiplication: an expensive call can cost 10 units instead of 1.

The Atomic Redis Script

The script takes the emission interval, the burst size and the cost of the request. It reads the clock from Redis instead of trusting the caller, so clock skew between API instances cannot let one instance admit requests another would reject.

gcra.lualua
-- KEYS[1] = limiter key
-- ARGV[1] = emission interval in microseconds (period / limit)
-- ARGV[2] = burst (how many requests may be admitted at once)
-- ARGV[3] = cost of this request
-- Returns { allowed, remaining, retry_after_ms, reset_after_ms }

local key = KEYS[1]
local emission = tonumber(ARGV[1])
local burst = tonumber(ARGV[2])
local cost = tonumber(ARGV[3])

local t = redis.call("TIME")
local now = tonumber(t[1]) * 1000000 + tonumber(t[2])

-- A missing key or a TAT in the past both mean a full bucket.
local tat = tonumber(redis.call("GET", key))
if tat == nil or tat < now then
  tat = now
end

local new_tat = tat + emission * cost
local allow_at = new_tat - emission * burst

if now < allow_at then
  -- Rejected: state is left untouched, so rejected calls cost nothing.
  return {
    0,
    0,
    math.ceil((allow_at - now) / 1000),
    math.ceil((tat - now) / 1000),
  }
end

local ttl_ms = math.ceil((new_tat - now) / 1000)
redis.call("SET", key, string.format("%.0f", new_tat), "PX", ttl_ms)

return {
  1,
  math.floor((now - allow_at) / emission),
  0,
  ttl_ms,
}

A few details in that script are deliberate:

  • Microsecond integers keep the arithmetic exact. Current epoch time in microseconds fits comfortably inside the 53 bits a Lua number can represent without rounding.
  • The TAT is written with an explicit integer format, because the default number-to-string conversion in Lua switches to exponent notation for values this large.
  • The TTL equals the time until the bucket is full. After that the key carries no information, so letting it expire is equivalent to keeping it.
  • Rejected requests do not move the TAT. A client hammering a limited key cannot extend its own penalty, and cannot cause write load either.
  • Calling TIME before a write is safe on Redis 5 and later, where scripts replicate their effects instead of being re-executed on replicas.

The TypeScript wrapper

The wrapper registers the script once, so ioredis sends it by SHA and only falls back to the full source when the script cache is cold. Policies are expressed the way a product person would say them: a limit per period, plus a burst.

rate-limiter.tsts
import { readFileSync } from "node:fs";
import { join } from "node:path";
import type Redis from "ioredis";
import type { Result } from "ioredis";

declare module "ioredis" {
  interface RedisCommander<Context> {
    gcra(
      key: string,
      emissionMicros: number,
      burst: number,
      cost: number,
    ): Result<[number, number, number, number], Context>;
  }
}

export interface RateLimitPolicy {
  name: string;
  limit: number; // sustained requests per period
  periodMs: number;
  burst: number; // requests admitted back to back from idle
}

export interface RateLimitDecision {
  policy: RateLimitPolicy;
  allowed: boolean;
  remaining: number;
  retryAfterMs: number;
  resetAfterMs: number;
}

export class RateLimiter {
  constructor(private readonly redis: Redis) {
    this.redis.defineCommand("gcra", {
      numberOfKeys: 1,
      lua: readFileSync(join(__dirname, "gcra.lua"), "utf8"),
    });
  }

  async consume(
    key: string,
    policy: RateLimitPolicy,
    cost = 1,
  ): Promise<RateLimitDecision> {
    if (cost > policy.burst) {
      // Such a request could never be admitted, so fail loudly instead.
      throw new RangeError("Request cost exceeds burst for policy " + policy.name);
    }

    const emissionMicros = Math.max(
      1,
      Math.round((policy.periodMs * 1000) / policy.limit),
    );

    const [allowed, remaining, retryAfterMs, resetAfterMs] = await this.redis.gcra(
      key,
      emissionMicros,
      policy.burst,
      cost,
    );

    return { policy, allowed: allowed === 1, remaining, retryAfterMs, resetAfterMs };
  }
}

Limit and burst are different promises

With burst equal to limit, a tenant starting from idle can send the full burst immediately and then keep receiving refill, so the first period can see close to twice the limit. The sustained rate is what GCRA guarantees. If the downstream cannot absorb that first spike, set burst well below limit.

Key Design: Tenant, Route Class, Credential

The algorithm decides how a limit behaves. The key decides who shares it, and this is where most multi-tenant limiters go wrong. A single counter per tenant lets a bulk export starve that tenant's own interactive traffic. A counter per endpoint multiplies the effective quota by the number of endpoints.

A layered scheme works better. Each request is checked against a small, fixed set of limits, narrowest first:

  1. Tenant and route class: groups endpoints by cost (read, write, search, export) instead of by URL, so adding an endpoint does not add quota.
  2. Tenant overall: the plan-level number that appears on the pricing page.
  3. Credential or IP: only for unauthenticated routes such as login, where no tenant is known yet.
rate-limit-policies.tsts
import type { RateLimitPolicy } from "./rate-limiter";

export type Plan = "free" | "team" | "enterprise";
export type RouteClass = "read" | "write" | "search" | "export";
export type FailMode = "open" | "closed";

export interface ResolvedLimit {
  key: string;
  policy: RateLimitPolicy;
  failMode: FailMode;
}

// Illustrative numbers. Real ones come from capacity tests, not from this file.
const TENANT_LIMITS: Record<Plan, RateLimitPolicy> = {
  free: { name: "tenant.free", limit: 120, periodMs: 60_000, burst: 30 },
  team: { name: "tenant.team", limit: 1_200, periodMs: 60_000, burst: 200 },
  enterprise: { name: "tenant.enterprise", limit: 6_000, periodMs: 60_000, burst: 600 },
};

const CLASS_SHARE: Record<RouteClass, { share: number; failMode: FailMode }> = {
  read: { share: 1, failMode: "open" },
  write: { share: 0.5, failMode: "open" },
  search: { share: 0.25, failMode: "open" },
  export: { share: 0.02, failMode: "closed" },
};

export function resolveLimits(input: {
  tenantId: string;
  plan: Plan;
  routeClass: RouteClass;
}): ResolvedLimit[] {
  const tenant = TENANT_LIMITS[input.plan];
  const { share, failMode } = CLASS_SHARE[input.routeClass];

  // The braces are a Redis Cluster hash tag: every key for one tenant
  // lands in the same slot, so they can be combined in one script later.
  const prefix = "rl:{" + input.tenantId + "}:";

  return [
    {
      key: prefix + "class:" + input.routeClass,
      policy: {
        name: tenant.name + "." + input.routeClass,
        limit: Math.max(1, Math.floor(tenant.limit * share)),
        periodMs: tenant.periodMs,
        burst: Math.max(1, Math.floor(tenant.burst * share)),
      },
      failMode,
    },
    { key: prefix + "all", policy: tenant, failMode: "open" },
  ];
}
  • Key on the tenant ID from the verified credential, never on a header the client controls.
  • Keep route classes few. Each class is a number someone has to justify and maintain.
  • For IP-based keys, trust the forwarded address only from your own proxy, and bucket IPv6 clients by /64 prefix so one host cannot rotate through addresses.
  • Checking the narrow limit first means a rejected export does not consume the tenant-wide budget. The reverse case, where the narrow check passes and the wide one rejects, charges one unit that was not used. That errs toward limiting and is usually acceptable; if it is not, check both keys inside one script.

The Response Contract

A limiter that only returns 429 teaches clients to retry in a tight loop. The response has to carry enough information for a well-behaved client to back off correctly without guessing.

  • 429 Too Many Requests for a rejected call, with a machine-readable error body.
  • Retry-After in whole seconds, rounded up. Rounding down invites a retry that is rejected again.
  • RateLimit-Limit, RateLimit-Remaining and RateLimit-Reset on every response, so clients can slow down before they hit the wall. Reset is a number of seconds from now, not an epoch timestamp.

The RateLimit header fields come from an IETF draft that has changed shape across revisions: earlier versions define the three separate fields used here, later ones fold them into structured RateLimit and RateLimit-Policy fields. Pick one form, document it, and do not switch silently, because client SDKs will parse whatever you ship first.

In NestJS this fits in a guard. It has to run after authentication, since it needs the tenant, and before anything expensive.

rate-limit.guard.tsts
import {
  CanActivate,
  ExecutionContext,
  HttpException,
  HttpStatus,
  Injectable,
} from "@nestjs/common";
import type { Request, Response } from "express";
import { RateLimitDecision, RateLimiter } from "./rate-limiter";
import { Plan, RouteClass, resolveLimits } from "./rate-limit-policies";

const LIMITER_TIMEOUT_MS = 50;

interface TenantRequest extends Request {
  tenant: { id: string; plan: Plan };
  routeClass: RouteClass;
}

function withTimeout<T>(work: Promise<T>, ms: number): Promise<T> {
  return new Promise<T>((resolve, reject) => {
    const timer = setTimeout(() => reject(new Error("rate limiter timeout")), ms);
    work.then(resolve, reject).finally(() => clearTimeout(timer));
  });
}

@Injectable()
export class RateLimitGuard implements CanActivate {
  constructor(
    private readonly limiter: RateLimiter,
    private readonly metrics: MetricsClient,
  ) {}

  async canActivate(context: ExecutionContext): Promise<boolean> {
    const http = context.switchToHttp();
    const req = http.getRequest<TenantRequest>();
    const res = http.getResponse<Response>();

    const limits = resolveLimits({
      tenantId: req.tenant.id,
      plan: req.tenant.plan,
      routeClass: req.routeClass,
    });

    let tightest: RateLimitDecision | undefined;

    for (const { key, policy, failMode } of limits) {
      let decision: RateLimitDecision;

      try {
        decision = await withTimeout(this.limiter.consume(key, policy), LIMITER_TIMEOUT_MS);
      } catch {
        this.metrics.increment("ratelimit.backend_error", { policy: policy.name, failMode });
        if (failMode === "closed") {
          res.setHeader("Retry-After", "5");
          throw new HttpException(
            { error: "rate_limiter_unavailable" },
            HttpStatus.SERVICE_UNAVAILABLE,
          );
        }
        continue; // fail open for this limit
      }

      this.metrics.increment("ratelimit.decision", {
        policy: policy.name,
        outcome: decision.allowed ? "allowed" : "limited",
      });

      if (!decision.allowed) {
        const retryAfterSeconds = Math.ceil(decision.retryAfterMs / 1000);
        this.writeHeaders(res, decision);
        res.setHeader("Retry-After", String(retryAfterSeconds));
        throw new HttpException(
          { error: "rate_limited", policy: policy.name, retryAfterSeconds },
          HttpStatus.TOO_MANY_REQUESTS,
        );
      }

      if (!tightest || decision.remaining < tightest.remaining) {
        tightest = decision;
      }
    }

    if (tightest) this.writeHeaders(res, tightest);
    return true;
  }

  private writeHeaders(res: Response, decision: RateLimitDecision) {
    res.setHeader("RateLimit-Limit", String(decision.policy.limit));
    res.setHeader("RateLimit-Remaining", String(decision.remaining));
    res.setHeader("RateLimit-Reset", String(Math.ceil(decision.resetAfterMs / 1000)));
  }
}

The headers report whichever limit is closest to exhaustion, because that is the one the client will hit next. MetricsClient stands in for whatever metrics library the service already uses.

Fail-Open vs Fail-Closed When Redis Is Down

Redis will be unavailable at some point: a failover, a network partition, a saturated connection pool. If the guard simply propagates the error, the rate limiter has turned a degraded cache into a full API outage. The decision has to be made in advance, per policy.

  • Fail open for ordinary reads and writes. The limiter protects fairness, and briefly losing fairness is better than rejecting every customer.
  • Fail closed for anything where the limit is a safety or cost control: login and OTP attempts, password reset, endpoints that send SMS or place calls, expensive exports, LLM-backed routes.
  • Fail closed with 503 and a short Retry-After, not 429. The client did nothing wrong, and dashboards should not show it as throttling.

Two things keep fail-open from being an unbounded hole. The first is the timeout in the guard: a slow Redis is more dangerous than a dead one, because every request waits on it. Pair a short per-call timeout with a circuit breaker so that after repeated failures the guard stops calling Redis for a few seconds. The second is a local fallback: an in-process token bucket per tenant, sized at the tenant limit divided by the number of instances, with some headroom. It is inaccurate, since load is never spread perfectly evenly, but it turns no limit at all into a rough one.

Count every fail-open decision

A limiter that fails open silently looks healthy on every dashboard while enforcing nothing. Emit a metric each time a request is admitted without a decision, and alert on it. It is also the earliest signal that Redis latency is climbing.

Outbound Limits to Third-Party Providers

The same problem exists in the other direction. Telephony, messaging, email and LLM providers all impose their own limits, and from their side your whole platform is one tenant. Two things change compared with inbound limiting.

First, the right response to being over the limit is usually to wait, not to reject. The work is typically a background job that can run a few hundred milliseconds later. Second, the quota is shared across all of your tenants, so without a per-tenant share one large campaign consumes the provider budget and every other tenant queues behind it.

outbound-limiter.tsts
import { setTimeout as sleep } from "node:timers/promises";
import type { RateLimitPolicy, RateLimiter } from "./rate-limiter";

export class OutboundBudgetExceededError extends Error {
  constructor(
    readonly key: string,
    readonly retryAfterMs: number,
  ) {
    super("Outbound budget exhausted for " + key);
  }
}

// Illustrative: 100 requests per second account-wide, 20 per tenant.
const PROVIDER: RateLimitPolicy = { name: "sms.provider", limit: 100, periodMs: 1_000, burst: 20 };
const PER_TENANT: RateLimitPolicy = { name: "sms.tenant", limit: 20, periodMs: 1_000, burst: 5 };

async function acquire(
  limiter: RateLimiter,
  key: string,
  policy: RateLimitPolicy,
  deadline: number,
) {
  for (;;) {
    const decision = await limiter.consume(key, policy);
    if (decision.allowed) return;

    // Jitter stops a group of waiting workers from waking at the same instant.
    const waitMs = decision.retryAfterMs + Math.floor(Math.random() * 25);
    if (Date.now() + waitMs > deadline) {
      throw new OutboundBudgetExceededError(key, decision.retryAfterMs);
    }
    await sleep(waitMs);
  }
}

export async function acquireSmsSlot(limiter: RateLimiter, tenantId: string, maxWaitMs = 2_000) {
  const deadline = Date.now() + maxWaitMs;

  // Tenant share first: a tenant over its share never touches the shared budget.
  await acquire(limiter, "out:sms:tenant:" + tenantId, PER_TENANT, deadline);
  await acquire(limiter, "out:sms:provider", PROVIDER, deadline);
}

When the wait budget runs out, the caller should put the job back on the queue with a delay equal to retryAfterMs instead of holding a worker. Sleeping callers are not served in arrival order, so a short in-process wait is fine for smoothing, and the queue is where ordering and long delays belong.

When the provider says no anyway

Your model of the provider limit is an estimate. Other systems may share the account, and providers adjust limits dynamically. When a 429 comes back with Retry-After, feed it into the limiter so every worker backs off, not only the one that received the response. With GCRA that is a matter of pushing the TAT far enough forward that nothing is admitted until the pause has elapsed.

gcra-penalize.lualua
-- KEYS[1] = limiter key
-- ARGV[1] = emission interval in microseconds
-- ARGV[2] = burst
-- ARGV[3] = pause in milliseconds (from the provider's Retry-After)

local emission = tonumber(ARGV[1])
local burst = tonumber(ARGV[2])
local pause = tonumber(ARGV[3]) * 1000

local t = redis.call("TIME")
local now = tonumber(t[1]) * 1000000 + tonumber(t[2])

-- A request is admitted once now >= tat + emission - emission * burst,
-- so this TAT blocks all admissions until the pause has elapsed.
local blocked_tat = now + pause + emission * burst

local tat = tonumber(redis.call("GET", KEYS[1]))
if tat == nil or tat < blocked_tat then
  local ttl_ms = math.ceil((blocked_tat - now) / 1000)
  redis.call("SET", KEYS[1], string.format("%.0f", blocked_tat), "PX", ttl_ms)
end

return 1

Testing the Limiter

Unit tests against a mocked Redis prove nothing here, because the property that matters is atomicity inside Redis. Run the script against a real instance, in a container in CI, and test properties instead of individual examples.

rate-limiter.spec.tsts
import { randomUUID } from "node:crypto";
import { setTimeout as sleep } from "node:timers/promises";
import Redis from "ioredis";
import { RateLimiter } from "./rate-limiter";

describe("RateLimiter (GCRA)", () => {
  const redis = new Redis(process.env.REDIS_URL ?? "redis://127.0.0.1:6379");
  const limiter = new RateLimiter(redis);

  afterAll(() => redis.quit());

  it("admits exactly the burst under concurrency", async () => {
    // 10 per minute means one refill every 6 seconds: none during the test.
    const policy = { name: "test", limit: 10, periodMs: 60_000, burst: 10 };
    const key = "rl:test:" + randomUUID();

    const decisions = await Promise.all(
      Array.from({ length: 200 }, () => limiter.consume(key, policy)),
    );

    expect(decisions.filter((d) => d.allowed)).toHaveLength(10);
  });

  it("reports a Retry-After that is actually sufficient", async () => {
    const policy = { name: "test", limit: 10, periodMs: 1_000, burst: 2 };
    const key = "rl:test:" + randomUUID();

    await limiter.consume(key, policy);
    await limiter.consume(key, policy);

    const rejected = await limiter.consume(key, policy);
    expect(rejected.allowed).toBe(false);
    expect(rejected.retryAfterMs).toBeGreaterThan(0);
    expect(rejected.retryAfterMs).toBeLessThanOrEqual(100);

    await sleep(rejected.retryAfterMs);
    expect((await limiter.consume(key, policy)).allowed).toBe(true);
  });

  it("does not let rejected calls extend the penalty", async () => {
    const policy = { name: "test", limit: 10, periodMs: 1_000, burst: 1 };
    const key = "rl:test:" + randomUUID();

    await limiter.consume(key, policy);
    const first = await limiter.consume(key, policy);
    for (let i = 0; i < 50; i++) await limiter.consume(key, policy);
    const last = await limiter.consume(key, policy);

    expect(last.retryAfterMs).toBeLessThanOrEqual(first.retryAfterMs);
  });
});
  • Concurrency: N parallel calls from idle admit exactly the burst, never burst plus one.
  • Sustained rate: over any long window, admitted requests do not exceed burst plus elapsed time divided by the emission interval.
  • Honest feedback: waiting for the reported Retry-After is always enough.
  • Failure injection: stop Redis mid-test and assert that open policies admit, closed policies return 503, and the backend-error metric increments.
  • Load test with one deliberately abusive tenant and confirm the latency of the others does not move.

Observability and Rollout

A rate limiter makes decisions customers can feel, so it needs the same visibility as any other customer-facing behaviour. The useful signals are few:

  • Decisions by policy and outcome. Label by policy and plan, not by tenant ID, to keep metric cardinality bounded.
  • Limiter latency (p50, p99) measured at the caller, including the network hop.
  • Fail-open and fail-closed counts, alerted on separately from ordinary throttling.
  • Outbound wait time and requeue count per provider.
  • A sampled structured log line for limited requests, carrying tenant ID, policy and retry-after. This is where per-tenant questions get answered.

New limits should ship in shadow mode first: run the full decision path, record what would have been rejected, and enforce nothing. A few days of that data shows which tenants a proposed number would actually hit, which is a much better basis for a conversation with them than a 429 in production. Then enforce per plan or per tenant behind a flag, lowest risk first.

The ratio of limited to total requests per policy is worth watching over time. A policy that never limits anyone is not protecting anything, and one that limits a steady share of traffic is either set too low or is absorbing a client retry loop that deserves a direct conversation.

What This Design Does and Does Not Give You

  • One Redis round trip and one small key per limit, with atomic decisions across any number of API instances.
  • Exact Retry-After values and standard headers, so clients can back off without guessing.
  • A declared failure mode per policy, with a timeout and a metric behind it.
  • The same primitive for outbound provider quotas, including per-tenant fairness and provider-driven backoff.

It does not replace capacity planning, load shedding or a WAF. Rate limiting enforces fairness between known callers at a rate you chose. It does not know whether the system behind it is healthy right now, and it is the wrong tool for absorbing a volumetric attack. Those need their own layers, in front of and behind this one.

Operational metric to watch

Track the fail-open rate and the limited-request ratio per policy side by side. The first tells you whether the limiter is working at all; the second tells you whether the numbers you chose still match how tenants use the API.