Limitar una API por plan en Rust: GCRA, Redis y tower-rate-tierRate limiting an API by plan in Rust: GCRA, Redis and tower-rate-tier
En esta páginaOn this page
El problema: limitar por plan#
Casi cualquier API de pago acaba teniendo planes: uno gratuito con un límite bajo, uno de pago con un límite más alto y, a veces, uno sin límite. Limitar las peticiones por IP no sirve para eso. Hay que saber quién hace la petición y qué plan tiene, y eso solo se sabe mirando la petición: una clave de API, un token o una consulta a la base de datos.
Y no todas las peticiones pesan lo mismo. Pedir un dato cuesta poco; generar un informe o una exportación cuesta mucho más, así que debería gastar más cuota. Además, en cuanto el servicio corre en varias instancias, cada una con su propio contador en memoria, el límite se multiplica por el número de instancias.
Para eso escribí tower-rate-tier, un middleware de Tower que limita las peticiones según el plan de cada usuario. Funciona con Axum, con Hyper y con cualquier servicio de Tower cuya respuesta se pueda construir a partir de un String (Tonic, de momento, no), y la versión 0.3.0, publicada el 2 de octubre de 2026, añade los límites compartidos a través de Redis. En esta entrada cuento cómo se usa, qué algoritmo hay debajo y cómo se comparte el límite entre instancias, con un servicio de ejemplo y salidas reales.
Un servicio con planes#
Este es el servicio de ejemplo de la entrada, completo. Tiene dos planes, free y pro, tres rutas con distinto coste y guarda los límites en Redis:
// Un servicio con planes. La clave de API decide el plan, Redis guarda
// los límites (los comparten todas las instancias) y cada ruta tiene su
// coste. Uso: cargo run --example server -- <puerto>
use std::sync::Arc;
use axum::http::StatusCode;
use axum::{Router, routing::get};
use tower_rate_tier::{
LimitEvent, OnMissing, Quota, RateTier, RedisStorage, TierIdentity,
TierLimitLayer,
};
// En un servicio real, esto saldría de la base de datos.
fn plan_of(api_key: &str) -> Option<&'static str> {
match api_key {
"key_ana" => Some("free"),
"key_bea" => Some("pro"),
"key_eva" => Some("gold"), // un plan que el limitador no conoce
_ => None,
}
}
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let port = std::env::args().nth(1).unwrap_or("3000".into());
let conn = redis::Client::open("redis://127.0.0.1:6379")?
.get_connection_manager()
.await?;
let tiers = RateTier::builder()
.tier("free", Quota::per_minute(10))
.tier("pro", Quota::per_minute(100))
.default_tier("free")
// Sin clave, o con una clave que no existe: 401.
.on_missing(OnMissing::Deny(StatusCode::UNAUTHORIZED))
.storage(Arc::new(RedisStorage::new(conn)))
.build();
let layer = TierLimitLayer::new(tiers)
.identifier_fn(|headers| {
let key = headers.get("x-api-key")?.to_str().ok()?;
Some(TierIdentity::new(key, plan_of(key)?))
})
.cost_fn(|req| match req.uri.path() {
"/report" => 5,
"/backup" => 50, // más de lo que el plan free da en un minuto
_ => 1,
})
.on_event(|event| {
if let LimitEvent::UnknownTier { user_id, tier, .. } = event {
eprintln!("unknown tier {tier:?} for {user_id}");
}
});
let app = Router::new()
.route("/data", get(|| async { "data\n" }))
.route("/report", get(|| async { "report\n" }))
.route("/backup", get(|| async { "backup\n" }))
.layer(layer);
let addr = format!("127.0.0.1:{port}");
let listener = tokio::net::TcpListener::bind(addr).await?;
axum::serve(listener, app).await?;
Ok(())
}Por partes:
- Los planes se declaran con un nombre y una cuota:
Quota::per_minute(10)son 10 peticiones por minuto.default_tier("free")es el plan de reserva para los casos que veremos más abajo. identifier_fnrecibe las cabeceras y devuelve quién es el usuario y su plan, oNonesi no lo sabe. Aquí la clave de API decide el plan; en un servicio real saldría de la base de datos, y para eso existe el traitTierIdentifier, que es asíncrono.cost_fndecide cuánto cuesta cada petición:/datagasta 1 de la cuota,/reportgasta 5 y/backup, 50.on_eventavisa de lo que el limitador resuelve por su cuenta, como un plan desconocido. Sirve para logs y métricas.
cost_fn se ejecuta dentro del middleware, y eso importa con Axum. El crate también tiene una capa tier_cost(n), pero si se pone en una ruta con .layer(...) mientras el limitador se añade con Router::layer, la capa de la ruta se ejecuta después del limitador: el coste llega tarde y se ignora. Hasta la 0.2, el README y el ejemplo de Axum hacían justo eso, así que todas las peticiones costaban 1 sin que nadie lo notara. La 0.3 lo corrige con cost_fn.
GCRA: por qué no una ventana fija#
La forma más sencilla de limitar es una ventana fija: un contador por minuto que vuelve a cero cuando cambia el minuto. El problema está justo en ese cambio. Este programa lo compara con lo que hace tower-rate-tier, con el mismo límite de 10 por minuto y las mismas 20 peticiones: 10 en el segundo 59 y otras 10 en el 60.
// Ventana fija frente a GCRA, con el mismo límite: 10 por minuto.
// Llegan 10 peticiones en el segundo 59 y otras 10 en el 60.
use std::time::Duration;
use tower_rate_tier::clock::FakeClock;
use tower_rate_tier::{Quota, RateTier};
#[tokio::main]
async fn main() {
let arrivals = [(59, 10), (60, 10)];
// Ventana fija: un contador por minuto que vuelve a 0 al cambiar.
let (mut minute, mut used, mut fixed) = (0, 0, 0);
for (t, count) in arrivals {
if t / 60 != minute {
(minute, used) = (t / 60, 0);
}
let ok = count.min(10 - used);
used += ok;
fixed += ok;
}
// GCRA, con tower-rate-tier y un reloj falso.
let clock = FakeClock::new();
let limiter = RateTier::builder()
.clock(clock.clone())
.tier("free", Quota::per_minute(10))
.build();
let (mut now, mut gcra) = (0, 0);
for (t, count) in arrivals {
clock.advance(Duration::from_secs(t - now));
now = t;
for _ in 0..count {
if limiter.check("ana", "free", 1).await.unwrap().is_ok() {
gcra += 1;
}
}
}
println!("fixed window: {fixed} of 20 allowed in 2 seconds");
println!("GCRA: {gcra} of 20 allowed in 2 seconds");
}Esta es la salida real, en Linux con Rust 1.94:
fixed window: 20 of 20 allowed in 2 seconds
GCRA: 10 of 20 allowed in 2 secondsCon la ventana fija pasan las 20 en dos segundos, el doble del límite, porque las 10 primeras cuentan en un minuto y las 10 siguientes en otro. Es la ráfaga en el borde de la ventana.
tower-rate-tier usa GCRA (Generic Cell Rate Algorithm), un algoritmo que viene del mundo de las redes ATM. En vez de contar peticiones, guarda un solo número por usuario: el TAT (theoretical arrival time), el momento en el que el cubo del usuario volvería a estar lleno. Con 10 peticiones por minuto:
- cada petición mueve el TAT 6 segundos hacia el futuro (60 / 10), multiplicados por su coste;
- la petición se acepta si, después de moverlo, el TAT no queda más de 60 segundos por delante de ahora;
- si no hay TAT guardado, o ya ha pasado, se parte de ahora.
Así se comporta, paso a paso, con un reloj falso:
// GCRA paso a paso: 10 peticiones por minuto, con un reloj falso.
use std::time::Duration;
use tower_rate_tier::clock::FakeClock;
use tower_rate_tier::{Quota, RateTier};
#[tokio::main]
async fn main() {
let clock = FakeClock::new();
let limiter = RateTier::builder()
.clock(clock.clone())
.tier("free", Quota::per_minute(10))
.build();
// (segundos que espera, peticiones que lanza después)
let mut t = 0;
for (wait, count) in [(0, 11), (6, 2), (30, 6)] {
clock.advance(Duration::from_secs(wait));
t += wait;
for _ in 0..count {
match limiter.check("ana", "free", 1).await.unwrap() {
Ok(info) => {
let left = info.remaining;
println!("t={t:>2}s 200 remaining {left}")
}
Err(limited) => {
let wait = limited.retry_after_secs();
println!("t={t:>2}s 429 retry after {wait} s")
}
}
}
}
}Esta es la salida real, en Linux con Rust 1.94:
t= 0s 200 remaining 9
t= 0s 200 remaining 8
t= 0s 200 remaining 7
t= 0s 200 remaining 6
t= 0s 200 remaining 5
t= 0s 200 remaining 4
t= 0s 200 remaining 3
t= 0s 200 remaining 2
t= 0s 200 remaining 1
t= 0s 200 remaining 0
t= 0s 429 retry after 6 s
t= 6s 200 remaining 0
t= 6s 429 retry after 6 s
t=36s 200 remaining 4
t=36s 200 remaining 3
t=36s 200 remaining 2
t=36s 200 remaining 1
t=36s 200 remaining 0
t=36s 429 retry after 6 sLas 10 primeras peticiones llevan el TAT del segundo 0 al 60. La undécima lo llevaría al 66, más de 60 segundos por delante, así que se rechaza: tiene que esperar 6 segundos. En el segundo 6 cabe justo una más, y el TAT pasa al 66. En el segundo 36 el TAT sigue en el 66, solo 30 segundos por delante: quedan 30 segundos libres, y a 6 segundos por petición son 5 peticiones.
Es decir, el usuario puede gastar todo su límite de golpe, pero luego la cuota vuelve poco a poco, al ritmo del plan, y no de una vez al cambiar el minuto.
GCRA guarda un solo número por usuario, el momento en el que su cuota estaría llena otra vez, y con él decide cada petición: no hay contadores que se reinicien ni ráfagas en el cambio de minuto.
El reloj falso, FakeClock, es parte del crate. Con él, los tests no dependen de la hora real ni tienen que esperar: avanzan el reloj 6 segundos y comprueban el resultado al instante, siempre igual.
Varias instancias, un solo límite: Redis#
Con la feature redis, RedisStorage guarda el TAT de cada usuario en Redis en vez de en memoria, así que todas las instancias comparten el mismo límite. Para comprobarlo, arranqué el servicio dos veces, en los puertos 3000 y 3001, contra el mismo Redis, y Ana, con el plan free, va alternando entre las dos:
# Ana (plan free, 10 por minuto) alterna entre las dos instancias.
for i in $(seq 1 11); do
port=$((3000 + i % 2))
curl -s -o /dev/null -w "$port %{http_code}\n" \
-H "x-api-key: key_ana" "localhost:$port/data"
doneEsta es la salida real, en Linux con Redis 7.0:
3001 200
3000 200
3001 200
3000 200
3001 200
3000 200
3001 200
3000 200
3001 200
3000 200
3001 429Diez peticiones aceptadas entre las dos instancias y la undécima rechazada: el límite es uno, no uno por instancia. La respuesta completa de un rechazo es esta:
curl -si -H "x-api-key: key_ana" localhost:3000/dataEsta es la salida real, en Linux con Redis 7.0:
HTTP/1.1 429 Too Many Requests
content-type: application/json
retry-after: 6
x-ratelimit-limit: 10
x-ratelimit-remaining: 0
x-ratelimit-reset: 1790947155
content-length: 61
date: Fri, 02 Oct 2026 13:18:14 GMT
{"error":"rate limit exceeded","tier":"free","retry_after":6}Retry-After dice cuántos segundos esperar, y el cuerpo lo repite en JSON. X-RateLimit-Reset es la hora Unix en la que la cuota estará llena otra vez. Retry-After se redondea hacia arriba. En la 0.2 se redondeaba hacia abajo, y un cliente que esperaba exactamente lo que le decían volvía a recibir un 429.
¿Qué queda en Redis? Una clave por usuario y plan:
# La clave del cubo de Ana: el plan y el SHA-1 del identificador.
key="trt:free:$(printf key_ana | sha1sum | cut -c1-40)"
redis-cli --scan --pattern 'trt:*'
redis-cli PTTL "$key"Esta es la salida real, en Linux con Redis 7.0:
trt:free:33d2315cbeb866b8b237240a98d605f2e6e41b8a
59841La clave es trt:<plan>:<sha1(usuario)>, así que el identificador no se guarda en claro. Si los identificadores tienen poca entropía, como direcciones IP o correos, el SHA-1 se puede revertir probando valores, y para eso existe .key_secret(secreto), que calcula la clave con HMAC-SHA1 y un secreto. El segundo número es PTTL: la clave caduca en unos 60 segundos, justo cuando el cubo de Ana vuelve a estar lleno. No hay que limpiar nada.
Cada comprobación es un script de Lua que Redis ejecuta de forma atómica: mientras corre, ningún otro comando toca la clave, así que dos instancias no pueden leer el mismo TAT y aceptar las dos la última petición. Esta es la parte final de gcra.lua, donde se decide:
-- A key of another type makes GET fail; pcall turns that into an error
-- table, which tonumber() reads as nil. Not an integer (this includes NaN)
-- means a corrupted value too: start fresh.
local stored = redis.pcall("GET", KEYS[1])
if type(stored) == "table" then
stored = nil
end
local tat = tonumber(stored)
local capped = false
if tat == nil or tat ~= math.floor(tat) or tat < now then
tat = now
elseif tat > now + burst_offset then
tat = now + burst_offset
capped = true
end
local increment = emission_interval * cost
if increment > MAX_EXACT - tat then
return fail("time values exceed the exact range of Lua numbers")
end
local new_tat = tat + increment
local allow_at = new_tat - burst_offset
if allow_at > now then
if capped then
redis.call("SET", KEYS[1], tat, "PX", math.ceil((tat - now) / 1000))
end
return { 0, 0, allow_at - now, tat - now }
end
if new_tat > now then
redis.call("SET", KEYS[1], new_tat, "PX", math.ceil((new_tat - now) / 1000))
end
local remaining = math.floor((burst_offset - (new_tat - now)) / emission_interval)
return { 1, remaining, 0, new_tat - now }El script es la misma lógica que el GCRA en Rust, con tres detalles que solo aparecen al llevarlo a Redis:
- El reloj es el de Redis. El script lee la hora con
TIME, así que da igual que los relojes de las instancias no coincidan. - La caducidad se redondea hacia arriba (
math.ceilen elPX). Si se redondeara hacia abajo, la clave podría desaparecer un milisegundo antes de tiempo y regalar peticiones. - Si el reloj de Redis va hacia atrás (por ejemplo, al pasar a una réplica con el reloj atrasado), el TAT se recorta a un cubo lleno y se guarda así. Sin eso, el usuario se quedaría bloqueado todo el tiempo que el reloj retrocedió.
El script trabaja en microsegundos, no en nanosegundos como la parte de Rust. Los números de Lua son doubles, exactos para enteros hasta 253: en microsegundos desde 1970 eso llega hasta el año 2255; en nanosegundos no llegaría ni a cubrir la fecha de hoy.
El script se prueba dos veces: en un Lua 5.1 embebido, la misma versión que usa Redis, sin necesidad de un servidor, y contra un Redis real en CI. Si Redis tarda más de 100 ms en responder, la comprobación cuenta como un error de almacenamiento, del que hablo más abajo.
Peticiones que cuestan más#
Bea tiene el plan pro, 100 por minuto. Pide un dato y luego un informe:
# Bea (plan pro, 100 por minuto): /data cuesta 1 y /report, 5.
for path in data report; do
curl -s -o /dev/null -H "x-api-key: key_bea" \
-w "/$path %{http_code} remaining %header{x-ratelimit-remaining}\n" \
"localhost:3000/$path"
doneEsta es la salida real, en Linux con Redis 7.0:
/data 200 remaining 99
/report 200 remaining 94/data gasta 1 y deja 99; /report gasta 5 y deja 94. Para GCRA, una petición de coste 5 es como cinco peticiones a la vez: mueve el TAT 5 intervalos de golpe.
Valores por defecto seguros#
Lo más delicado de un limitador no es el algoritmo, sino lo que hace cuando algo no encaja. Eva tiene una clave válida, pero su plan, gold, no está configurado:
# Eva tiene un plan que el limitador no conoce.
curl -s -o /dev/null -H "x-api-key: key_eva" \
-w "/data %{http_code} limit %header{x-ratelimit-limit}\n" \
localhost:3000/data
curl -s -o /dev/null -H "x-api-key: key_eva" \
-w "/backup %{http_code}\n" localhost:3000/backup
# Sin clave de API.
curl -s -o /dev/null -w "/data %{http_code}\n" localhost:3000/dataEsta es la salida real, en Linux con Redis 7.0:
/data 200 limit 10
/backup 403
/data 401Y esto es lo que escribió el servicio por la salida de error, desde on_event:
unknown tier "gold" for key_eva
unknown tier "gold" for key_evaTres casos, tres decisiones:
- Un plan desconocido no da acceso ilimitado. Eva recibe la cuota del plan por defecto (
limit 10) en su propio cubo, yon_eventavisa de cada petición. Un plan desconocido puede ser una errata, un plan nuevo que el limitador aún no conoce o, peor, un valor que el cliente controla. Hasta la 0.2, una petición con un plan desconocido pasaba sin ningún límite; en la 0.3 está corregido y anotado como fallo de seguridad en el changelog. Si no hay plan por defecto, la respuesta es403, y también se puede elegirDenyo unAllowexplícito conOnUnknownTier. - Una petición que nunca podría pasar no recibe un 429.
/backupcuesta 50 y el planfreeda 10 por minuto: ningún tiempo de espera la haría posible. Por eso responde403 Forbidden, sinRetry-Aftery sin tocar Redis. - Sin identificar, se decide explícitamente. Aquí, con
OnMissing::Denyy el estadoUNAUTHORIZED, una petición sin clave recibe401. Por defecto (OnMissing::UseDefault) recibiría el plan por defecto.
Si Redis falla o tarda más de 100 ms, por defecto la petición pasa (OnStorageError::Allow): es mejor no limitar durante una caída que tumbar la API entera. Si prefieres lo contrario, OnStorageError::Deny responde 503. En los dos casos, on_event recibe un LimitEvent::StorageError, así que una caída de Redis no pasa en silencio.
Lo que viene#
La hoja de ruta de la 0.4 se centra en la operación:
- cambiar los planes en caliente, sin reiniciar el servicio;
- métricas compatibles con Prometheus;
- soporte para Tonic y gRPC;
- seguir limitando en memoria local mientras Redis está caído, y un circuit breaker para que un Redis caído no cueste a cada petición sus 100 ms de espera.
Si alguna de estas versiones trae algo que merezca la pena contar, lo contaré aquí.
Lo que me llevo#
- Limitar por plan es, sobre todo, identificar bien. El algoritmo es la parte fácil; saber quién pide y con qué plan, y qué hacer cuando no se sabe, es lo que decide si el límite sirve.
- GCRA hace mucho con muy poco. Con un solo número por usuario reparte la cuota al ritmo del plan y hasta dice cuánto esperar.
- Con varias instancias, el estado y el reloj tienen que ser compartidos. Un script atómico en Redis y la hora de Redis resuelven las dos cosas a la vez.
- Los valores por defecto son decisiones de seguridad. Un plan desconocido, una petición sin identificar o un Redis caído necesitan una respuesta pensada, no la que salga por accidente.
- Un reloj inyectable hace que el tiempo se pueda testear. Los tests no esperan a que pase un minuto: mueven el reloj y comprueban.
tower-rate-tier es un crate mío, con licencia MIT o Apache 2.0. La 0.1.0 salió el 11 de marzo de 2026 y la 0.3.0, el 2 de octubre de 2026. Todo lo de esta entrada es la 0.3.0 descargada de crates.io: el servicio de ejemplo y las salidas se ejecutaron en Linux, con Rust 1.94 y Redis 7.0.
The problem: limiting by plan#
Almost every paid API ends up with plans: a free one with a low limit, a paid one with a higher limit and, sometimes, one with no limit at all. Limiting requests by IP doesn't help with that. You need to know who is making the request and which plan they're on, and you only learn that by looking at the request: an API key, a token or a database lookup.
And not every request weighs the same. Fetching a piece of data is cheap; generating a report or an export costs much more, so it should use up more quota. On top of that, as soon as the service runs on several instances, each with its own in-memory counter, the limit gets multiplied by the number of instances.
That's what I wrote tower-rate-tier for: a Tower middleware that limits requests according to each user's plan. It works with Axum, with Hyper and with any Tower service whose response can be built from a String (not Tonic, for now), and version 0.3.0, released on 2 October 2026, adds limits shared through Redis. In this post I cover how to use it, the algorithm underneath, and how the limit is shared between instances, with an example service and real outputs.
A service with plans#
This is the post's example service, in full. It has two plans, free and pro, three routes with different costs, and it keeps the limits in Redis:
// A service with plans. The API key decides the plan, Redis keeps
// the limits (every instance shares them) and each route has its
// cost. Usage: cargo run --example server -- <port>
use std::sync::Arc;
use axum::http::StatusCode;
use axum::{Router, routing::get};
use tower_rate_tier::{
LimitEvent, OnMissing, Quota, RateTier, RedisStorage, TierIdentity,
TierLimitLayer,
};
// In a real service, this would come from the database.
fn plan_of(api_key: &str) -> Option<&'static str> {
match api_key {
"key_ana" => Some("free"),
"key_bea" => Some("pro"),
"key_eva" => Some("gold"), // a plan the limiter doesn't know
_ => None,
}
}
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let port = std::env::args().nth(1).unwrap_or("3000".into());
let conn = redis::Client::open("redis://127.0.0.1:6379")?
.get_connection_manager()
.await?;
let tiers = RateTier::builder()
.tier("free", Quota::per_minute(10))
.tier("pro", Quota::per_minute(100))
.default_tier("free")
// No key, or a key that doesn't exist: 401.
.on_missing(OnMissing::Deny(StatusCode::UNAUTHORIZED))
.storage(Arc::new(RedisStorage::new(conn)))
.build();
let layer = TierLimitLayer::new(tiers)
.identifier_fn(|headers| {
let key = headers.get("x-api-key")?.to_str().ok()?;
Some(TierIdentity::new(key, plan_of(key)?))
})
.cost_fn(|req| match req.uri.path() {
"/report" => 5,
"/backup" => 50, // more than the free plan gives in a minute
_ => 1,
})
.on_event(|event| {
if let LimitEvent::UnknownTier { user_id, tier, .. } = event {
eprintln!("unknown tier {tier:?} for {user_id}");
}
});
let app = Router::new()
.route("/data", get(|| async { "data\n" }))
.route("/report", get(|| async { "report\n" }))
.route("/backup", get(|| async { "backup\n" }))
.layer(layer);
let addr = format!("127.0.0.1:{port}");
let listener = tokio::net::TcpListener::bind(addr).await?;
axum::serve(listener, app).await?;
Ok(())
}Piece by piece:
- Plans are declared with a name and a quota:
Quota::per_minute(10)means 10 requests per minute.default_tier("free")is the fallback plan for the cases we'll see further down. identifier_fnreceives the headers and returns who the user is and their plan, orNoneif it can't tell. Here the API key decides the plan; in a real service it would come from the database, and that's what theTierIdentifiertrait is for, since it's asynchronous.cost_fndecides what each request costs:/datauses 1 from the quota,/reportuses 5 and/backup, 50.on_eventreports what the limiter resolves on its own, such as an unknown plan. It's there for logs and metrics.
cost_fn runs inside the middleware, and that matters with Axum. The crate also has a tier_cost(n) layer, but if you put it on a route with .layer(...) while the limiter is added with Router::layer, the route's layer runs after the limiter: the cost arrives too late and is ignored. Up to 0.2, the README and the Axum example did exactly that, so every request cost 1 without anyone noticing. 0.3 fixes it with cost_fn.
GCRA: why not a fixed window#
The simplest way to limit is a fixed window: a counter per minute that goes back to zero when the minute changes. The trouble is right at that change. This program compares it with what tower-rate-tier does, with the same limit of 10 per minute and the same 20 requests: 10 at second 59 and another 10 at second 60.
// Fixed window versus GCRA, with the same limit: 10 per minute.
// 10 requests arrive at second 59 and another 10 at second 60.
use std::time::Duration;
use tower_rate_tier::clock::FakeClock;
use tower_rate_tier::{Quota, RateTier};
#[tokio::main]
async fn main() {
let arrivals = [(59, 10), (60, 10)];
// Fixed window: a counter per minute that resets when it changes.
let (mut minute, mut used, mut fixed) = (0, 0, 0);
for (t, count) in arrivals {
if t / 60 != minute {
(minute, used) = (t / 60, 0);
}
let ok = count.min(10 - used);
used += ok;
fixed += ok;
}
// GCRA, with tower-rate-tier and a fake clock.
let clock = FakeClock::new();
let limiter = RateTier::builder()
.clock(clock.clone())
.tier("free", Quota::per_minute(10))
.build();
let (mut now, mut gcra) = (0, 0);
for (t, count) in arrivals {
clock.advance(Duration::from_secs(t - now));
now = t;
for _ in 0..count {
if limiter.check("ana", "free", 1).await.unwrap().is_ok() {
gcra += 1;
}
}
}
println!("fixed window: {fixed} of 20 allowed in 2 seconds");
println!("GCRA: {gcra} of 20 allowed in 2 seconds");
}This is the real output, on Linux with Rust 1.94:
fixed window: 20 of 20 allowed in 2 seconds
GCRA: 10 of 20 allowed in 2 secondsWith the fixed window all 20 get through in two seconds, twice the limit, because the first 10 count towards one minute and the next 10 towards another. That's the burst at the window boundary.
tower-rate-tier uses GCRA (Generic Cell Rate Algorithm), an algorithm that comes from ATM networking. Instead of counting requests, it keeps a single number per user: the TAT (theoretical arrival time), the moment the user's bucket would be full again. With 10 requests per minute:
- each request moves the TAT 6 seconds into the future (60 / 10), multiplied by its cost;
- the request is allowed if, after moving it, the TAT is no more than 60 seconds ahead of now;
- if there's no stored TAT, or it's already in the past, it starts from now.
This is how it behaves, step by step, with a fake clock:
// GCRA step by step: 10 requests per minute, with a fake clock.
use std::time::Duration;
use tower_rate_tier::clock::FakeClock;
use tower_rate_tier::{Quota, RateTier};
#[tokio::main]
async fn main() {
let clock = FakeClock::new();
let limiter = RateTier::builder()
.clock(clock.clone())
.tier("free", Quota::per_minute(10))
.build();
// (seconds it waits, requests it sends afterwards)
let mut t = 0;
for (wait, count) in [(0, 11), (6, 2), (30, 6)] {
clock.advance(Duration::from_secs(wait));
t += wait;
for _ in 0..count {
match limiter.check("ana", "free", 1).await.unwrap() {
Ok(info) => {
let left = info.remaining;
println!("t={t:>2}s 200 remaining {left}")
}
Err(limited) => {
let wait = limited.retry_after_secs();
println!("t={t:>2}s 429 retry after {wait} s")
}
}
}
}
}This is the real output, on Linux with Rust 1.94:
t= 0s 200 remaining 9
t= 0s 200 remaining 8
t= 0s 200 remaining 7
t= 0s 200 remaining 6
t= 0s 200 remaining 5
t= 0s 200 remaining 4
t= 0s 200 remaining 3
t= 0s 200 remaining 2
t= 0s 200 remaining 1
t= 0s 200 remaining 0
t= 0s 429 retry after 6 s
t= 6s 200 remaining 0
t= 6s 429 retry after 6 s
t=36s 200 remaining 4
t=36s 200 remaining 3
t=36s 200 remaining 2
t=36s 200 remaining 1
t=36s 200 remaining 0
t=36s 429 retry after 6 sThe first 10 requests take the TAT from second 0 to 60. The eleventh would take it to 66, more than 60 seconds ahead, so it's rejected: it has to wait 6 seconds. At second 6 exactly one more fits, and the TAT moves to 66. At second 36 the TAT is still at 66, only 30 seconds ahead: that leaves 30 free seconds, and at 6 seconds per request that's 5 requests.
In other words, a user can spend the whole limit at once, but then the quota comes back little by little, at the plan's rate, rather than all at once when the minute changes.
GCRA keeps a single number per user, the moment their quota would be full again, and uses it to decide every request: no counters that reset and no bursts at the minute boundary.
The fake clock, FakeClock, is part of the crate. With it, tests don't depend on the real time and don't have to wait: they move the clock forward 6 seconds and check the result instantly, the same every time.
Several instances, one limit: Redis#
With the redis feature, RedisStorage keeps each user's TAT in Redis instead of in memory, so every instance shares the same limit. To check it, I started the service twice, on ports 3000 and 3001, against the same Redis, and Ana, on the free plan, alternates between them:
# Ana (free plan, 10 per minute) alternates between both instances.
for i in $(seq 1 11); do
port=$((3000 + i % 2))
curl -s -o /dev/null -w "$port %{http_code}\n" \
-H "x-api-key: key_ana" "localhost:$port/data"
doneThis is the real output, on Linux with Redis 7.0:
3001 200
3000 200
3001 200
3000 200
3001 200
3000 200
3001 200
3000 200
3001 200
3000 200
3001 429Ten requests allowed across both instances and the eleventh rejected: there's one limit, not one per instance. This is the full response to a rejection:
curl -si -H "x-api-key: key_ana" localhost:3000/dataThis is the real output, on Linux with Redis 7.0:
HTTP/1.1 429 Too Many Requests
content-type: application/json
retry-after: 6
x-ratelimit-limit: 10
x-ratelimit-remaining: 0
x-ratelimit-reset: 1790947155
content-length: 61
date: Fri, 02 Oct 2026 13:18:14 GMT
{"error":"rate limit exceeded","tier":"free","retry_after":6}Retry-After says how many seconds to wait, and the body repeats it in JSON. X-RateLimit-Reset is the Unix time at which the quota will be full again. Retry-After is rounded up. In 0.2 it was rounded down, and a client that waited exactly as long as it was told got another 429.
What's left in Redis? One key per user and plan:
# Ana's bucket key: the plan and the SHA-1 of the user id.
key="trt:free:$(printf key_ana | sha1sum | cut -c1-40)"
redis-cli --scan --pattern 'trt:*'
redis-cli PTTL "$key"This is the real output, on Linux with Redis 7.0:
trt:free:33d2315cbeb866b8b237240a98d605f2e6e41b8a
59841The key is trt:<plan>:<sha1(user)>, so the identifier isn't stored in the clear. If identifiers have little entropy, like IP addresses or emails, the SHA-1 can be reversed by trying values, and that's what .key_secret(secret) is for: it computes the key with HMAC-SHA1 and a secret. The second number is PTTL: the key expires in about 60 seconds, exactly when Ana's bucket is full again. There's nothing to clean up.
Every check is a Lua script that Redis runs atomically: while it runs, no other command touches the key, so two instances can't read the same TAT and both allow the last request. This is the last part of gcra.lua, where the decision is made:
-- A key of another type makes GET fail; pcall turns that into an error
-- table, which tonumber() reads as nil. Not an integer (this includes NaN)
-- means a corrupted value too: start fresh.
local stored = redis.pcall("GET", KEYS[1])
if type(stored) == "table" then
stored = nil
end
local tat = tonumber(stored)
local capped = false
if tat == nil or tat ~= math.floor(tat) or tat < now then
tat = now
elseif tat > now + burst_offset then
tat = now + burst_offset
capped = true
end
local increment = emission_interval * cost
if increment > MAX_EXACT - tat then
return fail("time values exceed the exact range of Lua numbers")
end
local new_tat = tat + increment
local allow_at = new_tat - burst_offset
if allow_at > now then
if capped then
redis.call("SET", KEYS[1], tat, "PX", math.ceil((tat - now) / 1000))
end
return { 0, 0, allow_at - now, tat - now }
end
if new_tat > now then
redis.call("SET", KEYS[1], new_tat, "PX", math.ceil((new_tat - now) / 1000))
end
local remaining = math.floor((burst_offset - (new_tat - now)) / emission_interval)
return { 1, remaining, 0, new_tat - now }The script is the same logic as the GCRA in Rust, with three details that only show up when you move it to Redis:
- The clock is Redis's. The script reads the time with
TIME, so it doesn't matter that the instances' clocks disagree. - The expiry is rounded up (
math.ceilin thePX). If it were rounded down, the key could vanish a millisecond early and give away requests. - If Redis's clock goes backwards (for example, on failover to a replica whose clock is behind), the TAT is capped at a full bucket and stored that way. Without that, the user would stay locked out for as long as the clock went back.
The script works in microseconds, not nanoseconds like the Rust side. Lua numbers are doubles, exact for integers up to 253: in microseconds since 1970 that reaches the year 2255; in nanoseconds it wouldn't even cover today's date.
The script is tested twice: in an embedded Lua 5.1, the same version Redis uses, with no server needed, and against a real Redis in CI. If Redis takes longer than 100 ms to answer, the check counts as a storage error, which I cover further down.
Requests that cost more#
Bea is on the pro plan, 100 per minute, and fetches some data and then a report:
# Bea (pro plan, 100 per minute): /data costs 1 and /report, 5.
for path in data report; do
curl -s -o /dev/null -H "x-api-key: key_bea" \
-w "/$path %{http_code} remaining %header{x-ratelimit-remaining}\n" \
"localhost:3000/$path"
doneThis is the real output, on Linux with Redis 7.0:
/data 200 remaining 99
/report 200 remaining 94/data uses 1 and leaves 99; /report uses 5 and leaves 94. For GCRA, a request costing 5 is like five requests at once: it moves the TAT 5 intervals in one go.
Safe defaults#
The trickiest part of a limiter isn't the algorithm, it's what it does when something doesn't fit. Eva has a valid key, but Eva's plan, gold, isn't configured:
# Eva has a plan the limiter doesn't know.
curl -s -o /dev/null -H "x-api-key: key_eva" \
-w "/data %{http_code} limit %header{x-ratelimit-limit}\n" \
localhost:3000/data
curl -s -o /dev/null -H "x-api-key: key_eva" \
-w "/backup %{http_code}\n" localhost:3000/backup
# No API key.
curl -s -o /dev/null -w "/data %{http_code}\n" localhost:3000/dataThis is the real output, on Linux with Redis 7.0:
/data 200 limit 10
/backup 403
/data 401And this is what the service wrote to standard error, from on_event:
unknown tier "gold" for key_eva
unknown tier "gold" for key_evaThree cases, three decisions:
- An unknown plan doesn't grant unlimited access. Eva gets the default plan's quota (
limit 10) in Eva's own bucket, andon_eventreports every request. An unknown plan might be a typo, a new plan the limiter doesn't know yet or, worse, a value the client controls. Up to 0.2, a request with an unknown plan went through with no limit at all; 0.3 fixes it and the changelog lists it as a security fix. If there's no default plan, the answer is403, andOnUnknownTiercan also chooseDenyor an explicitAllow. - A request that could never succeed doesn't get a 429.
/backupcosts 50 and thefreeplan gives 10 per minute: no amount of waiting would make it possible. So it answers403 Forbidden, with noRetry-Afterand without touching Redis. - Unidentified requests get an explicit decision. Here, with
OnMissing::Denyand theUNAUTHORIZEDstatus, a request with no key gets401. By default (OnMissing::UseDefault) it would get the default plan.
If Redis fails or takes longer than 100 ms, by default the request goes through (OnStorageError::Allow): better not to limit during an outage than to take the whole API down. If you prefer the opposite, OnStorageError::Deny answers 503. Either way, on_event receives a LimitEvent::StorageError, so a Redis outage doesn't go unnoticed.
What's next#
The 0.4 roadmap focuses on operations:
- changing plans on the fly, without restarting the service;
- Prometheus-compatible metrics;
- support for Tonic and gRPC;
- keeping the limit in local memory while Redis is down, and a circuit breaker so a down Redis doesn't cost every request its 100 ms timeout.
If any of these releases brings something worth telling, I'll tell it here.
What I take away#
- Limiting by plan is mostly about identifying well. The algorithm is the easy part; knowing who's asking and on which plan, and what to do when you don't know, is what decides whether the limit is any use.
- GCRA does a lot with very little. With a single number per user it spreads the quota at the plan's rate and even tells you how long to wait.
- With several instances, both the state and the clock have to be shared. An atomic script in Redis and Redis's own time solve both at once.
- Defaults are security decisions. An unknown plan, an unidentified request or a Redis outage need a deliberate answer, not whatever happens by accident.
- An injectable clock makes time testable. Tests don't wait for a minute to pass: they move the clock and check.
tower-rate-tier is a crate of mine, licensed under MIT or Apache 2.0. 0.1.0 came out on 11 March 2026 and 0.3.0 on 2 October 2026. Everything in this post is 0.3.0 as downloaded from crates.io: the example service and the outputs were run on Linux, with Rust 1.94 and Redis 7.0.
ReferenciasReferences
- tower-rate-tier on crates.iocrates.io · crates.io
- tower-rate-tier documentationdocs.rs · docs.rs
- tower-rate-tier: source codeGitHub · github.com
- tower-rate-tier: changelogGitHub · github.com
- tower-rate-tier: roadmapGitHub · github.com
- Generic cell rate algorithmWikipedia · en.wikipedia.org
- Scripting with LuaRedis · redis.io
- TIMERedis · redis.io
- RFC 6585: Additional HTTP Status Codes (429 Too Many Requests)IETF · rfc-editor.org
- RFC 9110: HTTP Semantics (Retry-After)IETF · rfc-editor.org
CompartirShare
¿Necesitas desarrollo en Rust?Need Rust development?
Desarrollo software seguro y de alto rendimiento con Rust para tu empresa. Consultoría, formación y proyectos a medida.Secure, high-performance software development with Rust for your business. Consulting, training, and custom projects.
HablemosLet's talk→
ComentariosComments
Para comentar necesitas una cuenta de GitHub. Los comentarios se guardan en GitHub Discussions (SoftDryzz/blog-comments) y se moderan: no se publican enlaces promocionales.Commenting requires a GitHub account. Comments live in GitHub Discussions (SoftDryzz/blog-comments) and are moderated: no promotional links.