Yes, the right design principles are:
- to push as much adversarial logic to the edge, out of the app core
- to make it scale independently of the app itself
- to make it scale horizontally, so you can dynamically add more rate limiter instances under peak load
- to adopt a stateless model if possible: auth token checks and rate limits should avoid db accesses
- to make to async and fuzzy: trade 1-2 requests over quota slipping in occasionally to eliminate a sync path
Rate limiting inside application code is a classic trap. By the time your Node, Python, or Go process parses the incoming HTTP request, runs middleware, and executes a Redis check to reject a bot, the attacker has already consumed your application's CPU and memory allocations.