Rate Limits
Per-endpoint request throttling with distributed sliding-window rate limiting.
Overview
GaaS enforces rate limits to prevent abuse, ensure fair resource allocation, and maintain system stability. Rate limits are applied per API key (organization) and vary by endpoint.
When you exceed a rate limit, the API returns an HTTP 429 Too Many Requests response with a Retry-After header indicating when you can retry.
GAAS_DATABASE_URL is set), GaaS uses Postgres-backed distributed rate limiting with atomic sliding-window counters. This ensures accurate rate limiting across multiple API server instances.
Per-Endpoint Rate Limits
Different endpoints have different rate limits based on their resource intensity:
| Endpoint Category | Limit | Window | Examples |
|---|---|---|---|
| Intent Submission | 120 requests | Per minute | POST /v1/intentsPOST /v1/intents/batch |
| Backtest | 5 requests | Per minute | POST /v1/learning/backtest |
| Onboarding | 10 requests | Per minute | POST /v1/onboarding/quickstartPOST /v1/onboarding/intake |
| Authentication (IP-based) | 5 requests | Per minute | Dashboard login, signup, MFA verification |
| General Endpoints | 120 requests | Per minute | All other API endpoints (decisions, escalations, audit, webhooks, etc.) |
429 Response Format
When you exceed a rate limit, the API returns a 429 Too Many Requests response:
POST https://api.gaas.is/v1/intents
X-API-Key: your_api_key
# Response (when rate limit exceeded):
HTTP/1.1 429 Too Many Requests
Retry-After: 42
RateLimit-Limit: 120
RateLimit-Remaining: 0
RateLimit-Reset: 1707824460
Content-Type: application/json
{
"error": {
"code": "rate_limit_exceeded",
"message": "Rate limit exceeded. Please try again later.",
"details": [{"retry_after_seconds": 42}],
"request_id": "9f1c2e7a4b3d4c21a0e5f6b7c8d9e0f1"
}
}
Rate Limit Headers
Every response from a rate-limited route includes these headers (IETF draft names, no X- prefix):
RateLimit-Limit— Maximum requests allowed in the windowRateLimit-Remaining— Requests remaining in current windowRateLimit-Reset— Unix timestamp when the rate limit resets
On 429 responses, the Retry-After header and error.details[0].retry_after_seconds give the number of seconds to wait before retrying. The SDKs raise their base error type (GaaSError in Python and TypeScript) with status_code / statusCode 429 and the same details.
Rate Limit Scoping
Per API Key
Rate limits are counted per API key: every request that carries the same X-API-Key counts toward the same limit, regardless of which agent or service sends it. Each key has its own budget, so separate keys for separate services do not share a limit. Requests without an API key are counted per client IP address.
Per IP Address
Authentication endpoints (login, signup, MFA verification) are rate-limited per IP address (not per API key) to prevent brute-force attacks. This applies to the dashboard authentication routes only.
Handling Rate Limits
Python (gaas_sdk)
import asyncio
from gaas_sdk import GaaSClient
from gaas_sdk.exceptions import GaaSError
async with GaaSClient("https://api.gaas.is", headers={"X-API-Key": api_key}) as client:
try:
response = await client.submit_intent(intent)
except GaaSError as e:
if e.status_code != 429:
raise
retry_after = (e.details or [{}])[0].get("retry_after_seconds", 1)
print(f"Rate limited. Retrying after {retry_after} seconds...")
await asyncio.sleep(retry_after)
response = await client.submit_intent(intent) # Retry
TypeScript (@governancehq/sdk)
import { GaaSClient, GaaSError } from '@governancehq/sdk';
const client = new GaaSClient({ baseUrl: 'https://api.gaas.is', headers: { 'X-API-Key': apiKey } });
try {
const response = await client.submitIntent(intent);
} catch (error) {
if (error instanceof GaaSError && error.statusCode === 429) {
const retryAfter = Number(error.details[0]?.retry_after_seconds ?? 1);
console.log(`Rate limited. Retrying after ${retryAfter}s...`);
await new Promise(resolve => setTimeout(resolve, retryAfter * 1000));
const response = await client.submitIntent(intent); // Retry
} else {
throw error;
}
}
Java (is.gaas.sdk)
The Java SDK installs from https://maven.gaas.is (see SDKs). A 429 arrives as GaaSRateLimitException, carrying the wait the API sent:
try {
handle(client.submitIntent(intent).getData(), intent);
} catch (GaaSRateLimitException e) {
Integer wait = e.getRetryAfterSeconds(); // the API sends it; null only if it did not
System.out.println("Rate limited. Retrying after " + wait + " seconds...");
Thread.sleep((wait != null ? wait : 1) * 1000L);
handle(client.submitIntent(intent).getData(), intent); // retry once
}
Best Practices
1. Implement Exponential Backoff
When a request is rate-limited, wait for the duration specified in Retry-After, then retry. If you hit the limit again, double the wait time with each subsequent retry:
async def submit_with_backoff(client, intent, max_retries=3):
for attempt in range(max_retries):
try:
return await client.submit_intent(intent)
except GaaSError as e:
if e.status_code != 429 or attempt == max_retries - 1:
raise # Not a rate limit, or final attempt failed
retry_after = (e.details or [{}])[0].get("retry_after_seconds", 1)
await asyncio.sleep(retry_after * (2 ** attempt))
2. Batch Requests Where Possible
Use the bulk intent submission endpoint (POST /v1/intents/batch) to submit up to 50 intents in a single request. This counts as 1 request toward the rate limit instead of 50.
3. Monitor Rate Limit Headers
Check RateLimit-Remaining on every response. If it's low (e.g., <10), slow down your request rate proactively to avoid hitting the limit.
# With a raw HTTP client (for example httpx):
remaining = int(response.headers['RateLimit-Remaining'])
if remaining < 10:
print("Approaching rate limit. Slowing down requests...")
time.sleep(5) # Pause before next request
4. Use Shadow Mode for Load Testing
Shadow mode intent submissions (?mode=shadow) are still subject to rate limits. When load testing, ensure your test traffic doesn't exceed the 120 req/min intent submission limit. Consider using multiple API keys to distribute load.
5. Cache Read-Only Data
Responses for read-only endpoints (e.g., GET /v1/intents/{intent_id}/decision) include ETag headers. Cache these responses locally and use If-None-Match headers to get 304 Not Modified responses, which don't count toward rate limits.
Upgrading Rate Limits
Rate limits are not tied to pricing tiers—all organizations (including Developer tier) have the same rate limits. This is intentional to ensure fair access.
If you need higher rate limits for high-volume production use cases, contact sales@gaas.is to discuss custom Enterprise rate limits. Enterprise plans can negotiate:
- Intent submission limits above the default 120 req/min
- Dedicated API endpoints with isolated rate limits
- Burst allowances (temporary rate limit increases)
Troubleshooting
Getting 429 Errors Despite Low Traffic
Cause: Multiple services or agents sharing the same API key are collectively exceeding the limit.
Solution: Monitor RateLimit-Remaining headers. Consider using separate API keys for different services (requires separate organizations).
Rate Limit Not Resetting
Cause: Sliding window means the limit resets gradually as old requests age out of the 1-minute window, not all at once.
Solution: Wait at least 60 seconds from your last successful request before retrying at full speed.
Auth Endpoints Blocked by IP Rate Limit
Cause: Dashboard authentication endpoints (login, signup, MFA) are limited to 5 req/min per IP address.
Solution: If multiple users share a single outbound IP (e.g., corporate NAT), requests may be rate-limited collectively. Contact support for IP allowlisting (Enterprise only).
Related Pages
- Billing & Quotas — Monthly action limits
- Advanced Features — Bulk submission to reduce rate limit impact
- Authentication — API key setup
- API Reference — Complete endpoint documentation