Skip to content

OmnAPI Rate Limits

Reference

OmnAPI applies a shared customer allowance across all of your API keys. Optional key sublimits let you control individual applications. A separate IP limit protects the service from excessive traffic from a single source.

TierOperations / secondTask-status queries / secondRequests / UTC dayRequests / UTC month
FREE555,000100,000
PREMIUM25100No default hard capNo default hard cap
ENTERPRISE200300No default hard capNo default hard cap

Only GET /api/v1/tasks/{id} uses the task-status query allowance. Task creation, lists, opening task event streams, Suno queries, and other endpoints use the operations allowance. Polling task status does not consume operation-rate capacity. Both classes count toward any configured daily or monthly allowance.

Rate capacity replenishes continuously. You can send up to one second’s allowance at once, then continue as capacity replenishes. Requests rejected by rate limiting do not consume the daily or monthly allowance. Requests that pass the limit checks but later fail validation or processing still count as received requests.

Daily and monthly allowances reset at UTC calendar boundaries. Paid tiers continue to record usage even when there is no daily or monthly hard cap. Customer-specific settings and preserved existing allowances may differ from the defaults above.

When managing API keys, rateLimit is a requests-per-second sublimit across both request classes. dailyLimit and monthlyLimit set optional key sublimits. A null value removes the key override; 0 disables that key sublimit. Neither value removes the shared customer allowance. When updating a key, omitted fields keep their current values; send null to clear an existing setting.

The default IP allowance is 60,000 requests per 60 seconds, shared by requests from the same public IP. This protection applies in addition to customer allowances. GET /api/v1/pricing/catalog does not require an API key but still uses IP protection.

Signed live media URLs on stream.omnapi.com have separate connection limits:

ScopeDefault simultaneous requests
Playback ticket4
Clip, across tickets and listeners4
Public source IP, across clips and listeners16
Shared service capacity1,000

All four checks apply. These limits count active media requests, not opens per minute or total listens per Generation. Each new GET occupies one slot until the request ends or is cancelled. Cancel the previous media request when switching candidates. Shared networks, including carrier NAT, share the IP limit. Obtaining a fresh URL does not reset the clip or IP count. Completed asset URLs are separate from this live playback limit.

A playback capacity rejection returns HTTP 429 with a media-specific JSON body:

{
"error": "playback-capacity-exceeded",
"allowed": false,
"scope": "clip",
"active": 4,
"limit": 4,
"retryAfter": 3
}
HeaderMeaning
Retry-AfterMinimum seconds before retrying; currently 3 for playback capacity
X-RateLimit-ScopeThe blocking scope: ticket, clip, ip, or global
X-RateLimit-Typeconcurrency
X-RateLimit-LimitThe active-request limit for that scope
X-RateLimit-Remaining0 on a capacity rejection
X-Request-IdReference to include when reporting the failed playback request

Active connections may continue beyond the retry delay, so it does not promise capacity will be available then. Use bounded, increasing backoff and re-query the Generation for its current playback URL. Temporary capacity-service failures use HTTP 503 with playback-capacity-unavailable. If a native player does not expose the HTTP response, refresh the Generation after a bounded delay instead of polling the media URL. Provide the request ID and failure time to support; exclude API keys and complete signed URLs from diagnostics.

In addition to request rate limits, each tier has a cap on non-terminal tasks. These caps prevent a single account or key from filling the worker queues with long-running generations.

TierAccount in-flight capPer-key in-flight cap
FREE2020
PREMIUM100100
ENTERPRISE1,0001,000

Queued and actively executing tasks count toward these caps. Completed, failed and cancelled tasks release capacity. A Suno task waiting for a submitted result stops occupying this allowance 15 minutes after creation. This exception applies only to the result-waiting stage; queued tasks and tasks still executing continue to count. Public task status may show PROCESSING for either stage. Increasing the allowance does not guarantee that all tasks execute simultaneously.

An admission-cap rejection also returns HTTP 429 with error.code: "RATE_LIMITED". Its details contains inFlight, limit, and tier; a per-key rejection also includes scope: "apiKey". This spelling differs from the request-rate scope api-key. A recovery time is not known, so Retry-After may be absent. Wait for existing tasks to finish, use bounded backoff, and inspect the body rather than treating the last request-rate headers as an in-flight capacity counter.

New website accounts using only a claimed trial can have one in-flight task and create five tasks per UTC day, in addition to the request limits above. Trial access is limited to eligible text/lyrics, subtitle, and Suno simple/custom song-generation operations; other features require purchased credits.

Trial usage counts gross credit deductions against the granted allowance; refunds do not restore that usage allowance. Unclaimed trials, unsupported features or an exhausted usage allowance return 403 FORBIDDEN. Trial task capacity returns 429 RATE_LIMITED: wait for the active task to finish or the next UTC day, as applicable. Existing accounts and accounts with purchased credits keep their ordinary task-capacity rules. See Credits.

Responses that reach the customer/key limit check describe its applicable allowance. Earlier rejections may report IP protection; headers can be absent when the check does not run. These headers also apply when opening an event stream. When several allowances apply, a successful response reports the enabled allowance with the smallest remaining fraction. An all-unlimited customer can report a limit of 0. Unauthenticated responses describe IP protection.

HeaderMeaning
X-RateLimit-Scopeuser, api-key, or ip
X-RateLimit-Policyoperation or query for customer/key checks
X-RateLimit-Typerps, daily, monthly, or the IP window
X-RateLimit-LimitLimit for the reported allowance; 0 means no hard cap
X-RateLimit-RemainingRemaining capacity for that allowance
X-RateLimit-ResetUnix timestamp for calendar reset or rate-capacity recovery
Retry-AfterMinimum seconds before retrying a rate-limited request

On a rate-limit rejection, headers describe the blocking allowance with the longest recovery time. Follow Retry-After; new traffic may still consume available capacity before your next attempt.

{
"success": false,
"error": {
"code": "RATE_LIMITED",
"message": "user operation rps limit exceeded; retry after 1s",
"details": {
"scope": "user",
"policy": "operation",
"type": "rps",
"limit": 5,
"current": 6,
"retryAfterSec": 1
}
}
}

details.type can be rps, daily, or monthly for customer and API-key limits. IP-limit responses may omit details and should be handled from Retry-After.

async function callWithRateLimitRetry(url: string, apiKey: string) {
for (let attempt = 0; attempt < 5; attempt++) {
const response = await fetch(url, { headers: { "x-api-key": apiKey } });
if (response.status !== 429) return response.json();
const waitSec = Number(response.headers.get("Retry-After") ?? 1);
await new Promise((resolve) => setTimeout(resolve, waitSec * 1000));
}
throw new Error("Rate limit retry budget exhausted");
}

For long-running generation tasks, avoid aggressive polling. Use the polling schedule in The Task Model or subscribe to Webhook Events.

Contact support before production traffic exceeds your tier. Customer-level allowances can be adjusted after capacity review. Creating additional API keys does not increase your customer allowance. Existing customer allowances may be preserved during policy transitions.