AI Rate Limits & Quotas
Hitting a 429 is almost always a tier problem, not a code problem. Every page on this hub documents the exact RPM, TPM, batch cap, and tier-unlock path for one provider — sourced from the provider's live limits dashboard.
If you're picking a provider, pick the limits page that matches your traffic profile. If you're already on one, bookmark the unlock path so your next scale-up doesn't take a week.
25 pages · updated 2026
Anthropic RSP ASL Levels Explained (2026): ASL-1 to ASL-5
Anthropic's Responsible Scaling Policy AI Safety Levels (ASL-1 to ASL-5) explained in depth. What each level requires for deployment + security; what current Claude models are rated at; what mitigations Anthropic states it has not yet developed. Sourced from anthropic.com/rsp + Capability/Safeguards Reports.
ReadAnthropic Tool Use Limits 2026: Max Tools, Token Costs & Parallel Calls
Complete guide to Anthropic tool use limits: 64-tool ceiling, token cost math, parallel calling, JSON schema constraints, and caching strategies for Claude agents.
ReadAnthropic Zero Data Retention Criteria (2026): Default + Enterprise
Anthropic's Zero Data Retention posture in 2026 — the API default no-long-term-retention behavior, the Enterprise contractual reinforcement, prompt caching lifecycle, sub-processor flow-down, and the AWS Bedrock / Google Vertex paths that inherit equivalent posture.
ReadAzure OpenAI Data Handling Tiers (2026): Abuse Monitoring, Opt-Out, ZDR-Equivalent
Azure OpenAI's data handling tiers in 2026 — default 30-day abuse monitoring, the modified content filter / abuse monitoring opt-out process, no-training-on-inputs commitment, region scope, and the BAA-covered HIPAA path. Sourced from Microsoft Learn Azure OpenAI documentation.
ReadCursor Pro Included Fast Requests (2026): The Fast Pool / Slow Pool Mechanics
How Cursor Pro's $20/mo fast-request pool actually works in 2026 — per-model quotas (gpt-5.5 ~500/mo, Opus 4.7 ~150/mo, Sonnet 4.6 ~500/mo), what happens when the fast pool exhausts (slow pool, 30-90s latency, possible auto-downgrade), the Composer multiplier (one Composer task fans out to multiple premium requests), Agent mode multipliers, and Pro vs Business tier delta. Sourced from cursor.com/pricing + cursor.com/docs, fetched June 2026.
ReadDevin ACUs Explained (2026): What An Agent Compute Unit Actually Buys
Cognition's Devin meters on ACUs — Agent Compute Units. 1 ACU ≈ ~1 hour of agent compute. How a trivial bug fix burns 0.1-0.3 ACU vs a full feature at 2-5 ACU. Per-plan included ACUs (Pro $20 ≈ 10, Max $200 ≈ 100, Teams 100 + 40/user). Overage at ~$2.25/ACU. The pause-and-resume mechanic, why long sessions burn unexpectedly fast, and how to forecast Devin spend. Sourced from devin.ai/pricing + Cognition docs, fetched June 2026.
ReadLangSmith Trace Quotas 2026: Plans, Limits, Pricing & Alternatives
Complete guide to LangSmith trace quotas: Developer 5k/month, Plus 50k/month, retention periods, 20 MB trace limits, overage pricing, and LangSmith vs Langfuse vs Helicone.
ReadOpenAI Assistants API v2 Rate Limits & Quotas 2026: Files, Threads, Runs
Complete guide to OpenAI Assistants API v2 limits: max files per vector store, thread retention, run expiry, RPM/TPM by tier, code_interpreter cost, and production strategies.
ReadOpenAI Embeddings Rate Limits 2026: All Tiers, Batch API, Matryoshka Dimensions
Complete guide to OpenAI embedding model rate limits across Tier 1-5: RPM, TPM, TPD, Batch API 50% discount, Matryoshka truncation, and ada-002 migration. Verified June 2026.
ReadOpenAI ZDR-Eligible Models (2026): Which Models + Endpoints
Which OpenAI models and endpoints are eligible for Zero Data Retention (ZDR) in 2026, what's excluded, how the eligibility list updates, and how to verify your project's ZDR status. Sourced from OpenAI Trust Portal and Enterprise documentation.
ReadPinecone Quota Tiers & Limits (2026): Starter, Standard, Enterprise Compared
Full breakdown of Pinecone's Starter, Standard, and Enterprise plan limits: vectors, namespaces, metadata caps, inactivity rules, serverless billing units, and BYOC. Verified June 2026.
ReadQdrant Cloud Quotas & Limits (2026): Free, Standard, Hybrid Cloud, Enterprise
Complete guide to Qdrant Cloud resource limits: RAM-based capacity planning, free-tier inactivity, quantization options, Hybrid Cloud architecture, and self-hosted vs managed trade-offs. Verified June 2026.
ReadReplit Agent Monthly Credits (2026): What $25 Core Actually Buys
Replit Core ($25/mo) includes a monthly Agent credit allotment. How credits convert to Agent actions (1 credit ≈ one substantive model call + tool invocations), how an MVP burns 5-15 credits vs a full-stack auth app at 20-50, overage billing at ~$0.25/credit, and when Core's monthly bucket runs out mid-month. Sourced from replit.com/pricing, fetched June 2026.
ReadAnthropic Message Batches API Limits (2026): 100k Requests, 256MB, 24h, 50% Off
Exact limits for Anthropic's Message Batches API in 2026: 100,000 requests per batch, 256MB max payload, 24-hour processing SLA (most batches finish in under 1 hour), 29-day results retention, 50% discount on input + output + cache writes. Separate quota pool from real-time Messages. Sourced from Anthropic's batch-processing docs.
ReadAzure OpenAI Quota Management 2026: TPM, PTU, Regional Caps & Increase Requests
Canonical 2026 reference for Azure OpenAI quota: per-subscription/per-region/per-model TPM, deployment SKUs (Standard, Global Standard, Data Zone, Regional Provisioned, Global Provisioned), Quota Tiers (0-6), default gpt-5.5 / gpt-5.4 allocations, PTU sizing and hourly billing, the quota increase request form, Dynamic Quota + spillover, and Azure vs OpenAI direct migration math. Sourced from Microsoft Learn, June 2026.
ReadClaude API Rate Limits 2026: RPM, ITPM, OTPM by Tier and Model
Exact Claude API rate limits in 2026 across Tier 1, 2, 3, 4, and Custom. Per-tier RPM, ITPM (input tokens per minute), and OTPM (output tokens per minute) for Claude Fable 5, Opus 4.7, Sonnet 4.6, and Haiku 4.5. Why Anthropic splits ITPM/OTPM instead of using combined TPM, how prompt caching multiplies effective throughput, Message Batches as a separate quota pool, 429 vs 529 handling, and the Tier 4 unlock path. Sourced from Anthropic's official rate-limits documentation.
ReadDALL·E 3 Rate Limit by Tier (2026): Full IPM Table + Workarounds
Exact DALL·E 3 rate limits at every OpenAI usage tier in 2026: Free → Tier 5, images per minute, per-image prices by resolution and quality, batch and concurrency workarounds, and what to do when you hit the cap. Sourced from OpenAI's live model documentation.
ReadFireworks AI Rate Limits 2026: Developer, Enterprise, On-Demand Deployments
Exact Fireworks AI rate limits in 2026 — Developer spending-tier ladder ($50 → $50,000 monthly caps), the 6,000 RPM account-wide ceiling, per-model serverless defaults for Llama 3.3 70B, DeepSeek V3/R1, Qwen 2.5, FireFunction V2, FLUX.1, the on-demand deployment alternative (per-GPU-hour, no rate limit), Business + Enterprise upgrade path, and the 429 vs 503 distinction. Sourced from Fireworks' live docs.
ReadGemini API Rate Limits 2026: Free Tier, Paid Tiers, Per-Model Quotas
The canonical 2026 reference for Google Gemini API rate limits. Free, Tier 1, Tier 2, Tier 3 thresholds; per-model RPM, TPM, RPD on Gemini 2.5 Pro, 2.5 Flash, 2.5 Flash-Lite; AI Studio vs Vertex AI quota systems; 429 / RESOURCE_EXHAUSTED handling; Batch API + Context Caching levers. Sourced from Google's live rate-limits documentation.
ReadGroq API Rate Limits 2026: RPM, TPM, RPD, TPD per Model — Free vs Dev vs Enterprise
Exact Groq API rate limits in June 2026 across Free, Developer, and Enterprise tiers. RPM, TPM, RPD, and TPD per model for Llama 3.3 70B Versatile, Llama 3.1 8B Instant, DeepSeek R1 Distill 70B, Qwen 2.5 32B, GPT-OSS 120B, and Whisper Large v3 Turbo. Which dimension binds first, how to upgrade, and how Groq's LPU speed translates to actual throughput. Sourced from console.groq.com/docs.
ReadOpenAI Batch API Limits 2026: Per-Tier Enqueued Tokens, 200MB Files, 24h SLA
Exactly what the OpenAI Batch API allows in 2026. Per-tier enqueued-token caps (Tier 1 → 5), 200MB max file size, 50,000 requests per batch, 24-hour SLA, the 50% discount on input + output, JSONL + custom_id schema, partial-completion behavior, and the Batch-vs-real-time decision tree. Sourced from OpenAI's live Batch API documentation.
ReadOpenAI Tier 1 vs Tier 5 (2026): What Each Usage Tier Unlocks
What every OpenAI usage tier unlocks in 2026 — Free → Tier 5. Monthly caps, indicative gpt-5.5 RPM/TPM, image limits, fine-tuning, batch quotas, prompt cache eligibility, priority routing. Sourced from OpenAI's rate-limits doc.
ReadOpenAI Tier 5 Unlock Requirements (2026): The Canonical Doc
Exactly what it takes to unlock OpenAI usage Tier 5 in 2026. $1,000 paid + 30 days since first payment. Full thresholds for every tier (Free → Tier 5), monthly usage caps, rate-limit gains per tier, payment-history strategies, common stuck-at-Tier-4 traps. Sourced from OpenAI's official rate-limits page.
ReadReplicate Rate Limits 2026: Predictions/Sec, Concurrency & Cold Starts
Exact Replicate rate limits in 2026: 600 predictions/min default, per-model concurrency caps, and the 30-90s cold-start problem that dominates latency. When to switch to always-on dedicated deployments, GPU class pricing (A100 vs H100 vs L40S), webhooks for long-running predictions, and self-hosted Cog. Sourced from Replicate's live docs.
ReadTogether AI Rate Limits 2026: Build, Scale, Enterprise — Per-Model Ceilings
Exact Together AI rate limits in 2026 across Build, Scale, and Enterprise tiers. Per-model RPM/TPM for Llama 3.3 70B/8B, DeepSeek R1, Qwen 2.5, FLUX.1, BGE embeddings. When to switch from serverless to dedicated endpoints (per-GPU-hour math). 429 handling, embedding + fine-tuning quotas, batch API. Sourced from Together's live docs.
Read
Stop guessing your AI bill.
Digital Dashboard Hub turns your real spend across OpenAI, Anthropic, and Google into one live dashboard — usage, cost, budget alerts, model mix. 14 days free.
Try DDH free