Stop runaway agents before they drain your API budget
AI Circuit Breaker is a drop-in OpenAI-compatible reverse proxy that sits between your agents and upstream LLM providers. It detects recursive agent loops using an in-memory sliding-window fingerprint check and enforces hard hourly and daily spend caps in volatile RAM, without persisting prompt payloads to disk.
Prompts are fingerprinted (SHA-256) and matched in a 3x sliding deque so repeated tool calls or retries are interrupted at the third identical attempt, before remaining requests are sent upstream.
Requests are checked against configurable hourly and daily account ceilings before being forwarded, returning HTTP 429 when a budget limit is breached.
Prompt bodies are processed for loop detection in volatile memory and never written to disk; usage logs and metadata exclude prompt text, and raw upstream keys are not logged.
Works with any client or framework that can point at a custom base URL, including OpenAI SDKs, CrewAI, LangChain, LangGraph, AutoGen, and direct REST clients — you change the base URL, keep the SDK.
Upstream API keys are encrypted with AES-256 (Fernet) before database insertion; protected key material is hashed for lookup.
HTTP 400 for a stopped infinite loop, HTTP 429 for a reached budget ceiling, and HTTP 502 for upstream provider errors, with structured JSON block payloads.
You configure your agent's OpenAI client with the Circuit Breaker base URL (e.g. https://circuit-breaker-api.onrender.com/v1) and your Circuit Breaker key instead of hitting the provider directly.
Each request passes through a RAM hash and spend-cap engine: a sliding-window prompt fingerprint check and hourly/daily spend checks run before the request is forwarded upstream.
If a prompt fingerprint crosses the repeat threshold (attempt 3 of an identical sequence), the request is stopped with an HTTP 400 'infinite_loop_killed' payload; if a cap is exceeded, HTTP 429 is returned instead of sending the request upstream.
Two plans: a free Developer tier with up to $15/mo monitored spend, and a $29/mo Production (Pro) tier with unlimited monitored spend, higher burst ceilings, and emergency webhook alerts.
$0/mo
For development and initial workloads: up to $15/mo monitored spend, 3x sliding-window loop detection, hourly and daily budget caps, and the OpenAI-compatible proxy endpoint.
$29/mo
For production agents needing higher budget ceilings: unlimited monitored spend, $50/hr and $200/day burst ceilings, emergency Slack and Discord webhooks, and priority European routing.
It's a reverse proxy that sits between your agents and LLM providers, detects recursive agent loops using an in-memory sliding-window fingerprint check, and enforces hourly and daily spend caps — all in volatile memory without retaining prompt payloads.
The proxy is designed for under 35ms of average overhead. Actual latency depends on network distance, provider response time, and deployment conditions.
No. Prompt bodies are processed for loop detection in memory and not written to disk; request metadata and usage logs do not include prompt text.
Any client or framework that can send OpenAI-compatible requests to a custom base URL, including CrewAI, LangChain, LangGraph, AutoGen, and direct REST clients.
There's a free Developer plan ($0/mo, up to $15/mo monitored spend) and a Production Pro plan at $29/mo with unlimited monitored spend, $50/hr and $200/day burst ceilings, and emergency Slack/Discord webhooks.
HTTP 400 indicates an infinite loop was stopped, HTTP 429 indicates a budget ceiling was reached, and HTTP 502 indicates an upstream provider error.
The site links to a GitHub repository and states it's built in Python and FastAPI by Roan de Jager; the repository link is provided on the homepage.