You run the bots. You pay the token bills. PromptCrunch cuts up to 75% off the input cost of every long client conversation, support, companion, tutoring, and hands you one dashboard, a spend cap per client, and a printable report to prove it. One key per client, one invoice. Bill the savings back or keep them.
Every client bot, under one dashboard.
Each client bot runs on its own API key, and the dashboard breaks input spend down by key, by model, and by day. Know exactly what every client costs you, without untangling one blended provider invoice at month end.
Set a hard spend limit on every client key, matched to the budget you agreed. No single bot can blow past it, so a runaway loop on one client's product stays that client's problem, not a hole in your margin.
The longer a chat runs, the more history you resend, and the bigger the bill. PromptCrunch strips that weight before it hits the provider, identical responses, a fraction of the input cost. Support, companion, and tutoring bots are where it hits hardest.
Every client key gets a monthly report you can print or export: what the bot spent, what compression saved, how much traffic it handled. Hand it over at renewal and your value is right there in dollars, not a promise, a receipt.
Your name is on the engagement, your numbers on the page. We stay out of the way.
Take a studio running 8 client bots. Swap in your own roster, the story is the same.
Your mix drives the number: the more conversation your bots run, the bigger your margin. Run your real roster through the savings estimator.
Route the whole roster through one base URL. Every bot gets the per-client dashboard and the spend cap, that runs on all of them. Then the long, prose-heavy conversations (support, companion, tutoring) hand back up to 75% of their input bill, and that is where your margin comes from.
Cached and code-heavy requests pass straight through, never a cent worse than sending direct, we never touch a cache hit to chase a number. Everything conversational, we crush. The comparison page shows exactly where each one wins.
Point everything at one dashboard, watch every client from one screen, and let the conversational bots pay for the whole plan several times over.
Any of the 20 API keys on the Agency plan. Give each client bot its own key: every key carries its own spend limit and its own printable report, so per-client accounting never blends together. One key per client is how the whole dashboard is wired.
Agency caps at 20 API keys. If your roster is bigger than that, email [email protected] and we will work out an enterprise arrangement. The self-hosted build, which has unlimited keys, is on the same path.
Only in the report footer. The proxy is invisible to end users: same models, same answers, no PromptCrunch branding anywhere in the bot itself. The printable client report carries one small PromptCrunch line in its footer; everything else on it is your engagement.
Requests that use cache_control pass through completely untouched, breakpoints intact, never a cent worse than sending direct. Compression goes to work on every conversational bot beside it. Caching wins when you resend the same prefix; compression wins when the conversation grows, and they stack. Details on the vs prompt caching page.
Start free: $5 credit, 100 requests a day, no card. Two lines of setup on any SDK with a configurable base URL. Point your first client bot at it today, move to Agency when the roster follows.