Cut every client's LLM bill. Pocket the difference.

You run the bots. You pay the token bills. PromptCrunch cuts up to 75% off the input cost of every long client conversation, support, companion, tutoring, and hands you one dashboard, a spend cap per client, and a printable report to prove it. One key per client, one invoice. Bill the savings back or keep them.

See it. Cap it. Crunch it.

Every client bot, under one dashboard.

Per-client spend at a glance

Each client bot runs on its own API key, and the dashboard breaks input spend down by key, by model, and by day. Know exactly what every client costs you, without untangling one blended provider invoice at month end.

Per-key spend limits

Set a hard spend limit on every client key, matched to the budget you agreed. No single bot can blow past it, so a runaway loop on one client's product stays that client's problem, not a hole in your margin.

Up to 75% fewer input tokens

The longer a chat runs, the more history you resend, and the bigger the bill. PromptCrunch strips that weight before it hits the provider, identical responses, a fraction of the input cost. Support, companion, and tutoring bots are where it hits hardest.

The client report

A printable report for every client, every month

Every client key gets a monthly report you can print or export: what the bot spent, what compression saved, how much traffic it handled. Hand it over at renewal and your value is right there in dollars, not a promise, a receipt.

Your name is on the engagement, your numbers on the page. We stay out of the way.

  • Spend. Input spend for the month, broken down by model.
  • Savings. Tokens and dollars cut by compression, counted per request from real token counts.
  • Volume. Requests handled and tokens processed, so the client sees exactly what their budget bought.

Where the margin comes from

Take a studio running 8 client bots. Swap in your own roster, the story is the same.

8 client bots at roughly $150/mo of input spend each
~$1,200/mo
6 of them run real conversations (support, companion, tutoring). That is where compression goes to work.
~$900/mo
Up to 75% off that spend, 60-75% across models in our benchmarks, 75% on Sonnet
$540 to $675/mo saved
The other 2 bots run code or cached prompts. They pass through billed exactly as today.
billed as-is
One flat Agency subscription, the only thing PromptCrunch ever charges
-$199/mo
Net margin back to you
roughly $340 to $475/mo

Your mix drives the number: the more conversation your bots run, the bigger your margin. Run your real roster through the savings estimator.

Every bot covered. The conversations pay the most.

Route the whole roster through one base URL. Every bot gets the per-client dashboard and the spend cap, that runs on all of them. Then the long, prose-heavy conversations (support, companion, tutoring) hand back up to 75% of their input bill, and that is where your margin comes from.

Cached and code-heavy requests pass straight through, never a cent worse than sending direct, we never touch a cache hit to chase a number. Everything conversational, we crush. The comparison page shows exactly where each one wins.

Point everything at one dashboard, watch every client from one screen, and let the conversational bots pay for the whole plan several times over.

Agency questions

What counts as a client key?

Any of the 20 API keys on the Agency plan. Give each client bot its own key: every key carries its own spend limit and its own printable report, so per-client accounting never blends together. One key per client is how the whole dashboard is wired.

What happens past 20 keys?

Agency caps at 20 API keys. If your roster is bigger than that, email [email protected] and we will work out an enterprise arrangement. The self-hosted build, which has unlimited keys, is on the same path.

Does the client see PromptCrunch?

Only in the report footer. The proxy is invisible to end users: same models, same answers, no PromptCrunch branding anywhere in the bot itself. The printable client report carries one small PromptCrunch line in its footer; everything else on it is your engagement.

What if a bot uses prompt caching?

Requests that use cache_control pass through completely untouched, breakpoints intact, never a cent worse than sending direct. Compression goes to work on every conversational bot beside it. Caching wins when you resend the same prefix; compression wins when the conversation grows, and they stack. Details on the vs prompt caching page.

Turn every client's token bill into your margin

Start free: $5 credit, 100 requests a day, no card. Two lines of setup on any SDK with a configurable base URL. Point your first client bot at it today, move to Agency when the roster follows.

Get your free key Estimate your margin See what Claude charges