---
title: TokenBouncer | Organizational AI usage and cost control
description: TokenBouncer controls usage and cost by user, group, and model in an OpenAI-compatible request path while tracing every AI request.
image: /og/lawnect-preview.jpg
---

# TokenBouncer

## Control organizational AI through policy and cost

TokenBouncer applies limits by user, group, and model in the OpenAI-compatible request path between internal AI services and model providers. It records requests, tokens, cost, latency, blocking, fallback, and errors for operations.

## Organization-aware usage policy

Set monthly request, token, and budget limits per user or across an entire group. Layer department policies and individual exceptions over global defaults, with separate limits and priorities for exact model names or patterns. When several policies match, operators can inspect their evaluation order and effective result.

Connect groups from Microsoft Entra ID (Active Directory), Keycloak, and other internal identity systems to real request identities to inspect inherited policy, individual overrides, and shared group budgets in one view. A separate self-service view lets regular users check utilization, remaining limits, and the next reset time without exposing administration.

## End-to-end AI cost optimization

TokenBouncer goes beyond reporting spend. It combines provider and model pricing with actual input, output, and cached tokens to establish a cost baseline, then analyzes which teams and users rely on each model.

Using actual usage distribution, it calculates model-budget combinations for savings targets from 10% to 90%. Operators compare projected total spend, spend per user, and affected users, then turn a selected scenario directly into user, group, and model limits plus fallback order. Premium models can remain available for critical work while repeatable tasks move to more efficient alternatives.

After rollout, TokenBouncer continues comparing model mix, confirmed spend, and cache savings. Latency and status events show whether savings create slower responses or excessive blocking, making cost optimization an ongoing operating process rather than a one-time report.

## Cost and usage observability

Combine model pricing with actual input and output tokens to calculate confirmed cost per request, while tracking tokens and spend saved through cache hits separately. Compare concentration and budget burn rate by provider, model, user, and group to find where policy should change.

Every request retains latency, status, block reason, and effective policy. When fallback runs, the event includes the requested and actual model, rule, reason, and attempt order. Support cases connect to an exact request through an event ID, while event CSV exports and administrator changes to settings, keys, and MCP connections remain separately auditable.

## Resilient request handling

When the original model reaches its quota, reevaluate approved alternatives in order within the same provider. Retry upstream 5xx and transport failures according to operating policy, and safely repair known request-shape errors before one additional attempt. Every retry, repair, and fallback result remains attached to the request event.

Customize messages for quota exhaustion, policy blocks, authentication failures, and upstream outages around organizational language and support procedure. Users receive a clear explanation and traceable event ID, while operators retain the original error and complete handling history.

## Identity, access, and agents

Require a Virtual Key before requests reach upstream model providers. Separate keys for workspaces, batch jobs, and services, with allowed models, budgets, expiration, and last-use tracking.

Combine the custom user header an existing AI application already sends with OIDC and service-account synchronization from Microsoft Entra ID (Active Directory), Keycloak, and other internal identity systems so existing users, groups, and memberships become the basis of policy. Configure how missing or unknown identities are handled for each environment. MCP servers and tokens available to AI agents retain connection and usage audit records.

## Deployment

TokenBouncer keeps the request format used by existing OpenAI-compatible clients. The proxy, administration interface, and database can be deployed in your own infrastructure and connected to the organization’s identity and permission model.

## Contact

Tell us about your AI services, model providers, and organizational structure. We will propose the integration, policy scope, deployment, and operating model.

Company location: Seocho-gu, Seoul, Republic of Korea

Email: contact(at)lawnect.com
