AI Cost Simulation
1. Overview
This document provides a practical way to estimate monthly AI usage cost for Centrify 360 AI.
|
AI usage requires a separate Centrify 360 AI add-on license. |
|
All figures in this document are illustrative only. |
2. Baseline Assumptions
This page uses a baseline of 500 queries per month.
For estimation purposes, assume the following per AI request:
-
Average input tokens per request: 90K
-
Cached portion of input: 70%
-
Non-cached portion of input: 30%
-
Average output tokens per request: 1K
This means a typical request behaves approximately like:
-
63K cached input tokens
-
27K non-cached input tokens
-
1K output tokens
At 500 queries per month, that becomes:
-
31.5M cached input tokens / month
-
13.5M non-cached input tokens / month
-
0.5M output tokens / month
These assumptions reflect the fact that a large portion of the semantic-model and instruction context is repeated across requests.
3. Why Caching Changes the Economics
Without caching, every request would be billed as a full 90K-token input interaction.
With caching, repeated context can often be reused, which materially lowers cost.
In many real-world usage patterns, this can reduce effective AI cost by around 70% to 80%.
4. Recommended Models for Evaluation
The following models are reasonable starting points for evaluation:
-
GPT-5.4
-
Claude Sonnet 5
-
DeepSeek V4 Pro
-
Gemini 3.6 Flash
5. Baseline Monthly Cost Comparison (500 Queries)
The table below applies publicly listed pricing to the baseline assumptions above.
| Model | Cached input rate | Standard input rate | Output rate | Illustrative monthly cost at 500 queries |
|---|---|---|---|---|
GPT-5.4 |
$0.25 / 1M |
$2.50 / 1M |
$15.00 / 1M |
≈ $49.13 / month |
Claude Sonnet 5 |
$0.30 / 1M |
$3.00 / 1M |
$15.00 / 1M |
≈ $64.95 / month |
DeepSeek V4 Pro |
$0.003625 / 1M |
$0.435 / 1M |
$0.87 / 1M |
≈ $6.85 / month |
Gemini 3.6 Flash |
$0.15 / 1M |
$1.50 / 1M |
$7.50 / 1M |
≈ $29.48 / month |
|
The monthly figures above are based on the formula:
All rates are expressed per 1 million tokens. |
6. Pricing Sources Used
The estimates above were aligned to public pricing pages available online and checked in July 2026:
-
GPT-5.4: public model catalog pricing showing $2.50 / 1M input, $0.25 / 1M cached input, and $15 / 1M output.
-
Claude Sonnet 5: Anthropic pricing page showing $3 / 1M input, $0.30 / 1M cache hits, and $15 / 1M output.
-
DeepSeek V4 Pro: DeepSeek pricing page showing $0.435 / 1M cache miss, $0.003625 / 1M cache hit, and $0.87 / 1M output.
-
Gemini 3.6 Flash: Google Gemini pricing page showing $1.50 / 1M input, $0.15 / 1M context caching, and $7.50 / 1M output.
Because vendors update pricing frequently, treat this page as a planning aid rather than a contractual price sheet.
7. Monthly Query Simulation
The table below shows how the same token assumptions scale with monthly usage.
| Monthly Queries | Average Input Tokens per Query | Cached Input Tokens | Non-Cached Input Tokens | Output Tokens |
|---|---|---|---|---|
500 |
90K |
31.5M |
13.5M |
0.5M |
1,000 |
90K |
63.0M |
27.0M |
1.0M |
2,500 |
90K |
157.5M |
67.5M |
2.5M |
5,000 |
90K |
315.0M |
135.0M |
5.0M |
10,000 |
90K |
630.0M |
270.0M |
10.0M |
8. Practical Recommendations
-
Start with the 500-query baseline and adjust once you have real usage telemetry.
-
Prefer models and providers with strong prompt caching support.
-
Compare the recommended premium-quality and cost-efficient options against your own workload.
-
Revisit pricing regularly because model pricing changes frequently.
-
Include the separate Centrify 360 AI add-on license in all commercial planning.