AI Infrastructure
Find your AI infrastructure fit.
The right AI infrastructure depends on your team size. Let's find yours.
5
Total users
—
Cloud cost / mo
—
H100 payback period
Interactive Calculator
How is your team using AI?
Configure your team
Enter the number of users in each category
Interactive
~3M tokens / user / mo
Chat, code assistance, document Q&A
users
Agent-assisted
~20M tokens / user / mo
Interactive + automated pipelines & workflows
users
Full automation
~75M tokens / user / mo
Heavy 24/7 agent workloads, batch inference
users
⚠ Above 150 total seats — Enterprise plan required by Anthropic
Hardware amortization period
Spreads the upfront hardware cost across this many years — sets the monthly amortization figures and the cost-chart horizon. Does not change break-even (payback is upfront ÷ monthly savings, independent of this).
years
⚠
Context window trade-off. Claude Sonnet supports up to 200K tokens of context. Local models on Entry Hardware top out at ~32K; the H100 running Llama 3.3 70B INT4 reaches ~128K. For long documents, large codebases, or extended reasoning chains, model selection on-prem is a critical factor — not just a technical detail.
Cloud API
Claude Sonnet API
Pay-per-token · No hardware · Infinitely elastic
Monthly cost
—
—
Upfront
€0
No hardware investment
Break-even
—
Reference line
Team Premium (≤150 seats)
—
Entry Hardware
Entry Server
Dual RTX 4090 · 2 × 24 GB VRAM (no NVLink) · 512 GB RAM
Runs Qwen2.5 32B Q4 · ~70 tok/s · ~32K context
Runs Qwen2.5 32B Q4 · ~70 tok/s · ~32K context
Monthly cost
€700
Fixed — power + ops (hardware shown separately)
Upfront investment
€20,000
5-yr amortization = €333/mo
Break-even vs cloud
—
—
Capacity
—
⚠
Above 70% capacity — queue latency increasing. Approaching hardware limit.
✕
Token demand exceeds Entry Hardware throughput — step up to H100.
H100 Tier
H100 Server
NVIDIA H100 80 GB · 1 TB RAM
~1,500 tok/s batch · ~128K context
Monthly cost
€1,800
Fixed — power + ops (hardware shown separately)
Upfront investment
€55,000
5-yr amortization = €917/mo
Break-even vs cloud
—
—
Capacity
—
⚠
Above 70% capacity — queue latency increasing. Approaching hardware limit.
✕
Token demand exceeds single-node capacity. Contact us to spec a multi-GPU setup.
60-Month Cumulative Cost
Hover to inspect monthly totals
Cloud
Entry Server
H100 Server
Model Assumptions
Sonnet pricing: $3/M input + $15/M output (60/40 mix → $7.80/M blended)
Interactive: ~3M tokens/user/mo · Agent-assisted: ~20M · Full automation: ~75M
Cloud cost (≤150 users): Team Premium $100/seat/mo + excess API tokens beyond plan
Estimated included in Team Premium: ~20M tokens/seat/mo (Anthropic states "6.25× Pro" — exact limits not published)
Enterprise (>150 seats): $20/seat platform fee + full API consumption — custom contract
USD/EUR conversion: $1 = €0.92 (approximate)
Entry Hardware: Qwen2.5 32B on single RTX 4090 · ~70 tok/s · 181M tok/mo capacity
H100 · Qwen2.5 72B INT4: ~1,200 tok/s · 3,110M tok/mo capacity · ~128K context
H100 · Qwen2.5 32B: ~500 tok/s · 1,296M tok/mo capacity · ~32K context
Capacity zones: ≤70% = comfortable · 70–100% = queue degrading · >100% = overloaded
Hardware amortized over 60 months (5 yr) · H100 electricity ~€90/mo included in ops figure
All figures in EUR (hardware) and USD-converted (API) excluding VAT