What the ~$3.0K/mo production floor buys (pilot scale)
| Component | $ / mo | What it covers |
|---|---|---|
| Compute — AKS / Container Apps | ~$600 | Hosts the ~12 .NET services; orchestration in-process over the event spine |
| Azure Data Explorer (ADX) | ~$1,200 | Production time-series cluster — every event + the LLM / decision audit trail |
| Event Hubs (Standard) | ~$100 | The event spine — POS, traffic, inventory, sensors, video metadata |
| Cosmos DB (provisioned) | ~$250 | Operational records, agent memory, tenant / store context |
| Cache for Redis (Standard C2) | ~$115 | Hot state, cooldowns, the statistical anomaly baselines |
| Azure AI Search (S1) | ~$250 | Vector index for RAG — store & enterprise knowledge base |
| Blob storage | ~$80 | Canonical audit — full LLM prompts / responses retained (SOC 2) |
| Platform overhead | ~$400 | Monitoring, Key Vault, identity, networking, egress, contingency |
| Production floor | ~$3.0K | ≈ $36K / yr at pilot |
What the floor carries — at near-zero marginal cost
The event spine already ingests POS, inventory, people-counters, sensors, weather, supply-chain & staffing. These structured streams are a rounding error next to video and agent reasoning — they ride essentially free on the same infrastructure.
Scales step-wise, not linearly
Floor grows by tier bumps (AI Search, ADX, Event Hubs Premium) as you add stores — so the per-store share of the floor keeps falling.
Not in this estimate — budget separately
Total cost of ownership — scaling with store count
| 1 storepilot | 5 storesexpansion | 10 storespilot target | → 100fleet · curve | |
|---|---|---|---|---|
| Platform floor /mo (slide 1) | $3.0K | $3.3K | $4.0K | $18K |
| Per-store variable (~$685/store·mo) | $0.7K | $3.4K | $6.8K | $68K |
| Total / month | ~$3.7K | ~$6.7K | ~$10.8K | ~$86K |
| Pilot basis / year hosted LLM / frame | ≈ $44K | ≈ $81K | ≈ $130K | ≈ $1.0M |
| ↓ At scale, cheap CV / year | ≈ $40K | ≈ $58K | ≈ $85K | ≈ $580K |
Per-store cost stack — pilot (~$685/mo)
Video cost — pilot vs. scale ($/store·mo)
A hosted LLM on each frame is the simplest pilot. At ~10+ stores, continuous detection moves to computer vision (one GPU fills up across the fleet) → video ~4× lower.