Every prompt your team writes is a description of how your business works.
Private AI in a Box is enterprise AI that runs inside your Azure subscription. Deployed, benchmarked, and signed off in five days. $25,000 fixed.
Three things you give up when your AI belongs to somebody else
Your data is valuable — and you already know where it is, who can reach it, and how it is classified. What sits on top of it has never been written down at all: how you price, how you underwrite, how you diagnose a failure, why one exception gets escalated and another doesn’t. Using AI means putting that into words, precisely, for the first time.
You would never put that in a document and send it to an outsider. Every prompt does exactly that.
To get useful work out of a model, your people have to describe the business to it in detail — the process, the exceptions, the judgement calls, the reasoning that has never been documented anywhere. That description is the asset. It goes to a third party under terms you did not write.
The evidence ↓A hosted model has a usage policy, and that policy decides what your organization is allowed to do with AI. Not your risk committee. Not your CISO. The vendor’s. And in some functions — security operations above all — that boundary sits directly across work you need to do.
The evidence ↓Someone else decides what your AI costs, what a seat includes, and what quietly moves from included to metered. Your finance team is not budgeting for a tool. It is budgeting for a price list it does not control.
The evidence ↓Everything below is about closing all three without giving up capability. The solution first — the evidence for each argument is further down, for whoever has to make the case internally.
Enterprise AI that runs inside your Azure subscription
An open-weight model, a private ChatGPT-style workspace your team signs into with corporate credentials, and the plumbing that makes it real — private networking, corporate sign-on, monitoring, cost controls.
It deploys into your subscription. Prompts, documents, and answers never cross your network boundary.
Built here, your accumulated AI capability sits in your subscription, in open formats, on hardware you control — and the model itself is a parameter you can swap without rebuilding anything around it.
Every major provider offers assurances about your data. Those assurances are real — and they are theirs to revise. Ours is a network topology with no outbound path. Only one of those is yours to verify, from your own logs.
Nine things, delivered and verified
Private endpoints, network isolation, corporate sign-on, public access disabled. All infrastructure-as-code.
gpt-oss-120b served on vLLM. Apache 2.0, US-origin, released by OpenAI as open weights — permissive commercial use, no copyleft exposure, nothing for legal to negotiate with a model vendor. Swappable for Qwen, Mistral, or whatever is best in twelve months.
A ChatGPT-style interface for the whole team. Corporate sign-in. Conversation history stored in your tenant.
Upload a policy library and ask questions against it. Vector database and embeddings run in-cluster.
Live results with citations, searched and scraped inside your cluster. No query reaches a third party.
Speech transcribed in-cluster. Audio never leaves.
Screenshots and scans extracted in-cluster. Text in images — not visual interpretation. See What this is not.
Latency, throughput, and spend dashboards. Budget alerts pre-wired.
Your capacity, measured on your hardware, against criteria agreed before day one.
Deployed, proved, measured, handed over
Template into your subscription; model serving live, TLS valid.
Sign-on configured, joint security review, boundaries agreed.
Every capability demonstrated and verified in the service logs.
Load benchmark at 8, 16, and 32 concurrent users.
Training, documentation, benchmark report, signed acceptance.
The clock starts when prerequisites are confirmed. We run those checks with you first, so the five days are five days.
A model narrates success in confident language whether or not anything happened. Every acceptance criterion is proved by a command against the component’s own log.
One fixed price, and a bill you control
| Deployment engagement — one time, to Data Meaning | $25,000 |
| Azure infrastructure — to Microsoft, your own bill | from ~$2,900 / moconfirmed during the prerequisite check |
| Box Managed — optional, to Data Meaning | $2,000 / mo |
We do not mark up your Azure spend and we do not receive it.
One commitment up front: the GPU runs continuously. A100 capacity is scarce enough that releasing the machine risks not getting it back — a fixed cost, not a variable one.
Available on Azure Marketplace. The solution template is published and certified by Microsoft, so it can be procured through the same channel as the rest of your Azure estate.
No per-seat licence. Cost is driven by capacity, not headcount — so the more your organization uses it, the better the economics get. Per-seat AI licensing works the other way around.
We would rather lose the deal here than in week three
Corporate sign-on is installed, not designed. Group structure and access policy are your identity team’s decisions.
Retrieval-ready. Connecting it to Snowflake or your document systems is accelerator work, quoted separately.
Text in a screenshot is extracted and read. Visual interpretation needs a second GPU node — we measured it, and it does not fit on one.
It is network-isolated inside your Azure subscription with no outbound path from the workload. If your requirement is a physically disconnected environment, say so on the first call and we will tell you honestly whether this is the right shape.
The base configuration ships with the model’s own safety behaviour intact. Purpose-selected configurations for security operations are a deliberate, scoped decision made with your risk function — not a switch we flip on day one.
It is a production environment. That is the point.
The honest test, before anyone books anything
This fits when
- Your people are describing genuinely proprietary process, judgement, or IP to an AI tool
- Legal, risk, or client contracts limit what you can send to a third-party provider
- A function in your business — security, fraud, compliance, investigations — keeps running into refusals
- You are in financial services, insurance, healthcare, or manufacturing
- You want a production environment, not another pilot that ends in a slide deck
- You have an Azure subscription and someone who can approve GPU quota
This does not fit when
- You need frontier-model reasoning at the very top of the benchmark tables
- You need a physically air-gapped or classified deployment
- You have no Azure footprint and no intention of building one
- Fewer than 20 people would use it
The four prerequisites. Deployability comes down to four checks: GPU quota on your subscription, GPU capacity in your region, subscription permissions, and availability zone. Thirty minutes on a call answers all four — before any commitment, any paperwork, or any pitch.
Everything below is the evidence behind the argument at the top — written for whoever has to make the case to a risk committee. If you are already convinced, the button above is the whole next step.
The five options, side by side
| Time to production | Cost model | What you disclose | What you’re permitted to do | Model control | |
|---|---|---|---|---|---|
| Private AI in a Box | 5 days | $25,000 fixed, then flat capacity | Nothing leaves your subscription | Your policy | Open weights, swappable |
| Per-seat AI (ChatGPT Enterprise, Copilot) | Days | Per seat, rises with headcount, vendor-set | Prompts and documents, under vendor terms | Vendor usage policy | None |
| Managed AI service in your cloud region | Weeks | Per token, rises with usage | In-region, vendor-operated service | Vendor usage policy | Vendor catalogue |
| On-prem AI appliance (Dell, HPE, Nutanix) | Months | Six figures capex, plus facilities | Nothing | Yours | Varies |
| Build it yourself | 3–6 months | GPU + salaries | Nothing | Yours | Full |
Building it yourself is the option most often underestimated: the licence is free, but GPU-experienced ML infrastructure engineers run $180,000–$250,000 and take three to six months to hire. And this is not either/or — plenty of clients run Copilot for productivity and the box for anything they will not disclose outside their boundary.
The three arguments above, each in one fact
Full treatment in the linked articles — written for whoever has to make the case to a risk committee.
Disclosure
Trade secret protection can be forfeited by disclosure. In 2025 a US federal court held that voluntarily disclosing claimed trade secrets to a consumer-tier AI platform — one whose terms imposed no confidentiality obligation — defeated trade secret protection, regardless of who owned the output. Counsel now advise confining confidential material to deployments where retention, training, and access are technically controlled, not only contractually promised.
Read: The trade secret you gave away in a text box →Permission
Malware analysis: 34.3%. Stating that you are on the blue team made refusal more likely, not less — because defenders and attackers use identical language. Attackers have adapted: malware is now seeded with content designed purely to trip an analysing model’s safety filter so the file goes unread.
Read: Your safety guardrail is a defender’s problem →Cost
Four pricing changes in one month, none of them a sticker price. In June and July 2026: GitHub Copilot moved chat and agent features to usage-based credits, Copilot Cowork added per-task billing on top of the existing seat, volume discounts expired, and the Microsoft 365 base suites underneath every Copilot seat rose. You cannot build a three-year business case on a price list a vendor can restructure in a quarter. A dedicated GPU has one price, whatever the volume.
Read: Nobody raised a price and everybody’s bill went up →Where it goes after week one
Proved end to end: the workspace queried a live Snowflake schema and returned real consolidated P&L figures, writing the query itself.
A model selected and deployed for defensive security work — malware analysis, threat intelligence, log and incident reasoning — where general-purpose guardrails refuse the task. Scoped to named users and named use cases, behind your identity provider, with full audit logging. Your risk function sets the policy; we build the environment that enforces it.
Model updates and evaluations, cost optimization, quarterly reviews. Priced below the 20–30% of a senior engineer this otherwise consumes.
Same subscription, same box, same guarantees. The pilot is the first slice of production — not a throwaway.
What your security, legal, and procurement teams will ask
Where does our data go?
Our enterprise AI agreement already says they won’t train on our data. Isn’t that enough?
What are we agreeing to for the model, and where does it come from?
gpt-oss-120b, released by OpenAI as open weights under Apache 2.0 — permissive commercial use, no copyleft exposure, US-origin. OpenAI’s gpt-oss usage policy sits alongside the licence. There is no model vendor agreement to negotiate, and where provenance is a procurement question, the objection disappears without a debate.Does this make us compliant with the EU AI Act, DORA, GLBA, or HIPAA?
Who operates it after handover, and what happens if we stop working with you?
We already have Copilot. Why this?
Somebody has to carry the operating burden
Sovereign AI platforms are arriving from every major vendor, and each one hands the operating burden to a partner. Where that bench is thin, it lands back on a customer who cannot carry it.
We are already inside your data estate — Snowflake, BI, analytics engineering, 20+ years and 200+ implementations. That is why your private model connects to the platforms your business actually runs on, and why the week ends in a signed acceptance document rather than a handover deck.
Microsoft partner · Governance practice built on Collibra and Alation · Delivery in English and Spanish across the Americas
Start with the prerequisites, not the pitch.
Four checks decide whether this is deployable in your subscription: GPU quota, GPU capacity in your region, subscription permissions, availability zone. Thirty minutes, before any commitment.