Private AI in a Box · SKU PAIB-01

Every prompt your team writes is a description of how your business works.

Private AI in a Box is enterprise AI that runs inside your Azure subscription. Deployed, benchmarked, and signed off in five days. $25,000 fixed.

01CONSUMEPrivate chat workspace · Corporate sign-on · Chat history in-tenantMicrosoft Entra ID[ chat workspace ]Azure Storage02SERVEModel serving · OpenAI-compatible endpoint · Dedicated GPU · Model swappablegpt-oss-120bApache 2.0vLLMNVIDIA A100 80GB/v1 OpenAI-compatible03AUGMENTRetrieval · Self-hosted search · Speech-to-text · Image text extraction[ vector database ][ embeddings model ][ search engine ][ STT model ][ OCR engine ]04GOVERNSecrets · Latency, throughput and spend dashboards · Budget alertsAzure Key VaultAzure MonitorAzure Cost Management05BOUNDARYNetwork isolation · No outbound path from the workloadAzure Private LinkAzure VNetAzure Private DNSTLSpublic access disabledYOUR AZURE SUBSCRIPTION[ bracketed, dashed ]  confirmed in the build, product not yet namedPUBLIC INTERNET  ·  THIRD-PARTY AIno prompts  no documents  no IP  no telemetry
The argument

Three things you give up when your AI belongs to somebody else

Your data is valuable — and you already know where it is, who can reach it, and how it is classified. What sits on top of it has never been written down at all: how you price, how you underwrite, how you diagnose a failure, why one exception gets escalated and another doesn’t. Using AI means putting that into words, precisely, for the first time.

You would never put that in a document and send it to an outsider. Every prompt does exactly that.

The first is disclosure.

To get useful work out of a model, your people have to describe the business to it in detail — the process, the exceptions, the judgement calls, the reasoning that has never been documented anywhere. That description is the asset. It goes to a third party under terms you did not write.

The evidence ↓
The second is permission.

A hosted model has a usage policy, and that policy decides what your organization is allowed to do with AI. Not your risk committee. Not your CISO. The vendor’s. And in some functions — security operations above all — that boundary sits directly across work you need to do.

The evidence ↓
The third is the meter.

Someone else decides what your AI costs, what a seat includes, and what quietly moves from included to metered. Your finance team is not budgeting for a tool. It is budgeting for a price list it does not control.

The evidence ↓

Everything below is about closing all three without giving up capability. The solution first — the evidence for each argument is further down, for whoever has to make the case internally.

What this is

Enterprise AI that runs inside your Azure subscription

An open-weight model, a private ChatGPT-style workspace your team signs into with corporate credentials, and the plumbing that makes it real — private networking, corporate sign-on, monitoring, cost controls.

It deploys into your subscription. Prompts, documents, and answers never cross your network boundary.

Built here, your accumulated AI capability sits in your subscription, in open formats, on hardware you control — and the model itself is a parameter you can swap without rebuilding anything around it.

A promise is a policy. A boundary is a fact.

Every major provider offers assurances about your data. Those assurances are real — and they are theirs to revise. Ours is a network topology with no outbound path. Only one of those is yours to verify, from your own logs.

What’s in the box

Nine things, delivered and verified

Hardened environment

Private endpoints, network isolation, corporate sign-on, public access disabled. All infrastructure-as-code.

Licence-vetted model on dedicated GPU

gpt-oss-120b served on vLLM. Apache 2.0, US-origin, released by OpenAI as open weights — permissive commercial use, no copyleft exposure, nothing for legal to negotiate with a model vendor. Swappable for Qwen, Mistral, or whatever is best in twelve months.

Private chat workspace

A ChatGPT-style interface for the whole team. Corporate sign-in. Conversation history stored in your tenant.

Document upload and retrieval

Upload a policy library and ask questions against it. Vector database and embeddings run in-cluster.

Self-hosted web search

Live results with citations, searched and scraped inside your cluster. No query reaches a third party.

Voice input

Speech transcribed in-cluster. Audio never leaves.

Reads text in images

Screenshots and scans extracted in-cluster. Text in images — not visual interpretation. See What this is not.

Monitoring and cost controls

Latency, throughput, and spend dashboards. Budget alerts pre-wired.

Benchmark report, documentation, training

Your capacity, measured on your hardware, against criteria agreed before day one.

The five days

Deployed, proved, measured, handed over

1Deploy

Template into your subscription; model serving live, TLS valid.

2Secure

Sign-on configured, joint security review, boundaries agreed.

3Prove

Every capability demonstrated and verified in the service logs.

4Measure

Load benchmark at 8, 16, and 32 concurrent users.

5Hand off

Training, documentation, benchmark report, signed acceptance.

The clock starts when prerequisites are confirmed. We run those checks with you first, so the five days are five days.

We verify at the service, not at the screen.

A model narrates success in confident language whether or not anything happened. Every acceptance criterion is proved by a command against the component’s own log.

What it costs

One fixed price, and a bill you control

Deployment engagement — one time, to Data Meaning$25,000
Azure infrastructure — to Microsoft, your own billfrom ~$2,900 / moconfirmed during the prerequisite check
Box Managed — optional, to Data Meaning$2,000 / mo

We do not mark up your Azure spend and we do not receive it.

One commitment up front: the GPU runs continuously. A100 capacity is scarce enough that releasing the machine risks not getting it back — a fixed cost, not a variable one.

Available on Azure Marketplace. The solution template is published and certified by Microsoft, so it can be procured through the same channel as the rest of your Azure estate.

No per-seat licence. Cost is driven by capacity, not headcount — so the more your organization uses it, the better the economics get. Per-seat AI licensing works the other way around.

What this is not

We would rather lose the deal here than in week three

Not governed on arrival.

Corporate sign-on is installed, not designed. Group structure and access policy are your identity team’s decisions.

Not connected to your data estate.

Retrieval-ready. Connecting it to Snowflake or your document systems is accelerator work, quoted separately.

It reads images; it does not see them.

Text in a screenshot is extracted and read. Visual interpretation needs a second GPU node — we measured it, and it does not fit on one.

Not air-gapped.

It is network-isolated inside your Azure subscription with no outbound path from the workload. If your requirement is a physically disconnected environment, say so on the first call and we will tell you honestly whether this is the right shape.

Not an ungoverned model by default.

The base configuration ships with the model’s own safety behaviour intact. Purpose-selected configurations for security operations are a deliberate, scoped decision made with your risk function — not a switch we flip on day one.

Not a proof of concept.

It is a production environment. That is the point.

Is this for you?

The honest test, before anyone books anything

This fits when

  • Your people are describing genuinely proprietary process, judgement, or IP to an AI tool
  • Legal, risk, or client contracts limit what you can send to a third-party provider
  • A function in your business — security, fraud, compliance, investigations — keeps running into refusals
  • You are in financial services, insurance, healthcare, or manufacturing
  • You want a production environment, not another pilot that ends in a slide deck
  • You have an Azure subscription and someone who can approve GPU quota

This does not fit when

  • You need frontier-model reasoning at the very top of the benchmark tables
  • You need a physically air-gapped or classified deployment
  • You have no Azure footprint and no intention of building one
  • Fewer than 20 people would use it

The four prerequisites. Deployability comes down to four checks: GPU quota on your subscription, GPU capacity in your region, subscription permissions, and availability zone. Thirty minutes on a call answers all four — before any commitment, any paperwork, or any pitch.

Everything below is the evidence behind the argument at the top — written for whoever has to make the case to a risk committee. If you are already convinced, the button above is the whole next step.

How it compares

The five options, side by side

Time to productionCost modelWhat you discloseWhat you’re permitted to doModel control
Private AI in a Box5 days$25,000 fixed, then flat capacityNothing leaves your subscriptionYour policyOpen weights, swappable
Per-seat AI (ChatGPT Enterprise, Copilot)DaysPer seat, rises with headcount, vendor-setPrompts and documents, under vendor termsVendor usage policyNone
Managed AI service in your cloud regionWeeksPer token, rises with usageIn-region, vendor-operated serviceVendor usage policyVendor catalogue
On-prem AI appliance (Dell, HPE, Nutanix)MonthsSix figures capex, plus facilitiesNothingYoursVaries
Build it yourself3–6 monthsGPU + salariesNothingYoursFull

Building it yourself is the option most often underestimated: the licence is free, but GPU-experienced ML infrastructure engineers run $180,000–$250,000 and take three to six months to hire. And this is not either/or — plenty of clients run Copilot for productivity and the box for anything they will not disclose outside their boundary.

The evidence

The three arguments above, each in one fact

Full treatment in the linked articles — written for whoever has to make the case to a risk committee.

Disclosure

Trade secret protection can be forfeited by disclosure. In 2025 a US federal court held that voluntarily disclosing claimed trade secrets to a consumer-tier AI platform — one whose terms imposed no confidentiality obligation — defeated trade secret protection, regardless of who owned the output. Counsel now advise confining confidential material to deployments where retention, training, and access are technically controlled, not only contractually promised.

Read: The trade secret you gave away in a text box →

Permission

43.8%Refusal rate for system hardening requests across 2,390 real prompts from a sanctioned cyber-defence competition.

Malware analysis: 34.3%. Stating that you are on the blue team made refusal more likely, not less — because defenders and attackers use identical language. Attackers have adapted: malware is now seeded with content designed purely to trip an analysing model’s safety filter so the file goes unread.

Read: Your safety guardrail is a defender’s problem →

Cost

Four pricing changes in one month, none of them a sticker price. In June and July 2026: GitHub Copilot moved chat and agent features to usage-based credits, Copilot Cowork added per-task billing on top of the existing seat, volume discounts expired, and the Microsoft 365 base suites underneath every Copilot seat rose. You cannot build a three-year business case on a price list a vendor can restructure in a quarter. A dedicated GPU has one price, whatever the volume.

Read: Nobody raised a price and everybody’s bill went up →
Grow from the box

Where it goes after week one

Accelerators · from $25K each · 2–4 weeksConnect it to your data estate

Proved end to end: the workspace queried a live Snowflake schema and returned real consolidated P&L figures, writing the query itself.

Security operations configurationFor work general guardrails refuse

A model selected and deployed for defensive security work — malware analysis, threat intelligence, log and incident reasoning — where general-purpose guardrails refuse the task. Scoped to named users and named use cases, behind your identity provider, with full audit logging. Your risk function sets the policy; we build the environment that enforces it.

Box Managed · from $2,000 / monthSomeone else owns the upkeep

Model updates and evaluations, cost optimization, quarterly reviews. Priced below the 20–30% of a senior engineer this otherwise consumes.

Scale · 2–4 weeksPilot to enterprise

Same subscription, same box, same guarantees. The pilot is the first slice of production — not a throwaway.

Questions

What your security, legal, and procurement teams will ask

Where does our data go?
Nowhere. The workload sits behind private endpoints in your VNet with public access disabled. Retrieval, embeddings, web search, and transcription all execute in-cluster. The absence of an outbound path is verifiable in your own network logs — you do not have to take our word for it.
Our enterprise AI agreement already says they won’t train on our data. Isn’t that enough?
It may be. The questions worth putting to your counsel: what rights are reserved over aggregated or de-identified usage data, what happens if the terms are revised, what happens if the provider becomes party to litigation, and what proportion of your organization’s actual AI usage runs through that agreement rather than through personal accounts. This is not an argument that your provider is untrustworthy. It is an argument that the protection is contractual rather than architectural — and that some material warrants the architectural version, particularly where trade secret status depends on demonstrating technical control.
What are we agreeing to for the model, and where does it come from?
gpt-oss-120b, released by OpenAI as open weights under Apache 2.0 — permissive commercial use, no copyleft exposure, US-origin. OpenAI’s gpt-oss usage policy sits alongside the licence. There is no model vendor agreement to negotiate, and where provenance is a procurement question, the objection disappears without a debate.
Does this make us compliant with the EU AI Act, DORA, GLBA, or HIPAA?
No product makes you compliant. What changes is your evidentiary position: you hold the logs, you control the data path, and you can document what the system did with which inputs. Those frameworks require documentation, risk classification, human oversight, and audit trails — all substantially easier when the audit trail is generated inside your own boundary.
Who operates it after handover, and what happens if we stop working with you?
You operate it, by default — everything is infrastructure-as-code, documented, with a benchmark report and training delivered in week one. If we part ways, nothing happens: your subscription, your infrastructure, your data, your open-format assets. There is no proprietary layer holding it together. Box Managed exists if you would rather not own model updates and cost tuning, but the environment does not degrade without us.
We already have Copilot. Why this?
Different jobs. Copilot lives where your documents and meetings are. The box exists for the work you will not disclose outside your boundary, and for functions where a general-purpose usage policy blocks legitimate work.
Why Data Meaning

Somebody has to carry the operating burden

Sovereign AI platforms are arriving from every major vendor, and each one hands the operating burden to a partner. Where that bench is thin, it lands back on a customer who cannot carry it.

We are already inside your data estate — Snowflake, BI, analytics engineering, 20+ years and 200+ implementations. That is why your private model connects to the platforms your business actually runs on, and why the week ends in a signed acceptance document rather than a handover deck.

Microsoft partner · Governance practice built on Collibra and Alation · Delivery in English and Spanish across the Americas

Start with the prerequisites, not the pitch.

Four checks decide whether this is deployable in your subscription: GPU quota, GPU capacity in your region, subscription permissions, availability zone. Thirty minutes, before any commitment.