Metadata Governance Agent

Every column documented, classified, and evidenced.

Inside the catalog you already own. We deploy an agent into your own cloud project that reads your data estate and, for every column, writes a plain-English description, a sensitivity classification, a business domain and the observed value domain — as suggestions your stewards approve, with an auditable record of how each was reached.

civilian_labor_forceUS Census data
Owner wrote

Population in Civilian Labor Force. The number of civilians 16 years and over…

Agent wrote

People aged 16 and older who are either employed or actively seeking employment.

Rule

Not null · >= 0 · must not exceed the population aged 16 and over in the same row

Real output — the agent saw only name, type and sample values.

The problem

Your estate is undocumented, and it has stopped being harmless

Thousands of columns named cust_dt_ltv_fl were tolerable when only analysts used them. They asked a colleague, or worked it out from the data, and the knowledge lived in people’s heads.

An AI agent cannot ask a colleague. Pointed at an undocumented warehouse it either refuses or guesses wrong. Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls.

Source: Gartner press release, 25 June 2025. Third-party research, not a Data Meaning result.

TodayAfter one pass
No description, no owner — the column name is all anyone has
A plain-English description a business person can read
Nothing says whether a column holds personal data
Every column classified — personal, health, financial, or none
Quality rules that never get written — three minutes a column
Observed value domains become your data quality rules
Why it never gets fixed

The arithmetic is impossible

~50 hours

of machine time — one full pass over a 40,000-column estate, running in your project.

2,000 hours

of steward time for the same estate at three minutes a column — which is why it never gets done.

Nobody has 2,000 spare hours, so it goes on the roadmap, gets deferred, and the catalog stays empty. Every attempt to fix it by policy fails the same way: the people who know what a column means have the least time to write it down.

The obvious automation has not worked either.

Regex flags FLOAT_PROJECT_ID as personal data because it ends in _ID, and misses USER_PSEUDO_ID because “pseudo” reads as anonymous. Generic AI writes fluent text that restates the column name. What both lack is evidence — and a steward approving four thousand descriptions without it is not reviewing, they are rubber-stamping.

What lands

Four things for every column, in the catalog you already have

Description and classification

A plain-English description and a sensitivity class, for every column, written as proposals.

Business and value domains

Finance, customer, supply chain, HR — plus the observed domain that defines a valid value.

Rules your stewards approve

All four land in your catalog as proposals. Nothing is overwritten; approvals teach the agent.

How it runs

Beside your warehouse, learning from your stewards

1Read

Deployed into your Google Cloud project, it profiles the estate — never credentials, IDs or free text.

2Propose

Description, classification, domain and rule land in your catalog as proposals, each with its evidence.

3Review

Your stewards approve in the catalog you already own — Dataplex, or directly against BigQuery column metadata. No migration.

4Learn

Every approved description becomes an example the agent learns from on the next run.

Nothing leaves your project.

No long-lived credentials, no copy stored. We enrich the catalog you already own rather than replacing it. Connectors for Collibra and Alation are on the roadmap; we will tell you exactly where they stand before you buy.

What you are paying for

You are buying judgment, not infrastructure

Google’s — included, not billed by us
  • Vertex AI — the models that read your columns
  • BigQuery — your estate, and the evidence trail
  • Cloud Run — runs the agent, scales to zero
  • Dataplex — where metadata is written to by default
Ours — this is the product
  • Classification taxonomy — tuned to your glossary, not a template
  • Routing and prompt library — decides meaning, and when to say “unclear”
  • Profiler — samples safely; never reads credentials or free text
  • Evidence writer — the record behind every proposal, in your project
Commercials

One engagement, priced by the size of your estate

$18,000

under 10,000 columns · 2 weeks

  • 1 business domain mapped
  • 1 taxonomy tuning cycle
  • One 2-hour steward session
  • Run summary
Most common

$39,000

10,000–50,000 columns · 5 weeks

  • Up to 4 business domains mapped
  • 2 taxonomy tuning cycles
  • Two steward sessions
  • Governance readout

$75,000

over 50,000 columns · 10 weeks

  • Up to 8 business domains mapped
  • 3 tuning cycles, domain by domain
  • Enablement program and run-book
  • Executive readout and priorities

The bands differ by business domains, not column count — the agent itself is identical at every band. Additional business domains are $7,500 each. No recurring fee: the agent is yours afterwards and keeps running in your project. Google Cloud usage is billed to you directly and runs under $100 for a full pass over a 40,000-column estate; incremental runs are a few dollars, and it scales to zero between runs.

Why us

Everyone can document an estate. Nobody else can show you why

  • Your data never leaves your project — no long-lived credentials, no copy stored
  • Every entry carries its evidence — which model, what it looked at, how confident
  • We enrich the catalog you already own — nothing to migrate, nothing to replace
  • Nothing is overwritten — a column someone documented is left exactly as it is

What is your catalog coverage today?

If the honest answer is a number you report every quarter and it never moves, that is the conversation.