A data marketplace turns governed data from a cost center into a revenue and efficiency engine — an internal exchange where data products are listed, discovered, and consumed with the same discipline as any commercial marketplace. For enterprises, the marketplace is how data assets get reused at scale: teams stop rebuilding the same pipelines, data products reach more consumers, and monetization — internal chargeback or external sale — becomes possible. This article explains the governance model that makes data marketplaces work without creating new risks.
What is a data marketplace, and what is it not?
A data marketplace is a governed exchange where data products are published by the teams that produce them and consumed by the teams that need them, with a consistent mechanism for discovery, access request, entitlement, and usage measurement. That definition has four parts, and most initiatives that claim to have built a marketplace have implemented only the first.
Discovery is the catalogue: what exists, what it contains, how fresh it is, who owns it. Access request is the workflow that gets a consumer from "I found it" to "I can query it" without a series of emails. Entitlement is the permission model that determines who may see which rows and columns, applied automatically rather than by manual provisioning. Usage measurement is the instrumentation that records who consumed what, how often, and for which purpose — which is what makes governance auditable and chargeback possible.
What it is not: it is not a data catalogue with a nicer interface, not a shared folder with a naming convention, and not a data lake with an access request form bolted on. Those artefacts answer "where is the data?" A marketplace answers "can I use this, under what terms, and how quickly?"
The distinction matters commercially, because the difference between the two is the difference between a cost centre and a revenue line. A catalogue reduces the time analysts spend searching. A marketplace enables an enterprise to sell data products to partners, to charge internal consumers in a way that funds data quality, and to demonstrate to a regulator exactly who accessed what.
It is also worth separating internal from external marketplaces. An internal marketplace is about reuse and cost allocation; its success metric is consumption breadth. An external marketplace is about monetisation; its success metric is revenue, and it introduces contractual, licensing, and delivery obligations that most internal data teams are not structured to meet. The governance requirements differ substantially, and conflating the two is a common cause of stalled programmes.
Why do internal data marketplaces fail?
Supply without demand, or demand without supply. Marketplaces are two-sided, and both failure modes are common. Publishing teams produce data products nobody consumes because they were specified by producers rather than by consumers. Or consumers arrive, find nothing usable, and go back to requesting extracts from the data team. Seeding supply from demonstrated demand — the ten extracts the data team is asked for most often — is the reliable way to break the deadlock.
No producer incentive. Publishing a well-documented, well-governed data product is real work, and if the producing team receives nothing for it, they will do the minimum. The incentive does not have to be money; it can be consumption metrics that count toward team objectives, or a chargeback model that funds the producer's roadmap. What does not work is asking for it as a favour.
Governance as a gate rather than a service. If requesting access takes three weeks and two approvals, consumers route around the marketplace entirely — usually by asking someone to export the data to a spreadsheet. The marketplace must be the fastest path to data, not the most compliant one.
Quality debt inherited silently. A data product that is published without a quality contract transfers the debugging burden to every consumer. Once two consumers have been burned, the marketplace's reputation is set, and rebuilding it is far harder than establishing it correctly the first time.
Measuring the wrong thing. Counting published data products rewards volume. The metrics that matter are active consumers, repeat consumption, and time-to-first-query for a new consumer.
What are the models for data monetisation?
There are four distinct models, and they carry very different operational requirements. Choosing one late, or drifting between them, is a common source of confusion.
Internal chargeback or showback. Business units are charged for the data they consume, either as real budget transfers or as visible cost attribution. The purpose is behavioural: when consumers see the cost, they stop requesting redundant extracts and producers get funded for quality. This is the lowest-risk model and the right starting point for most enterprises.
Data-as-a-product sold externally. Packaged data products licensed to customers or partners — benchmark datasets, market intelligence, aggregated industry metrics. This requires contractual infrastructure, service levels, delivery mechanisms, and usually a dedicated commercial function. It is genuinely a new business line, not an extension of the data team.
Data sharing for reciprocal value. Exchanging data with partners, suppliers, or industry consortia where the benefit is access rather than payment. Common in supply chain, insurance, and financial services. The governance burden is high because you are accountable for how the counterparty handles what you send, and the commercial benefit is diffuse.
Enabling revenue rather than generating it. Using data to improve a product the customer already buys — better recommendations, faster onboarding, more accurate risk assessment. This is where most enterprises capture the most value, and it requires no monetisation infrastructure at all.
The strategic mistake is treating external monetisation as the goal because it sounds more ambitious. For most organisations, the value in a data marketplace comes from eliminating duplicated effort and improving decisions, and the monetisation case should be built on evidence of consumption, not on ambition.
How should governance be designed for a marketplace?
Marketplace governance has to do three things simultaneously: protect the data, enable the consumer, and leave an audit trail. Designs that optimise one at the expense of the others fail.
Classify at publication, not at consumption. Every data product is classified when it is published — public, internal, restricted, confidential — and the classification travels with it. Consumers then request access to a classification tier rather than negotiating per dataset, which dramatically reduces friction. Trying to classify at the point of each access request is what produces three-week waits.
Automate entitlement. Permissions should be derived from the consumer's role and purpose, applied by the platform, and revocable centrally. Manual provisioning creates two failure modes: delays that push people around the system, and orphaned access that nobody remembers granting.
Make purpose explicit. Consumers declare why they need the data, and the declared purpose is recorded and auditable. This is increasingly a regulatory requirement rather than a best practice, and it is also operationally useful — it tells you which purposes are underserved and which data products are being used for something they were never designed for.
Enforce at query time, not at export time. Entitlement applied to a row or column at the moment of query means consumers can self-serve safely. Entitlement applied only when data leaves the platform means the only safe pattern is a mediated extract, which reintroduces the data team as a bottleneck.
Build for revocation. Someone will eventually need to know who has ever touched a given dataset, and to cut access across every consumer at once. Design the entitlement model as a graph of grants rather than a list of permissions, so that both questions are answerable in minutes rather than weeks.
What does the product model look like — who publishes, who consumes?
Treat every data product as a product with an owner, a roadmap, and users. That framing, borrowed from data mesh, is what makes a marketplace sustainable rather than a dumping ground.
Each data product should carry, at minimum: a named owner with a backup; a plain-language description of what it contains and what it is good for; a schema with column-level descriptions; a freshness and availability commitment — how current the data is and how often it updates; a quality contract covering completeness, uniqueness, and known caveats; a classification; and a support channel. A product missing any of these generates support burden rather than value.
On the consumer side, the marketplace needs three things to convert interest into usage. A sample or preview, so a consumer can determine suitability without requesting access. A self-service request flow with a published service level — ideally automated for non-restricted tiers. And usage documentation: not just what the columns mean but what the data is known to be wrong about, which is the single most valuable thing a producer can publish and the thing most often omitted.
The operating rhythm matters as much as the artefacts. Producers need a regular review of consumption, quality incidents, and consumer feedback, with the authority to deprecate. Consumers need a visible route to request new products. Without both, the marketplace becomes a one-way publication channel and supply drifts away from demand within a year.
How do you price and charge for data products?
Pricing internal data is uncomfortable, which is why many programmes avoid it and then wonder why producers lack incentive. The mechanics are simpler than they appear if the objective is behavioural rather than revenue-maximising.
The standard approach is cost recovery with a consumption driver: allocate the platform's running cost, plus an allocation of the producing team's effort, across consumers in proportion to usage. The unit can be compute consumed, rows scanned, queries issued, or seats — pick the one that correlates most closely with the behaviour you want to influence, and keep it to a single unit, because multi-factor allocation models are impossible to explain and therefore impossible to act on.
Start with showback before chargeback. Publish the cost without moving budget for two or three quarters. This surfaces the behavioural response — consumers consolidate redundant extracts, producers see which products justify investment — without triggering the political fight that real budget transfers provoke. Move to chargeback only once the numbers are trusted.
Price external data products differently and more carefully. Value-based pricing, tiered by volume or freshness, usually outperforms cost-plus, because the cost of production bears no relation to the value a customer derives. But external pricing requires competitive awareness and contractual flexibility that internal allocation does not, and it should be owned by a commercial function rather than by the data team.
One warning: do not price to recover transformation cost from the first consumer. Early data products have one consumer and full cost, which produces a price that guarantees there is never a second one. Spread cost across expected consumption at maturity, and accept a subsidy during the adoption curve.
How do you measure whether the marketplace is working?
Measure both sides of the market and the friction between them.
Supply-side metrics: number of published products meeting the full documentation standard, coverage of the highest-demand domains, and the proportion of products with a freshness commitment actually being met.
Demand-side metrics: monthly active consumers, repeat consumption rate — the share of consumers who come back — and consumption concentration, which tells you whether a handful of products carry the marketplace or whether value is broadly distributed.
Friction metrics are the ones that predict whether the marketplace survives: median time from access request to granted access; median time from first visit to first successful query; and the share of requests that are abandoned before completion. A marketplace with growing supply and flat active consumers almost always has a friction problem rather than a content problem.
Outcome metrics are what justify the programme: reduction in duplicated datasets, reduction in ad-hoc extract requests to the central team, and — where external monetisation is in scope — revenue and pipeline attributable to data products. Track the first two from the start; they are achievable within two quarters and they fund the third.
Report supply and demand together. A marketplace that reports only publication counts will optimise for volume, and volume is the metric that is easiest to game and least correlated with value.
What should enterprises do first?
Find the demand that already exists. Every data team maintains an informal queue of recurring extract requests; that queue is the specification for your first ten data products. Converting the ten most frequent requests into properly documented, properly governed, self-service products is the fastest route to demonstrated value, because you are serving demand that has already proven itself.
Build the minimum governance that makes self-service safe: classification at publication, automated entitlement by role and purpose, query-time enforcement, and an access log. Do not attempt a comprehensive data classification programme across the whole estate first — classify what you publish, and expand coverage as the catalogue grows.
Pick one domain to prove the model end to end, preferably one where the producing team is enthusiastic and the consuming demand is high. A single domain taken from publication through to chargeback teaches you more than a cross-domain pilot that never reaches entitlement automation.
Start with showback rather than chargeback, and publish the numbers where both producers and consumers can see them. The transparency does most of the behavioural work, and it builds the trust needed before any actual budget moves.
Finally, set the incentive for producers before launch rather than after. If nothing in a producing team's objectives rewards publication, consumption, and quality, the marketplace will be populated with undocumented extracts and the programme will be remembered as a catalogue project that did not work.
What Are the Key Takeaways on Data Marketplaces?
A data marketplace is discovery, access request, entitlement, and usage measurement — not a catalogue with a better interface. The governance model is what separates a cost centre from a revenue line.
- Seed supply from demonstrated demand: the ten extracts your data team is asked for most often.
- Classify at publication, automate entitlement by role and purpose, enforce at query time, and design for revocation.
- Treat each data product as a product: owner, description, schema, freshness commitment, quality contract, and known caveats.
- Start with showback before chargeback, use a single consumption driver, and do not recover full cost from the first consumer.
- Measure supply and demand together, and watch friction metrics — time-to-access predicts survival better than publication count.
- Set producer incentives before launch, not after the marketplace fails to populate.