AI Economics Wiki
Data Moats
A data moat exists when a company has durable, lawful access to hard-to-replicate data and a repeatable system that turns it into measurably better customer outcomes, which in turn helps the company obtain more or better data.
Snapshot
What it is
A data moat is a durable competitive advantage created by exclusive or preferential access to useful data and the organizational ability to convert that data into better outcomes. The advantage compounds when better outcomes attract more usage, transactions, integrations, or expert feedback that competitors cannot obtain as quickly or cheaply.
What it is not
A large database, a customer list, public web data, raw product telemetry, or the sentence "our model gets smarter with every user" is not automatically a moat. Data must be lawful to use, relevant to a valuable decision, sufficiently accurate and representative, operationally usable, and hard to substitute. The product must also capture part of the value created.
Core mechanism
privileged observations -> reliable labels or outcomes -> product learning -> better customer result -> more qualified observations
Every arrow requires evidence. A loop that never reaches a verified outcome is data collection, not learning.
Founder rule
Measure the moat with a counterfactual. Compare the data-enhanced product with the best realistic product a competitor could build from public, licensed, customer-provided, or synthetic data. Then show the incremental customer value, not only a model metric.
Practical heuristic
Data moat strength ≈ access durability x data utility x learning yield x value capture x reinforcement
Score each factor from 0 to 1 for internal planning. This is a diagnostic, not an accounting measure or scientific law. The multiplicative form is deliberate: if the company has no usable rights, no learning, no customer value, or no reinforcing loop, the claimed moat approaches zero.
On this page11 sections
What exactly is a data moat?#
A data asset is information the company can access. A data advantage exists when that information lets the company perform a useful task better, faster, or more cheaply than a realistic alternative. A data moat is the durable version: competitors cannot readily reproduce the access, learning system, outcome improvement, and accumulation loop at an acceptable cost and within a commercially relevant time.
The term is most useful when it identifies a causal chain rather than an inventory. Ask four questions:
- What valuable decision or workflow does the data improve?
- What exact data and ground truth cause the improvement?
- Why can the company keep accessing and using them?
- Why will competitors remain behind after they see the opportunity?
If the answer stops at "we have more records," the company has described stock, not defensibility.
Data stock, data flow, and data system#
Three layers are often confused:
| Layer | Meaning | What a founder should prove |
|---|---|---|
Data stock | The historical corpus already held | Coverage, uniqueness, provenance, rights, quality, freshness, and replacement cost |
Data flow | New observations arriving through normal use or partnerships | Volume of eligible and useful new examples, label latency, churn, consent, and cost |
Data system | The process that turns raw observations into product improvements | Schemas, quality checks, labeling, evaluation, deployment, monitoring, governance, and cycle time |
A static stock can decay. A high-volume flow can be mostly duplicates or noise. A strong system can sometimes create more advantage from a smaller, better-labeled corpus than a competitor creates from a larger one.
Controlled experiments in DataComp-LM illustrate the point. The researchers compared extraction, deduplication, filtering, and data-mixing strategies under standardized training and evaluation. Their DCLM-baseline dataset reached 64% five-shot MMLU accuracy training a 7B model on 2.6 trillion tokens, surpassing the prior open-data state of the art with 40% less compute; the paper identifies model-based filtering as a key step. The lesson is not that every startup should train a language model. It is that dataset design can change performance and compute efficiency, so row count alone is a poor proxy for data value. (DataComp-LM paper)
Meta's Llama 3 work provides the complementary lesson that scale can matter when the data and training system support it. Meta reported pretraining on more than 15 trillion tokens and described substantial investment in data filtering, deduplication, and data-mix experiments. A founder citing this evidence should keep both halves: quantity can help, but curation and evaluation determine whether additional data are productive. (Meta, The Llama 3 Herd of Models)
Common sources of potentially defensible data#
| Source | Potential advantage | Main weakness to test |
|---|---|---|
Workflow exhaust | Observations created while users complete a job in the product | Activity may not contain the correct outcome or label |
Transaction outcomes | Purchases, repayments, returns, conversions, failures, or resolutions | Outcomes may be delayed, censored, or influenced by the product itself |
Sensor and device data | High-frequency operational measurements | Portability rights, hardware access, drift, and storage cost |
Expert labels | Diagnoses, classifications, rankings, annotations, or adjudications | Expensive, slow, inconsistent, and difficult to scale |
Customer-contributed data | Domain context supplied during onboarding or operation | Rights may be tenant-limited and customers may export the same data to rivals |
Licensed or partner data | Specialized corpus unavailable in public sources | Non-exclusivity, renewal risk, price increases, and use restrictions |
Synthetic data | Coverage of rare cases, privacy support, or low-cost augmentation | Model errors and assumptions can be amplified; independent ground truth is still needed |
Public data | Cheap baseline and broad coverage | Rivals can usually access the same source |
The UK Competition and Markets Authority distinguishes public, synthetic, third-party proprietary, and first-party proprietary data in its foundation-model market review. It reports mixed views on how important first-party proprietary data are for different tasks and notes that no single data source was then regarded as indispensable — a useful antidote to the claim that any private dataset is automatically strategic. (CMA, AI Foundation Models technical update)
Why do data moats matter to founders?#
They can improve product performance without competing only on model access#
Foundation models, APIs, open-source models, and standard software components make many baseline capabilities widely available. A startup can still differentiate through the context, outcomes, exception cases, and workflow signals a general-purpose supplier does not possess. The defensible layer may be a domain-specific retrieval corpus with verified provenance, a ranked set of expert decisions and their eventual outcomes, an error taxonomy and correction history, rare failure cases underrepresented in public data, or a cross-customer benchmark customers cannot build individually. The model may be replaceable while the data pipeline, evaluation set, and embedded workflow are not.
They change pricing only when the data changes customer economics#
A higher accuracy score is not a pricing metric. Connect the improvement to an economic outcome:
Incremental customer value = eligible decision volume x absolute error reduction x value per avoided error
For a revenue use case, replace avoided error cost with incremental contribution profit per successful decision. Then test capture:
Value capture rate = incremental price attributable to the improvement / incremental customer value
This supports value-based pricing, outcome pricing, or a higher usage tier. It also exposes a common failure: the data may improve the model but not enough to alter willingness to pay, retention, conversion, or cost to serve.
They influence fundraising and valuation narratives#
Investors often hear "more customers create more data, which makes the product better." A credible pitch replaces that circle with evidence: the scarce observation and the right to use it; the ground-truth label and the time to obtain it; a learning curve against incremental eligible data; a holdout or prospective test showing the improvement survives outside the training set; the customer KPI affected; the fully loaded cost; and the reason competitors cannot buy, request, or infer an equivalent dataset.
Do not inflate the story. In a 2024 enforcement action, the SEC charged investment advisers for "AI washing" — making false or misleading statements about their use of AI. The direct rules applied to regulated advisers, but the founder lesson is simple: claims about proprietary AI and customer-data learning should match the deployed capability and supporting evidence. (SEC, AI-washing enforcement action)
They affect product design and go-to-market sequencing#
If defensibility depends on labeled outcomes, the product must be designed to capture them. A free tool that generates many anonymous queries may produce less useful learning than a narrow paid workflow that records the action taken and the eventual result. That changes early choices: target customers who hit the problem frequently, pick use cases where outcomes arrive in days or weeks, negotiate data-use rights before integration, make corrections part of the workflow, and preserve an evaluation holdout instead of training on every record. The goal is not "collect everything" but "capture the minimum data needed to improve a valuable decision, with clear rights and a measurable return."
How do you test a data moat? The five-link chain#
1. Durable, lawful access#
Identify the exact source and the operative rights. Ownership is not the only path; a company might rely on a license, customer authorization, a partner agreement, public-domain status, or a documented legal basis. What matters commercially is whether the access and permitted uses survive customer churn, contract renewal, product changes, deletion requests, and regulatory review. Build a rights matrix for each material source covering who supplied the data, what the company may do with it, at what scope, retention limits, what happens on termination, whether a customer can give the same data to a rival, and whether provenance can be proven.
Do not treat a privacy-policy edit as a substitute for permission. The FTC has warned that quietly changing terms to permit new AI-training uses may be unfair or deceptive, and notes prior remedies have included deletion of models built with unlawfully obtained data. (FTC, changing terms for AI data use) For personal data in the EEA, the European Data Protection Board's Opinion 28/2024 states a model trained on personal data is not automatically anonymous and describes a case-specific assessment. Involve qualified counsel for the actual jurisdiction; the commercial point is that uncertain or revocable rights weaken the moat. (EDPB, Opinion 28/2024)
2. Relevant observations and credible ground truth#
The data must predict or explain the outcome the customer values. A click is not the same as satisfaction. An agent's draft is not the same as an accepted answer. A clinician opening a recommendation is not the same as a correct diagnosis. Define the unit of observation, the target label, the label source, the label latency, the coverage, the noise, and the representativeness. Count economically useful examples, not raw events: one thousand duplicates from one configuration may add less value than ten verified failures from rare configurations.
3. Repeatable data operations#
Raw data do not improve a product by themselves. The company needs a system for ingestion, normalization, access control, quality checks, labeling, lineage, versioning, evaluation, retraining or retrieval updates, deployment, and monitoring. NIST's Generative AI Profile recommends documenting data origin and lineage, testing data flows, and evaluating output against known ground truth — controls that make the claimed learning system auditable. (NIST AI 600-1, Generative AI Profile) Track:
Learning-loop cycle time = median days from eligible event to validated production improvement
Label yield = validated labels / eligible production events
A startup can receive millions of events while producing almost no labels or deployable improvements. That is a broken loop hidden by volume.
4. Measurable learning and product impact#
Use a versioned evaluation set and a realistic baseline — a public or open-source model, the same model without proprietary context, the customer's current rules or workflow, a competitor's product, or the previous production version. Measure absolute improvement, not only relative. Moving an error rate from 10% to 8% is a 20% relative reduction but a 2 percentage-point absolute reduction, and the economic formula uses the absolute change. Useful evidence includes learning curves against cumulative unique labeled examples, ablations with and without each data source, temporal and customer holdouts, and prospective pilots. Do not train on the test set, and do not treat corrections caused by a bad model as independent proof the model is improving.
5. Value capture and reinforcement#
First, check whether the company captures enough value to fund the loop:
Incremental contribution margin = incremental revenue - collection - labeling - model/retrieval - monitoring/support cost
Second, measure whether success produces more useful data:
Reinforcement rate = additional validated examples caused by product adoption / period
The word caused matters. If the company would receive the same dataset from a partner regardless of product success, that source may be valuable but it is not a product-driven flywheel. The CMA's review cautions that user feedback is not automatically fed into a model, may require rigorous manual review, and that retraining can be expensive. Feedback loops can become barriers to entry, but the loop must actually operate. (CMA, AI Foundation Models technical update)
Worked example: predictive maintenance#
A startup sells diagnostic software to industrial maintenance operators. A widely available baseline predicts the correct failed component on the first attempt for 62% of work orders. Using 80,000 lawfully pooled, normalized, and technician-validated historical work orders, the data-enhanced system reaches 75% on a time-based holdout from the same target environment.
For one large customer, assume 20,000 eligible diagnostic decisions per month, each avoided wrong first diagnosis saves ~$24 in repeat labor and parts handling, the startup charges a $9,000 monthly premium, and incremental compute, QA, labeling, and monitoring cost $2,400 per month.
Value created:
Avoided errors = 20,000 x (0.75 - 0.62) = 2,600 per month
Monthly customer value = 2,600 x $24 = $62,400
Value captured:
Value capture rate = $9,000 / $62,400 = 14.4%
The customer keeps about $53,400 of modeled monthly value. That is a stronger pricing claim than "accuracy improved by 21%," the relative increase, which does not describe money saved.
Contribution economics:
Incremental monthly contribution margin = $9,000 - $2,400 = $6,600
If the reusable ingestion, labeling, and evaluation system cost $132,000 to build, simple payback is $132,000 / $6,600 = 20 months with one customer, or $132,000 / (5 x $6,600) = 4 months if the same improvement transfers to five comparable customers. That transfer assumption is the critical moat test. If each customer's equipment taxonomy requires an independent dataset and another $132,000 build, the startup has customer-specific switching costs but no cross-customer data scale advantage. Note the holdout result establishes association under the test design, not permanent causality; a prospective controlled rollout gives stronger evidence.
Key Facts
Curation beats raw scale
DataComp-LM's DCLM-baseline reached 64% five-shot MMLU training a 7B model on 2.6 trillion tokens with 40% less compute than the prior open-data state of the art, showing dataset design — not row count — drives value
DataComp-LMScale still needs a system
Meta pretrained Llama 3 on more than 15 trillion tokens alongside heavy filtering and data-mix work — quantity helped only because curation and evaluation did too
Meta, Llama 3Data problems are pervasive
In interviews with 53 practitioners on high-stakes AI, Google researchers found 92% experienced one or more "data cascades" — compounding downstream failures from upstream data issues
Google ResearchRaw-data exclusivity is eroding
The EU Data Act began applying on 12 September 2025, giving users rights to access and share raw data generated by connected products — pushing defensibility toward derived insight and workflow
European CommissionOverclaiming has teeth
The SEC's 2024 "AI washing" actions penalized advisers for misstating their use of AI and client data — claims must match the deployed capability
SECCommon mistakes and misinterpretations#
"We have millions of data points"#
Rows are not units of information. Duplicates, correlated events, missing labels, stale observations, and dominant easy cases can make a large corpus economically thin. Report unique eligible cases, coverage of the target population, verified outcomes, and marginal learning.
"Customer usage is ground truth"#
Usage reveals behavior, not correctness or value. Users click the first answer, accept a default, or abandon a task silently. Capture the later outcome or expert adjudication when that is the relevant target.
"Every user makes the model smarter"#
Feedback does nothing until it is captured, permitted, validated, incorporated, evaluated, and deployed. Report cycle time and deployment yield. Research on data cascades found upstream data problems compounding downstream in 92% of high-stakes projects studied. (Google Research, Data Cascades)
"Proprietary means defensible"#
A private dataset can be small, irrelevant, nonexclusive, exportable, or legally unusable. A rival may buy an equivalent license, ask customers for the same export, generate synthetic substitutes, or solve the task with a stronger general model.
"We can fix the rights later"#
Later consent, licensing, or anonymization may be expensive or impossible. The U.S. Copyright Office's 2025 generative-AI training report describes fair use as fact-specific and notes illegal access can weigh against it. Establish provenance and obtain advice before the dataset becomes embedded in the product. (U.S. Copyright Office, Part 3)
When this breaks: limitations of a data moat#
The baseline catches up. General models, open-source models, public datasets, synthetic data, or a vendor feature can erase the lift. Re-run the counterfactual against today's best substitute, not the system used before the startup existed.
Marginal learning approaches zero. The dataset may cover common cases quickly; new records then add duplicates rather than information. If the learning curve flattens while acquisition and labeling costs continue, the moat stops compounding.
The product creates its own biased training distribution. Once a model changes which offers users see or which cases humans review, later data are no longer independent of the model. Preserve exploration, audits, and independent outcome measurement.
Rights narrow or disappear. Customers can revoke permission, contracts end, regulators require deletion. The EU Data Act's portability rights, applying since 12 September 2025, can also weaken exclusivity for IoT startups — so the defensible layer may be the derived insight and service rather than raw device data. (European Commission, EU Data Act)
The loop is customer-specific. If data cannot be pooled or the learning does not transfer, each implementation starts near zero. That may create a services business or high switching costs, but it is not a scalable cross-customer data network effect.
A founder's 90-day validation plan#
Days 1-30 — define the claim. Name one customer decision and one economic outcome. Inventory sources, rights, retention, deletion, and portability. Define the observation, label, eligibility rule, and realistic competitor baseline. Freeze a time-based or customer-based holdout. Estimate collection, labeling, compute, monitoring, and support cost.
Days 31-60 — test learning and transfer. Build a learning curve using unique validated examples. Run ablations with and without the claimed proprietary source. Slice performance by customer, cohort, time, and edge case. Test a customer or environment excluded from training.
Days 61-90 — test economic reinforcement. Run a prospective pilot against the existing workflow where feasible. Convert absolute performance lift into customer value. Test a price or expansion offer tied to that value. Reconcile incremental revenue with fully loaded data and serving cost. Write the strongest competitor substitution plan and update the moat score.
At day 90, choose among four honest conclusions: (1) moat forming — durable access, transferable learning, measurable value, and reinforcement are visible; (2) useful data advantage — product lift exists, but durability or accumulation is unproven; (3) customer-specific advantage — the data improve each deployment but do not transfer; or (4) no material advantage yet. The last three are not failures — they prevent a false moat story from distorting product, pricing, and fundraising decisions.
Frequently asked questions
01Is a big proprietary dataset a moat?
Not by itself. A moat needs durable lawful access, data that improves a valuable decision, a working learning system, measurable customer value, and a reinforcing loop competitors cannot cheaply reproduce. Size without those is stock, not defensibility.
02How is a data moat different from a network effect?
A network effect makes the product more valuable as more users join. A data moat makes the product perform better as it accumulates useful, labeled outcomes. They can coexist, but the mechanisms and evidence differ — test each separately.
03How do I prove the data actually improves the product?
Use a versioned evaluation set and a realistic baseline, show learning curves against unique labeled examples, run ablations with and without the proprietary source, and validate on temporal and customer holdouts. Report absolute improvement, not just relative.
04Does "more data" always help?
No. Additional data can be redundant, shifted, mislabeled, adversarial, or generated by your own system. Curation, coverage of rare cases, and label quality often matter more than volume.
05What's the single most common founder mistake here?
Pitching data volume instead of the causal chain. Investors and buyers care about eligible cases, label yield, measured lift, customer value, gross-margin effect, and how fast the loop turns — not the row count.
Related concepts#
- TAM, SAM, SOM: estimate the reachable market that can generate the required data density.
- Startup Valuation Methods: connect defensibility claims to valuation without treating data volume as an asset appraisal.
- Bootstrapping vs. Venture Capital: decide whether the cost and speed of building the data system justify external capital.
- Value-Based Pricing: price the outcome created by the data advantage.
- Economic Value Estimation: quantify avoided costs or incremental profit in pilots.
- Customer Use Cases: start with the decision and outcome before designing the dataset.
- Ideal Customer Profile: target customers with sufficient event frequency and feedback quality.
Sources#
- DataComp-LM: In Search of the Next Generation of Training Sets for Language Models (2024)
- Meta, The Llama 3 Herd of Models (2024)
- Google Research, Data Cascades in High-Stakes AI (2021)
- UK Competition and Markets Authority, AI Foundation Models technical update report (2024)
- NIST, AI Risk Management Framework: Generative AI Profile (NIST AI 600-1, 2024)
- European Data Protection Board, Opinion 28/2024 on AI models and personal data
- Federal Trade Commission, Quietly Changing Terms for AI Data Use Could Be Unfair or Deceptive (2024)
- U.S. Copyright Office, Copyright and Artificial Intelligence, Part 3: Generative AI Training (2025 pre-publication version)
- European Commission, EU Data Act gives users control over connected-device data (2025)
- SEC, AI-washing enforcement action (2024)
This page is educational and does not constitute legal or financial advice. Data rights, privacy, and copyright questions are fact- and jurisdiction-specific; consult qualified counsel before relying on a specific data source.
Author
Dr. Sarah Zou
Independent economist · EconNova
Commercial strategy for technical products, with a focus on pricing, unit economics, and the operating choices behind the model.
About SarahTopics
Cite this page
Canonical URL
https://sarahzou.com/wiki/ai-economics/data-moatsSuggested citation
Zou, S. (2026). Data Moats: How Proprietary Data Becomes Defensible. In AI Economics. Pricing & Monetization Wiki. https://sarahzou.com/wiki/ai-economics/data-moats
Open license
Reuse with attribution
This content is available for reuse. When referencing or republishing it, please credit Dr. Sarah Zou and link back to the original source.
Licensed under Creative Commons Attribution 4.0 International. You may share and adapt the material with appropriate credit.