Methodology
How the preliminary estimate works
The online assessment combines self-reported answers with fixed model assumptions to produce a preliminary score and, where supported, an indicative range. It uses neither buyer bids nor completed-deal comparisons. This page explains its rules and limits.
- Inputs
- 26 multiple-choice questions about one dataset
- Factors
- 16, each scored 0 to 100
- Outputs
- Score, range, confidence, and rights risk
- Model
- epistemic-labs-valuation-2026.10.2
01
What the score measures
Volume alone does not establish value. The model uses your answers to estimate potential usefulness to an AI developer and readiness for licensing. Its heaviest value factor is the reported chain from input to human decision to real-world outcome. No records or agreements have been inspected.
Value to AI buyers
How useful the model assumes the data could be to an AI developer. These factors make up 70% of the score.
- 01
Decision to outcome pairing
Reported links from a situation to a human decision and its outcome. The model gives this structure more weight when estimating potential AI usefulness.
- 02
Uniqueness
How much comparable public data you report. The model assumes less public availability means greater distinctiveness.
- 03
Expert involvement
Reported human decisions, corrections, and demonstrations. The model assumes these preserve useful judgment signals.
- 04
Labels and outcomes
Reported links to outcomes, ratings, rankings, or feedback. The model gives more weight to consistent links.
- 05
Domain scarcity
Reported specialization combined with a fixed industry assumption about scarcity. This is not measured market availability.
- 06
Cost to reproduce
Your estimate of the time and expertise needed to recreate comparable records. Greater reported difficulty scores higher.
- 07
Scale
Reported record count and media volume. The model gives scale less weight than decision and outcome structure.
- 08
Historical depth
Reported years of history. The model assumes longer histories may preserve more change over time.
- 09
Assumed buyer fit
A fixed industry prior combined with reported AI signals. This is a model assumption, not observed buyer demand.
- 10
Freshness
Reported update frequency. The model assumes frequent updates could support ongoing deliveries, subject to verified rights.
- 11
Modality mix
Reported kinds of data, such as text, images, audio, or sensor readings. The model gives additional weight to a mix of modalities.
Ability to transact
How ready the answers suggest the data is for licensing and delivery. These make up 30%.
- 01
Rights confidence
Reported licensing confidence, confidentiality terms, and customer or vendor data. The model discounts unresolved issues without verifying rights.
- 02
Privacy simplicity
Reported personal information combined with fixed industry assumptions about privacy work. This does not establish compliance.
- 03
Provenance
Reported internal creation and third-party material. The model assumes internal creation makes origin easier to trace.
- 04
Structure and cleanliness
Reported structure and outcome links. The model assumes structured, linked records need less preparation.
- 05
Exclusivity potential
Distinctiveness and origin as proxies for possible licensing flexibility. The model does not establish exclusivity rights or a premium.
Factors are listed in order of influence within each group. Higher scores reflect more favorable model assumptions, not verified rights, quality, or demand.
02
From score to range
Indicative aggregate licensing value across one or more licenses over roughly the first two to three years, not a sale price. The range combines reported scale, assumed usefulness for AI work, assumed buyer demand, and reported licensing readiness. It is not a price, offer, bid, or appraisal. Rights, uniqueness, quality, scale, freshness, provenance, permitted uses, and buyer demand need separate review.
- 01
Volume sets a starting point
The model assigns each record band a dollar starting point that rises more slowly than record count. Confirmed volumes may receive higher starting points for terabyte-scale media. Unknown record counts use no more than the 10,000 to 100,000 band, with no media uplift. These are model assumptions, not observed license prices.
- 02
Quality multiplies it
The value factors combine into an index that multiplies the starting point. The model gives more weight to paired decisions and outcomes in specialized domains than to generic record volume. This does not establish what a buyer would pay.
- 03
Demand assumptions adjust it
Fixed assumptions about each industry and the reported signals adjust the estimate. They are not live buyer requests or evidence of demand for your dataset. Buyer fit must be tested separately.
- 04
Rights and readiness discount it
The readiness factors apply 35% to 100% of the modeled amount. Unclear rights, personal information, and third-party material can reduce it. This adjustment is a scoring rule, not a determination that licensing is permitted.
- 05
Freshness adjusts it
The model increases the amount by about 15% when updates are reported as daily or continuous. This assumes possible ongoing deliveries; it does not establish a subscription premium or licensing rights.
- 06
High rights risk reduces it further
High rights risk applies 50% of the amount, multiplied by the capped score divided by the uncapped score. The range is closed and limited to 25% of the record-band ceiling, rounded down. The result lists the applied caps and policies.
- 07
Uncertainty sets the width
The high end starts at about 3.5 times the low end. Each "Not sure" answer lowers the low end, up to a ratio of 7, without raising the high end. Bounds are rounded to familiar figures within the applicable ceiling. Unknown volume also keeps the range closed.
03
Ceilings, and when we withhold a number
Each record band has a display ceiling set by the model. For known volumes without high rights risk, when the calculated range exceeds that limit, we show the ceiling with a plus sign. The + is a limit of this calculator, not evidence that a buyer would pay that amount or more. These ceilings are assumptions, not market benchmarks.
| Records | Highest figure shown |
|---|---|
| Under 10,000 | $150K |
| 10,000 to 100,000 | $400K |
| 100,000 to 1 million | $1M |
| 1 to 10 million | $2.5M |
| 10 to 100 million | $4M |
| Over 100 million | $5M |
| Not sure | $400K |
We show the score but no figure when
- You indicated the business likely does not hold the rights to license the data.
- The data is mostly simulated or test activity, outside this calculator's focus on operating records.
- The overall score is below 35.
- The modeled high end would fall below $25K, this calculator's minimum display threshold, not a measured cost of licensing.
Withholding a range reflects this calculator's scope and input rules, not a judgment on your business or on all synthetic data. The model also caps the score for some answers: likely lacking rights caps it at 45, mostly simulated data at 50, and extensive or sensitive personal information without confident rights at 65.
04
Confidence and rights risk
Confidence is never high. Every answer is self-reported and nothing has been inspected, so the best we will say is moderate. It drops to low when the record count is unknown, three or more answers are “Not sure”, or rights risk is high. These labels describe model rules, not statistical probabilities or validation against completed deals.
Rights risk reads the rights and privacy factors together with confidentiality terms and third-party material. It is a prompt for the conversation with your counsel, not a legal conclusion. Epistemic Labs is not a law firm and does not give legal advice.
05
What a questionnaire cannot know
- Whether your contracts, privacy notices, and customer terms permit the proposed uses. Your counsel must review those rights; ownership alone does not establish them.
- How accurate, complete, and consistently recorded the data is, and whether its collection history is documented.
- Whether outcomes really join to records at the volume described, and how much work the join takes.
- Which buyers need this exact structure right now, and what they have already licensed elsewhere.
- What de-identification and redaction will cost, and how much useful signal they remove.
- Market conditions at the time a deal is negotiated.
Diligence under NDA may refine or rule out an opportunity. Any license fee is negotiated separately after rights, privacy, quality, and buyer fit have been reviewed. Every assessment records the model version that produced it, so estimates can be re-run when the model is recalibrated.
See where your data lands
About five minutes, dataset metadata only. You will see a preliminary score, an indicative range where supported, and what would need to be verified.