Public Data and Open Information Toolkit: Sources, Licences, Derived Rights, and Attribution
By Casey Scott McKay ·
Public data is the most misunderstood input in modern product development, because "publicly available" describes an access condition and says nothing about what may lawfully be done with the material. This toolkit assembles the working material for building on it. It distinguishes the four categories that get collapsed into one — works outside copyright, works licensed openly, works accessible under access legislation, and works merely visible on a website — each carrying different obligations. It covers the licence families and their share-alike and attribution terms, the derived-data analysis that determines whether a product is encumbered, the access and scraping questions that turn on contract rather than copyright, and the diligence a purchaser will run on a product built from other people's information.
IP and Technology > General IP | Toolkit | Published 30 December 2025 - Updated 24 January 2026 | Casey Scott McKay - marksy.us
Summary. Public data is the most misunderstood input in product development, because "publicly available" describes an access condition and says nothing about what may lawfully be done with the material. This toolkit distinguishes the four categories that get collapsed into one — works outside copyright, works licensed openly, works accessible under access legislation, and works merely visible on a website — each carrying different obligations. It covers licence families and their share-alike terms, the derived-data analysis, the access and scraping questions that turn on contract rather than copyright, and the diligence a purchaser will run.
Keywords: public data · open data licences · government works · freedom of information · derived data rights · attribution obligations · share-alike terms · database rights · web scraping · terms of use · refresh obligations · provenance records · public records privacy · data product diligence · source verification
Start Here
The phrase "publicly available data" conceals four distinct legal positions, and almost every failure in this area is a confusion between them.
Material outside copyright. Works of the federal government are generally not subject to copyright under 17 U.S.C. § 105. Facts are never protected, per Feist Publications v. Rural Telephone Service. Works whose term has expired are in the public domain. This material may be used freely, subject only to any contractual conditions attached to how it was obtained.
Material licensed openly. A dataset published under an open licence is copyrighted material made available on terms. Those terms bind, and they commonly require attribution and sometimes require that derivative works be released under the same licence.
Material accessible under access legislation. Records obtainable under freedom of information regimes are accessible; accessibility is not a licence. State and local records may be copyrighted, may carry use restrictions, and may be supplied under terms.
Material merely visible on a website. Publicly readable does not mean publicly usable. The site's terms of use, its access controls, and the copyright in whatever is displayed all apply.
Four questions organise the practice.
Which category is this source in? The threshold question, asked per source, recorded per source.
What obligations attach — attribution, share-alike, restrictions, refresh?
Is the output a derivative work, and does it inherit anything?
Can the provenance be proved? Because a purchaser or a defendant will ask.
See Information the Government Holds for the doctrinal treatment and Building a Product on Public Data for the sequence.
Government works and the limits of the exception
The rule that federal government works are outside copyright is narrower than practitioners assume, and the exceptions matter commercially.
It covers works prepared by officers or employees as part of their official duties. It does not cover work produced by contractors, which is the majority of technical output in many programmes.
State and local government works are not covered. Many are copyrighted, several jurisdictions assert those rights, and some charge for access.
Edicts of government — statutes, judicial opinions, and official annotations — are outside copyright on a distinct doctrine, and the boundary around privately produced annotations and compilations has been litigated.
Standards incorporated by reference into law raise a live question about whether the standard remains protected once compliance with it is legally required.
Government-funded but privately produced material may be protected, with the funding agreement determining what rights the government obtained.
Foreign government works are not covered by the domestic exception at all and are frequently protected.
Trademarks and official seals are separately restricted regardless of copyright status.
Contractual conditions can attach anyway. A portal supplying uncopyrighted data under terms of use has created a contractual obligation independent of any intellectual property right, which is a point that repeatedly surprises people.
Open licence families and what they require
Open data licences are short, are rarely read, and impose real obligations.
Public domain dedications waive rights to the extent possible and impose no conditions. The cleanest option and the least common for institutional publishers.
Attribution-only licences permit any use subject to crediting the source in a specified manner. The obligation is easy and is routinely breached because nobody built the attribution into the product.
Share-alike licences additionally require that adaptations be released under the same terms. This is the provision that can encumber a commercial product, and the boundary of what counts as an adaptation is the whole question.
Non-commercial licences exclude commercial use, which is defined variably and generously by the licensor and narrowly by the user.
No-derivatives licences permit redistribution but not adaptation, which excludes most product uses.
Database-specific licences address the compilation separately from the contents, which matters in jurisdictions with a distinct database right and which is drafted into some licences regardless.
Custom government terms frequently combine attribution with restrictions on implying endorsement, on redistributing personal data, and on representing derived outputs as official.
Version matters. Licence families have versions with materially different terms, and the version in force when the data was obtained governs.
Multiple licences in one product is the norm, and the obligations aggregate rather than average.
The derived-data question
Whether a product built from open data inherits obligations is the commercially decisive question, and it is answered by looking at what the product contains rather than at how it was made.
Facts extracted from a source carry nothing. If the output is a set of facts, and facts are unprotectable, the licence has nothing to attach to.
Substantial reproduction of a protected compilation is different. Where the output reproduces the selection and arrangement that made the source protectable, it is a copy.
Adaptation triggers share-alike. Cleaning, correcting, enriching, and restructuring a dataset may produce an adaptation, and share-alike terms then reach the result.
Aggregation is usually not adaptation. A product combining a share-alike dataset with independent material may be a collection rather than a derivative, which is the standard structural answer — and the boundary depends on the licence's own definitions.
Model training is unsettled. Whether training on a licensed dataset produces a derivative of it, and whether outputs inherit obligations, is contested. Assume the answer may be unfavourable and structure accordingly.
Attribution survives everything. Even where share-alike does not reach the output, attribution obligations usually do.
Document the boundary decision. A recorded analysis of why the product is not an adaptation is the difference between a defensible position and an assumption.
See the Data Licensing and Rights Toolkit and the Data Licensing Checklist.
Access, scraping, and the contract layer
Much public data is obtained by automated collection, and the legal analysis there is mostly not about copyright.
Terms of use bind where they are agreed, and the formation question — whether a visitor assented — determines whether the terms apply at all. See the Online Terms and Consumer Contracts Toolkit.
Technical measures matter. Circumventing authentication, rate limits, or blocks moves the analysis toward unauthorised access under 18 U.S.C. § 1030, as construed in Van Buren v. United States.
Publicly accessible pages occupy contested ground, with courts differing on whether accessing material available to anyone without authentication can be unauthorised.
Copyright still applies to what is copied. Scraping reproduces the page, and whether that is fair use under 17 U.S.C. § 107 depends on purpose and on what is retained.
Server load and interference support trespass and related theories in some jurisdictions.
Robots directives are not law but are evidence about permission and expectation.
Official APIs are the safe route where they exist, subject to their own terms, rate limits, and versioning.
Third-party data suppliers may themselves have scraped, which pushes the diligence upstream.
Personal data hiding in public records
Public records contain information about people, and the fact that a record is public does not exempt the user from privacy obligations.
Comprehensive privacy statutes reach personal information regardless of whether it was publicly available, subject to exemptions that vary and that are narrower than users assume.
Aggregation creates sensitivity. Individually innocuous public records combined into a profile is a different product from any of its inputs.
Consumer reporting characterisation attaches where information about individuals is compiled and supplied for employment, credit, insurance, or housing decisions — which can convert a public-records product into a regulated one under 15 U.S.C. § 1681a.
Court records carry their own access rules, sealing regimes, and increasingly restrictions on bulk redistribution.
Property and voter records are subject to jurisdiction-specific use restrictions that are frequently contractual conditions of bulk access.
Removal and correction requests will arrive, and a product built on public records needs a process even where no statute compels one.
Downstream use restrictions should be passed through to customers, since the supplier's obligations do not disappear on resale.
See the State Privacy Compliance Toolkit and the State Privacy Law Applicability and Readiness Checklist.
Provenance, and the record a buyer will ask for
A data product's value in a transaction depends on whether its inputs can be proved lawful, and most cannot.
Record the source per dataset: where obtained, when, under what terms, at what version.
Preserve the licence text as it stood, since publishers change terms and the version that governed the acquisition is the one that matters.
Record the access method — download, API, scrape — and the terms in force at the time.
Record any registration or agreement signed to obtain bulk access, since those are contractual obligations independent of any licence.
Record refresh obligations, since some licences require using current data or removing withdrawn records.
Record the derived-work analysis for anything share-alike touched.
Record the attribution implemented and where it appears in the product.
Make it a build artefact, generated rather than compiled by hand, on the model of a software bill of materials.
This record is the single most valuable document a data business produces, and its absence is the finding that most reliably reduces a valuation.
Attribution, done properly
Attribution obligations are the easiest to satisfy and the most commonly breached, because nobody designs for them.
Read what the licence actually requires: the name of the source, a link, a licence identifier, an indication of changes, and sometimes specific wording.
Decide where it appears. A product with a hundred sources cannot credit each on every screen, and the accepted practice is a dedicated attributions page linked from the product with per-source detail.
Indicate modifications where required, since many licences require a statement that changes were made.
Do not imply endorsement, which most government terms expressly prohibit and which a prominent official logo implies.
Carry attribution through to customers where the product is redistributed, since the obligation follows the data.
Automate it. Attribution generated from the provenance record stays current; attribution written once by hand goes stale within a release.
Audit it on a cadence, because sources are added by engineers who do not know the obligation exists.
Building the sourcing discipline
The organisational problem is that data acquisition is done by analysts and engineers who reasonably believe that public means free.
Install an intake gate. Any new external data source is registered before use, with source, terms, and category recorded. One form, one owner, five minutes.
Give a decision rule, not a doctrine. Four categories, four answers, and an escalation route for anything share-alike or restricted.
Train on the failure modes, which are: assuming public means free, assuming government means uncopyrighted, missing the terms of use behind a download button, and adding a share-alike source to a commercial product.
Review at release. The attribution page and the provenance record are checked before a version ships.
Re-check terms periodically, since publishers change them and continued use is under the current terms in some regimes.
Keep a substitution plan for any source whose terms could become problematic, because dependency on a single restricted source is a business risk.
Audit annually, and treat the audit output as diligence preparation rather than as a compliance chore.
The freedom of information route, and what it does and does not give you
Access legislation is a source of data rather than a licence, and the difference causes recurring problems for businesses that build on requested records.
Access is not permission. A record obtained under an access regime has been disclosed; nothing about the disclosure grants rights in whatever copyright subsists in it. A privately authored document held by a government body and released on request remains the author's work.
Exemptions shape what arrives. Commercially confidential information, personal data, law enforcement material, and deliberative content are commonly withheld or redacted, and the redactions are frequently the interesting part.
Third parties may object. Many regimes provide notice to the submitter of commercial information before release, with an opportunity to object — which means a business requesting a competitor's submissions may find the process slow and contested.
Fees and bulk access differ. Individual requests are one thing; a product depending on continuous bulk supply needs a different arrangement, frequently a contractual one with its own terms.
Reuse regimes are separate. Some jurisdictions have distinct legislation governing what may be done with public sector information after it is obtained, imposing licensing, attribution, and non-discrimination requirements on the body rather than on the requester.
Timeliness is the practical constraint. A product requiring current data cannot depend on a request process measured in months, which is why bulk arrangements and published datasets are the realistic sources.
Requests are public in some regimes. The fact that a company asked for particular records may itself be disclosable, which is a competitive intelligence problem for the requester.
The same records may be published elsewhere on better terms. Before building a request programme, check whether the data is already available through an open portal, a commercial aggregator, or another jurisdiction's publication — which is faster, cheaper, and comes with clearer terms.
The practical instruction is to treat access legislation as one channel among several, to record what came from where, and never to assume that a document released under a statutory access regime may be redistributed commercially without further analysis.
Comparative notes
The framework above is domestic, and practitioners advising internationally should know the four points where it diverges most.
Crown and state copyright. Many jurisdictions assert copyright in government works rather than excluding them, and then license that copyright openly. The result looks similar in practice — the material is usable — but the legal basis is a licence with conditions rather than an absence of rights, which means the obligations bind and the terms can change.
Database rights. Several jurisdictions recognise a right in a database independent of copyright, protecting substantial investment in obtaining, verifying, or presenting contents. That reaches material which is unprotectable domestically, and it means a compilation of pure facts may be protected abroad and free at home. Extraction and re-utilisation of a substantial part is the infringing act, and repeated extraction of insubstantial parts can qualify.
Public sector information reuse regimes. Some jurisdictions impose obligations on public bodies to make information reusable on non-discriminatory terms, with charging limits and format requirements. That is a right running against the body rather than a licence to the user, but it shapes what terms are available.
Text and data mining exceptions. Some jurisdictions have specific exceptions permitting mining of lawfully accessible works, sometimes limited to research, sometimes with an opt-out mechanism for rights holders. Where such an exception exists, the analysis of a training corpus differs materially from the domestic fair use inquiry.
The practical instruction for a cross-border product is to run the source register against the markets in which the product is offered rather than only against the market in which it was built. A dataset lawfully compiled in one jurisdiction may infringe a database right in another, and the exposure attaches where the product is supplied.
Model training on public sources
The largest current use of public data is training, and it deserves separate treatment because the analysis is unsettled and the exposure is large.
Collection is reproduction. Assembling a corpus copies the material, whatever happens afterwards. That is the act with the clearest legal characterisation, and it precedes any argument about the model.
Fair use is the principal defence and turns on the four factors in 17 U.S.C. § 107, read through Warhol v. Goldsmith on purpose comparison and Google LLC v. Oracle America on functional material. The arguments are strong for genuinely transformative analysis and weaker where outputs substitute for the source.
Licensed sources carry their terms into the analysis. A share-alike dataset in a training corpus raises the question whether the model is an adaptation, which nobody has answered authoritatively and which most licences did not contemplate.
Non-commercial licences are a clear problem for a commercial model, and are frequently in corpora because they were free.
Attribution is impossible at corpus scale, which means an attribution-only licence is technically breached by most training uses even where the copyright analysis favours the user.
Facts remain unowned. A model that learns facts from public records has learned unprotectable material, and the argument is strongest where the outputs are factual rather than expressive.
Personal information in the corpus carries obligations independent of copyright, including rights requests that a trained model cannot satisfy.
Provenance is the practical defence. A developer who can say precisely what was in the corpus, on what terms, is in a materially better position than one who cannot — both in litigation and in a transaction.
The honest advice is that this area will change, that positions taken now may not survive, and that the mitigations available are recording provenance, excluding categories where the terms are clearly incompatible, and structuring so that a problematic source can be removed and the model retrained. See the AI Content and IP Toolkit and Deploying Generative AI Without Losing Your IP.
A short glossary
Publicly available. An access condition, not a permission. The phrase that causes most of the errors in this field.
Public domain. Material in which no copyright subsists, whether because it never did or because the term expired. Distinct from openly licensed material.
Government edicts doctrine. The rule placing statutes, judicial opinions, and certain official annotations outside copyright.
Open licence. A grant of permission on stated terms, commonly requiring attribution and sometimes requiring share-alike release.
Share-alike. The condition requiring adaptations to be released under the same licence. The provision that can encumber a commercial product.
Adaptation. A derivative of a licensed work, as defined by the licence rather than by intuition. The boundary question in every share-alike analysis.
Collection. An assembly of independent works, which under most open licences does not trigger share-alike as an aggregate.
Database right. A right recognised in some jurisdictions protecting investment in a compilation independently of copyright, reaching material unprotectable domestically.
Extraction and re-utilisation. The infringing acts under a database right, including repeated taking of insubstantial parts.
Access legislation. Statutes compelling disclosure of records. A supply channel, not a licence.
Reuse regime. Legislation obliging public bodies to make information reusable on non-discriminatory terms.
Text and data mining exception. A statutory permission, in some jurisdictions, to mine lawfully accessible works, sometimes subject to an opt-out.
Provenance record. The generated inventory of sources, terms, versions, and obligations. The asset a purchaser asks for.
Refresh obligation. A licence condition requiring use of current data or removal of withdrawn records.
Practitioners who keep those fourteen straight will avoid the field's characteristic errors: treating access as permission, treating government as uncopyrighted, and treating a licence obligation as satisfied because nobody has complained.
What the company itself owns
A data business spends its attention on what it may take and very little on what it holds, which is a mistake at exit.
The compilation. Selection, arrangement, and structure attract thin protection under 17 U.S.C. § 103, and registration in periodic batches makes that protection enforceable under 17 U.S.C. § 412.
The enrichment. Matching, cleaning, normalising, resolving entities, and correcting errors produce material the sources do not contain. This is where the genuine value is added and where a competitor cannot simply re-derive the product from the same inputs.
The methods. Matching algorithms, quality heuristics, and the accumulated knowledge of which sources are unreliable in which respects are trade secrets under 18 U.S.C. § 1839 if treated as such.
The pipeline. Ingestion, transformation, and delivery software, with the ordinary ownership questions for contractor-written code under 17 U.S.C. § 201 and 17 U.S.C. § 204.
The customer terms. What customers may do with the product, whether they may build derivatives, and what happens on termination are the terms that determine whether the business has recurring revenue or a one-time sale.
The provenance record itself, which is both a compliance artefact and a competitive advantage, since a customer choosing between two suppliers will prefer the one that can prove its inputs.
The brand, which in a data business signals reliability and is the reason a customer renews.
The practical instruction is to run the same discipline inward as outward: register the compilations, treat the methods as secret, assign the pipeline code, and be able to describe in one page what the company owns as distinct from what it licensed. A data business that can only describe its inputs is describing someone else's asset.
The first meeting
Six questions asked of a new data-product client surface almost everything.
Show me your source list. If it does not exist, that is the engagement. If it exists as a spreadsheet last updated eighteen months ago, that is also the engagement.
Which of your sources are share-alike? Most clients do not know, and the answer determines whether the product is encumbered.
Where does your attribution appear? If the answer is nowhere, obligations are being breached daily and the fix is a page and a build step.
Did you agree to terms to get any of this? Registration forms, bulk access agreements, and API terms are contractual obligations independent of licences, and clients rarely think of them as such.
Does the product contain information about identifiable people? If yes, a second regime applies regardless of the public status of the sources, and consumer reporting characterisation may apply on top.
What would you hand a buyer who asked to see your provenance? The honest answer is usually "we would put something together," which is exactly what a diligence team hears as a red flag.
Six questions, half an hour, and a work plan whose first item is the register, because everything else depends on it.
A closing observation
There is a persistent belief in technology companies that data found on the internet is a free input, comparable to air or sunlight. It is a natural belief, it is encouraged by how easy collection has become, and it is wrong often enough to be dangerous.
The accurate position is more nuanced and more useful. A great deal of public data genuinely is free to use, including some of the most valuable — facts, government edicts, and material whose term has expired. Some is free subject to conditions that are trivial to satisfy and are nonetheless routinely breached, which is a solved problem that nobody has solved. Some carries conditions that are genuinely incompatible with a commercial product, and using it anyway is a decision that should be taken knowingly rather than discovered in diligence. And some is not public at all in the relevant sense, having been obtained in breach of terms that a court will enforce.
Four categories, four answers, and a register recording which is which. The whole of this toolkit reduces to that, and the businesses that do it are not more cautious than their competitors — they are simply able to sell to customers who ask.
Building the register when there is no register
Most engagements begin with a product already built on sources nobody documented, and the reconstruction exercise has a workable method.
Start from the pipeline, not from memory. Ingestion code names its sources. Configuration files list endpoints. Scheduled jobs identify what runs and how often. An afternoon with an engineer produces a more complete list than a month of asking people.
Check the network egress. Where the pipeline calls out to is a definitive record of where data comes from, including sources nobody mentions because they were added by a departed contributor.
Look for the manual step. Every data business has one source that a person downloads and drops into a folder, and it is invariably the one with the restrictive terms.
Retrieve the terms as at acquisition where possible, using archived versions of the publisher's pages, and record where the reconstruction was uncertain.
Triage rather than boil the ocean. Classify sources by whether the product could survive their removal. Critical sources get full analysis; incidental ones get a category and a note.
Flag the three shapes of problem in order of seriousness: a share-alike source in a commercial product, a non-commercial source in a commercial product, and an agreement signed to obtain bulk access whose terms nobody has read.
Fix forward as you go. Attribution can be implemented while the register is being built, and it is the cheapest obligation to discharge.
Write down what could not be established. A register with honest gaps is far more useful than one with confident guesses, both internally and in diligence, where an admitted uncertainty is a manageable finding and a wrong answer is a credibility problem.
Two to four weeks for a mid-sized product, and the output is simultaneously a compliance artefact, a sales asset, and the document that makes the next financing straightforward.
Selling into regulated customers
A data product sold to banks, insurers, healthcare organisations, or government buyers meets a procurement process that asks questions no other customer asks, and the answers determine whether the sale closes.
Provenance warranties are standard. Regulated buyers require representations that inputs were lawfully obtained and that the supplier has rights to grant. A supplier that cannot support the warranty either declines the business or gives a representation it cannot stand behind.
Indemnities follow the warranties, and the cap negotiation is where the risk actually lands.
Audit rights are common, and a buyer entitled to inspect the source register will eventually inspect it.
Sub-processor and onward transfer terms apply where the product contains personal information, with flow-downs and localisation questions.
Model and automated decision questions arrive where the buyer will use the product in decisions about people — credit, employment, insurance, housing — which can convert the supplier's product into a regulated input.
Continuity and escrow requirements appear where the buyer's operations depend on the feed. See the Software Continuity and Escrow Toolkit.
Security assessment is a separate workstream with its own timeline. See the Cybersecurity Governance and Disclosure Toolkit.
Accuracy representations are the sharpest, because a data product's accuracy is a function of sources the supplier does not control, and the honest position is a described methodology and a correction process rather than a guarantee.
The commercial point worth making to a client early: the register is not overhead. It is the document that lets the company answer a regulated buyer's procurement questionnaire in a week rather than a quarter, and in this market that timing difference decides deals.
It is also the document that makes the difference between a warranty a founder can give honestly and one given because the deal required it — a distinction that matters enormously two years later, when the buyer's regulator asks the buyer where the data came from and the question arrives back at the supplier with a contract attached.
That chain — regulator to buyer to supplier — is the reason provenance discipline has moved from a compliance preference to a commercial requirement in this market, and it is why the register repays its cost several times over in the first enterprise deal it enables.
Make that argument to the commercial team rather than to the general counsel, because in a data business the commercial team controls the roadmap and the general counsel does not.
Framed as a sales enabler it gets built in a quarter; framed as a compliance project it gets deferred indefinitely, and the deferral ends at diligence.
That is the single most useful piece of practice advice in this toolkit, and it has nothing to do with the law of public data at all.
Which is frequently how it goes in this field: the doctrine is manageable and the organisational habit is the hard part.
Solve the habit and the doctrine takes care of itself; solve only the doctrine and nothing changes.
And the habit begins in exactly one place: start with the register.
A Suggested Reading Path
New to public data: Information the Government Holds, then Building a Product on Public Data, then the Public Data Use Checklist.
Data licensing generally: the Data Licensing and Rights Toolkit and the Data Licensing Checklist.
The copyright framework: the Copyright Fundamentals Toolkit, the Copyright Duration and Public Domain Toolkit, and the Public Domain Clearance Checklist.
Fair use where reproduction occurs: Running a Fair Use Analysis and the Fair Use Risk Assessment Checklist.
Open licensing analogues: Copyleft and Consequences and the Software, Data, and Open Source Toolkit, since share-alike reasoning is the same problem in a different medium.
Privacy overlay: the Privacy and Marketing Data Toolkit and the Biometric and Sensitive Data Toolkit.
Model training: the AI Content and IP Toolkit, the AI Procurement and Governance Toolkit, and the Generative AI IP Compliance Checklist.
Transactions: the IP Due Diligence Toolkit.
Primary Authorities
| Authority | Use | |---|---| | 17 U.S.C. § 102 | What is protected, and the exclusion of ideas and facts | | 17 U.S.C. § 103 | Compilations and the thin protection of selection and arrangement | | 17 U.S.C. § 105 | Government works, and its limits | | 17 U.S.C. § 106 | Reproduction and derivative rights engaged by collection | | 17 U.S.C. § 107 | Fair use where reproduction occurs | | 17 U.S.C. § 201 | Contractor-produced government material | | 17 U.S.C. § 204 | Transfers of rights in acquired datasets | | 17 U.S.C. § 302 | Duration, and what has entered the public domain | | 17 U.S.C. § 412 | Registration of a company's own compilations | | Feist v. Rural Telephone | Facts are not owned; compilations are thinly protected | | Georgia v. Public.Resource.Org, Inc. | Edicts of government and official annotations | | Banks v. Manchester | The foundation of the government edicts doctrine | | Harper & Row v. Nation Enterprises | Fair use limits on unpublished and newsworthy material | | Warhol v. Goldsmith | Purpose comparison in the first factor | | Google LLC v. Oracle America | Fair use applied to functional material | | 18 U.S.C. § 1030 | Unauthorised access in collection disputes | | Van Buren v. United States | What exceeding authorised access means | | 15 U.S.C. § 1681a | When a public-records product becomes a consumer report | | 15 U.S.C. § 45 | Accuracy and provenance representations to customers | | 15 U.S.C. § 1125 | Implying official endorsement | | 18 U.S.C. § 1839 | The company's own enrichment and matching methods | | FRCP 26 | Discovery into sourcing in a data dispute |
Search the underlying materials directly for government edicts doctrine annotations, open data share-alike derived product, web scraping terms of use enforceability, public records aggregation privacy statute, and open licence attribution requirement compliance.
Forms and Templates
A source register, one row per dataset: name, publisher, category, access method, date, licence and version, obligations, refresh requirement, and owner.
An intake form completed before any new source is used, feeding the register.
A four-category decision rule, one page, distinguishing outside copyright, openly licensed, access-legislation, and merely-visible sources, with an escalation route for share-alike and restricted terms.
A licence text archive capturing terms as they stood at acquisition.
A derived-work analysis template for any share-alike source, recording why the output is or is not an adaptation.
An attributions page specification, generated from the register, with per-source detail, modification statements, and no implied endorsement.
A refresh and withdrawal procedure for sources requiring current data or removal of withdrawn records.
A downstream terms rider passing source restrictions through to customers.
A removal and correction process for personal information appearing in the product.
A substitution plan identifying an alternative for each critical source.
A provenance export suitable for handing to a diligence team without further work.
For general drafting starting points, see the Draft License Agreement and the License Agreement Template.
Five recurring matters
Diligence finds a share-alike dataset in a commercial product. Establish first whether the output is actually an adaptation, which is frequently arguable and occasionally clearly not. Where it is, the options are replacing the source, restructuring the product so the dataset is aggregated rather than adapted, or complying. Compliance is sometimes acceptable and is rarely considered.
A government portal changes its terms. Check whether continued use is governed by the current terms or by those in force at acquisition, which differs by publisher. Then check whether the product's dependency on that source is critical, and activate the substitution plan if it is.
A person demands removal from a public-records product. There may be no statutory obligation and there is almost always a commercial and reputational one. Have a process, apply it consistently, and record the decision.
A competitor alleges the product's data was scraped from its site. The analysis is contract and access first, copyright second. Establish what was collected, from where, whether terms were agreed, whether any technical measure was circumvented, and what of the collected material survives in the product.
A customer asks for a provenance warranty. This is now a standard request in data purchasing, and a supplier that cannot answer it loses the deal. The register is the answer, which is why it is a commercial asset rather than a compliance artefact.
What good looks like
A source register exists, generated rather than hand-maintained, and current.
Every source has a recorded category and a licence text archived at the version acquired.
Share-alike sources have a written derived-work analysis.
Attribution is generated from the register and audited at release.
Refresh and withdrawal obligations are implemented, not merely noted.
Personal-information handling has a process, including removal requests.
A substitution plan exists for critical sources.
The provenance export can be produced in a day, because that is what a purchaser will ask for.
Businesses with those eight sell data products at full value. Businesses without them discover during diligence that their most important input has terms nobody read.
Related Documents
The core cluster is Information the Government Holds, Building a Product on Public Data, and the Public Data Use Checklist.
For sectors whose products are built substantially on public information, see the Space and Satellite IP Toolkit, the Logistics and Supply Chain Technology IP Toolkit, and the Mining, Energy, and Natural Resources IP Toolkit.
For the cultural and heritage side of open information, see the Museums, Libraries, and Cultural Heritage IP Toolkit, The Rights You Cannot Trace, and the Traditional Knowledge and Cultural Expressions Toolkit.
For the workforce and consumer-reporting overlay that public-records products frequently trigger, see Everything the Application Knows and the Recruitment and Workforce Data Toolkit.
Marksy is not a law firm and this toolkit is not legal advice. The status of government works, the enforceability of terms of use, the treatment of database rights, and the reach of privacy statutes over publicly available information all vary by jurisdiction. Advice on a specific product requires the source register and the terms.