Who Owns the Data? Scraping, Databases, and the Limits of Ownership
By Casey Scott McKay ·
Clients arrive with a firm conviction that they own their data, and the law's answer is that almost nobody owns data as such. This article explains what actually protects a dataset - a thin copyright in original selection and arrangement, a trade secret claim if it is genuinely secret, a contract if someone agreed to something, and technical controls - and why the United States, unlike the European Union, has no database right at all. It works through the scraping analysis in the order it should be run: is the material protected, was access authorized, was a contract formed and is it enforceable, was a technical measure circumvented, and does the output infringe. It covers what Van Buren did to the Computer Fraud and Abuse Act, what hiQ v. LinkedIn resolved and left open, and the personal-data overlay that has become the largest exposure in the area.
IP and Technology > Information Technology | Article | Published 29 December 2024 - Updated 21 January 2025 | Casey Scott McKay - marksy.us
Summary. Clients arrive with a firm conviction that they own their data, and the law's answer is that almost nobody owns data as such. This article explains what actually protects a dataset — a thin copyright in original selection and arrangement, a trade secret claim if it is genuinely secret, a contract if someone agreed to something, and technical controls — and why the United States, unlike the European Union, has no database right at all. It works through the scraping analysis in the order it should be run: is the material protected, was access authorized, was a contract formed and is it enforceable, was a technical measure circumvented, and does the output infringe. It covers what Van Buren did to the Computer Fraud and Abuse Act, what hiQ v. LinkedIn resolved and left open, and the personal-data overlay that has become the largest exposure in the area.
Keywords: web scraping legality · computer fraud and abuse act · van buren exceeds authorized access · hiq v linkedin · feist originality compilations · sweat of the brow · sui generis database right · browsewrap clickwrap enforceability · terms of service breach · trespass to chattels servers · dmca 1201 circumvention · robots txt · rate limiting · personal data in scraped sets · ccpa and scraped data · training data provenance · contractual scraping prohibitions · data licensing · hot news misappropriation · preemption of state data claims
This is premium Marksy content — the full document is available to subscribers.