AI, Content, and IP Toolkit: Training Data, Generated Works, and the Ownership Gap
By Casey Scott McKay ·
This toolkit is a curated research guide to what United States intellectual property law currently does — and refuses to do — with generative artificial intelligence, assembling the Marksy corpus on the subject into one ordered shelf. It opens with the organizing distinction that resolves most confused AI questions: the input problem (was the training corpus lawfully acquired and used) is legally separate from the output problem (does anyone own the generated asset, and does it infringe anything), and each of those splits again into a defensive and an offensive question. Thematic sections route the reader to the documents that work each stage, beginning with the human authorship requirement running from Burrow-Giles through Thaler v. Perlmutter and the Copyright Office's 2023 registration guidance and 2025 reports, and the training-data fair use litigation from Thomson Reuters/Ross through Bartz/Anthropic and Kadrey/Meta. Further sections cover output-similarity and memorization risk, the ownership paperwork that still governs the human contributions, trademark, right of publicity, and trade secret as the regimes that fill the gap copyright leaves, and the contract terms, indemnities, and insurance that allocate a risk no doctrine has yet resolved. It closes with a branching reading path for six different readers, a table of controlling authorities with one-line holdings, the Marksy templates that apply, and annotated pointers to neighbouring toolkits and checklists.
IP and Technology > General IP | Toolkit | Published 22 August 2024 - Updated 12 February 2025 | Casey Scott McKay - marksy.us
Summary. A curated tour of everything in the Marksy corpus about intellectual property and generative AI. It starts with the distinction that untangles most of the confusion — the input question (was the corpus lawfully acquired and lawfully used) is a different legal problem from the output question (does anybody own the generated asset, and does it infringe) — and then routes you to the documents that work each one. Sections cover the human authorship requirement from Burrow-Giles to Thaler v. Perlmutter, the Copyright Office's registration guidance and its three AI reports, the training-data fair use docket, output-similarity and memorization risk, the ownership paperwork that governs the human contributions, and the trademark, publicity, and trade secret regimes that fill the copyright gap. It ends with a branching reading path, a table of controlling authorities, the templates that apply, and pointers to the neighbouring toolkits.
Keywords: ai ip toolkit · generative ai · human authorship requirement · thaler v. perlmutter · copyright office ai guidance · training data fair use · bartz v. anthropic · kadrey v. meta · thomson reuters v. ross · output substantial similarity · the ownership gap · ai vendor indemnity · digital replicas · generated logo clearance · provenance and disclosure · 17 u.s.c. 1202 · registration disclaimer · trade secret fallback · ai contract clauses · eu ai act training summary
Start Here
Halyard Studios, a thirty-four-person branding agency in Charleston, delivered a rebrand in March to Tidewater Provisions, a Savannah hot-sauce maker: a wordmark, a pelican mascot, 140 packaging illustrations, and a thirty-second radio spot narrated by a synthetic voice built to sound like a retired local sportscaster. Every visual asset started in a diffusion model. The statement of work says Tidewater owns "all intellectual property in the Deliverables."
That sentence is partly false, and nobody at either company knows which part.
This toolkit is for whoever has to answer three questions about a matter like that:
- Does anyone own the output? Often no one does — not the user, not the vendor, not the model. That is the ownership gap, and it is a feature of the statute rather than a temporary uncertainty.
- Was the input lawful, and who bears that risk? Training-data infringement is a separate case from output infringement, with separate defendants, defenses, and appellate posture.
- What can go out the door, and what paper allocates the residue? The doctrine will not settle for years, so the operative instrument in most AI matters is a contract, not a case.
If you read only one thing, read Who Owns What the Machine Made: Copyright Authorship in the Age of Generative AI. It is the doctrinal spine of everything below — human authorship, Thaler, the March 2023 registration guidance, the selection-and-arrangement carve-out, the disclosure duty on the application form, and an honest account of what a vendor indemnity actually covers. The rest of this toolkit assumes it. Everything here is United States law; a client who trains or deploys abroad gets a different answer, noted below.
The Field, Mapped
Copyright law has a human being at its center and always has. Section 102(a) protects "original works of authorship," and the Supreme Court reads "author" as the person "to whom anything owes its origin; originator; maker." Burrow-Giles Lithographic Co. v. Sarony, 111 U.S. 53, 58 (1884). Originality means independent creation plus a modicum of creativity, not effort. Feist Publications, Inc. v. Rural Telephone Service Co., 499 U.S. 340, 345 (1991). A macaque cannot sue under the Act. Naruto v. Slater, 888 F.3d 418 (9th Cir. 2018). And the D.C. Circuit has now held squarely that a machine cannot be listed as an author, affirming the Register's refusal of Stephen Thaler's application. Thaler v. Perlmutter, 687 F. Supp. 3d 140 (D.D.C. 2023), aff'd, 130 F.4th 1039 (D.C. Cir. 2025). The Patent Act answers the same way about inventors. Thaler v. Vidal, 43 F.4th 1207 (Fed. Cir. 2022).
The administrative rule is shorter than the case law and matters more day to day. Copyright Registration Guidance: Works Containing Material Generated by Artificial Intelligence, 88 Fed. Reg. 16,190 (Mar. 16, 2023), says four things: a work must be the product of human authorship; material generated in response to a prompt is not; an applicant must disclose more-than-de-minimis AI material and exclude it from the claim; and a human's own selection, arrangement, and modification remain claimable. See also Compendium of U.S. Copyright Office Practices §§ 306, 313.2, 621 (3d ed. 2021). The Zarya of the Dawn correspondence, Reg. No. VAu001480196 (Feb. 21, 2023), set the template — text and arrangement registered, Midjourney images excluded — and the Copyright Review Board then refused Théâtre D'opéra Spatial (Sept. 5, 2023) after 624 prompt iterations and SURYAST (Dec. 11, 2023) even though the applicant supplied his own base photograph. Three Office reports followed: digital replicas (2024), copyrightability (2025), and generative AI training (2025).
That is the output side. The input side is a different case, and confusing the two is the commonest analytic error here. A total defense win on training does nothing for the copyrightability of the output, because the obstacle is the absence of a human author, not the provenance of the corpus. A model trained only on public-domain material still produces unownable pictures.
Four questions, worth seeing at once:
| | Input — the corpus, the ingestion, the weights | Output — the generated asset | |---|---|---| | Defensive: can somebody stop me? | Training infringement and fair use; CMI removal, 17 U.S.C. § 1202(b); breach of site terms; scraping claims | Substantial similarity and memorization; trademark confusion and dilution; publicity and false endorsement, 15 U.S.C. § 1125(a) | | Offensive: can I stop anyone else? | Trade secret in the corpus, tags, and weights, 18 U.S.C. § 1839(3); contract and access controls | The ownership gap — usually no copyright; thin copyright in human selection, arrangement, and modification, § 103(b); trademark; trade secret |
"AI training" is also not one act. It is at least four — acquisition, reproduction into a durable corpus, training itself, and output — and courts have begun answering them differently. Judge Bibas held that training a competing, non-generative legal research tool on Westlaw headnotes was not fair use. Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc., No. 1:20-cv-00613 (D. Del. Feb. 11, 2025), interlocutory appeal taken to the Third Circuit. Judge Alsup held that training on lawfully purchased books was exceedingly transformative but that retaining a permanent library of pirated downloads was not fair use at all. Bartz v. Anthropic PBC, No. 3:24-cv-05417 (N.D. Cal. June 23, 2025); the case settled on publicly reported terms near $1.5 billion. Judge Chhabria granted Meta partial summary judgment while saying plainly that the result reflected an evidentiary failure, and identified market dilution — a flood of machine-made substitutes suppressing demand for human work — as the theory that should decide these cases. Kadrey v. Meta Platforms, Inc., No. 3:23-cv-03417 (N.D. Cal. June 25, 2025). The consolidated OpenAI litigation in the Southern District of New York, the Andersen v. Stability AI induced-infringement theory, and the music publishers' lyric case against Anthropic remain live. There is no controlling appellate authority on generative training. Anyone who tells you otherwise is selling something.
Four other bodies of law sit around the copyright question, and copyright specialists trip on all of them. Trademark does not care who drew the logo — a generated mark that identifies source is protectable through use and registration, 15 U.S.C. §§ 1051, 1127 — the best answer available to the ownership gap for brand assets. The right of publicity governs synthetic voices and faces and is moving faster than copyright, led by Tennessee's ELVIS Act and California's 2024 digital replica statutes, Cal. Lab. Code § 927 and Cal. Civ. Code § 3344.1. Trade secret protects confidential outputs without asking who authored them. And contract is where the residual risk lands.
Finally: fair use does not exist abroad. The European Union runs on the text-and-data-mining exceptions of Directive (EU) 2019/790, arts. 3-4, with a machine-readable opt-out rights holders are now exercising, plus the transparency and copyright-policy duties for general-purpose models in Regulation (EU) 2024/1689, art. 53(1)(c)-(d). Japan's Copyright Act art. 30-4 is broader than anything American law offers; the United Kingdom's exception is narrower. No multinational deployment clears on a United States analysis.
Theme 1 — The Output Question: Authorship, Disclosure, and the Gap
Start here whenever the client asks "do we own this." The doctrine is settled at the extremes and open in the middle: a raw generation has no author, a substantially repainted one carries a copyright in the human changes, and nobody has drawn the line between them.
- Who Owns What the Machine Made: Copyright Authorship in the Age of Generative AI — the full doctrine, from Sarony through Thaler and the Review Board refusals, with the clearest statement in the corpus of why "you own it" and "you can stop anyone" are different sentences. Read it before the first client call, and again before you sign an opinion.
- Deploying Generative AI Without Losing Your IP — twelve stages converting the doctrine into a file: tool inventory and output tiering, a policy whose operative clause manufactures registrable subject matter rather than merely prohibiting things, a provenance record that survives a deposition, field-by-field disclaimer language, and a vendor redline with ask, fallback, and walk-away positions. Work from this once you know there is a program to build.
- Generative AI IP Compliance Checklist — the same workflow as dated, checkable actions, each naming the form, fee, field, or office. Hand it to in-house counsel at kickoff.
Registration is where the disclosure duty bites. Filing late forfeits statutory damages and fees under 17 U.S.C. § 412, and an application that quietly omits AI material can be attacked under 17 U.S.C. § 411(b) after Unicolors, Inc. v. H&M Hennes & Mauritz, L.P., 595 U.S. 178 (2022).
- What Copyright Registration Actually Buys You — the two gates, § 411(a) after Fourth Estate Public Benefit Corp. v. Wall-Street.com, LLC, 586 U.S. 296 (2019), and § 412's three-month grace window, with the arithmetic. Read it when a client asks whether registering a mixed work is worth $65.
- Registering a Copyright: A Practitioner's Guide — the limitation-of-claim mechanics every AI-assisted application needs, the group options at 37 C.F.R. § 202.4, $800 special handling under § 201.3(d), and the refusal ladder to the Copyright Review Board. Reach for it the first time you type a disclaimer.
- Copyright Registration Checklist: From Deposit to Certificate — eleven phases with the four dates every registered work needs on a docket. Use it for volume filings, where the disclosure must be right a hundred times.
Trap. Supplementary registration under 17 U.S.C. § 408(d) costs $100 and repairs an application that failed to disclose AI material. Fixing it before a defendant finds it is housekeeping; fixing it after is an admission. Audit the back catalogue first.
Theme 2 — The Input Question: Training Data and Fair Use
Every model builder and fine-tuner has this problem, and so does every company whose vendor has it. The four-factor analysis is the one that governs documentaries and quotation; what differs is scale, provenance, and factor four.
- Fair Use After Warhol — the doctrine after Andy Warhol Foundation v. Goldsmith, 598 U.S. 508 (2023), with a part separating the four copying events in training and reporting what each district court actually held. Read it before forming a view about Bartz or Kadrey from a headline.
- Running a Fair Use Analysis: A Practitioner's Guide for Content, Software, and AI Training — thirteen stages from the use statement through scoring, mitigations, permission, insurance, and litigation, with a playbook for the model developer training on scraped data and model language for a training-data warranty and indemnity. It turns a judgment call into a defensible file.
- Fair Use Risk Assessment Checklist — ten phases and a twelve-tab clearance file ending in a signed, dated decision — proceed, change, license, or drop — with a scoring grid and an escalation matrix keyed to who signs.
Two things practitioners get wrong. First, provenance is an independent liability question: Bartz separated how the corpus was obtained from what was done with it, and torrenting a shadow library is not cured by a transformative downstream use. Second, the claims that ride along are often the dangerous ones — CMI removal under 17 U.S.C. § 1202(b), at $2,500 to $25,000 per violation under § 1203(c)(3)(B), multiplies terrifyingly across a corpus, subject to the double-scienter rule of Stevens v. CoreLogic, Inc., 899 F.3d 666 (9th Cir. 2018), and to an unresolved identicality objection. Add breach of site terms and the contract claims that come with scraping.
Otterbein Labs, a Columbus SaaS company, fine-tuned an open-weights model on 2.1 million support tickets, product manuals, and forum threads. The tickets are its own. The manuals include eleven OEM documents obtained under a distributor agreement carrying a no-derivative-works clause — a contract problem fair use does not touch. The threads came from a site whose terms prohibit automated collection. Otterbein's copyright exposure is the least of its three problems, and counsel spent the first week on the wrong one.
Theme 3 — When the Output Looks Like Something
Output claims are ordinary infringement claims, won and lost on the ordinary elements. Style is not protectable; specific expression is. A plaintiff who can put a verbatim regurgitation next to the source has a real case; one who can only say "it looks like my work" usually does not.
- Proving Copyright Infringement: Access, Substantial Similarity, and the Idea-Expression Divide — originality, § 102(b), merger and scenes a faire, copying in fact versus unlawful appropriation, the competing circuit tests, and abstraction-filtration-comparison. Read it before drafting either side of an output complaint; corpus membership makes access nearly automatic, which throws the case onto similarity.
- Filing a Copyright Infringement Complaint in Federal Court — thirteen stages from the registration gate, with the works-in-suit schedule that makes § 412, and therefore the damages election, a per-work question. Essential for a rights holder assembling a multi-work claim against a model developer.
- Copyright Infringement Complaint Checklist — ten phases through filing and service, including web capture built to authenticate under Fed. R. Evid. 902(13) — exactly how you preserve a prompt-and-output pair before the model updates and the output stops reproducing.
Practice tip. Capture generations with the model version, system prompt, seed, and timestamp, somewhere a litigation hold can reach. A defendant who cannot reproduce the complained-of output six months later has lost the argument that it was a fluke; a plaintiff who cannot has lost the exhibit.
Theme 4 — The Human Parts Still Need Paperwork
Nothing about generative tools changes who owns the human contributions, and those are the only copyrights in the file. The contractor who repaints a generated illustration authors the repaint. Halyard's pelican was finished by a freelance illustrator, and Halyard cannot convey to Tidewater what she never assigned to Halyard.
- Who Owns the Work: Employees, Contractors, Joint Authors, and Work Made for Hire — the two exclusive routes to work-made-for-hire status, the Community for Creative Non-Violence v. Reid, 490 U.S. 730 (1989), agency factors, and why most contractor work — logos, websites, software, illustrations — fails the commissioned-work route entirely. The most consequential document here for agencies and studios.
- Transfers, Licenses, and Termination Rights — the execution manual: model contractor language pairing a work-for-hire recital with a present assignment, the third-party and open-source representations that should now cover generated material too, and the § 205 recordation race. Use its Stage 4 language and add a tool-disclosure representation.
- Copyright Ownership and Chain-of-Title Checklist — eleven phases building a title package a buyer or lender will accept, including a phase auditing the stock, fonts, open source, and machine-generated material buried inside deliverables. Run it before a financing, not during one.
- Assignments vs. Licenses: What's the Difference? — short, and it settles the question under every vendor output clause. An assignment of "all right, title, and interest" in uncopyrightable output conveys nothing; the covenant not to assert is the part with content.
Theme 5 — Generated Brands, Faces, and Voices
This is where the money is lost, because most AI policies address copyright and stop. Halyard's pelican is a trademark problem, its wordmark is a clearance problem, and its synthetic sportscaster is a right-of-publicity problem no copyright license can fix.
Marks. Trademark asks whether a designation identifies source, not who authored it, which makes registration the practical answer to the ownership gap for brand assets. Two cautions: generative tools produce near-misses of existing marks with cheerful regularity, and a generated product mockup is not a specimen of use, 37 C.F.R. § 2.56; TMEP § 904 — signing a declaration on goods that are not sold flirts with the deliberate-intent standard of In re Bose Corp., 580 F.3d 1240, 1245 (Fed. Cir. 2009).
- Trademark Clearance Searching — why a free database search is screening, not clearance, and how constructive notice under 15 U.S.C. § 1072 makes "the model suggested it" worth nothing.
- Running a Full Trademark Clearance Search and the Trademark Clearance Search Checklist — the protocol and the working paper. Run them on generated candidates as you would on human ones, and budget for a higher knockout rate.
- Choosing a Strong Trademark — models produce names that sound like the category, which is a machine for generating descriptive marks. Read it before the naming session, not after the refusal.
- Specimen Refusals — the refusal a generated mockup earns, and the one genuinely expensive to cure.
Third-party marks in generated copy. Models write comparative copy naming competitors. Route it through Descriptive and Nominative Fair Use for the two doctrines and the circuit split over which test applies, the Trademark Fair Use Audit Checklist for the working review, and the Expressive Use and Parody Risk Checklist for creative work that riffs on a brand now that Jack Daniel's Properties, Inc. v. VIP Products LLC, 599 U.S. 140 (2023), has removed the Rogers shortcut for source-identifying uses.
Faces and voices. Not a copyright problem, and filing it as one wastes a year.
- Your Face Is Not Public Domain — the patchwork mapped, from Midler v. Ford Motor Co., 849 F.2d 460 (9th Cir. 1988), and Waits v. Frito-Lay, Inc., 978 F.2d 1093 (9th Cir. 1992), through the ELVIS Act's untested tool-distribution theory. Read the digital replica part before any synthetic voice ships.
- Clearing and Licensing Name, Image, and Likeness — fifteen stages including the separate digital replica consent Cal. Lab. Code § 927 now requires. A 2019 release permitting "use of the recordings in advertising" does not authorize training a voice model on them.
- Name, Image, and Likeness Clearance Checklist — term, media, sell-off, and post-mortem treatment, itemized. Put "does this asset depict, evoke, or imitate a real, identifiable person?" on the intake form.
Audio. Generated music needs two clearances because recorded music carries two copyrights: see Two Copyrights, One Song, Clearing a Track, and the Music Clearance Checklist.
Theme 6 — What Copyright Will Not Hold: Trade Secret and Contract
When there is no copyright, two regimes still work, and both are indifferent to authorship.
- Trade Secrets and the DTSA: Protecting What You Cannot Register — the two elements, what courts count as reasonable measures, and the whistleblower notice at 18 U.S.C. § 1833(b) whose omission silently forfeits fees and exemplary damages. The right frame for a corpus, a tagging taxonomy, weights, prompt libraries, and unreleased design candidates.
- Building a Trade Secret Program That Survives Litigation — building the reasonable-measures record before you need it, including the onboarding hygiene that keeps a competitor's data out of your corpus.
- Trade Secret Protection and Departure Checklist — the inventory register and the seventy-two-hour departure protocol. The ML engineer who leaves with a corpus is the fact pattern of the next five years.
On the contract side, the house licensing materials transfer cleanly once you know what to change. How to Draft a Trademark License Agreement and the Trademark License Agreement Template supply the architecture — scope, territory, term, quality control, royalties, termination. An AI output license needs that skeleton plus three additions: an express acknowledgment that some output may not be protectable, a covenant not to assert rather than a bare assignment, and a training carve-out over the licensee's inputs.
Theme 7 — Platforms, Takedowns, and Shipping the Product
If AI output touches a product other people upload to, the safe harbor analysis is the one platforms have run since 1998, and the six-dollar designated-agent registration is still the highest-return filing in technology law.
- The DMCA Safe Harbor — the four harbors, the § 512(i) conditions, red-flag knowledge, and the repeat-infringer requirement where safe harbors actually die.
- Sending and Fighting a DMCA Takedown — fifteen stages including the documented fair use look Lenz v. Universal Music Corp., 815 F.3d 1145 (9th Cir. 2016), requires. Automated notices fired at fingerprint matches carry § 512(f) exposure, and an AI-detection false positive is the case that tests it.
- DMCA Takedown Notice Checklist — the six elements of § 512(c)(3)(A) and the put-back clock, from both sides.
- The Legal Layers of a Website, Launching a Website or App Without Legal Debt, and the Website and App Launch Legal Checklist — the pre-ship stack, into which AI disclosures slot: the training-rights term in your own terms of service, provenance metadata, and the state transparency statutes coming online.
Diligence deserves a warning. No buyer should accept a schedule of owned copyrights that silently includes generated material. The workflows are in Trademark Due Diligence in Mergers and Acquisitions and the Trademark Due Diligence Checklist; the question to add is "what did the human do," asked in writing, per asset.
A Suggested Reading Path
Pick the branch that matches your matter. Each is ordered.
Company using AI tools with no policy. (1) Who Owns What the Machine Made. (2) Generative AI IP Compliance Checklist, Phases 1-3, to inventory and tier. (3) Deploying Generative AI, Stages 2-6. (4) Copyright Registration Checklist for whatever is worth filing.
Training or fine-tuning a model. (1) Fair Use After Warhol. (2) Running a Fair Use Analysis, especially the model-developer playbook. (3) Fair Use Risk Assessment Checklist as the corpus register's cover sheet. (4) Trade Secrets and the DTSA, because the weights are the asset.
A demand letter about an output. (1) Deploying Generative AI, Stage 11, for the first fourteen days and the vendor tender. (2) Proving Copyright Infringement to test the claim. (3) Running a Fair Use Analysis, Stage 11. (4) Pre-Litigation Enforcement Checklist if you are the one sending it.
A rights holder whose work is in a corpus. (1) What Copyright Registration Actually Buys You — the § 412 calendar decides what your claim is worth. (2) Registering a Copyright, using group options for volume. (3) Filing a Copyright Infringement Complaint in Federal Court. (4) Copyright Infringement Complaint Checklist.
Launching a brand built with generative tools. (1) Choosing a Strong Trademark. (2) Trademark Clearance Searching, then the full search guide. (3) Who Owns What the Machine Made, for why you are filing a trademark rather than a copyright. (4) Specimen Refusals before the mockup becomes a specimen.
A synthetic voice, face, or persona in the pipeline. Read Your Face Is Not Public Domain first and the copyright material second; the mechanics are in Clearing and Licensing NIL and its checklist.
Primary Authorities
| Authority | Holding in one line | |---|---| | 17 U.S.C. §§ 102(a), 103(b) | Protection runs to original works of authorship, and in a derivative work only to the author's own contribution | | 17 U.S.C. §§ 411(a), 412 | No suit until the Register acts; no statutory damages or fees for infringement commenced before registration | | 17 U.S.C. §§ 1202(b), 1203(c)(3)(B) | Intentional CMI removal, at $2,500-$25,000 per violation | | 15 U.S.C. §§ 1051, 1127; 18 U.S.C. § 1839(3) | Trademark and trade secret protection turn on use and secrecy, not on who authored the thing | | Burrow-Giles Lithographic Co. v. Sarony, 111 U.S. 53 (1884) | An author is the human originator who gives visible form to a mental conception | | Feist Publications, Inc. v. Rural Telephone Service Co., 499 U.S. 340 (1991) | Originality needs independent creation plus minimal creativity; effort is not enough | | Thaler v. Perlmutter, 130 F.4th 1039 (D.C. Cir. 2025) | "Author" under the 1976 Act means a human being; a machine cannot be named as author | | Andy Warhol Foundation v. Goldsmith, 598 U.S. 508 (2023) | Factor one asks whether the particular use has a further purpose, weighed against commerciality and justification | | Thomson Reuters v. Ross Intelligence, No. 1:20-cv-00613 (D. Del. Feb. 11, 2025) | Training a competing non-generative research tool on headnotes was not fair use; on interlocutory appeal | | Bartz v. Anthropic PBC, No. 3:24-cv-05417 (N.D. Cal. June 23, 2025) | Training on lawfully acquired books was exceedingly transformative; retaining pirated copies was not | | Kadrey v. Meta Platforms, Inc., No. 3:23-cv-03417 (N.D. Cal. June 25, 2025) | Summary judgment for Meta on this record; market dilution identified as the live factor-four theory | | Stevens v. CoreLogic, Inc., 899 F.3d 666 (9th Cir. 2018) | Section 1202(b) requires knowledge that removal will induce a specific infringement | | Midler v. Ford Motor Co., 849 F.2d 460 (9th Cir. 1988) | Deliberate imitation of a distinctive voice to sell a product is actionable | | Tenn. Code Ann. §§ 47-25-1101 to -1108; Cal. Lab. Code § 927 | Voice is protected identity; digital replica contract terms without specificity are void | | 88 Fed. Reg. 16,190 (Mar. 16, 2023); Compendium §§ 306, 313.2, 621 | AI-generated material must be disclosed and excluded from a registration claim | | Directive (EU) 2019/790, arts. 3-4; Reg. (EU) 2024/1689, art. 53(1)(c)-(d) | EU text-and-data-mining exceptions with opt-out, plus training-summary and copyright-policy duties |
Forms and Templates
The Marksy template library is trademark-first, which is convenient here, because trademark is where AI-generated brand assets actually get protected.
- Trademark Assignment Agreement — Template — moves a generated wordmark with its goodwill, as the anti-assignment-in-gross rule requires. Pair it with a separate copyright assignment of the human contributions; one instrument does not do both jobs.
- Trademark License Agreement — Template — the architecture for licensing a generated brand asset, including the quality-control terms that keep the license from going naked. For an output license, add the three provisions in Theme 6.
- Trademark Cease-and-Desist Letter — Template, with Sending an Effective Cease-and-Desist Letter — what works when a generated logo is being copied and there is no copyright to assert.
- Trademark Portfolio Inventory — Template — repurpose as the generated-asset register: mark, first use, clearance date, tool, human contribution, registration status.
- Trademark Coexistence Agreement — Template — for two companies that independently generate similar marks in adjacent classes.
Related Toolkits and Checklists
- Fair Use and Permissions Toolkit — copyright, trademark, and publicity permissions in one workflow. Read it when the question is whether to license rather than whether to defend.
- Copyright Fundamentals Toolkit — ownership, registration, duration, and scope. Start here if § 203 termination and § 205 recordation are not yet familiar.
- Copyright Enforcement Toolkit — takedowns, demands, and litigation in sequence, including the Copyright Claims Board for small output claims.
- Right of Publicity and Personal Brand Toolkit — the right destination for a deepfake matter that arrives labelled as a copyright problem.
- Trade Secret Protection Toolkit — where a training corpus and a set of weights are actually defended.
- IP Due Diligence Toolkit — treats AI-generated material on a schedule of assets. Read it before signing a representation that the company owns its content.
- Website and App Launch IP Toolkit — the pre-ship sequence into which AI disclosures and provenance metadata slot.
- Trademark Clearance and Brand Selection Toolkit — the naming workflow, unchanged by the fact that a model produced the candidates.
- Music, Film, and Creative Industry IP Toolkit — generated audio and audiovisual work.
- Evidence and Expert Witness Toolkit — Rule 26(a)(2) timing and Fed. R. Evid. 702 as amended, which is how the machine-learning expert testifies or does not.
Related Documents
Articles
- Who Owns What the Machine Made — the doctrinal spine.
- Fair Use After Warhol — the factors and the training docket.
- Proving Copyright Infringement — the test an output claim must meet.
- Who Owns the Work — who owns the human contributions.
- What Copyright Registration Actually Buys You — the calendar beats the certificate.
- Your Face Is Not Public Domain — voices, faces, digital replicas.
- Trade Secrets and the DTSA — the fallback for corpora and weights.
- The DMCA Safe Harbor — hosting user-generated AI content.
- Assignments vs. Licenses — what an output clause conveys.
Guides
- Deploying Generative AI Without Losing Your IP — policy, provenance, vendor redline.
- Running a Fair Use Analysis — the clearance procedure.
- Registering a Copyright — deposits, disclaimers, group filings.
- Transfers, Licenses, and Termination Rights — the paperwork, with model clauses.
- Filing a Copyright Infringement Complaint in Federal Court — the rights holder's route.
- Sending and Fighting a DMCA Takedown — notices and § 512(f).
- Clearing and Licensing Name, Image, and Likeness — synthetic voice consents.
- How to Draft a Trademark License Agreement — the skeleton to adapt.
Checklists
- Generative AI IP Compliance Checklist — inventory to incident response.
- Fair Use Risk Assessment Checklist — the scored, signed file.
- Copyright Ownership and Chain-of-Title Checklist — with a generated-material audit.
- Copyright Registration Checklist — deposit to certificate.
- DMCA Takedown Notice Checklist — elements and put-back clock.
- Name, Image, and Likeness Clearance Checklist — releases for synthetic use.
Toolkits
- Fair Use and Permissions Toolkit — clearance across three regimes.
- Copyright Fundamentals Toolkit — the baseline this assumes.
- Right of Publicity and Personal Brand Toolkit — digital replicas and NIL.
- Trade Secret Protection Toolkit — indifferent to authorship.
- IP Due Diligence Toolkit — AI assets on a schedule.
- Website and App Launch IP Toolkit — the pre-ship stack.
Templates & Forms
- Trademark Assignment Agreement — Template — moving a generated mark with its goodwill.
- Trademark License Agreement — Template — the skeleton for an output license.
- Trademark Cease-and-Desist Letter — Template — the demand with no copyright behind it.
- Trademark Portfolio Inventory — Template — repurposed as the generated-asset register.
- Trademark Coexistence Agreement — Template — independently generated near-identical marks.
Across the Wider Corpus
The library now spans patents, trade secrets, data, and sector-specific practice. These sit outside this document's immediate subject and bear on it directly.
- Buying a Model: AI Vendor Contracts, Training Rights, Output Ownership, and the Indemnity That Is Not There — the doctrinal treatment of AI vendor contracts, training rights, output ownership, and the indemnity that is not there.
- Whose Invention Is It? Joint Development, Background IP, and the Ownership Default Nobody Wants — the doctrinal treatment of joint development, background IP, and the ownership default nobody wants.
- The Image Business: Photography, Stock Licensing, and Visual Content Rights — the doctrinal treatment of photography, stock licensing, and visual content rights.
- Negotiating an AI Vendor Agreement: A Practitioner's Guide to Training Rights, Output Ownership, Indemnities, and Model Governance — the operational steps for training rights, output ownership, indemnities, and model governance.
- Managing an Agency Relationship: A Practitioner's Guide to Pitch Material, Ownership, Third-Party Assets, and Termination — the operational steps for pitch material, ownership, third-party assets, and termination.
- Clearing and Contracting for Digital Replicas: A Practitioner's Guide to Consent, Scope, and Duration — the operational steps for consent, scope, and duration.
- Licensing and Clearing Visual Content: A Practitioner's Guide to Stock, Commissions, Releases, and Usage Terms — the operational steps for stock, commissions, releases, and usage terms.
- AI Procurement Checklist: Use Case Review, Data Rights, Output Terms, Indemnity, Evaluation, and Monitoring — the working sequence for use case review, data rights, output terms, indemnity, evaluation, and monitoring.
- Agency Engagement Checklist: Pitch and Spec Work Terms, Deliverable Ownership, Third-Party Asset Schedules, Approval Records, and Transition on Exit — the working sequence for pitch and spec work terms, deliverable ownership, third-party asset schedules, approval records, and transition on exit.
- AI Procurement and Governance Toolkit: Vendors, Models, Data, and Output — clause language and working templates for vendors, models, data, and output.
- Choosing Your Protection Toolkit: Patent, Copyright, Trademark, or Trade Secret — clause language and working templates for patent, copyright, trademark, or trade secret.
- Data Licensing and Rights Toolkit: Provenance, Scope, Derived Data, and Compliance — clause language and working templates for provenance, scope, derived data, and compliance.
This document is general information about the law, not legal advice, and does not create an attorney-client relationship. Trademark and copyright outcomes turn on specific facts. Marksy is not a law firm.