OUR SERVICES
What AI-assisted development means for who owns your code, what you can protect, what you can train on, and what you have to disclose.
Your engineers are shipping AI-written code. Your product may train on data you did not create, and your team is pasting who-knows-what into chatbots to debug it. Each of those habits creates legal risks that are no longer hypothetical. Courts have ruled on AI authorship and training-data fair use, regulators have imposed disclosure obligations, and litigants have lost trade secret rights over what went into a chatbot.
None of this means slowing down. It means knowing where the lines are, documenting the right things, and putting a handful of contracts and policies in place while they are still cheap. Here is the compliance picture for a software company whose software development runs on artificial intelligence.
Copyright ownership starts with a simple rule: copyright requires a human author. The courts have settled the pure case. In Thaler v. Perlmutter, the US Copyright Office’s refusal to register a work generated entirely by a machine was upheld on appeal, and the Supreme Court declined to disturb that result. Purely AI-generated code and other AI-generated output get no copyright at all. That means a competitor that copies it commits no copyright infringement. Contracts, trade secret law, and trademarks can still restrict what others do with it, but the automatic copyright protection your human-written code enjoys is simply absent.
For AI-assisted work, copyright protection follows the human contribution. AI-assisted code means code your developers write with an assistant’s help, and it is protectable where their human creativity and expressive choices shape it. Content the AI originated on its own is not. The uncomfortable end of that spectrum is what the industry calls vibe-coding. In vibe-coding, a program’s code is generated end to end from prompts, with minimal human oversight and little human authorship in the code itself. That program may be almost entirely unprotectable by copyright. In the Copyright Office’s view, prompts alone are human input that amounts to suggestions rather than authorship, no matter how many it took.
Three practical consequences follow. First, keep meaningful human authorship in the source code that matters most to you, and document the human contributions. Second, for AI-generated code, trade secret law becomes the default protection, because it has no human-creation requirement. The protection is conditional rather than automatic, lasting only as long as you actually keep the material secret and valuable. That makes confidentiality practices and contracts, not copyright, the fence around AI-generated code. Third, watch the other direction. AI-generated code can reproduce third-party code from public repositories, including copyleft-licensed code and other open source code. So open source license scanning and provenance checks on AI-written source code belong in the pipeline, to catch licensing violations before they ship.
The Copyright Office registers AI-assisted works, and works containing AI-generated material, under a framework set out in its artificial intelligence registration guidance. That framework distinguishes three situations. AI used as an assistive tool enhancing human expression does not limit copyright protection. Human-authored material that remains perceptible in the output stays protected. And copyright law can protect human selection and arrangement of AI-generated material as a compilation even where the pieces are not.
The obligations are concrete. Applications must disclose and disclaim AI-generated content that is more than minimal. Knowingly failing to do so can jeopardize the registration itself, up to cancellation or a court disregarding it in litigation. The registrations that succeed are the documented ones. Contemporaneous records of prompts, iterations, edits, and the human decisions that shaped the result have made the difference in close cases. For a software company, that argues for keeping authorship records the way you already keep commit history, and for auditing past registrations if AI-generated material went undisclosed.
If you train or fine-tune machine learning models, the early court decisions have drawn the lines that matter. In Bartz v. Anthropic, training on lawfully acquired books was held transformative fair use, but building a library of pirated copies was not. The piracy carried the liability, and the case ended in a settlement reported at $1.5 billion covering roughly half a million works, among the largest copyright settlements ever. Read the fine print of that settlement and you find the lesson for everyone else. It released only the input-side claims and expressly preserved claims over infringing model outputs. In Kadrey v. Meta, training won as fair use only because the plaintiff authors failed to prove market harm. The court warned that plaintiffs with developed market-dilution evidence may win the same fight. And in Thomson Reuters v. Ross, copying a competitor’s content to build a directly competing research tool was not fair use, though the court limited its holding to non-generative AI.
The variables that matter are acquisition, licensing, and market impact, and none of them is settled. The courts have even split on how much weight pirated acquisition carries on its own. What is consistent is that fair use runs out where training substitutes for the market of the works you trained on. Scraping adds a second layer, because platform terms of service may also support contract and unfair-competition claims regardless of whether the content was publicly visible.
For most software companies the exposure arrives through vendors rather than through training runs of their own. That makes training-data provenance a diligence item and a contract item: where the data came from, what was licensed, what opt-outs were honored. Back that with representations that continue past signing and indemnities that actually match the claims. A generic intellectual property indemnity is underfit for an AI deal. The risks worth allocating separately include training-data infringement, scraping and terms-of-service claims, output regurgitation, and disclosure-law violations.
Consumer-tier AI tools commonly reserve the right to use, and sometimes disclose, what users type into them. That sits at odds with the secrecy trade secret law requires. Feed confidential source code or data into one and you may have disclosed it in the legal sense. A federal court has dismissed a trade secret claim where the plaintiff had built her claimed secrets through a public chatbot, voluntarily disclosing them to a counterparty that owed her no confidentiality. And courts have ordered litigants to keep confidential material out of mainstream consumer AI tools entirely. Note the trap inside the trap: the tier that matters is defined by the contract terms on training, confidentiality, and retention, not by whether the account is paid. A vendor’s promise that you own the outputs does nothing to keep your inputs secret.
The fix costs little. Restrict confidential material to enterprise-tier tools with contractual no-training commitments, put the rule in a written AI-use policy, update employment and confidentiality agreements to cover AI tools and AI-generated output, and extend your ordinary security controls to AI workflows. One engineer debugging proprietary code in a personal chatbot account can undo years of trade secret accumulation. This is the cheapest serious risk to eliminate in the entire AI stack.
Disclosure and transparency obligations for artificial intelligence reach further than most teams assume, and more arrive every year:
If your product makes or supports decisions in those categories, or you offer a model others build on, map your features against these regimes. Date-stamp whatever compliance summary you rely on, because these laws have been amended mid-flight more than once. They require documentation: provenance records, dataset summaries, risk assessments. That documentation is far easier to produce if it was kept from the start, and it overlaps heavily with what good intellectual property hygiene requires anyway.
For companies that host user content or provide tools others might misuse, the Supreme Court’s decision in Cox Communications v. Sony Music reset the standard. Contributory copyright liability requires affirmatively inducing infringement or offering a service tailored to it, and merely knowing that some users infringe is not enough. That is meaningful protection for neutral tools and platforms. It is not a reason to relax DMCA hygiene, because the safe harbor is a defense that must be earned. A registered agent, expeditious takedowns, and a repeat-infringer policy you actually enforce still preserve it and still get cases dismissed early. And if you deploy AI agents that act autonomously, govern them as if their actions are your company’s actions, because that is how a court is likely to see it.
The common thread across all of this is documentation: of human authorship, of data provenance, of tool tiers, of policies enforced. Those records are cheap to create now and expensive to reconstruct later. They are exactly what a court, a regulator, or an acquirer will ask for. Sorting out which of them your company needs first is worth a conversation with experienced counsel before the next model ships, rather than after the letter arrives.
Not the portions the AI generated on its own. Copyright law requires human authorship, so copyright protection covers your developers’ human contributions: the code they wrote, the modifications they made, and creative selection and arrangement of AI output. A program generated almost entirely from prompts may be effectively unprotectable. That is why documented human involvement in your core code matters, and why trade secret protection carries more of the load for AI-generated code.
Sometimes, and the details decide it. Courts have held training transformative where the data was lawfully acquired, found liability where the copies were pirated, warned that proof of market harm could flip the outcome, and rejected fair use where the training built a directly competing product. Treat acquisition, licensing, and market impact as the live variables. Get provenance representations from any vendor whose models you build on.
Yes. Applications must disclose and disclaim AI-generated content that is more than minimal. Knowing failures can jeopardize the registration, up to cancellation or a court setting it aside in litigation. Works with AI-assisted but human-authored expression remain registrable, and past registrations that omitted required disclosures can be corrected by supplementary registration.
Increasingly, yes. California’s AB 2013 requires public training-dataset summaries for generative AI made available in the state, with no trade-secret exception, and the EU AI Act requires documentation and a public training-content summary from general-purpose model providers. State automated-decision laws add developer-to-deployer documentation duties on their own timelines. Build the provenance record before someone with leverage asks for it.
No. Using AI as an assistive tool does not reduce copyright protection for what your developers author, and the Copyright Office says so expressly. The line falls between AI-assisted human expression, which is protected, and AI-originated content, which is not. A policy preserves both your copyright and your trade secrets if it keeps human authorship in critical code, records how AI was used, and routes confidential material only through enterprise-tier AI tools.