I found out that companies may be using personal data and online content to train AI models, possibly without clear consent. I’m trying to understand the laws around AI training data, privacy, copyright, and whether people can opt out or take legal action.
Publicly visible does not mean legally free to use. In the U.S., AI training can involve separate privacy, contract, and copyright issues, and there is no universal right to opt out. A company may argue copyrighted material was used fairly, but that depends on how it was copied, what the model produces, and market harm. Privacy rights vary by state; California residents, for example, can request access or deletion and opt out of certain sales or sharing, though those rights do not automatically ban every form of AI training. Save the relevant terms and privacy policy, submit written deletion and objection requests, and document any output that closely reproduces your work or exposes personal information. Legal action is possible, but usually requires identifiable misuse, actual harm, or ownership of the copied material, not merely suspicion that it entered a dataset.
Don’t assume deleting a post later removes it from an already trained model. @rocketlogic’s written-request advice is useful, but the stronger issue may be how the company obtained the data and what it promised at the time. Scraping behind a login or violating a stated “no training” policy can create a clearer contract or consumer-protection claim than merely using a public post. The FTC has warned that breaking privacy promises may even require deletion of models built from unlawfully obtained data. Save the terms that applied when you posted, since companies can rewrite policies later.
If the material was created within your job duties or posted under a license broad enough to cover machine learning, you may not be the person with the strongest copyright claim. Your employer, client, publisher, or the platform may hold some of the relevant rights. You could still have privacy rights in personal information contained in the material, but ownership and privacy are separate questions.
That distinction gets lost because “my data” can mean several different things:
- Original expression, such as an article, illustration, photo, or substantial piece of code
- Personal information, such as your name, location, account history, or private messages
- Sensitive information, such as biometrics, health details, precise location, or children’s data
- Confidential material, such as an unreleased product plan or customer file
Each category follows a different legal path. Copyright generally protects original expression, not facts, ideas, methods, or a person’s general style. Privacy law may cover identifiable information even when the post is not copyrightable. Confidentiality may depend on an agreement or on whether reasonable steps were taken to keep the material secret.
I agree with @smartminer that the terms in effect when the data was collected are more useful than whatever the company’s policy says today. I would not treat a login screen as the dividing line, though. Restricted access can strengthen a contract or unauthorized-access argument, but it does not automatically make every scrape illegal. Likewise, public access does not automatically authorize every later use.
“AI training” is really a chain of activities: obtaining the material, retaining it, making training copies, processing it, creating model weights, and generating outputs. A company can have a defensible position at one stage and a problem at another. For example, it might lawfully possess a photo but still create an infringing output from it. Or it might produce no recognizable copy while having collected the underlying personal information through deceptive privacy practices. The FTC has specifically warned companies that secret changes in data use or broken promises about training can lead to enforcement, including orders affecting models derived from improperly obtained data.
Deletion has a technical limitation too. Removing your original record from a database is different from reversing its effect on a completed model. Privacy laws may require deletion of identifiable source records, copies, or inferences in certain circumstances, but they do not create a universal “untrain me” command. If information has genuinely been deidentified, a business may not be required to reconstruct your identity merely to locate it. Publicly available information and various internal-use exceptions can limit deletion rights as well.
A useful request should therefore be narrow and concrete. Identify the account, files, posts, dates, and email addresses involved. Ask whether the company collected the material directly, bought it from a broker, received it through a partner, or scraped it. Then ask whether it was used for pretraining, fine-tuning, evaluation, retrieval, moderation, or human review. Those are technically different uses, and a company may answer “not used to train” while still retaining the content for evaluation or retrieval.
For a copyright issue, preserve the original work, publication date, metadata, and any model output that reproduces distinctive portions. Registration matters more than many creators realize. Copyright exists automatically when an original work is fixed, but registration is ordinarily needed to sue over a U.S. work, and timely registration can affect access to statutory damages and attorney fees. A vague claim that the model “sounds like me” is usually much weaker than evidence of repeated text, unusual errors, watermarks, characters, or other protected elements.
For privacy, focus on identifiability and sensitivity. A dataset containing your public username and forum comments presents a different case from one connecting those comments to your home address, faceprint, medical condition, or private messages. State of residence matters heavily. California residents now have the DROP mechanism for deletion requests to registered data brokers, with brokers required to begin retrieving and processing those requests as of August 1, 2026, subject to exceptions. That still does not cover every platform or every AI developer.
So yes, companies can legally use some personal data and online content for AI training without obtaining a separate checkbox consent every time. That is not blanket permission. The practical question is which rights attach to the specific material, what the company represented when it collected the material, how it obtained and used it, and whether you can document an actual violation rather than infer one from the existence of the model.
Realistically, most of these cases die at the proof stage, not the law stage. @stacktigersync is right that ‘my data’ splits into different claims, but unless you can show an actual output tracing back to your stuff or a broken promise in writing, you’re left arguing that it probably got scraped. Probably isn’t a case.
A DMCA notice is a poor tool for demanding that a model “forget” your work. It is designed to identify removable infringing material, while a privacy request concerns identifiable information and a contract complaint concerns promises the company made. Those routes overlap far less than people assume.
I partly disagree with @aieagle9663hq on proof. A matching output is strong evidence, but it is not the only useful kind. The company’s own response may confirm collection, retention, or a training use that conflicts with its privacy policy. Regulators have specifically warned AI companies about secretly using customer data contrary to their commitments.
Send separate, narrowly worded requests rather than one broad “remove my data” demand. Ask privacy staff what identifiable information they hold and how it was used. Send copyright staff exact URLs or outputs containing copied expression. Send any broken-promise complaint with the policy language and date attached. Using the wrong process often gets you a technically correct but practically useless rejection.
A model trained on your vacation photo and a model that uses your profile to deny you a loan may involve the same data, but the second situation gives you a much clearer target. Training disputes can get stuck on how the data was collected and whether it can be traced. A real decision affecting your credit, job, housing, or insurance creates a separate issue, even if the original training was legal.
That is where I’d push back a little on @aieagle9663hq. You do not always need an output that copies your post. If a lender uses an AI system to reject you, it still has to provide specific reasons rather than blame an unexplained black box. Other consumer-protection and discrimination rules may apply depending on the decision and location.
So if the AI actually affected you, don’t focus only on getting “untrained.” Save the rejection or decision notice, ask what information was used and where it came from, and dispute anything inaccurate. A challenge to the concrete decision may be faster and stronger than trying to prove your content existed somewhere in a giant training set.
If the terms you accepted contain a binding arbitration clause and class-action waiver, even a decent claim may have to be pursued individually instead of in court. That does not make questionable AI training legal, but it can make enforcement too expensive to be practical, so check the dispute section before paying a lawyer to analyze the training issue.