SSS · AI Data Sets, Copyright & KVKK/GDPR

Can we train a model with data we collect from the internet?

Conditionally: the scope of copyright exceptions, opt-out records, website terms of use, and personal data rules must be analysed on a source-by-source basis. “Everyone does it” is not a le…

Updated · July 20261 min readCategory · AI Data Sets, Copyright & KVKK/GDPR
Short answer

Conditionally: the scope of copyright exceptions, opt-out records, website terms of use, and personal data rules must be analysed on a source-by-source basis. “Everyone does it” is not a legal basis.

Conditionally: the scope of copyright exceptions, opt-out records, website terms of use, and personal data rules must be analysed on a source-by-source basis. “Everyone does it” is not a legal basis.

For EU-facing training there is a specific gate: the text-and-data-mining exception in Article 4 of the DSM Directive applies only where the rightsholder has not reserved use in a machine-readable way, so honouring opt-out signals is part of having a lawful basis at all. And where a source contains personal data, you need a KVKK or GDPR basis for it as well: clearing copyright does not clear data protection, which is a separate gate. That is why the analysis is genuinely source-by-source, and why a blanket scrape is not defensible.

Shall we apply this matter to your situation?

Tell us your specific situation in a few sentences; we'll assess it with the right team.

Get in touch
This content is for general information only and does not constitute legal advice. Please contact our team for an assessment of your specific circumstances.
Categories
AI Data Sets, Copyright & KVKK/GDPR

The right start means a predictable process.

From the first meeting to completion of the work; let's plan every step transparently.