Conditionally: the scope of copyright exceptions, opt-out records, website terms of use, and personal data rules must be analysed on a source-by-source basis. “Everyone does it” is not a legal basis.
For EU-facing training there is a specific gate: the text-and-data-mining exception in Article 4 of the DSM Directive applies only where the rightsholder has not reserved use in a machine-readable way, so honouring opt-out signals is part of having a lawful basis at all. And where a source contains personal data, you need a KVKK or GDPR basis for it as well: clearing copyright does not clear data protection, which is a separate gate. That is why the analysis is genuinely source-by-source, and why a blanket scrape is not defensible.
Shall we apply this matter to your situation?
Tell us your specific situation in a few sentences; we'll assess it with the right team.