this post was submitted on 27 Nov 2024
65 points (100.0% liked)

TechTakes

1480 readers
342 users here now

Big brain tech dude got yet another clueless take over at HackerNews etc? Here's the place to vent. Orange site, VC foolishness, all welcome.

This is not debate club. Unless it’s amusing debate.

For actually-good tech, you want our NotAwfulTech community

founded 1 year ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] Breve@pawb.social 1 points 3 weeks ago* (last edited 3 weeks ago) (1 children)

My point is that you can't talk about usage rights of a dataset without talking about a specific use case. The suggested use case was to provide a static test dataset for systems developed to use the firehose API, but the dataset could be used for literally anything from making funny memes (fair use) to training a LLM model (arguably not fair use). Does the existence of an illegal use case automatically mean the dataset itself should be illegal though?

As a collorary, a photocopier can be used to create unauthorized reproductions of copyrighted works. Should making and disturbing photocopiers be illegal because they are capable of and used in the process of violating copyright law, or should we accept the photocopier absent of a use case isn't breaking any laws and go after the people who use them to illegally create unauthorized reproductions?

[–] davidagain@lemmy.world 5 points 2 weeks ago

A data set isn't like a photocopier in any meaningful way.

It's not a tool, it's information, and some of it counts as personal data under EU data protection laws.