Hacker News (curated)new | past | comments | ask | show | jobs| show hidden

To be clear, making pages with bad training data for bots won't make the bots go away.

It'll just punish the bad actors running the scrapers. As the original poster mentions that they are using TVs as proxies to get residential IPs, one really can't think of these bots as criminal enterprises.

Sadly, if the bad actors has two cents for brain, they'll limit how much importance each domain name can have on training data. To mitigate impact of bad data like this.

(Note I'd suggest only linking to them from robot.txt as pages to not be indexed, that way no human or well behaved not ever will see them, which is kind of the point).



> … making pages with bad training data for bots won't make the bots go away. It'll just punish the bad actors running the scrapers.

Exactly. I can't hope to keep them all at bay, but I can at least have the petty little victory of making their visit less convenient than it might otherwise be.

> if the bad actors has two cents for brain

I suspect that a majority of them are little better than the script kiddies of yore, running tools with minimal understanding of what is actually going on.

> I'd suggest only linking to them from robot.txt as pages to not be indexed

Agreed. Blocking all bots from all pages, well those that bother to listen to robots.txt. All bots because pretty much all of them are scraping for AI and similar these days, even googlebot. If I want people to see my stuff they'll get a link, and maybe they'll pass it on further, but all indexers/trainers can get stuffed. I'll likely make an exception for archive.org and similar.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact | github