Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Given that we have AGI around the corner cough, building a model which discriminates SEO-spam from original content might be viable, don't you think so? Or for every new domain in the index you check whether it's legit or not by human intervention, then repeat each year and build a blacklist, similar to how blacklisting IPs works for Email (how do these sites even get listed? There should be networks apparent!). One could also put a button: "report abuse" and make a process to unblacklist legit sites (the same process for email works very well with most providers... except Google.)

My theory, why this is not happening: I guess that most of these SEO-spam sites are actually including Google-funneled ads, so this means there's something wrong there too...



AGI may initially reduce the incidents of SEO spam, but I very much doubt that it can eliminate it. The same can be said for human vetting. The thing to keep in mind is that SEO is performed for different reasons and takes many forms, so developing a universal model will likely be impossible. When the current forms of SEO spam become less effective, more effective forms will be adopted. Any form of filtering also presents the problem of false positives. More aggressive filters would likely produce more false positives.

Blacklists are also problematic. I have an email account with a smaller provider that I never bother to use since there is a good chance that their servers are blacklisted at any given point in time. Getting their server removed from these lists is non-trivial in most cases. Once they are removed, they usually get relisted within a few months. The problem also runs in the opposite direction: I have corporate email accounts where every external email (including those from certain departments of their own organization) is labelled as such and as potentially suspicious simply because blacklists are not sufficient. The only reliable outcome of blacklisting is the reduced reliability of communications channels.

User reported abuse is even more problematic since it opens up avenues for abuse. While it may be relatively easy to filter out bad reports in situations where there is a minimal vested interest (e.g. finding something disagreeable), that won't be the case when there is a considerable vested interest (e.g. attacking the competition).


Can GPT-3 be flipped on its head, to remove such content, rather than create it?


wait, isn't the point of AGI that it reasons like a human? E.g. for a simple (SEO-plagued) query like "disassembly of consumer electronics X", I get actual, worthwhile content and not a 5min video-ad? Basically every mildly intelligent human can discern this content from fake content, AGI should be able to do the same imo...


Then you have hundreds of SEO companies putting their own AGI to beat google one.


I mean, can we really call it an intelligence if it wants to work in marketing?


Obviously, because it has figured out the highest return for the least effort.... Get that "agi" a contract, please.


What if the SEO companies keep improving their AGIs with the aim of beating Google's AGI and in the process inadvertently start producing high quality content? Will SEO companies morph into knowledge mining and organizing companies (the business Google is supposed to be in)?


Obligatory xkcd:

https://xkcd.com/810/


Sounds like a GAN with more steps


Browser extension uBlacklist* can "Blocks specific sites from appearing in Google search results". It also supports DuckDuckGo and Startpage. You can add subscription like those ad-blocking extension. Maybe this is what you want.

*: https://github.com/iorate/uBlacklist




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: