There was a similar system that showed up a few months ago called blogsurf[1]. Their focus was on a new ranking algorithm, but both dato and blogsurf seem to use the RSS feed as a data aggregation API. This could be a big opportunity to create a search engine which avoids most of the SEO-spam, I'm not sure if you would need some sort of human-in-the-loop validation process or not. Either way, I can't wait to see the first group that offers something similar while indexing ~100k+ feeds and making an effort to cover the long-tail.
I've dabbled with using RSS feeds to inform the crawling of my search engine as well.
I think the big issue with leaning too heavily into RSS without much human intervention is that many CMS platforms generate an absurd number of RSS feeds, and it can be hard bordering on impossible to figure out which of them is canonical. Some (but not all) of these sites also tend to be fairly spammy low-value domains.
I understand that nitter serve a purpose. However, for OP who is trying to crawl high quality content such as blogs and news automatically, nitter is a nightmare.
> This could be a big opportunity to create a search engine which avoids most of the SEO-spam
As soon as it's worthwhile, there will be RSS spam targeting search engines.
Many companies already have near zero-useful-content blogs because having blobs of text that search engines consider fresh increases their ranking. If RSS-based search becomes a thing, you'll get that kind of content pushed to it.
Or, if that doesn't work because of feed moderation or something, expect getting sponsored entries in your favorite indie blog's feed. Sponsored content is already a thing, but mostly to get "yes-follow" links. Now it will be to get content pushed to RSS search engines.
> As soon as it's worthwhile, there will be RSS spam targeting search engines.
That can be solved with a web of trust.
I don't mind showing my friends which RSS feeds I subscribe to. But I wouldn't let all of them see my browsing history (if I even kept one, which I don't). I also might not mind them seeing some kind of normalized ratio of viewing time for articles from each feed.
I would love to be able to rank my search results based on my friends-friends-friends opinion of the trustworthiness of any particular feed.
Top-level domain isn't the right trust granularity. Feed probably is.
I think a lot of the problems with domain ranking is domain expiry and dead links. It's extremely susceptible to domain ownership changes. If you monitor high value domains for expiry and then immediately purchase them and fill them up with links to your websites, you will effectively inherit that website's ranking power.
HTML does not provide any information of when a link was created, and DNS and HTTP does not provide any information bout how long a website has been around or ownership changes. HTTPS isn't much use either. Maybe you can coax it out of a registrar, but the information is not readily available and not easy to interpret. It's really a bit of a design flaw.
The WWW was just never designed with this in mind, it was designed around long-lived academic and government institutions, and it really struggles to deal the ephemeral nature of the public and commercial web adequately. I'm not sure if RSS feeds are much better.
Not really sure what's a good and scalable fix to this that allows the linked resources to still be mutable by their owners.
You are suggesting that Google and Bing don't account for domain ownership changes in their "link juice" algorithm.
I find this to be a totally implausible claim.
The high-ranking domains all have enough data in their WHOIS to detect an ownership change. Even without that you can detect a missed renewal date and pause the emission of link juice to new links appearing on that domain.
Which one do you think it is, considering both the functionality of the program and the high possibility of the author not being a native English speaker?
For millions of people 'dato' simply means 'date', as in day of the month.
The geographical first reference from the name of the author (Davide Sant'Angelo) suggests 'dato' must be the singular for 'data' - "piece of data".
What some number of people decide in their spare time about the nicknames they will give to things - from cockney rhyming slang to extra-academical trends -, remains their own business, and should not bother us. A "private" business, not in the default terms of their right to keep it private to us, but of our right to remain uninvolved.
Given the number of words in all human languages, chances are someone, somewhere will at some time take exception to your name choice no matter how carefully you chose.
22 comments
[ 3.2 ms ] story [ 74.8 ms ] thread[1]: https://blogsurf.io/
I think the big issue with leaning too heavily into RSS without much human intervention is that many CMS platforms generate an absurd number of RSS feeds, and it can be hard bordering on impossible to figure out which of them is canonical. Some (but not all) of these sites also tend to be fairly spammy low-value domains.
As soon as it's worthwhile, there will be RSS spam targeting search engines.
Many companies already have near zero-useful-content blogs because having blobs of text that search engines consider fresh increases their ranking. If RSS-based search becomes a thing, you'll get that kind of content pushed to it.
Or, if that doesn't work because of feed moderation or something, expect getting sponsored entries in your favorite indie blog's feed. Sponsored content is already a thing, but mostly to get "yes-follow" links. Now it will be to get content pushed to RSS search engines.
That can be solved with a web of trust.
I don't mind showing my friends which RSS feeds I subscribe to. But I wouldn't let all of them see my browsing history (if I even kept one, which I don't). I also might not mind them seeing some kind of normalized ratio of viewing time for articles from each feed.
I would love to be able to rank my search results based on my friends-friends-friends opinion of the trustworthiness of any particular feed.
Top-level domain isn't the right trust granularity. Feed probably is.
HTML does not provide any information of when a link was created, and DNS and HTTP does not provide any information bout how long a website has been around or ownership changes. HTTPS isn't much use either. Maybe you can coax it out of a registrar, but the information is not readily available and not easy to interpret. It's really a bit of a design flaw.
The WWW was just never designed with this in mind, it was designed around long-lived academic and government institutions, and it really struggles to deal the ephemeral nature of the public and commercial web adequately. I'm not sure if RSS feeds are much better.
Not really sure what's a good and scalable fix to this that allows the linked resources to still be mutable by their owners.
I find this to be a totally implausible claim.
The high-ranking domains all have enough data in their WHOIS to detect an ownership change. Even without that you can detect a missed renewal date and pause the emission of link juice to new links appearing on that domain.
What some number of people decide in their spare time about the nicknames they will give to things - from cockney rhyming slang to extra-academical trends -, remains their own business, and should not bother us. A "private" business, not in the default terms of their right to keep it private to us, but of our right to remain uninvolved.
Since you pointed this out, can we assume you're in the latter group?