49 comments

[ 2.7 ms ] story [ 121 ms ] thread
Personally, I don't think this was a complete/valid test - Yandex could be censoring other things. The input test data primarily covered things that would fit a specific narrative, not one for the 'other side'. Has anyone done a more complete test?
Yes. I think the conclusion is that you would do well to use both Yandex and a typical western one for anything political.
I think the validity is mostly scoped the current major world-wide events. It would be interesting to have some sort of topic censorship ranking website; give it a topic, and it ranks various search engines.
Can you give examples of queries for the "other side"? I'd like to test them out.
Zelensky is a hero -- should return five results in which the author considers Zelensky to be a hero.

Russia thousands arrested -- should return five results about arrests at recent war protests in Russia

Russian casualties in ukraine -- should return five results listing figures in the thousands.

This is why the index shouldn't be full of crap. A smaller collection, curated by the user, is where it's at.

Maybe you should do a pull request!

This is a little silly when you consider volume. A search engine should crawl millions of pages and sites, and over time develop quality indicators so the crap can be filtered out. Anything that works equivalently to Google's algorithm will suffer from the same crap infection as Google results, and to a lesser degree any of the other top 20 engines and meta-search tools.

Something like ublock lists and p2p sharing in closed communities could provide semantic indexing and gradual manual curation, so with a few thousand people you might get novelty + quality in an engine. If you stick to a small, curated list of sites, you won't get site level novelty, and quality will eventually suffer relative to searches that scan over new things.

Even then, consider reddit - build a reddit search that's as good as Google. That updates once an hour.

Sharing bookmarks isn't feasible when you want the functionality of a general purpose search engine.

Oddly enough, Yandex pass these tests with flying colors.
I expect any search engine worth using to filter out spam and low effort crap. Yandex does neither.
neither does google these days
But this gets me thinking.

"Censored" is definitely the wrong word here; it implies that there is some objective pool of information that we could all be getting to if it weren't for the search engines getting it wrong, and that's just not it.

It's definitely much closer to -- imagine 5 hypothetical bookstores owned by different people. Each has to make decisions on what to choose to carry based on their business interests, and so they do. But there's nothing I would call "censorship" about that.

And if we're deciding it is, then we have to be much more serious and intelligent in what to do about it. I'm thinking something like the EPA or FDA for search engines. If they have to tell us whats in the ketchup, then they have to tell us what's in the "search engine algorithms."

When a company gets large enough, whether in the US or Russia, "good relations with the government" becomes one of your aforementioned "business interests". Call it whatever you want.
>"Censored" is definitely the wrong word here; it implies that there is some objective pool of information that we could all be getting to if it weren't for the search engines getting it wrong, and that's just not it.

1. I think that when it comes to "search engine censorship", most people know it's about search engines downranking/hiding certain results, not them literally going around burning books or blocking/taking down websites.

2. The definition for "censorship" is "The use of state or group power to control freedom of expression or press, such as passing laws to prevent media from being published or propagated". When it comes to the internet, getting downranked/hidden from search engines does a pretty good job at preventing such information from being "propagated", even though you could still access it via the deep web. As a thought experiment, suppose the government banned 1984 from being sold/read anywhere, save for you going to the library of congress in person. You can theoretically still read it, but at great difficulty. It's certainly more work than reading harry potter or whatever. Would you say such measures are not "censorship"?

I think there are pretty clear examples of censorship that would cross the line.

The most widely recognized example is when major search engines delisted pictures and results for the Tiananmen Square massacre last year.

The bookstore analogy would be if you came in and asked for a book on the massacre, and the owner said "sorry, I don't know what you are talking about" when they have a whole shelf in stock.

yes great insight! A better word might be 'curated' - we are getting access to preselected set of information from which we can further select
This is interesting because I tried Yandex for the first time after this DDG debacle, and I found it solved a major problem I have with every other popular search engine - Yandex respected my actual query and gave me the correct results every time, instead of trying to mind read me and mangle the results like the other ones. Google always does it, DDG started doing it some time ago, and unfortunately Brave search also seems to do it.

I understand there's probably a reason others do it - maybe that that's what's needed to serve a general audience, they're optimising the results for my 70 yr old dad and 10 yr old niece. But that does mean that using them is a frustrating experience for me half of the time, of quoting everything and micromanaging the query terms so that they understand that I do mean this specific thing rather than this overgeneralised idea they seem to parse from it. Yandex was a breath of fresh air in that respect.

In search engines and in so many other applications, we desperately need a checkbox that says "I know what I'm doing, don't assume anything, give it to me straight". Expert modes, literal modes, they need to be everywhere.
Google has sort of added this - but they hide it so well I didn't even know it existed until a few weeks back.

If you do a search, click tools, click the "more" dropdown, and choose "Verbatim" then the results are actually pretty close to "don't fucking try to read my mind".

They claim the following about Verbatim (although this info is from 2011, and it may be different now)

----

This verbatim search removes personalized, corrected, suggested, related, and non-inclusive results.

On the verbatim page, users will see only results that:

Include all their search terms.

Match their exact spelling.

Use the same tense (e.g., “is” and “was” will be seen as distinct).

Use the same verb form (e.g., “swimming” and “swim” will be seen as distinct).

Use the same plural vs singular form (e.g., “hat” and “hats” will be seen as distinct).

> click tools, click the "more" dropdown

It's under Tools -> "All results" dropdown for me.

I tried to see if it was possible to bookmark to get a verbatim search directly, and while I couldn't decipher the obfuscated query they seem to be using these days, the overall query length was much shorter than with the original non-verbatim search. I wonder if that has anything to do with the:

> This verbatim search removes personalized, corrected, suggested, related, and non-inclusive results.

claim. The original search had a very long 64 bit encoded parameter, which perhaps had some context information - it seems plausible it had some typing history information (perhaps taken from their Google Docs code), maybe how fast I typed (for error probability), some kind of personally identifying ID about my current session perhaps. Looking at these query strings and how opaque they are, feels a bit creepy that all this information is being sent from my computer without my knowledge or significant control. Makes me glad I rarely use Google these days.

Using quotes in DDG used to work for me. Unfortunately, for a couple months now it doesn't. The results are garbage. Often they don't even include most of the words I typed, only the first one.
I figure to be a more complete test you'd have to include Baidu the Chinese search engine, which is in no doubt censored. Or old school engines like ask.com, and Yahoo even though they serve Bing results to my understanding it would be nice to include them. Either way, good showcase, great example queries. Although I do think some should be catering the MSM narrative to see if the results give contrary results included.
I think in the future people will just resort to checking yandex/baidu/duckduckgo every time they think the results might be censored. The opposing regimes are likely to censor complimentary parts of the web.
I've been doing this for awhile. Most importantly, I do this with my news. If there's a war in Europe, I don't even bother to check the BBC anymore, or German outlets. I try to grok Al Jazeera, India Times, Channel News Asia, see what's common between them.
> Results for "evidence of election fraud 2020" - at least 5 results showing major election fraud took place.

Shouldn't an "objective" search engine return results that are true (that there was no systemic election fraud) instead of results that fit his personal biases?

(comment deleted)
That query is biased to begin with. Let’s say I want to understand the flat earth arguments. Should the search engine fill my feed with “earth is not flat” when I’m trying to search for “proof that earth is flat”. This is different from searching “is earth flat” which is an attempt to discover the truth.

Your claim that the search engine should show results which indicate that there was no systematic election fraud, is a wishful thinking that you don’t want people to believe in it, not a good decision making regarding what search engines should display.

> Should the search engine fill my feed with “earth is not flat” when I’m trying to search for “proof that earth is flat”.

What should it do? Article titled "There is no proof that the earth is flat" also satisfies the query completely from a keyword perspective, although not from a semantic. But if there is no actual proof that earth is flat, any article claiming there is would satisfy the query only from a keyword perspective but technically not semantically. It is a good question.

An objective search engine should return the results that are most relevant to your query without attempting to be the arbiter of "truth". An objective search engine doesn't try to weigh in on disputed events like the existence of systemic election fraud (which has been a regular occurrence in every election since before any one of us was born, despite the exhortations of the DC blob and legacy media outlets). For example, a Special Counsel in Wisconsin recently did a report on 2020 election fraud. If you Google "Wisconsin report on election fraud" the front page is comprised entirely of DC blob legacy media outlet editorials about why the Wisconsin report is debunked without a single link to the report itself. You are within your rights to think that the Wisconsin report is total garbage, but if you go to a search engine and search for "Wisconsin report on election fraud" an objective search engine will have a link to the actual report on the front page, if not as the first result.
DDG for me is a mixed bag on that query, some links on how it’s been debunked but more than half promoting the report…

It I mean it would be nice if it surfaced whatever actual report you’re talking about, but maybe it’s just not the best query for that? I’m sure I could find it by clicking through any of the links

A search engine should return what I'm searching.

If I type "the daily stormer" (a very unique name which corresponds to something very specific, i.e. the website for the far-right online newspaper by that name), I expect to find that thing in the first 5 results. That's not the case with Google; it's just not there at all. You have articles from other papers talking about it, you've got a wikipedia page talking about it, you've got sites like the ADL and the SPLC talking about it, but not the actual thing itself. Bing, Yahoo, Yandex all return it as the first result.

This.

When I “search” for something, I am searching for it. I mean, it’s the definition of the word “search” for crying out loud. Somehow over time this idea of a “search engine” got perverted into something that either tries to predict what I don’t know I really want or what someone else thinks I need to see.

How can a search engine know what is true?
> No search engine, not even the Russian based Yandex, returned an objective article that included the reasons laid out in Putin's invasion speech. [...] The level of disinformation I observed for this test was terrifying. Based on what I saw, I don't believe it is possible for an uninformed person to learn the actual reasons Russia invaded from a web search on any search engine.

Really, you consider Putin's stated reasons to be the actual reasons for the invasion? And the linked article in that paragraph is just a pile of misinformation (even calling it subjective would be a complement).

a search engine is meant for searching text about whatever I ask it. I don’t need it to vet the result based on geopolitics. If I search for “Putin invasion speech” I want articles that show me what Putin said about the invasion ranked by various technical metrics but not by whether the world thinks Putin is a liar and then giving me articles with what others said about Putin’s speech
I do believe that Google’s ‘quality’ of results has been declining in recent years, but:

This article is biased, cherry-picked nonsense. Is anyone actually looking at the results of his example queries and evaluating the quality themselves before responding?

To go with your example: the first three results a Google search returns for me on ‘Putin Invasion Speech’ are two NY Times articles analyzing the speech, and then a full translated transcript from Bloomberg. It’s hardly hidden.

On the author’s own example, “Why did Russia invade Ukraine? 2022”, the first article is from the BBC and tries to lay out the full historical context - and mentions Putin’s speech.

The rest are from mainstream news sources of various descriptions, again laying out historical context.

The author may not _like_ that these are trusted, mainstream, popular sources - but that’s what Google has _always_ strived to return. Pagerank was never an ‘unbiased’ thing that searched solely for page content - and, wow, can you imagine the toxic SEO nightmare we’d be in if it was…

> The author may not _like_ that these are trusted, mainstream, popular sources

Because that's not what I'm searching for. Even if we all agree all the Russian sites are propaganda, I still want to be able to search through it. It's a political decision by Google to filter out these results. It's not algorithmic because no matter how precise your search query is, you will never get results from certain forbidden domains. Google just directly removed rt.com from results recently, so even if you search for "rt.com" you will not see it. I'm sure they have similar filtering based on certain keywords.

But that wasn't the author's search query, it was "Why did Russia invade Ukraine? 2022" - and Putin's speach is not a good answer to the for that, even ignoring the the inherent bias.

I specifically didn't call out the complaint about missing results when searching for direct quotes from an article (event thought they are also blatant misinformation), because that indeed is an issue. If I search for some pronoun X then I expect to find a known website with the name X. But if I search for a general term, then the only reasonable way to thing that a search engine can return is websites that fit the consensus view around that topic.

A Google search for “rt.com”, or ”Russia Today”, returns rt.com as the first result for me. On the google.com home page, starting to type ‘rt’ will have rt.com as the first autocomplete suggestion.

Did you try this out?

Yes, I found articles on the first page of DDG results providing Putin's own reasoning in his own words. The author just doesn't like the analysis and commentary surrounding it.

Take a look at who the author of this "comparison" of search engines is. His next post is describing Zelensky as a neo-Nazi. He claims Putin will exit Ukraine once these "Nazis" have been removed. He says he trusts Putin more than Biden and that the US was responsible for the invasion.

Actually, if you open the post and scroll down enough to see one of the other posts published by the author (look for the post titled "Why did Russia attack Ukraine..."), then you would realize that the entire website has piles of misinformation.
Search engines turn searches into money. Some do this by selling ads, some by selling your data. Unless you’re paying for it you’re bound with their database as the product and are mostly irrelevant to them besides spending more time there. They have an incentive to delay you finding what you want.
Yet here we are with no usable alternatives because competitors get crushed
I suggest anyone who read this post to scroll down the webpage, and once past the comment section - see the first article that was previously posted by the author, titled "Why did Russia attack Ukraine...".

Reading that post should tell you how smoothly you're being manipulated.

I’ve found in the past that Yandex’s reverse image search is vastly better than Google’s. I am almost always able to find things with Yandex that I cannot with Google.
Is deranking considered censorship? On what basis does RT need to be the top result?

Now of course I know its been purposely deranked. But couldn’t google claim its not censored if it exists but buried in results?

I want to know what legal basis we have for asserting certain sites must have the highest rank on google for certain searches. I’m assuming there isn’t one.