Sang's Blog

Search engines are broken

Search used to be simple. You typed words into a box, and a machine gave you a list of pages that contained those words, ranked by how many other pages linked to them. You clicked a result, left the search engine, and read what someone else had written. The search engine was a map. The web was the territory. That arrangement is gone. Google does not want to send you to a website anymore. It wants to answer your question directly, using AI trained on the very websites it no longer sends you to. The AI Overview now appears at the top of many searches, synthesising dozens of pages into a single paragraph. You get the information without ever visiting the source. Google gets your attention and your data. The website that did the original research gets nothing. The map has consumed the territory.

But the AI Overview problem, while real, is a symptom of a deeper structural condition. The search engine market appears competitive on the surface — dozens of alternatives, each with its own pitch about privacy, neutrality or European values. Yet almost all of them share a single dependency that undermines their claim to independence: they do not own their own index. A search engine without its own index does not have a product. It has a user interface. It submits queries to Google or Bing, reformats the results, and presents them as its own. The privacy improvements are genuine. The independence is not. You cannot compete with a company you depend on for your raw material.

The reason almost no one owns their own index is that building one is one of the hardest problems in computing. You need to crawl hundreds of billions of pages continuously, keep them fresh, parse their content, extract meaning, rank them against every conceivable query, and return results in under a second. Google has spent over twenty-five years and tens of billions of dollars building its index. Microsoft has invested comparable amounts in Bing. No independent organisation has come close in breadth, freshness or relevance. The barriers are not just financial — they are temporal. You cannot buy a twenty-five-year head start in crawling, ranking and relevance tuning. You have to earn it, and earning it takes time that new entrants do not have.

This constraint creates a hierarchy with three tiers. At the top sit Google and Bing, the only two companies that own web-scale indexes built from their own crawlers. They are the infrastructure layer. Everything else sits on top of them. The second tier consists of services that maintain their own index but fall back to Google or Bing for queries their index cannot handle. Brave Search is the only general-purpose engine in this tier — after acquiring the failed search engine Cliqz, Brave built its own primary index and relies on Bing fallback for roughly seven percent of queries. The third tier is everything else: services that have no index of their own and function entirely as API clients. DuckDuckGo draws primarily from Bing, supplemented by its own crawler for certain queries. Startpage submits queries to Google on behalf of its users, returns Google’s results without tracking, and does not pretend to have an independent index. Qwant and Ecosia rely on Bing and Google while working on a shared European index that is years from being operational. Kagi operates a small index called Teclis but still depends on Google and Bing for the majority of results.

This is the same structural problem that appears everywhere in modern internet infrastructure. The browser engine market collapsed to Chromium because maintaining a modern engine requires thousands of engineers. The cloud market collapsed to three providers because building global data centres requires capital expenditure almost no one else can afford. The search index market collapsed to Google and Bing because continuously crawling the entire web requires resources that only they possess. In each case, the surface appears competitive — many browsers, many search engines, many cloud providers — but underneath, the infrastructure is a duopoly or monopoly. The competition is in the packaging, not the foundation.

The implications for how you choose a search engine are straightforward once the hierarchy is clear. Your choice operates on two axes: who owns the infrastructure, and who pays the bills. On the infrastructure axis, there are exactly three independent crawlers — Google, Bing and Brave — and Brave covers only a fraction of the web that Google does. On the payment axis, you either pay with money (Kagi, SearXNG self-hosted) or with data (everything free). The privacy-focused services improve the payment axis — you pay less in data — but they do not change the infrastructure axis. DuckDuckGo makes your queries private, but it still sends them to Bing. The privacy gain is real. The infrastructure dependency is unchanged.

The one solution that sidesteps both axes is SearXNG. It is a metasearch engine you host yourself, aggregating results from whichever backends you choose — Google, Bing, DuckDuckGo, Brave, Wikipedia — while stripping tracking from every request. It has no index and makes no claim to independence. Its value is honesty: it tells you exactly where each result comes from and lets you decide which backends to trust. For the technically inclined, this is the cleanest answer. For everyone else, the question is whether privacy is worth paying for — and if so, whether you trust a paid service like Kagi or a funded one like DuckDuckGo more than you trust Google or Bing directly.

The deeper truth is that there is no complete escape from this dependency through consumer choice alone. Building a competitive web index requires resources that no privacy-focused organisation possesses, and the gap widens every year as AI training demands more data and more compute. The German court ruling that Google bears full legal responsibility for AI-generated search content is a meaningful regulatory step, but it addresses a symptom. The structural problem is that one company controls the index most of the world uses to find anything, and that company has decided its interests are better served by keeping users on its own pages than by sending them elsewhere. Replacing the faucet does not change who controls the water. The plumbing is still owned by the same people who built the old pipes.

← Prev Post Next Post →