Google must share anonymized search data with competitors


The European Commission has adopted two binding decisions that force Google to share anonymized search data with competing search engines and open parts of Android to competing AI assistants.

Search Data Metrics gives eligible vendors, including AI chatbots with search capabilities, access to anonymized query, click, view, and results position data that they can use to build their own retrieval and ranking systems.

The Commission outlined the two decisions under the Digital Markets Act, six months after the opening of the procedure which gave them.

We addressed the research data proposal in April, when it was preliminary conclusions submitted for public consultation. The version adopted this month is final.

What the decision requires

Google is required to share anonymized data on rankings, queries, clicks and views free and paid search results under fair and non-discriminatory conditions. This includes information such as search queries, metadata such as language and device type, URLs viewed, user interactions and result positions.

However, it does not include Google’s ranking algorithms. Some sensitive data, such as account details, search histories, timestamps, and rare or long queries, are deleted to protect individuals.

The Commission said Google’s current data sharing approach has failed. The new ruling details what effective sharing entails, specifying who is eligible and how data is priced, with costs based on recovery rather than free market rates.

AI chatbots that qualify as online search engines under the DMA are eligible to use data to improve their systems, but not to train general AI models or replicate Google results. These requirements are binding under the DMA but do not involve fines, unlike a separate DMA case related to self-preference and antitrust cases pending before European courts.

Why data matters for AI research

This decision extends beyond search engines to AI responses, as it involves grounding. AI chatbots use recent web data to ensure the accuracy of their responses, and the quality of this data depends on the research information behind it. A 2025 explainer on AI mode explained that Google bases its models with a system called FastSearch, which relies on its own search ranking signals.

The decision does not give this to competitors. This does not require Google to share FastSearch or its search algorithms and technologies. Instead, Google must share anonymized data about queries, clicks, views, and result rankings that eligible parties can use to develop their own retrieval and ranking systems, with grounding being one of the approved uses.

In February we suggested that this European process could have greater long-term implications for AI research than the US antitrust case, because the question of whether Google search data feeds competing AI tools affects the entire system of AI answers, citations and references. The decision marks the point where this issue begins to be resolved operationally. A chatbot with access to lots of anonymized interaction data from Google Search starts with a different baseline than a chatbot without that data.

Who can actually use it

Which companies benefit first depends on who can already use the data effectively, not just their eligibility. All applicants must have at least 50,000 monthly users in the EU and pass either a two-year operating history or, for new entrants, an investment test. A security review and independent audit is required before Google shares data.

Established search engines like Bing and DuckDuckGo will likely reach these thresholds quickly and act on new data rights. New entrants must develop the ability to leverage data before it influences their offerings.

In the short term, the impact on traffic will be limited by a benchmark that we have followed throughout the year. AI chatbots still represent a small part of the references. According to SE Rankingall AI platforms combined accounted for approximately 0.24% of global internet traffic in January. Better access to this data could influence what competing engines and chatbots can develop, but it alone won’t determine where researchers go.

The Android AI track

The second decision concerns Android. Google needs to open up a set of operating system features to compete with AI assistants so that a person can activate a competing assistant by voice, similar to the “Hey Google” command, and let it act in applications, like booking a taxi or writing a response.

Google is expected to add most of these features in the next major Android release, Android 18, and no later than August 1, 2027. Concurrent voice activation, which lets more than one assistant respond to different wake words, has a later deadline of August 1, 2028. Google’s own Gemini assistant already has this level of access on Android, which is the asymmetry the decision is meant to close.

Google’s response

Google disagrees with both decisions. Kent Walker, president of global affairs at Google and Alphabet, wrote that they “risk undermining vital privacy and security safeguards” for millions of Europeans. He also mentioned that Google has repeatedly proposed solutions aimed at achieving the DMA’s goals. Regarding measures relating to research data, his concern is the disclosure of European research data to unknown companies without proper anonymization and users’ knowledge and consent.

The Commission explains that anonymization involves a multi-level technical process combined with contractual guarantees, developed with internal and external confidentiality experts. This process allows Google to review an applicant on cybersecurity or data protection criteria before sharing data, and the measures can be re-evaluated if independent testing reveals the safeguards are insufficient.

Why it matters

Once vendors successfully complete the access process, competing search engines and AI chatbots gain access to anonymized search data, similar to what Google has accumulated on a large scale. A wider range of providers using this data could allow more search engines and chatbots to cite sources and drive referral traffic, moving away from the current dominance of a few platforms. However, this does not guarantee this result. It depends on who qualifies through eligibility and audits, and how the data performs in actual product development.

Looking to the future

Initially, researchers and editors will not notice any changes. Google will spend the remainder of 2026 developing the dataset and establishing terms, with its pricing proposal expected no later than January 2027. Each eligible vendor will then access the data on its own schedule, after obtaining a license and agreeing to the price. Major Android changes are expected by August 1, 2027 and simultaneous voice activation by August 1, 2028.

The Commission plans to review these measures every two years and may reopen them if independent testing indicates that anonymization is inadequate. It remains unclear whether this will increase the number of engines and chatbots vying for visibility, and the outcome will only be clear when eligible providers start mining the data.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *