
Google's AI intellectual property policy: text-and-data mining exceptions, opt-out rights, and what builders need to know
Published by AINave Editorial • Reviewed by Ramit
Google's President of Global Affairs Kent Walker delivered a keynote at the Global Forum on Intellectual Property during Singapore IP Week 2026, laying out the company's position on AI intellectual property policy. The core argument: AI training on publicly accessible data should be allowed under text-and-data mining exceptions, paired with machine-readable opt-out rights for creators. For AI builders, this framework directly affects how training data can be sourced and what safeguards are expected.
The policy framework: text-and-data mining exceptions and opt-out rights
Walker argued that text-and-data mining exceptions for training on publicly accessible content are essential, similar to approaches in Singapore, Japan, and the EU. He compared AI training to a student reading books in a library and then writing original novels: the learning process itself is not infringement. Without such exceptions, requiring a commercial license for every piece of publicly available data would, in Google's view, end AI innovation.
To balance creator rights, Google has implemented opt-out controls like Google-Extended and robots.txt, allowing rights holders to signal that their content should not be used for training or grounding. The company also updated Search Console protocols to let website owners manage how their content appears in generative AI search features. Walker described this combination of a broad right to train and a machine-readable opt-out as a "reasonable middle way."
Why this matters for AI builders
The practical impact is direct. If text-and-data mining exceptions become widely adopted, builders can train on publicly accessible web data without negotiating individual licenses, reducing friction for model development. However, the opt-out mechanisms mean that respecting robots.txt and Google-Extended signals becomes a compliance requirement, not just a courtesy. Builders who ignore these signals risk legal exposure or losing access to data sources.
Google also highlighted safeguards for AI outputs. SynthID embeds imperceptible watermarks into AI-generated images, audio, text, or video, helping users identify synthetic content. Likeness detection tools on YouTube scan for videos that may contain a creator's face. For builders shipping products that generate media, integrating similar watermarking or detection may become an expectation for trust and compliance.
Practical implications for training data and content verification
Builders should audit their data sourcing pipelines to ensure they respect opt-out signals. If you scrape the web for training data, you need to check robots.txt and Google-Extended headers. The policy also opens the door for commercial partnerships: Google noted that a balanced framework with clear exceptions does not preclude negotiations with rights holders, and the company is exploring new value-exchange models.
On the regulatory side, Google supports the NO FAKES Act of 2025 and the TAKE IT DOWN Act in the U.S., which target unauthorized digital replicas. These signal a direction for deepfake regulation that could affect how builders handle likeness generation. Walker argued that copyright law is not the right tool for deepfakes, and that misappropriation laws are a better fit.
Caveats and what remains unclear
This is Google's position, not settled law. The keynote is a policy advocacy document, and the legal landscape varies by jurisdiction. The claim that AI models are 300 times more efficient than two years ago is a vendor claim without independent verification. The Oxford Economics projection of up to $6 trillion in GDP impact over 10 years was commissioned by Google. There is no independent benchmarking of SynthID's effectiveness in preventing deception. Builders should monitor actual legislation and court rulings rather than relying solely on this framework.






















