
Study reveals only 30% of South African Publishers Block AI Crawlers despite widespread fair use concerns
Press Releases
A new study, Protocol Gap: South Africa, reveals that only a third of publishers in South Africa are actively aiming to limit AI crawler access to their content using robots.txt; these are files that allow a site owner to tell search engine crawlers which URLs they can access on the site, and which ones they should avoid. The new data from South Africa reveals that these AI-blocking measures are concentrated among larger and better-resourced publishers while smaller outlets remain widely exposed to scraping – the same outlets whose sustainability is most at risk due to declining traffic and referrals.
Over the last year, South Africa’s leading news sites have lost an average of 20% of their daily page views, mirroring international dynamics. As extraction and use of publisher content by AI companies increases, often without compensation or attribution, it accelerates visibility issues that threaten the survival of newsrooms – in South Africa and beyond.
In late 2025, the South African Competition Commission concluded a landmark Media and Digital Platforms Market Inquiry (MDPMI), which found that publishers are increasingly being forced to make trade-offs between maintaining visibility and protecting their content. In this context, use of robots.txt is one indicator of how publishers are responding to growing pressure to decide whether and how their content can be accessed by automated crawlers. The Protocol Gap report’s findings contribute to ongoing debates in South Africa’s media sector on application of competition law, copyright reform, and the development of a national AI policy framework.
Robots.txt is a text file that can be added to the root domain of a website that contains instructions to search engines and web crawlers about what content they can access. This includes directives to AI bots that systematically browse and collect data to train LLMs and power AI assistants and search engines. Robots.txt is one of the few free tools available to content creators to signal how they want their data to be used by AI companies. While not legally binding, robots.txt has been used in legal complaints against unauthorised crawling and scraping in Canada, US, and UK. The file can serve as leverage in licensing negotiations, especially where AI models rely heavily on large volumes of unconsented, uncompensated data. However, effectiveness depends on voluntary compliance by AI developers, and its role within broader governance frameworks remains contested.
Based on analysis of 263 news websites, Protocol Gap: South Africa finds that while 74.1% of publishers have a robots.txt file in place, only 30.4% block at least one AI crawler as of April 2026. Although the percentage of publishers that use robots.txt to block AI crawlers rose modestly from 29.7% in December 2025 to 30.4%, the findings suggest that most publishers remain open to AI data collection.
Below these broad findings, the study identified a gap in adoption of AI-scraping policies between large and small publishers that reflects disparities in resources, technical capacity, referral traffic dependency, and strategic positioning. This raises important questions about bargaining power, content ownership and sustainability, as AI systems become increasingly reliant on news content. In the absence of clear legal and regulatory frameworks, the decisions publishers make now about how AI crawlers access their content could have long-lasting consequences.

The report, released by the Media Leadership Think Tank (MLTT) at the GIBS Business School (Gordon Institute of Business Science), the Journalism Relay Project, and the International Fund for Public Interest Media (IFPIM) is part of a broader research collaboration focused on Brazil, Indonesia, and South Africa. Through technical and qualitative research, it aims to increase awareness among media organizations on how AI systems access their content and help to inform strategies and policies to manage AI crawlers in ways that best serve newsroom interests. The report includes a detailed methodology for all those who may want to replicate this research on robots.txt in other countries and regions.
The study also provides a replicable methodology for researchers and publisher associations interested in conducting similar audits and concludes with next steps for research and engagement with ongoing efforts to strengthen publisher agency, bargaining power, and transparency in emerging AI licensing ecosystems.
Full suite of Protocol Gap reports: ctrl-j.info/publications



%20(1000%20x%20627%20px).jpg)
