robots.txt
robots-txt-blocked-domains-englishfineweb-robots-txt-files-compressedrobots-txt-blocked-domains-multilingualaurora-robots-txtFetched robots.txt for creation of the Aurora dataset.
Columns are
url: The URL of the robots.txt file fetched
content: The content of the response. Pages that responded with 404 Not Found or 410 Gone were assumed to allow all. Pages that responded with 403 Forbidden were assumed to disallow all. Pages that were tried but gave other errors contain a null value.
fetched: The timestamp of the visit.
global-robots-txt-ai-blocking-stats
Global robots.txt AI Bot Blocking Statistics (2026)
Aggregated sector-by-sector statistical index of AI Crawler blocking policies (opt-out rate) based on analysis of robots.txt directives across 50,000 global target domains.
Published by Pixel Office EU.
Included Sectors
News & Publishing
SaaS & Tech
Finance & Banking
E-Commerce
Healthcare & Medical
Travel & Hospitality
Blogs & Personal Sites
🚨 Is Your Domain Blocked or Misconfigured for AI… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/global-robots-txt-ai-blocking-stats.Robots-TXT-Des-Domaines-FRCe DataSet est concu pour que nous puissions crawler des sites, mais tout en respectant les regles en vigueur,
Ainsi grace a un mapping des robots.txt nous pourrons savoir quel sites autorise ou refuse l'utlisation des données du site pour l'entrainement des modeles, et ce, pour la totalité des domaine en .fr
Et d'ici pas long je l'espere nous aurons un dataset opensource a > 1T token , PUREMENT FRANCAIS !
serieux c'est normal que les gros dataset FR ont plus de 50% d'anglais ???
On… See the full description on the dataset page: https://huggingface.co/datasets/Data-Gouv-ML/Robots-TXT-Des-Domaines-FR.
