CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-index /ccrawl-recrawl-domains Common Crawl Domain Recrawl Live fetches of the home page of every ranked domain in Common Crawl's web graph, rendered to Markdown as they are fetched What is it? Common Crawl's web graph ranks domains by how central they are, but it does not tell you what those domains actually serve today. This dataset walks that ranking from the top and fetches each domain's home page now, storing the response as one Parquet row with the body, the headers, the timing and the… See the full description on the dataset page: https://huggingface.co/datasets/open-index/ccrawl-recrawl-domains.tabulartext-generation1M<n<10M0 likes1.2k downloads1mo agoHugging Face02open-index /ccrawl-domains Common Crawl Domain Ranks Web domains ranked by harmonic centrality and PageRank, ready to prioritize a crawl What is it? This dataset is the domain-level ranking from Common Crawl's hyperlink web graph, republished as clean Parquet. Common Crawl builds a graph of which domains link to which, then scores every domain by harmonic centrality and PageRank. A high rank means many other well-connected domains link to it, which is a solid proxy for importance when you… See the full description on the dataset page: https://huggingface.co/datasets/open-index/ccrawl-domains.tabulargraph-ml100M<n<1B0 likes722 downloads2mo agoHugging Face03Blackroot /Tiny-Open-Domain-BooksA tiny example dataset consisting of four books dedicated to the open domain in JSONL format: Alice in Wonderland - Lewis Caroll Dracula - Bram Stoker The Wonderful Wizard of Oz - L. Frank Baum The Count of Monte Cristo - Alexandre Dumas & Auguste Maquet All works are open domain, thus this dataset is also dedicated to the open domain. The dataset has been made to have extremely long context lengths, ideally as close to 2048 at possible without cutting off chunks in strange places. Each… See the full description on the dataset page: https://huggingface.co/datasets/Blackroot/Tiny-Open-Domain-Books.textn<1K5 likes120 downloads3y agoHugging Face04nthakur /mkqa-open-domaintext10K<n<100K1 likes103 downloads2y agoHugging Face05CyberMax-tools /open-domain-ranks Linkheft Open Domain Ranks: free domain authority data for 10.3 million domains Try the paid tool: Linkheft on Apify: score any domain list via API, with 4-month trends. First try costs cents; pay only for results. An open alternative to proprietary "domain authority" scores. For each of the top 10 million domains on the web (plus every Majestic Million domain) this dataset gives: column meaning domain registrable domain, lowercase (stripe.com) cc_rank… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/open-domain-ranks.tabulartabular-regression10M<n<100M0 likes74 downloads17h agoHugging Face06Lines /Open-Domain-Oral-Disease-QA-Dataset Open-Domain-Oral-Disease-QA-Dataset Dataset Details Dataset Description This dataset is meticulously designed to evaluate the diagnostic capabilities of Large Language Models (LLMs) in the domain of oral disease. We currently offer a suite of evaluation datasets encompassing models such as GPT-3.5, GPT-4, Palm2, and Llama2-70B. More data is under reviewed. This dataset is meticulously designed to evaluate the diagnostic capabilities of Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/Lines/Open-Domain-Oral-Disease-QA-Dataset.textn<1K5 likes47 downloads2y agoHugging Face07vietnqw /en-mT5-ner-open_domaintext100K<n<1M0 likes46 downloads2y agoHugging Face08Finnish-NLP /OrcaAgentInstruct-opendomainqatabular100K<n<1M0 likes43 downloads10mo agoHugging Face09Kaballas /open_domain_qatext100K<n<1M0 likes39 downloads1y agoHugging Face10vietnqw /raw-llama3-generated-open_domain_NERtext10K<n<100K0 likes23 downloads2y agoHugging Face11ivan604 /open-domain-authority-index open-domain-authority-index Domain-level authority metrics over the global Common Crawl link graph, produced by the openhrefs pipeline. Columns: domain, open_authority, open_volume, window_id. Best-effort snapshot, not a maintained service. This is an on-demand byproduct of the openhrefs pipeline, provided as-is. Run the pipeline yourself for current or custom data. Source terms Derived from Common Crawl and composite-domain-rating. Source terms apply; you are… See the full description on the dataset page: https://huggingface.co/datasets/ivan604/open-domain-authority-index.tabular100M<n<1B0 likes19 downloads4mo agoHugging Face12philipfourie /Tiny_Open-Domain-Books-Morsetext10K<n<100K0 likes16 downloads1y agoHugging Face13Bedru /open-domain-qa-prompt-settextn<1K0 likes16 downloads1mo agoHugging Face14tomrb /open_domain_biorxivtext10K<n<100K0 likes14 downloads1y agoHugging Face15vietnqw /vi-mt5-ner-open_domain-rawtext100K<n<1M0 likes12 downloads2y agoHugging Face16vietnqw /en-mT5-ner-open_domain-remaketext100K<n<1M0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.