CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OmriShtayer /Website_Traffic_and_Engagementtabulartable-question-answeringn<1K0 likes45 downloads1y agoHugging Face02openbenchmarks /OB-Company-Websearch OB Company Websearch 10 multi-constraint company-discovery questions with a frozen, hand-labelled reference set. The public set accompanies the OpenBenchmarks Multi Turn Company Search Benchmark. The benchmark holds the research agent and budgets fixed while varying the web search provider. This release is for the search-only condition: the agent can search and use result snippets but is not given a page-fetch tool. Dataset contents Each question asks for the… See the full description on the dataset page: https://huggingface.co/datasets/openbenchmarks/OB-Company-Websearch.tabularquestion-answeringn<1K0 likes45 downloads1mo agoHugging Face03davidsmandrade /ileoro-pt-web Ileoro-pt-web Dataset de verificacao externa para recuperacao em Historia da Africa em portugues. O arquivo principal preserva o schema tabular do davidsmandrade/Ileoro-pt; provenance e auditoria sao persistidos como artefatos sidecar do pipeline. Fontes UNESCO/UNESDOC e paginas relacionadas aos livros da Historia Geral da Africa sao rejeitadas pelo pipeline. tabularquestion-answeringn<1K0 likes40 downloads5mo agoHugging Face040xrphl /USCIS-knowledge-base-full-website A comprehensive dataset of 99,489 content chunks from 4,666 pages on the USCIS website, with pre-computed OpenAI text-embedding-ada-002 embeddings (1536 dimensions). Built for RAG (Retrieval-Augmented Generation), semantic search, and GraphRAG applications focused on U.S. immigration law and policy. 🔗 GitHub: github.com/0xrphl/USCIS-knowledge-base-full-website🔥 Scraped with: Firecrawl — The open-source web scraping API for AI🍎 Visualized with: Embedding Atlas — Interactive embedding… See the full description on the dataset page: https://huggingface.co/datasets/0xrphl/USCIS-knowledge-base-full-website.tabulartext-retrieval100K<n<1M0 likes29 downloads3mo agoHugging Face05OptiTransferData /swiss-web-premium-chgated *.ch Swiss Web Premium (A+) Grade A+ -- 112.4M tokens -- 110,491 records -- 29 languages -- Full provenance -- PII-redacted -- RAG-ready -- SFT-formatted A production-grade Swiss web corpus from the .ch TLD namespace. 110,491 documents independently quality-scored (avg 93.3/100, minimum 90), PII-redacted, and SHA256-verified. Built for LLM training, RAG pipelines, SFT fine-tuning, and multilingual NLP. OptiTransferData Portfolio Premium sovereign web corpora for… See the full description on the dataset page: https://huggingface.co/datasets/OptiTransferData/swiss-web-premium-ch.documenttext-generation10K<n<100K0 likes10 downloads6mo agoHugging Face06weblab-GENIAC /OpenBookQA-Japanese-maskedgated OpenBookQA-Japanese-masked 与えられた問題に対して4つの選択肢から答えを選択するデータセット allenai/openbookqaをcyberagent/calm3-22b-chatで翻訳 5,957件 train split: 4,956件(4,957件の内1件削除) validation split: 500件 test split: 499件(500件の内1件削除) ルールベース・機械学習ベースのフィルタリング処理の後、目視確認を行い、個人情報が含まれ得るサンプルについては全件削除対応を実施 Format データセットの構成は以下 { "idx": ID, "id": 元ID, "question_stem_en": 英語の質問文, "choices_en": { "text": 選択肢の文章, "label": 選択肢の記号, }… See the full description on the dataset page: https://huggingface.co/datasets/weblab-GENIAC/OpenBookQA-Japanese-masked.tabulartext-generation1K<n<10K0 likes6 downloads2y agoHugging Face07BYU-Idaho /Web-Content BYU-Idaho Web Content Dataset (NLP-Enhanced) State-of-the-art university web content dataset with full NLP enrichment: entity extraction, acronym detection, domain terminology, and semantic features. Enterprise-ready for advanced RAG, semantic search, and AI applications. Dataset Description Records: 2,448 ultra-high-quality pages Source: byui.edu and subdomains Format: Markdown + NLP metadata (JSON fields) Quality: 40.2% filtered + 91.5/100 avg score + Full NLP… See the full description on the dataset page: https://huggingface.co/datasets/BYU-Idaho/Web-Content.tabularquestion-answering1K<n<10K0 likes6 downloads9mo agoHugging Face08OptiTransferData /swiss-web-premium-ch-fullgated *.ch Swiss Web Premium (A+) -- Full Dataset Grade A+ -- 112.4M tokens -- 110,491 records -- 29 languages -- 22 production files -- 554.1 MB The complete production release of OptiTransfer's Swiss web corpus. 110,491 documents from the .ch TLD namespace, independently quality-scored (avg 93.3/100, minimum 90), PII-redacted, and SHA256-verified. Delivered in Parquet, JSONL, language splits, and pre-built RAG chunks. This is the full commercial dataset. For evaluation, see the free… See the full description on the dataset page: https://huggingface.co/datasets/OptiTransferData/swiss-web-premium-ch-full.tabulartext-generation100K<n<1M0 likes5 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.