CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01permutans /wdc-common-crawl-embedded-jsonldtext10B<n<100B4 likes5.7k downloads2y agoHugging Face02logiover /json-ld-schema-meta-tag-extractor-sample-data JSON-LD Schema & Meta Tag Extractor Extract JSON-LD/Schema.org structured data, Meta tags, OpenGraph and Twitter Cards from any URL. Get page title + meta description with a clean JSON output for SEO audits, validation, competitor research and AI datasets. Proxy-ready for large crawls. What the actor scrapes 🧩 JSON-LD Schema & Meta Tag Extractor — Scrape Schema.org, OpenGraph & Meta Tags Extract structured data and SEO metadata from any webpage in seconds. This… See the full description on the dataset page: https://huggingface.co/datasets/logiover/json-ld-schema-meta-tag-extractor-sample-data.textn<1K0 likes34 downloads4mo agoHugging Face03UWV /wim-instruct-signaalberichten-to-jsonld-agent-steps Dataset Card for UWV/wim_instruct_signaalberichten_to_jsonld_agent_steps Dataset Summary This dataset contains 116,056 instruction-following examples from a production pipeline that converts Dutch municipality complaint messages (signaalberichten) into structured JSON-LD knowledge graphs using Schema.org vocabulary. Each example represents an atomic LLM call from a 4-stage agent pipeline designed to extract semantic information from citizen-government communications. The… See the full description on the dataset page: https://huggingface.co/datasets/UWV/wim-instruct-signaalberichten-to-jsonld-agent-steps.text100K<n<1M0 likes33 downloads1y agoHugging Face04UWV /wim-instruct-wiki-to-jsonld-agent-steps Dataset Card for UWV/wim_instruct_wiki_to_jsonld_agent_steps Dataset Summary This dataset contains instruction-following examples for training language models on the task of converting unstructured Wikipedia text to structured JSON-LD format using Schema.org vocabulary. Each example represents a single step in a multi-agent pipeline where different specialized models handle entity extraction, schema retrieval, knowledge graph transformation, and validation. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/UWV/wim-instruct-wiki-to-jsonld-agent-steps.tabular100K<n<1M1 likes27 downloads1y agoHugging Face05moo3030 /jsonl-datasettextn<1K0 likes9 downloads11mo agoHugging Face06DannyAI /jsonl-datasettextn<1K0 likes8 downloads11mo agoHugging Face07supergoose /buzz_sources_356_jsonldtextn<1K0 likes5 downloads2y agoHugging Face08ozgurkrkrt /jsonl-datatext10K<n<100K0 likes4 downloads3y agoHugging Face09MohamedQiqa /jsonl-dataset Dataset Name jsonl-dataset Fields instruction: The task or question input: Optional context or input output: The expected response Data Splits Training: X examples Validation: Y examples Usage from datasets import load_dataset dataset = load_dataset("MohamedQiqa/jsonl-dataset") texttext-generationn<1K0 likes4 downloads3mo agoHugging Face10Nimsprod /jsonl-datasettextn<1K0 likes3 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.