CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01permutans /wdc-common-crawl-embedded-jsonldtext10B<n<100B4 likes5.4k downloads2y agoHugging Face02UWV /wim-instruct-signaalberichten-to-jsonld-agent-steps Dataset Card for UWV/wim_instruct_signaalberichten_to_jsonld_agent_steps Dataset Summary This dataset contains 116,056 instruction-following examples from a production pipeline that converts Dutch municipality complaint messages (signaalberichten) into structured JSON-LD knowledge graphs using Schema.org vocabulary. Each example represents an atomic LLM call from a 4-stage agent pipeline designed to extract semantic information from citizen-government communications. The… See the full description on the dataset page: https://huggingface.co/datasets/UWV/wim-instruct-signaalberichten-to-jsonld-agent-steps.text100K<n<1M0 likes39 downloads1y agoHugging Face03logiover /json-ld-schema-meta-tag-extractor-sample-data JSON-LD Schema & Meta Tag Extractor Extract JSON-LD/Schema.org structured data, Meta tags, OpenGraph and Twitter Cards from any URL. Get page title + meta description with a clean JSON output for SEO audits, validation, competitor research and AI datasets. Proxy-ready for large crawls. What the actor scrapes 🧩 JSON-LD Schema & Meta Tag Extractor — Scrape Schema.org, OpenGraph & Meta Tags Extract structured data and SEO metadata from any webpage in seconds. This… See the full description on the dataset page: https://huggingface.co/datasets/logiover/json-ld-schema-meta-tag-extractor-sample-data.textn<1K0 likes33 downloads4mo agoHugging Face04UWV /wim-instruct-wiki-to-jsonld-agent-steps Dataset Card for UWV/wim_instruct_wiki_to_jsonld_agent_steps Dataset Summary This dataset contains instruction-following examples for training language models on the task of converting unstructured Wikipedia text to structured JSON-LD format using Schema.org vocabulary. Each example represents a single step in a multi-agent pipeline where different specialized models handle entity extraction, schema retrieval, knowledge graph transformation, and validation. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/UWV/wim-instruct-wiki-to-jsonld-agent-steps.tabular100K<n<1M1 likes23 downloads1y agoHugging Face05pyrihtm /text_to_schema.org_json-ldtextn<1K1 likes12 downloads2y agoHugging Face06moo3030 /jsonl-datasettextn<1K0 likes9 downloads11mo agoHugging Face07DannyAI /jsonl-datasettextn<1K0 likes8 downloads11mo agoHugging Face08linges0103 /jsonldatasettextn<1K0 likes6 downloads3y agoHugging Face09supergoose /buzz_sources_356_jsonldtextn<1K0 likes5 downloads2y agoHugging Face10ozgurkrkrt /jsonl-datatext10K<n<100K0 likes4 downloads3y agoHugging Face11Nimsprod /jsonl-datasettextn<1K0 likes4 downloads7mo agoHugging Face12MohamedQiqa /jsonl-dataset Dataset Name jsonl-dataset Fields instruction: The task or question input: Optional context or input output: The expected response Data Splits Training: X examples Validation: Y examples Usage from datasets import load_dataset dataset = load_dataset("MohamedQiqa/jsonl-dataset") texttext-generationn<1K0 likes4 downloads2mo agoHugging Face13butlerj /toy_data_jsonld_testtextn<1K0 likes2 downloads1y agoHugging Face14PhongGoldFish /jsonl_dhspHN2textn<1K0 likes1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.