CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01albertvillanova /tests-raw-jsonltext10K<n<100K1 likes40k downloads5y agoHugging Face02hf-internal-testing /raw_jsonltext10K<n<100K0 likes24k downloads5y agoHugging Face03kaczmarj /wsinfer-model-zoo-jsonThis is the registry of models in the WSInfer Model Zoo. See https://wsinfer.readthedocs.io/en/latest/ and https://github.com/SBU-BMI/wsinfer-zoo for more information. textn<1K1 likes10k downloads3y agoHugging Face04introspector /ocaml-opam-ppxlib-json-astversion https://git-lfs.github.com/spec/v1 oid sha256:b797e216eb8d6720d794ca5d07a01019cac2d79df456fbcc69b6407497cc267c size 473 0 likes8.7k downloads2y agoHugging Face05chupei /format-jsonltextn<1K0 likes8.5k downloads2y agoHugging Face06chupei /format-jsontextn<1K0 likes8.5k downloads2y agoHugging Face07hf-internal-testing /ner-jsonltext10K<n<100K0 likes8.2k downloads1y agoHugging Face08permutans /wdc-common-crawl-embedded-jsonldtext10B<n<100B4 likes5.5k downloads2y agoHugging Face09epfl-dlab /JSONSchemaBench JSONSchemaBench JSONSchemaBench is a benchmark of real-world JSON schemas designed to evaluate structured output generation for Large Language Models (LLMs). It contains approximately 10,000 JSON schemas, capturing diverse constraints and complexities. import datasets from datasets import load_dataset def main(): # Inspect the available subsets of the datasetall_subsets = datasets.get_dataset_config_names("epfl-dlab/JSONSchemaBench") print("Available subsets:"… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/JSONSchemaBench.texttext-generation10K<n<100K12 likes4.4k downloads1y agoHugging Face10endomorphosis /Caselaw_Access_Project_JSON The Caselaw Access Project In collaboration with Ravel Law, Harvard Law Library digitized over 40 million U.S. court decisions consisting of 6.7 million cases from the last 360 years into a dataset that is widely accessible to use. Access a bulk download of the data through the Caselaw Access Project API (CAPAPI): https://case.law/caselaw/ Find more information about accessing state and federal written court decisions of common law through the bulk data service documentation here:… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/Caselaw_Access_Project_JSON.text-generation1M<n<10M3 likes4.2k downloads2y agoHugging Face11datasets-examples /doc-formats-jsonl-1 [doc] formats - jsonl - 1 This dataset contains one jsonl file at the root. textn<1K0 likes2.4k downloads3y agoHugging Face12daydreamer-json /temporalHlsRawDataStorage0 likes2k downloads1y agoHugging Face13jsonhash /LLVIP LLVIP 数据集 [中文] [English] 这里存储了[LLVIP 数据集]的备份 和 其 [COCO标注格式的标注] 下载 数据集:https://huggingface.co/datasets/UserNae3/LLVIP/blob/main/LLVIP.zip COCO格式标注:https://huggingface.co/datasets/UserNae3/LLVIP/blob/main/coco_annotations.7z 版权 版权链接: https://github.com/bupt-ai-cz/LLVIP?tab=readme-ov-file#license 1 likes1.7k downloads2y agoHugging Face14w601sxs /gsm8k-json Dataset Card for "gsm8k-json" More Information needed text1K<n<10K0 likes1.6k downloads3y agoHugging Face15Arun63 /sharegpt-quizz-generation-json-output ShareGPT-Formatted Dataset for Quizz Generation in Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate quizz in structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-quizz-generation-json-output.texttext-generationn<1K1 likes1.5k downloads2y agoHugging Face16flax-sentence-embeddings /stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering100K<n<1M12 likes1.5k downloads4y agoHugging Face17scilons /SciLaD-all-json-v1 SciLaD (JSON) SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing over 10 million scientific publications and a multilingual, unfiltered TEI XML split including more than 35 million publications. We also publish the extensible pipeline for generating SciLaD. Dataset Details In this repository we share the full… See the full description on the dataset page: https://huggingface.co/datasets/scilons/SciLaD-all-json-v1.text10M<n<100M0 likes1.5k downloads1mo agoHugging Face18cpral /forums_pol_json_zst0 likes1.4k downloads6mo agoHugging Face19Arun63 /sharegpt-structured-output-json ShareGPT-Formatted Dataset for Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.texttext-generationn<1K7 likes1.4k downloads2y agoHugging Face20flax-sentence-embeddings /stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M8 likes1.2k downloads4y agoHugging Face21dixantp /desktop-accessibility-screenshot-json-dumpsimagen<1K1 likes1k downloads1y agoHugging Face22minpeter /hermes-function-calling-v1-jsonl Hermes Function-Calling V1 This dataset is the compilation of structured output and function calling data used in the Hermes 2 Pro series of models. This repository contains a structured output dataset with function-calling conversations, json-mode, agentic json-mode and structured extraction samples, designed to train LLM models in performing function calls and returning structured output based on natural language instructions. The dataset features various conversational scenarios… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/hermes-function-calling-v1-jsonl.texttext-generation10K<n<100K1 likes993 downloads2y agoHugging Face23NousResearch /json-mode-evaltextn<1K44 likes907 downloads3y agoHugging Face24Obscure-Entropy /conceptual_captions_jsonimage1M<n<10M0 likes806 downloads2y agoHugging Face25flax-sentence-embeddings /stackexchange_title_body_jsonljsonl.gz format from https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml Each line contains a dict in the format: {"text": ["title", "body"], "tags": ["tag1", "tag2"]} The following parameters have been used for filtering: min_title_len = 20 min_body_len = 20 max_body_len = 4096 min_score = 0 If a stackexchange contained less than 10k questions (after filtering), it is written to the small_stackexchanges.jsonl.gz file. This is a dump of the files from… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl.text1M<n<10M2 likes804 downloads5y agoHugging Face26gary2oos /cs-net-json3d10K<n<100K0 likes740 downloads2mo agoHugging Face27flax-sentence-embeddings /stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M9 likes737 downloads4y agoHugging Face28tianzl66 /Sheetpedia_json_1005text100K<n<1M0 likes690 downloads1y agoHugging Face29cpral /step35-en2pl-conv-pass4-jsonlconversations: 1,251,034 chat-template tokens (role+content, incl. special tokens): 2,664,206,408 reasoning_content tokens (not covered by chat template, counted separately): 6,662,763,429 avg tokens/conversation: 2129.6 used tokenizer: APT4 100K<n<1M0 likes687 downloads2mo agoHugging Face30jsonhash /shitjournal-backupimage1K<n<10K0 likes604 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.