CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lance-format /fineweb-edu FineWeb-Edu (Lance Format) A Lance-formatted version of FineWeb-Edu — over 1.5 billion educational web passages with cleaned text, source metadata, language detection signals, and 384-dim text embeddings — available directly from the Hub at hf://datasets/lance-format/fineweb-edu/data/train.lance. Key features Cleaned passage text in the text column with the source url and title carried alongside. Language detection signals (language, language_probability) for filtered… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/fineweb-edu.tabulartext-retrieval1B<n<10B8 likes8.4k downloads4mo agoHugging Face02ackermans26 /LLaVA-OneVision-1.5-Instruct-Data-qwen-formattext1M<n<10M0 likes6.8k downloads2mo agoHugging Face03chupei /format-texttextn<1K0 likes6.7k downloads2y agoHugging Face04chupei /format-jsontextn<1K0 likes6.4k downloads2y agoHugging Face05chupei /format-jsonltextn<1K0 likes6.4k downloads2y agoHugging Face06jeggers /logiqa2_formatted Dataset Card for "logiqa2_formatted" More Information needed tabular10K<n<100K2 likes6.3k downloads2y agoHugging Face07chupei /format-listjsontextn<1K0 likes6.1k downloads2y agoHugging Face08EleutherAI /filtering-pretraining-mix-arrow-formattabular100M<n<1B0 likes3k downloads2y agoHugging Face09lance-format /Openvid-1M OpenVid Dataset (Lance Format) Lance format version of the OpenVid dataset with 937,957 high-quality videos stored with inline video blobs, embeddings, and rich metadata. Why Lance? Lance is an open-source format designed for multimodal AI data, offering significant advantages over traditional formats for modern AI workloads. Blazing Fast Random Access: Optimized for fetching scattered rows, making it ideal for random sampling, real-time ML serving, and interactive… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/Openvid-1M.tabulartext-to-video100K<n<1M8 likes2.8k downloads8mo agoHugging Face10datasets-examples /doc-formats-jsonl-1 [doc] formats - jsonl - 1 This dataset contains one jsonl file at the root. textn<1K0 likes2.7k downloads3y agoHugging Face11iamroot /chat_formatted_examplestextn<1K0 likes2.3k downloads2y agoHugging Face12datasets-examples /doc-formats-parquet-1textn<1K0 likes2.2k downloads2y agoHugging Face13datasets-examples /doc-formats-csv-1 [doc] formats - csv - 1 This dataset contains one csv file at the root: data.csv kind,sound dog,woof cat,meow pokemon,pika human,hello The YAML section of the README does not contain anything related to loading the data (only the size category metadata): --- size_categories: - n<1K --- textn<1K0 likes2.2k downloads3y agoHugging Face14heegyu /glaive-function-calling-v2-formatted original dataset: glaiveai/glaive-function-calling-v2 {'system_message': 'You are a helpful assistant with access to the following functions. Use them if required -', 'function_description': '{\n "name": "get_random_quote",\n "description": "Get a random quote",\n "parameters": {}\n}', 'conversations': [{'content': 'Hi, can you help me with something?', 'role': 'user'}, {'content': "Of course! I'm here to assist you. What do you need help with?", 'role': 'assistant'}… See the full description on the dataset page: https://huggingface.co/datasets/heegyu/glaive-function-calling-v2-formatted.text100K<n<1M12 likes1.9k downloads3y agoHugging Face15Merserk /Krea-2-Turbo-Checkpoint-Format-Benchmark Krea 2 Turbo ComfyUI Format Fidelity Benchmark This release is a paired, deterministic comparison of eight Krea 2 Turbo checkpoint formats in ComfyUI: BF16, FP8 Scaled, INT8 ConvRot, MXFP8, NVFP4, INT4 ConvRot W4A4, GGUF Q8_0, and GGUF Q4_K_M. It contains 240 scored 1024×1024 images, saved float32 decoded tensors and final latents, every denoising trajectory, raw metric tables, telemetry, statistical comparisons, and reproduction code. Main result BF16 is the… See the full description on the dataset page: https://huggingface.co/datasets/Merserk/Krea-2-Turbo-Checkpoint-Format-Benchmark.imagetext-to-imagen<1K5 likes1.8k downloads2mo agoHugging Face16ysn-rfd /text-dataset-tiny-code-script-py-format USED of tahamajs/medicine_ds_persian for .parquet file USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file USED of Abirate/english_quotes for .jsonl file NEW FILES (05/12/2025) NEW FILES (12/26/2025) NEW FILES (02/15/2026) texttext-generation10K<n<100K3 likes1.7k downloads4mo agoHugging Face17wissamantoun /fineweb-edu-format-topic FineWeb-Edu w/ Topic and Format Annotations FineWeb-Edu dataset consists of 1.3T tokens annotated for Topic and Format using wissamantoun/WebOrganizer-TopicClassifier-ModernBERT and wissamantoun/WebOrganizer-FormatClassifier-ModernBERT classifiers. Similar to WebOrganizer/Corpus-200B but using FineEdu instead of DCLM. Topic Labels: Adult Art & Design Software Dev. Crime & Law Education & Jobs Hardware Entertainment Social Life Fashion & Beauty Finance & Business Food & Dining… See the full description on the dataset page: https://huggingface.co/datasets/wissamantoun/fineweb-edu-format-topic.texttext-generation1B<n<10B5 likes1.6k downloads1y agoHugging Face18lance-format /openvid-lance OpenVid (Lance Format) A Lance-formatted version of the OpenVid-1M corpus — 937,957 high-quality clips with inline MP4 bytes, 1024-dim video embeddings, captions, and rich per-clip quality signals — available directly from the Hub at hf://datasets/lance-format/openvid-lance/data/train.lance. Key features Inline MP4 bytes in the video_blob column, stored in a side blob file and surfaced as lazy BlobFile handles via take_blobs — metadata scans, search, and filtering… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/openvid-lance.tabulartext-to-video100K<n<1M2 likes1.5k downloads4mo agoHugging Face19Lyric1010 /repo_format3 Dataset: repo_format3 This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train/repo_format3/no-curriculum/tmp. text0 likes1.4k downloads11mo agoHugging Face20KaiChen1998 /coda-lm-llava-format CODA-LM Dataset Card CODA-LM is the multi-modal version of the CODA dataset, used in the CODA-LM paper. Both English and Chinese annotations are available. Check detailed usage in our Github repo. This repo contains the CODA-LM dataset, which has been reorganized in the LLaVA data format. You are also welcome to check the original CODA-LM data which contains more metadata vanilla annotations. Usage from datasets import load_dataset # name can be selected from… See the full description on the dataset page: https://huggingface.co/datasets/KaiChen1998/coda-lm-llava-format.imageimage-to-text10K<n<100K3 likes1.3k downloads2y agoHugging Face21alexwww94 /Rexverse-2M-formattedimage1M<n<10M0 likes1.3k downloads8mo agoHugging Face22togethercomputer /glaive-function-calling-v2-formatted Dataset Card for "glaive-function-calling-v2-formatted" More Information needed text100K<n<1M37 likes1.2k downloads3y agoHugging Face23ksterx /hle-no-img-prompt-completion-formatimage1K<n<10K0 likes1.2k downloads1y agoHugging Face24justus27 /math-hendrycks-genesys-formattext1K<n<10K0 likes1.1k downloads1y agoHugging Face25vanloc1808 /pico-banana-smolvlm-format-with-rejected-answer pico-banana-smolvlm-format-with-rejected-answer Balanced image-level tampering detection dataset in SmolVLM-style format with chosen/rejected answer pairs, derived from the pico-banana MCQ pipeline. Suitable for preference learning (e.g. DPO) and RLHF-style training. Dataset overview Same as vanloc1808/pico-banana-smolvlm-format, but each example includes a rejected_answer field: the answer from the counterpart sample (same edited/original image pair, opposite… See the full description on the dataset page: https://huggingface.co/datasets/vanloc1808/pico-banana-smolvlm-format-with-rejected-answer.image100K<n<1M1 likes1.1k downloads7mo agoHugging Face26jeggers /gpqa_formattedgated Dataset Card for GPQA Formatted version of original GPQA dataset. This removes most columns and adds single columns options and answer to contain a list of the possible answers and the index of the correct one. GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy… See the full description on the dataset page: https://huggingface.co/datasets/jeggers/gpqa_formatted.textn<1K4 likes1k downloads2y agoHugging Face27vidore /vidore_v3_finance_en_mteb_format Vidore3FinanceEnRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_finance_en How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3FinanceEnRetrieval") evaluator… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_finance_en_mteb_format.imagevisual-document-retrieval10K<n<100K1 likes995 downloads11mo agoHugging Face28vidore /vidore_v3_computer_science_mteb_format Vidore3ComputerScienceRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_computer_science How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes956 downloads11mo agoHugging Face29vidore /vidore_v3_industrial_mteb_format Vidore3IndustrialRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_industrial How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3IndustrialRetrieval")… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_industrial_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes938 downloads11mo agoHugging Face30lance-format /droidimage10M<n<100M0 likes889 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.