CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ackermans26 /LLaVA-OneVision-1.5-Instruct-Data-qwen-formattext1M<n<10M0 likes6.8k downloads2mo agoHugging Face02jeggers /logiqa2_formatted Dataset Card for "logiqa2_formatted" More Information needed tabular10K<n<100K2 likes6.3k downloads2y agoHugging Face03datasets-examples /doc-formats-parquet-1textn<1K0 likes2.2k downloads2y agoHugging Face04heegyu /glaive-function-calling-v2-formatted original dataset: glaiveai/glaive-function-calling-v2 {'system_message': 'You are a helpful assistant with access to the following functions. Use them if required -', 'function_description': '{\n "name": "get_random_quote",\n "description": "Get a random quote",\n "parameters": {}\n}', 'conversations': [{'content': 'Hi, can you help me with something?', 'role': 'user'}, {'content': "Of course! I'm here to assist you. What do you need help with?", 'role': 'assistant'}… See the full description on the dataset page: https://huggingface.co/datasets/heegyu/glaive-function-calling-v2-formatted.text100K<n<1M12 likes1.9k downloads3y agoHugging Face05ysn-rfd /text-dataset-tiny-code-script-py-format USED of tahamajs/medicine_ds_persian for .parquet file USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file USED of Abirate/english_quotes for .jsonl file NEW FILES (05/12/2025) NEW FILES (12/26/2025) NEW FILES (02/15/2026) texttext-generation10K<n<100K3 likes1.7k downloads4mo agoHugging Face06wissamantoun /fineweb-edu-format-topic FineWeb-Edu w/ Topic and Format Annotations FineWeb-Edu dataset consists of 1.3T tokens annotated for Topic and Format using wissamantoun/WebOrganizer-TopicClassifier-ModernBERT and wissamantoun/WebOrganizer-FormatClassifier-ModernBERT classifiers. Similar to WebOrganizer/Corpus-200B but using FineEdu instead of DCLM. Topic Labels: Adult Art & Design Software Dev. Crime & Law Education & Jobs Hardware Entertainment Social Life Fashion & Beauty Finance & Business Food & Dining… See the full description on the dataset page: https://huggingface.co/datasets/wissamantoun/fineweb-edu-format-topic.texttext-generation1B<n<10B5 likes1.6k downloads1y agoHugging Face07KaiChen1998 /coda-lm-llava-format CODA-LM Dataset Card CODA-LM is the multi-modal version of the CODA dataset, used in the CODA-LM paper. Both English and Chinese annotations are available. Check detailed usage in our Github repo. This repo contains the CODA-LM dataset, which has been reorganized in the LLaVA data format. You are also welcome to check the original CODA-LM data which contains more metadata vanilla annotations. Usage from datasets import load_dataset # name can be selected from… See the full description on the dataset page: https://huggingface.co/datasets/KaiChen1998/coda-lm-llava-format.imageimage-to-text10K<n<100K3 likes1.3k downloads2y agoHugging Face08alexwww94 /Rexverse-2M-formattedimage1M<n<10M0 likes1.3k downloads8mo agoHugging Face09togethercomputer /glaive-function-calling-v2-formatted Dataset Card for "glaive-function-calling-v2-formatted" More Information needed text100K<n<1M37 likes1.2k downloads3y agoHugging Face10ksterx /hle-no-img-prompt-completion-formatimage1K<n<10K0 likes1.2k downloads1y agoHugging Face11justus27 /math-hendrycks-genesys-formattext1K<n<10K0 likes1.1k downloads1y agoHugging Face12vanloc1808 /pico-banana-smolvlm-format-with-rejected-answer pico-banana-smolvlm-format-with-rejected-answer Balanced image-level tampering detection dataset in SmolVLM-style format with chosen/rejected answer pairs, derived from the pico-banana MCQ pipeline. Suitable for preference learning (e.g. DPO) and RLHF-style training. Dataset overview Same as vanloc1808/pico-banana-smolvlm-format, but each example includes a rejected_answer field: the answer from the counterpart sample (same edited/original image pair, opposite… See the full description on the dataset page: https://huggingface.co/datasets/vanloc1808/pico-banana-smolvlm-format-with-rejected-answer.image100K<n<1M1 likes1.1k downloads7mo agoHugging Face13jeggers /gpqa_formattedgated Dataset Card for GPQA Formatted version of original GPQA dataset. This removes most columns and adds single columns options and answer to contain a list of the possible answers and the index of the correct one. GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy… See the full description on the dataset page: https://huggingface.co/datasets/jeggers/gpqa_formatted.textn<1K4 likes1k downloads2y agoHugging Face14vidore /vidore_v3_finance_en_mteb_format Vidore3FinanceEnRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_finance_en How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3FinanceEnRetrieval") evaluator… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_finance_en_mteb_format.imagevisual-document-retrieval10K<n<100K1 likes995 downloads11mo agoHugging Face15vidore /vidore_v3_computer_science_mteb_format Vidore3ComputerScienceRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_computer_science How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes956 downloads11mo agoHugging Face16ducido /calvin_task_D_D_scale_100_lerobo_formatThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "panda", "total_episodes": 5124, "total_frames": 303794, "total_tasks": 389, "total_videos": 0, "total_chunks": 6, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:5124" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ducido/calvin_task_D_D_scale_100_lerobo_format.imagerobotics100K<n<1M1 likes944 downloads1y agoHugging Face17vidore /vidore_v3_industrial_mteb_format Vidore3IndustrialRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_industrial How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3IndustrialRetrieval")… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_industrial_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes938 downloads11mo agoHugging Face18vidore /vidore_v3_hr_mteb_format Vidore3HrRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_hr How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3HrRetrieval") evaluator = mteb.MTEB([task])… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_hr_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes859 downloads11mo agoHugging Face19vidore /vidore_v3_finance_fr_mteb_format Vidore3FinanceFrRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_finance_fr How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3FinanceFrRetrieval") evaluator… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_finance_fr_mteb_format.imagevisual-document-retrieval10K<n<100K1 likes849 downloads11mo agoHugging Face20vidore /vidore_v3_pharmaceuticals_mteb_format Vidore3PharmaceuticalsRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_pharmaceuticals How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_pharmaceuticals_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes831 downloads11mo agoHugging Face21HuggingFaceH4 /rlaif-v_formattedfrom datasets import load_dataset, features def format(examples): """ Convert prompt from "xxx" to [{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": "xxx"}]}] and chosen and rejected from "xxx" to [{"role": "assistant", "content": [{"type": "text", "text": "xxx"}]}]. Images are wrapped in a list. """ output = {"images": [], "prompt": [], "chosen": [], "rejected": []} for image, question, chosen, rejected in zip(examples["image"]… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/rlaif-v_formatted.image10K<n<100K17 likes820 downloads2y agoHugging Face22vidore /vidore_v3_energy_mteb_format Vidore3EnergyRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_energy How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3EnergyRetrieval") evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_energy_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes801 downloads11mo agoHugging Face23vidore /vidore_v3_physics_mteb_format Vidore3PhysicsRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_physics How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3PhysicsRetrieval") evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_physics_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes794 downloads11mo agoHugging Face24sliuau /DeepScaleR-Preview-Dataset-verl-formattext10K<n<100K0 likes788 downloads11mo agoHugging Face25TimoImhof /TriviaQA-in-SQuAD-format Dataset Card for "TriviaQA-in-SQuAD-format" More Information needed text10K<n<100K6 likes760 downloads3y agoHugging Face26maelic /VG150-coco-format VG150 — Visual Genome 150 (COCO format) This dataset is the standard VG150 split of Visual Genome (Krishna et al., 2017), the most widely used benchmark for Scene Graph Generation, reformatted in standard COCO-JSON format. VG150 contains the top 150 object categories and 50 relations from the original Visual Genome dataset, selected by frequency in the Scene Graph Generation by Iterative Message Passing paper. This version in COCO format was produced as part of the… See the full description on the dataset page: https://huggingface.co/datasets/maelic/VG150-coco-format.imageobject-detection100K<n<1M0 likes740 downloads2mo agoHugging Face27open-paws /tool-use-llama-format Open Paws Tool Use Llama Format This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Tool Use Data Format: JSONL (JSON Lines) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning Organization: Open Paws License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/tool-use-llama-format.texttext-generation1M<n<10M3 likes618 downloads1y agoHugging Face28open-paws /visual-qa-llama-format Open Paws Visual Qa Llama Format This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Multimodal Data Format: JSONL (JSON Lines) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning Organization: Open Paws License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/visual-qa-llama-format.imagetext-generation1M<n<10M1 likes541 downloads1y agoHugging Face29openforcefield /descent-format-geom Dataset Card for Meta-OMol25 Descent Formatted GEOM v1.0 Dataset Details Dataset Description Meta-OMol25 provides molecular structures, coordinates, energies, and forces, and we derived mapped SMILES for broad OpenFF parameter fitting workflows. This release is designed for general fitting and evaluation of van der Waals and valence terms. Curated by: Jennifer A Clark; jaclark5 Funded by: Open Force Field Initiative Shared by: Open Force Field Initiative, Open… See the full description on the dataset page: https://huggingface.co/datasets/openforcefield/descent-format-geom.tabular10M<n<100M0 likes538 downloads6mo agoHugging Face30PerRing /coco_captioning_complete_formatimage100K<n<1M0 likes467 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.