CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yuanhezhang /DAG-MATH-Formatted-CoT Benchmark Overview This dataset card contains 2,894 gold-standard DAG-MATH formatted CoT from problems from Omni-MATH. Top‑Level Schema Each JSON file is a list with a single object describing the problem: problem_id: integer identifier of the problem. domain: list of strings describing the topic taxonomy. difficulty: numeric difficulty indicator from 1 (easiest) to 6 (hardest). problem_text: problem statement. sample_id: sample identifier for the solution trace.… See the full description on the dataset page: https://huggingface.co/datasets/yuanhezhang/DAG-MATH-Formatted-CoT.tabular1K<n<10K1 likes127 downloads11mo agoHugging Face02Nhanvi282 /openvivqa-formating-vlm OpenViVQA Formatting Dataset for VLM A Vietnamese multimodal instruction-format dataset for training Vision Language Models (VLMs) on Visual Question Answering (VQA) tasks. This dataset reformats OpenViVQA-style samples into conversational instruction-tuning format compatible with modern VLM training pipelines such as: Qwen2-VL LLaVA InternVL Phi-3 Vision Idefics SmolVLM Dataset Structure Each sample contains: image: input image conversations: multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/Nhanvi282/openvivqa-formating-vlm.imagevisual-question-answering10K<n<100K1 likes84 downloads5mo agoHugging Face03germtf /astrolabe-trace-format-test Astrolabe trace format test (temporary) tabularn<1K0 likes31 downloads3mo agoHugging Face04playra /numeric-format-catalog Numeric Format Catalog (Trinity S3AI / t27) Version: v3.0 (2026-06-13) -- supersedes v2 (count=81, withdrawn) and v1 (count=77, withdrawn). Count: 83 numeric formats across 13 clusters. License: CC-BY-4.0. Maintainer: Trinity S3AI -- admin@t27.ai -- ORCID 0009-0008-4294-6159. Companion dataset: playra/numeric-conformance-packs (bit-exact conformance vectors that instantiate this catalog). Linked papers: Catalog preprint (84-format ruler): arXiv:2606.09686 -- HF paper page… See the full description on the dataset page: https://huggingface.co/datasets/playra/numeric-format-catalog.tabularothern<1K0 likes29 downloads3mo agoHugging Face05Arabic-Clip-Archive /Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-SigLIP-512_validationimage1K<n<10K0 likes24 downloads3y agoHugging Face06spectralbranding /exp-compounding-format Compounding x Format: Specification Framing in Agentic Pipelines Dataset Summary Two experiments (1,440 LLM calls across 8 models) testing whether specification framing attenuates or amplifies dimensional collapse across multi-step agentic shopping pipelines. The dataset operationalises Section 5.16 ("Specification Paradox") of the companion R15 paper: Brand Function specification works in single-step contexts (reducing the Dimensional Collapse Index toward the… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-compounding-format.tabulartext-classification1K<n<10K0 likes22 downloads3mo agoHugging Face07spectralbranding /exp-bf-format Brand Function Format Optimization (Exp D) Dataset Summary 375 LLM responses (355 valid, 20 parse errors, 0 failures) testing which representational format of a brand function specification maximizes AI comprehension fidelity. Five formats (JSON structured, prose narrative, tabular minimal, ranked list, score-only vector) were crossed with five canonical SBT brands (Hermes, IKEA, Patagonia, Tesla, Erewhon), five model families (Claude Haiku 4.5, GPT-4o-mini… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-bf-format.tabulartext-generationn<1K0 likes15 downloads3mo agoHugging Face08Arabic-Clip-Archive /Arabic_dataset_1M_translated_jsonl_format_ViT-B-16-plus-240This translation done using https://huggingface.co/Helsinki-NLP/opus-mt-en-ar image100K<n<1M0 likes13 downloads3y agoHugging Face09Arabic-Clip-Archive /Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240This dataset repo contains the dataset (CC3M+CC12M+SBU) translated using opus-mt-en-ar and cleaned. Its size about 13M image1M<n<10M0 likes12 downloads3y agoHugging Face10Arabic-Clip-Archive /Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240-fulldata-v2DatasetDict({ train: Dataset({ features: ['index', 'embeddings', 'en_caption', 'ar_caption', 'nr_words', 'url'], num_rows: 12166802 }) }) image1M<n<10M0 likes12 downloads3y agoHugging Face11Arabic-Clip-Archive /ccs_synthetic_ar_1M-Arabic_dataset_1M_translated_jsonl_formatimage1M<n<10M0 likes9 downloads3y agoHugging Face12open-llm-leaderboard /ontocord__wide_3b_sft_stage1.2-ss1-expert_formatted_text-detailsgated Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.2-ss1-expert_formatted_text Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.2-ss1-expert_formatted_text The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.2-ss1-expert_formatted_text-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face13open-llm-leaderboard /ontocord__wide_3b_sft_stag1.2-lyrical_news_software_howto_formattedtext-merge-detailsgated Dataset Card for Evaluation run of ontocord/wide_3b_sft_stag1.2-lyrical_news_software_howto_formattedtext-merge Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stag1.2-lyrical_news_software_howto_formattedtext-merge The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stag1.2-lyrical_news_software_howto_formattedtext-merge-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face14open-llm-leaderboard /openbmb__MiniCPM-S-1B-sft-llama-format-detailsgated Dataset Card for Evaluation run of openbmb/MiniCPM-S-1B-sft-llama-format Dataset automatically created during the evaluation run of model openbmb/MiniCPM-S-1B-sft-llama-format The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openbmb__MiniCPM-S-1B-sft-llama-format-details.tabular10K<n<100K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.