CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kainecorneko /twaitch-txttabular1K<n<10K0 likes4.7k downloads8mo agoHugging Face02Kaij00 /MSVQAThis is a multimodal cross-scenario dataset for continual learning with MLLMs. We provide a simple script to split the dataset in multiple ways. The dataset format has been adjusted for Qwen. The coordinates in 'train_annfiles.json' and 'val_annfiles.json' are adjusted to Qwen2.5VL format. And 'train_annfiles_ori.json' and 'val_annfiles_ori.json' retain the original coordinates of the bounding box. You need to adjust the coordinates fit your format. Detailed information can refer to… See the full description on the dataset page: https://huggingface.co/datasets/Kaij00/MSVQA.imagevisual-question-answering10K<n<100K2 likes3.5k downloads9mo agoHugging Face03nips26anonymous159 /Kairos Kairos — Long-Form Video Annotation and Benchmark Kairos is an automated annotation pipeline for long-duration videos (10–30 minutes). This repository hosts a benchmark of 2,870 multiple-choice and 2,870 free-form (OpenQA) questions across 820 videos, spanning 17 fine-grained capabilities and 5 temporal tiers (T1: single moment, T2: 1–60 s, T3: 60–300 s, T4: 300–900 s, T5: >900 s). What's inside . ├── data/ │ ├── kairos_benchmark.jsonl # 2,870 MCQs (bilingual… See the full description on the dataset page: https://huggingface.co/datasets/nips26anonymous159/Kairos.tabularvideo-text-to-text1K<n<10K0 likes1.9k downloads5mo agoHugging Face04kainecorneko /twaitch-txt-2tabularn<1K0 likes1.1k downloads2mo agoHugging Face05kaivoss /system-one-270m-data system-one-270m-data 25,002 synthetic typed decisions: a piece of state, a question, a caller-supplied option set, and a soft target distribution over those options. Built to train kaivoss/system-one-270m, an open take on the System One model class (TypeSafe Jev, Laya). Schema Field Type Meaning prompt string the full rendered prompt, state + question + lettered options letters list[string] the option letters in play, ["A", "B", ...] target… See the full description on the dataset page: https://huggingface.co/datasets/kaivoss/system-one-270m-data.tabulartext-classification10K<n<100K0 likes169 downloads4d agoHugging Face06kai-os /carnice-glm5-hermes-traces Carnice GLM-5 Hermes Traces This dataset is a merged release bundle of GLM-5 traces collected through the Hermes Agent harness. It was generated by running the carnice_trace_prompt_bank_v4 prompt bank through Hermes Agent with: z-ai/glm-5 via OpenRouter local/file/terminal/code-execution tools for local tasks Hermes browser tools plus Tavily-backed web_search / web_extract for web tasks isolated disposable workspaces per prompt This release is prepared for Hugging Face upload and… See the full description on the dataset page: https://huggingface.co/datasets/kai-os/carnice-glm5-hermes-traces.tabulartext-generation1K<n<10K59 likes148 downloads6mo agoHugging Face07KaiWu123 /awesome-ai4ai Awesome AI4AI — the catalog behind the survey The structured catalog accompanying "AI4AI Survey: From Long-Horizon Agents to Recursive Self-Improvement — Definitions, Reliable Horizons, and Open Problems", by 23 authors across Tongji, SJTU, UC Berkeley, CASIA, NUS, NTU, and Simple Agent Lab. 📄 Paper: https://www.preprints.org/manuscript/202608.2108/v1 🔗 DOI: https://doi.org/10.20944/preprints202608.2108.v1 📥 PDF, original layout:… See the full description on the dataset page: https://huggingface.co/datasets/KaiWu123/awesome-ai4ai.tabularn<1K0 likes69 downloads24d agoHugging Face08kairawal /MultiLingual-SorryBench MLSFT Multilingual SORRY-Bench Evaluation Dataset ⚠️ CONTENT WARNING: This dataset contains adversarial prompts specifically designed to elicit harmful outputs from language models. It is intended for safety research and evaluation purposes only. Dataset Description A comprehensive multilingual safety evaluation dataset based on SORRY-bench for assessing model refusal rates and safety properties across 8 languages: Chinese (zh) Danish (da) Greek (el) Hindi (hi) Irish… See the full description on the dataset page: https://huggingface.co/datasets/kairawal/MultiLingual-SorryBench.tabular1K<n<10K0 likes44 downloads6mo agoHugging Face09open-llm-leaderboard /kaist-ai__janus-rm-7b-detailsgated Dataset Card for Evaluation run of kaist-ai/janus-rm-7b Dataset automatically created during the evaluation run of model kaist-ai/janus-rm-7b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/kaist-ai__janus-rm-7b-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face10Kai910 /EmoSet_15Kimage10K<n<100K0 likes25 downloads1y agoHugging Face11yellooot /kaia-timer-dataset1tabular1K<n<10K0 likes16 downloads1y agoHugging Face12kaizen9 /mcq_test_2tabular1K<n<10K0 likes11 downloads1y agoHugging Face13kaizen9 /rewire2-ultrafineweb-moonlighttabular1K<n<10K0 likes9 downloads1y agoHugging Face14kaizen9 /cynthiav2tabular1K<n<10K0 likes9 downloads1y agoHugging Face15kaizen9 /rewire2-ultrafineweb-moonlight2tabular1K<n<10K0 likes8 downloads1y agoHugging Face16kaizen9 /labeled_creativetabular1K<n<10K0 likes8 downloads1y agoHugging Face17kaizen9 /labeled_creative_newtabular1K<n<10K0 likes8 downloads1y agoHugging Face18open-llm-leaderboard /kaist-ai__mistral-orpo-capybara-7k-detailsgated Dataset Card for Evaluation run of kaist-ai/mistral-orpo-capybara-7k Dataset automatically created during the evaluation run of model kaist-ai/mistral-orpo-capybara-7k The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/kaist-ai__mistral-orpo-capybara-7k-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face19open-llm-leaderboard /kaist-ai__janus-dpo-7b-detailsgated Dataset Card for Evaluation run of kaist-ai/janus-dpo-7b Dataset automatically created during the evaluation run of model kaist-ai/janus-dpo-7b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/kaist-ai__janus-dpo-7b-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face20kaizen9 /diverse-qa-dclm-moonlight-texttabular100K<n<1M0 likes7 downloads1y agoHugging Face21kaizen9 /ernie_dclmpro_diverse_synthtabular100K<n<1M0 likes7 downloads1y agoHugging Face22kaizen9 /refine_codetabularn<1K0 likes7 downloads9mo agoHugging Face23kaizen9 /mcq-gen-testtabularn<1K0 likes6 downloads1y agoHugging Face24kaizen9 /synth-ultrafineweb-moonlight-diversetabular1K<n<10K0 likes6 downloads1y agoHugging Face25kaizen9 /diverse_synthtabular100K<n<1M0 likes6 downloads1y agoHugging Face26kaizen9 /ernie_dclmpro_synth_sample_17ktabular10K<n<100K0 likes6 downloads1y agoHugging Face27kaizen9 /refine_code_onetabular10K<n<100K0 likes6 downloads9mo agoHugging Face28open-llm-leaderboard /kaist-ai__janus-7b-detailsgated Dataset Card for Evaluation run of kaist-ai/janus-7b Dataset automatically created during the evaluation run of model kaist-ai/janus-7b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/kaist-ai__janus-7b-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face29kaizen9 /nemotron-diverse-qa-dclm-moonlighttabular100K<n<1M0 likes5 downloads1y agoHugging Face30kaizen9 /diverse-synth-moonlight-proctabular100K<n<1M0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.