datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twaitch-txtMSVQAThis is a multimodal cross-scenario dataset for continual learning with MLLMs. We provide a simple script to split the dataset in multiple ways.
The dataset format has been adjusted for Qwen. The coordinates in 'train_annfiles.json' and 'val_annfiles.json' are adjusted to Qwen2.5VL format. And 'train_annfiles_ori.json' and 'val_annfiles_ori.json' retain the original coordinates of the bounding box.
You need to adjust the coordinates fit your format.
Detailed information can refer to… See the full description on the dataset page: https://huggingface.co/datasets/Kaij00/MSVQA.Kairos
Kairos — Long-Form Video Annotation and Benchmark
Kairos is an automated annotation pipeline for long-duration videos (10–30 minutes).
This repository hosts a benchmark of 2,870 multiple-choice and 2,870 free-form
(OpenQA) questions across 820 videos, spanning 17 fine-grained capabilities and
5 temporal tiers (T1: single moment, T2: 1–60 s, T3: 60–300 s, T4: 300–900 s, T5: >900 s).
What's inside
.
├── data/
│ ├── kairos_benchmark.jsonl # 2,870 MCQs (bilingual… See the full description on the dataset page: https://huggingface.co/datasets/nips26anonymous159/Kairos.twaitch-txt-2system-one-270m-data
system-one-270m-data
25,002 synthetic typed decisions: a piece of state, a question, a
caller-supplied option set, and a soft target distribution over those options.
Built to train kaivoss/system-one-270m,
an open take on the System One model class (TypeSafe
Jev,
Laya).
Schema
Field
Type
Meaning
prompt
string
the full rendered prompt, state + question + lettered options
letters
list[string]
the option letters in play, ["A", "B", ...]
target… See the full description on the dataset page: https://huggingface.co/datasets/kaivoss/system-one-270m-data.carnice-glm5-hermes-traces
Carnice GLM-5 Hermes Traces
This dataset is a merged release bundle of GLM-5 traces collected through the Hermes Agent harness.
It was generated by running the carnice_trace_prompt_bank_v4 prompt bank through Hermes Agent with:
z-ai/glm-5 via OpenRouter
local/file/terminal/code-execution tools for local tasks
Hermes browser tools plus Tavily-backed web_search / web_extract for web tasks
isolated disposable workspaces per prompt
This release is prepared for Hugging Face upload and… See the full description on the dataset page: https://huggingface.co/datasets/kai-os/carnice-glm5-hermes-traces.awesome-ai4ai
Awesome AI4AI — the catalog behind the survey
The structured catalog accompanying "AI4AI Survey: From Long-Horizon Agents to
Recursive Self-Improvement — Definitions, Reliable Horizons, and Open Problems",
by 23 authors across Tongji, SJTU, UC Berkeley, CASIA, NUS, NTU, and Simple Agent Lab.
📄 Paper: https://www.preprints.org/manuscript/202608.2108/v1
🔗 DOI: https://doi.org/10.20944/preprints202608.2108.v1
📥 PDF, original layout:… See the full description on the dataset page: https://huggingface.co/datasets/KaiWu123/awesome-ai4ai.MultiLingual-SorryBench
MLSFT Multilingual SORRY-Bench Evaluation Dataset
⚠️ CONTENT WARNING: This dataset contains adversarial prompts specifically designed to elicit harmful outputs from language models. It is intended for safety research and evaluation purposes only.
Dataset Description
A comprehensive multilingual safety evaluation dataset based on SORRY-bench for assessing model refusal rates and safety properties across 8 languages:
Chinese (zh)
Danish (da)
Greek (el)
Hindi (hi)
Irish… See the full description on the dataset page: https://huggingface.co/datasets/kairawal/MultiLingual-SorryBench.kaist-ai__janus-rm-7b-details
Dataset Card for Evaluation run of kaist-ai/janus-rm-7b
Dataset automatically created during the evaluation run of model kaist-ai/janus-rm-7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/kaist-ai__janus-rm-7b-details.EmoSet_15Kkaia-timer-dataset1mcq_test_2rewire2-ultrafineweb-moonlightcynthiav2rewire2-ultrafineweb-moonlight2labeled_creativelabeled_creative_newkaist-ai__mistral-orpo-capybara-7k-details
Dataset Card for Evaluation run of kaist-ai/mistral-orpo-capybara-7k
Dataset automatically created during the evaluation run of model kaist-ai/mistral-orpo-capybara-7k
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/kaist-ai__mistral-orpo-capybara-7k-details.kaist-ai__janus-dpo-7b-details
Dataset Card for Evaluation run of kaist-ai/janus-dpo-7b
Dataset automatically created during the evaluation run of model kaist-ai/janus-dpo-7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/kaist-ai__janus-dpo-7b-details.diverse-qa-dclm-moonlight-texternie_dclmpro_diverse_synthrefine_codemcq-gen-testsynth-ultrafineweb-moonlight-diversediverse_synthernie_dclmpro_synth_sample_17krefine_code_onekaist-ai__janus-7b-details
Dataset Card for Evaluation run of kaist-ai/janus-7b
Dataset automatically created during the evaluation run of model kaist-ai/janus-7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/kaist-ai__janus-7b-details.nemotron-diverse-qa-dclm-moonlightdiverse-synth-moonlight-proc
