datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
adaption-experimentalcotqa
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-ExperimentalCoTQA
This dataset contains 1,000 question-and-answer pairs focusing on the evolution of African societies, including topics like the role of women, indigenous languages, traditional leadership, and polyrhythmic music. The content explores historical transitions from pre-colonial eras through colonialism to modern-day challenges and cultural innovations.… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/adaption-experimentalcotqa.tr-cc-experimental
Veri Seti Hakkında
Kaynak: Common Crawl (CC) Türkçe Ağ Verileri
Toplayan [Ben (Swag Victoria)]
Temizleme & Filtreleme: GLM 5.3 Flash yardımıyla HTML etiketleri, navigasyon gürültüleri ve spam metinler ayıklanmaya çalışılmıştır.
Format: .jsonl (JSON Lines)
Q&A Abi neden bu kadar yavaş geliyor CC kazıma işlemi? (Answer) Abi'nin amına koyim.
persona-belief-probeslm-eval-results-ChaoticNeutrals-Prima-LelantaclesV7-experimental-7b-private
Dataset Card for Evaluation run of ChaoticNeutrals/Prima-LelantaclesV7-experimental-7b
Dataset automatically created during the evaluation run of model ChaoticNeutrals/Prima-LelantaclesV7-experimental-7b
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-ChaoticNeutrals-Prima-LelantaclesV7-experimental-7b-private.experimental-optimizationHumanAgencyBench_Human_Annotations
Human annotations and LLM judge comparative Dataset
Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
Code: https://github.com/BenSturgeon/HumanAgencyBench/
Dataset Description
This dataset contains 60,000 evaluated AI assistant responses across 6 dimensions of behaviour relevant to human agency support, with both model-based and human annotations. Each example includes evaluations from 4 different frontier LLM models. We also provide… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Human_Annotations.RaifuWars-Warrior-Experimental-SFT
Raifu Wars — Warrior SFT, built-in AI on Arboretum (v0)
Experimental first cut — read the limitations before training on it.
Two facts the repository name does not carry:
the teacher is the game's built-in heuristic AI, not a human and not a strong model
every row is the same map, Arboretum, across 40 seeds
Treat this as something to be superseded rather than built on.
Supervised fine-tuning data for an LLM that plays a seat in Raifu Wars, a turn-based strategy
game, through… See the full description on the dataset page: https://huggingface.co/datasets/yotisstudios/RaifuWars-Warrior-Experimental-SFT.RP-logs-V2-Experimental-prefixedBlueSky-Experimental-sharegptPinkstack__Superthoughts-lite-1.8B-experimental-o1-details
Dataset Card for Evaluation run of Pinkstack/Superthoughts-lite-1.8B-experimental-o1
Dataset automatically created during the evaluation run of model Pinkstack/Superthoughts-lite-1.8B-experimental-o1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pinkstack__Superthoughts-lite-1.8B-experimental-o1-details.transduction_experimental_resultssethuiyer__LlamaZero-3.1-8B-Experimental-1208-details
Dataset Card for Evaluation run of sethuiyer/LlamaZero-3.1-8B-Experimental-1208
Dataset automatically created during the evaluation run of model sethuiyer/LlamaZero-3.1-8B-Experimental-1208
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sethuiyer__LlamaZero-3.1-8B-Experimental-1208-details.nbeerbower__Mistral-Nemo-Moderne-12B-FFT-experimental-details
Dataset Card for Evaluation run of nbeerbower/Mistral-Nemo-Moderne-12B-FFT-experimental
Dataset automatically created during the evaluation run of model nbeerbower/Mistral-Nemo-Moderne-12B-FFT-experimental
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nbeerbower__Mistral-Nemo-Moderne-12B-FFT-experimental-details.nbeerbower__Nemo-Loony-12B-experimental-details
Dataset Card for Evaluation run of nbeerbower/Nemo-Loony-12B-experimental
Dataset automatically created during the evaluation run of model nbeerbower/Nemo-Loony-12B-experimental
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nbeerbower__Nemo-Loony-12B-experimental-details.nbeerbower__SmolNemo-12B-FFT-experimental-details
Dataset Card for Evaluation run of nbeerbower/SmolNemo-12B-FFT-experimental
Dataset automatically created during the evaluation run of model nbeerbower/SmolNemo-12B-FFT-experimental
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nbeerbower__SmolNemo-12B-FFT-experimental-details.sometimesanotion__Lamarck-14B-v0.1-experimental-details
Dataset Card for Evaluation run of sometimesanotion/Lamarck-14B-v0.1-experimental
Dataset automatically created during the evaluation run of model sometimesanotion/Lamarck-14B-v0.1-experimental
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Lamarck-14B-v0.1-experimental-details.sethuiyer__Llama-3.1-8B-Experimental-1208-Instruct-details
Dataset Card for Evaluation run of sethuiyer/Llama-3.1-8B-Experimental-1208-Instruct
Dataset automatically created during the evaluation run of model sethuiyer/Llama-3.1-8B-Experimental-1208-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sethuiyer__Llama-3.1-8B-Experimental-1208-Instruct-details.uninstruct-v1-big-experimental-chatmlsethuiyer__Llama-3.1-8B-Experimental-1206-Instruct-details
Dataset Card for Evaluation run of sethuiyer/Llama-3.1-8B-Experimental-1206-Instruct
Dataset automatically created during the evaluation run of model sethuiyer/Llama-3.1-8B-Experimental-1206-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sethuiyer__Llama-3.1-8B-Experimental-1206-Instruct-details.uninstruct-v1-experimental-chatmlSubset of SlimPajama-6B where tokens associated with chatml prompt format are randomly added mid-text to make model forget how to do instruct and make it behave like a completion model which is not instruction following. Base Yi 1.5 models are contaminated on synthetic SFT data, hence the need for de-contamination attempts before further finetuning, if you don't want your end model to behave like ChatGPT. Should also work with Qwen 1.5 and Qwen 2 models, they are contaminated too.
hawky-ai-andromeda-cot-experimentalDdataset_tonto
