datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen35-4b-drpo-vs0f49th-trainer-logprobs
Qwen3.5 4B DRPO trainer logprobs from W&B run vs0f49th
This dataset contains the raw trainer-logprob JSONL shards saved by W&B run ai2-llm/open_instruct_internal/vs0f49th (qwen35_4b_drpo__42__1782345587).
Contents
Source run: https://wandb.ai/ai2-llm/open_instruct_internal/runs/vs0f49th
Source path: /weka/oe-adapt-default/allennlp/deletable_rollouts/
Filename pattern: qwen35_4b_drpo__42__1782345587_trainer_logprobs_step*_rank*.jsonl
Files: 4320 JSONL shards… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/qwen35-4b-drpo-vs0f49th-trainer-logprobs.alpaca-farm-davinci-003-2048-tokenbaligh-hamasa
Balīgh
Instruction-tuning data for classical Arabic, built from printed books of the
Arabic philological tradition. 40,145 records across two books.
taḥwīl (تحويل) — rewrite modern Arabic, MSA or dialect, into classical Arabic.
qa — one linguistic fact from the book, asked in a real voice (student,
reader, writer, teacher, editor, preacher, learner) and answered in the
author's words.
sharḥ_kāmil — a verse explained under fixed scholarly headings from
several of its claims at… See the full description on the dataset page: https://huggingface.co/datasets/omarabb315/baligh-hamasa.Hambobos-RandomNumbers_50M
Внимание!⚠️
этот датасет использует split в 50 секций для адекватного отправления на сервер
Детали⚙️
было созданно с помощью ChatGPT 5 mini
Использование✨
датасет состоит из 50 split деталей с названиями типа random_number.jsonl.part001, удачи в использовании!
AUTOSTT-ENG-correctionsAraLingBench
AraLingBench
📄 Paper: arXiv:2511.14295💻 GitHub: hammoudhasan/AraLingBench
AraLingBench is a 150-question Arabic multiple-choice benchmark that tests core linguistic competence of language models across five pillars:
النحو (Grammar)
الصرف (Morphology)
الإملاء (Spelling & Orthography)
فهم اللغة (Reading Comprehension)
التركيب اللغوي والأسلوبي (Syntax & Stylistics)
All questions are human-authored and validated, with a single correct answer and a difficulty label: Easy, Medium, or… See the full description on the dataset page: https://huggingface.co/datasets/hammh0a/AraLingBench.Hala-4.6M-SFT
Hala: Arabic-Centric Instruction & Translation Dataset
Paper: Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
Authors: Hasan Abed Al Kader Hammoud*, Mohammad Zbeeb*, Bernard Ghanem
Affiliation: King Abdullah University of Science and Technology (KAUST)
*Equal contribution
In Arabic, حلا (Hala) conveys sweetness and beauty—qualities long associated with the language itself. In this spirit, we extend Hala to datasets that aim to enrich… See the full description on the dataset page: https://huggingface.co/datasets/hammh0a/Hala-4.6M-SFT.NASA-EO-Bench
NASA-EO-Bench
A large-scale benchmark for geoscience dataset retrieval, derived from citation relationships in peer-reviewed NASA publications.
Paper: Bringing Agentic Search to Earth Observation Data Discovery — CIKM '26, 10.1145/3799682.3841109
Overview
Finding the right NASA Earth observation dataset for a given research need is hard even for domain experts. NASA-EO-Bench operationalises this task as an information retrieval problem: given a natural-language… See the full description on the dataset page: https://huggingface.co/datasets/HamiltonMYu/NASA-EO-Bench.Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1
Persian Civil Procedure QA Dataset
Dataset Description
این مجموعهداده شامل پرسشوپاسخهای حقوقی به زبان فارسی در حوزه آیین دادرسی مدنی است.
هر نمونه شامل سه فیلد اصلی است:
question: پرسش حقوقی
answer: پاسخ پرسش
evidence_quote: عبارت دقیق و مستند از دادهٔ منبع که پاسخ بر اساس آن استخراج شده است
هدف مجموعهداده، فراهمکردن دادهای ساختاریافته برای آموزش، ارزیابی و توسعه مدلهای زبانی فارسی در زمینه پرسشوپاسخ حقوقی است.
Dataset Structure
نمونهای… See the full description on the dataset page: https://huggingface.co/datasets/hamidsalimi/Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1.200k-tulu-2-unbalanced
Tulu 2 Unfiltered - 200k subset
This the 200k subset of the 'unfiltered' version of the Tulu v2 SFT mixture, created by collating the original Tulu 2 sources and avoiding downsampling.
This was used for the 200k-size experiments.
Details
The dataset consists of a mix of :
FLAN (Apache 2.0, we only sample 961,322 samples along with 398,439 CoT samples from the full set for this data pool)
Open Assistant 1 (Apache 2.0)
ShareGPT (Apache 2.0 listed, no official repo… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/200k-tulu-2-unbalanced.hexphitulu-2-unfiltered
Tulu 2 Unfiltered
This is an 'unfiltered' version of the Tulu v2 SFT mixture, created by collating the original Tulu 2 sources and avoiding downsampling.
Details
The dataset consists of a mix of :
FLAN (Apache 2.0, we only sample 961,322 samples along with 398,439 CoT samples from the full set for this data pool)
Open Assistant 1 (Apache 2.0)
ShareGPT (Apache 2.0 listed, no official repo found)
GPT4-Alpaca (CC By NC 4.0)
Code-Alpaca (CC By NC 4.0)
LIMA (CC BY-NC-SA)… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/tulu-2-unfiltered.Marathon
Dataset Card for Marathon
Release
[2024/05/15] 🔥 Marathon is accepted by ACL 2024 Main Conference.
Dataset Summary
Marathon benchmark is a new long-context multiple-choice benchmark, mainly based on LooGLE, with some original data from LongBench. The context length can reach up to 200K+. Marathon benchmark comprises six tasks: Comprehension and Reasoning, Multiple Information Retrieval, Timeline Reorder, Computation, Passage Retrieval, and Short Dependency… See the full description on the dataset page: https://huggingface.co/datasets/Hambaobao/Marathon.rds-sels-tulu-3-arena-hard-939k
RDS+ Selected Tulu 3 Arena Hard 939k
This is the dataset (and associated scores) selected by RDS+ when selecting 939k samples using Arena Hard samples.
For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning.
This was used to train this model.
This dataset is selected from Tulu 3 unfiltered, and please see that page for more information on sources.
License
This dataset is licensed under ODC-BY-1.0. It is intended for… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-tulu-3-arena-hard-939k.rds-sels-multitask-rrmax-top326k
RDS+ Selected Multitask 326k
This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples for multiple tasks at once.
For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning.
This was used to train this model.
This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources.
License
We are releasing this dataset under the terms of ODC-BY. By using this, you… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-multitask-rrmax-top326k.HealthyLife-Insurance-Charge-Prediction-v2tulu_mix_store
Tulu Mix Store
A bit of a dumping ground for files that we created as part of the tulu 2 mix.
Please see our official dataset page for details on the mix and associated licenses.
multiple-choice-questions
Questões de Múltipla Escolha - Base de dados (PT-BR)
Contextualização
Este repositório contém uma base de dados (data.json) com questões de múltipla escolha, a qual foi utilizada principalmente no desenvolvimento de modelos de recuperação de informação.
Descrição do conjunto de dados
O conjunto de dados é composto por questões de múltipla escolha, abrangendo uma variedade de temas dentro da área da Ciência da Computação. Cada questão é estruturada em formato… See the full description on the dataset page: https://huggingface.co/datasets/mateus-hamade/multiple-choice-questions.coco_2017_caption_trainhambobos-basicmath_10M
ВНИМАНИЕ
датасет состоит из 10 чанков!!!
использование
expression это сам типа 2+2, answer это ответ на этот самый примерный 2+2
информация
сделано с помощью ChatGPT 5 mini на телефоне
лицензия mit
Oxfaicorpusgsm8k-symbolic
GSM8k Symbolic
This is an uploaded form of the dataset from Diffusion of Thoughts,
adapted to follow the same format as Tulu datasets.
Citation
If you find this work useful, please cite the original work:
@article{ye2024diffusion,
title={Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models},
author={Ye, Jiacheng and Gong, Shansan and Chen, Liheng and Zheng, Lin and Gao, Jiahui and Shi, Han and Wu, Chuan and Li, Zhenguo and Bi, Wei and Kong… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/gsm8k-symbolic.fitness-qa
Fitness-QA
This is a synthetic dataset for fitness content based on "neuml/txtai-wikipedia" embedding index.
The generation of statements from context uses txtinstruct.
This dataset contains questions generated from contexts using the statement generator "flan-t5-base" trained on SQuAD dataset.
Each context includes generated questions with coherent relevant answers, and the irrelevant questions with (I don't have data on that).
Fitness data is pulled from wikipedia data stored… See the full description on the dataset page: https://huggingface.co/datasets/hammamwahab/fitness-qa.rds-sels-arena-hard-top326k
RDS+ Selected Arena Hard 326k
This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples using Arena Hard samples.
For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning.
This was used to train this model.
This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources.
License
We are releasing this dataset under the terms of ODC-BY. By using this, you… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-arena-hard-top326k.math_rlvr_mixture_dpodigitize-pid-ner
Digitize-PID: Pipeline numbers (NER)
Note: I am not the author of this dataset
Named Entity Recognition dataset for extracting pipeline numbers from full text of P&ID
(Piping and Instrumentation Diagram) documents.
Dataset Details
Dataset Description
Pipeline numbers are structured identifiers in engineering documents:
Example Format: A-123-BC (3-5 segments with a separator such as -, , or _)
Use case: Automated extraction from P&ID document text
Domain:… See the full description on the dataset page: https://huggingface.co/datasets/hamzas/digitize-pid-ner.Teeet
Demet Turkish Chat Dataset
Bu dataset, Türkçe sohbet modeli fine-tuning için hazırlanmış konuşma örnekleri içerir.
Dataset Bilgileri
Karakter: Demet
Yaş: 17
Şehir: Ankara
Dil: Türkçe
Format: ChatML (messages formatı)
Kullanım
from datasets import load_dataset
dataset = load_dataset("hammur/Teeet")
Format
Her örnek şu formatta:
{
"messages": [
{"role": "system", "content": "..."},
{"role": "user", "content": "..."},
{"role":… See the full description on the dataset page: https://huggingface.co/datasets/hammur/Teeet.MoroccanMedMCQA-FR
MoroccanMedMCQA-FR: A French-Language Moroccan Medical Multiple-Choice QA Benchmark
Dataset Description
MoroccanMedMCQA-FR is the first French-language medical multiple-choice question answering (MCQ) benchmark grounded in the Moroccan medical faculty curriculum. It comprises 6,771 officially sourced MCQs drawn from past examinations of the Faculty of Medicine and Pharmacy of Fès (FMPF), Sidi Mohammed Ben Abdellah University, Morocco… See the full description on the dataset page: https://huggingface.co/datasets/hamzaaouadi/MoroccanMedMCQA-FR.rds-sels-alpacafarm-top326k
RDS+ Selected AlpacaFarm 326k
This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples using AlpacaFarm samples.
For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning.
This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources.
When finetuning a Llama 2 7b model on this data using the associated codebase and evaluating with the same codebase, the expected… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-alpacafarm-top326k.advbench
