datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prop_logic_puzzledataioiemeHuffPostA dataset of approximately 200K news headlines from the year 2012 to 2018 collected from HuffPost.dataset-meds-84149598ff52f2cads-5540a63e610e2a90dataset-namedatameimagenet-ccybermetric-10000
Dataset Card for "cybermetric-10000"
More Information needed
multi-hop-qa-function-calling-format-V1.0This dataset is converted from khaimaitien/qa-expert-multi-hop-qa-V1.0 to OpenAI function calling format.
Each data point is a list of messages with role=user, assistant or function:
message that role=user, content is the question
message that role=assistant, content is not None, function_call is None: --> assistant responds with text only
message that role=assistant and function_call is not None --> assistant asks to execute a function call
function_call is of the form: {"name": "retrieve"… See the full description on the dataset page: https://huggingface.co/datasets/khaimaitien/multi-hop-qa-function-calling-format-V1.0.pairs_with_scores_v27kothar-dataset-it
Kothar fine-tuning datasets
This repo documents the instruction-tuning datasets built by this repository for a
protein language model. All datasets are seeded from the Neo4j protein knowledge graph
(docs/neo4j_schema.md) or from computed sequence features, built by the scripts under
scripts/python/, and published as HuggingFace Hub configs at
khairi/kothar-dataset-it.
What's here
Document
Covers
shared-conventions.md
Sequence encoding, the computed… See the full description on the dataset page: https://huggingface.co/datasets/khairi/kothar-dataset-it.asr-youtube-datasetdatasenamekhakas-russian-parallel-corpus
Khakas-Russian Parallel Corpus
The creation of this dataset is aimed at supporting the development of natural language processing (NLP) tools and
machine translation for the Khakas language, which is classified as a "Definitely Endangered" language. By providing
high-quality parallel data, this project helps preserve the linguistic heritage of the Khakas people.
Dataset Overlap:
The Khakas sentences in this corpus do not overlap with those in the Khakas… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-parallel-corpus.arocrbench_khattKITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding
This dataset is designed to evaluate the performance of Arabic OCR and document understanding systems. It includes a variety of document types and tasks.
Please see paper & code for more information:
GitHub Repository
Project Page
arXiv Paper
InternVL_Chat_V12_SFT_DatacyberQA
Dataset Card for "cyberQA"
More Information needed
ppe-benchmark-eval
PPE Benchmark Eval Set (v1)
A held-out, human-verified benchmark for evaluating vision-language models on
personal protective equipment (PPE) detection — specifically hardhat and
safety-vest presence — framed as a VQA-style classification task.
What this is
96 images, balanced 24/24/24/24 across the four hardhat × vest combinations
(yes/yes, yes/no, no/yes, no/no). Sourced from a forked, filtered subset of
the karabuk-university PPE dataset
on Roboflow Universe… See the full description on the dataset page: https://huggingface.co/datasets/khadijah00/ppe-benchmark-eval.KhanomTanLLM-pretrained-dataset
KhanomTanLLM pretrained dataset
This daataset collect all raw text for pretraining LLM.
Codename: numfa v2
Repository: https://github.com/pythainlp/KhanomTanLLM
Tokens
53,376,211,711 Tokens
English: 31,629,984,243 Tokens
Thai: 12,785,565,497 Tokens
Code: 8,913,084,300 Toekns
Parallel data: 190,310,686 Tokens
Based on Typhoon-7B (https://huggingface.co/scb10x/typhoon-7b) tokenizer
All subset
Thai
pythainlp/thai_food_v1.0
pythainlp/thailaw-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset.arabic-latin-invoices-synthetic
Synthetic Arabic/Latin Invoices — label-first multimodal dataset
A large, perfectly-labeled synthetic invoice dataset for training multimodal
invoice-parsing models, with first-class support for Arabic script, mixed Arabic/Latin
layouts, and all three numeral glyph systems.
Every sample is generated label-first: a structured invoice record is sampled, then
rendered to pixels via a headless browser, and bounding boxes are read back from the same
DOM — so labels and boxes are… See the full description on the dataset page: https://huggingface.co/datasets/KhalfounMehdi/arabic-latin-invoices-synthetic.openai_mmlu_arabic
Dataset Card
arabic-audio-collection-mohamed-khairy
Mohamed Khairy Arabic Speech Dataset
Dataset Summary
The Mohamed Khairy Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 430 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-mohamed-khairy.es-futures-1mkhakas-russian-dict
Khakas-Russian Dictionary (Dataset)
Dataset Developer: Vasily AdeshkinContact for inquiries: adeshkin.vi@phystech.edu
📌 Important Notice & Citation
When using this dataset, please be sure to cite this repository and the original dictionary: https://khakas.altaica.ru/dictionary/.
Please note: Some optical character recognition (OCR) errors may still be present in the data.
🛠 Contribution & Authorship
I am not the author of the original dictionary, but I have… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-dict.amc-ruler-qwen35-32k
AMC RULER 32k
This dataset contains frozen inputs for the RULER benchmark.
Generation metadata
Benchmark: RULER
Sequence length: 32,768 tokens
Tokenizer: Qwen/Qwen3.5-9B
Tokenizer revision: c202236
lm-eval version: 0.4.12
Task configurations: 13
Samples per configuration: 500
Deterministic generation: Yes. Each configuration resets Python, NumPy, and task random state to seed 42.
Task configurations
niah_single_1
niah_single_2
niah_single_3… See the full description on the dataset page: https://huggingface.co/datasets/khashazad/amc-ruler-qwen35-32k.amc-ruler-qwen35-16k
AMC RULER 16k
This dataset contains frozen inputs for the RULER benchmark.
Generation metadata
Benchmark: RULER
Sequence length: 16,384 tokens
Tokenizer: Qwen/Qwen3.5-9B
Tokenizer revision: c202236
lm-eval version: 0.4.12
Task configurations: 13
Samples per configuration: 500
Deterministic generation: Yes. Each configuration resets Python, NumPy, and task random state to seed 42.
Task configurations
niah_single_1
niah_single_2
niah_single_3… See the full description on the dataset page: https://huggingface.co/datasets/khashazad/amc-ruler-qwen35-16k.pp4avPP4AV is the first public dataset with faces and license plates annotated with driving scenarios.
P4AV provides 3,447 annotated driving images for both faces and license plates.
For normal camera data, dataset sampled images from the existing videos in which cameras were mounted in moving vehicles, running around the European cities.
The images in PP4AV were sampled from 6 European cities at various times of day, including nighttime.
This dataset use the fisheye images from the WoodScape dataset to select 244 images from the front, rear, left, and right cameras for fisheye camera data.
PP4AV dataset can be used as a benchmark suite (evaluating dataset) for data anonymization models in autonomous driving.cosmopedia-openstax-khanacademy-150k-sharegpt
Dataset Card for "cosmopedia-openstax-khanacademy-150k-sharegpt"
More Information needed
