datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openvino-arc140v-lunarlake
OpenVINO local-inference on an Intel Arc 140V (Lunar Lake) iGPU
Reference performance data for running local models on a single Intel Core Ultra 7 258V
(Lunar Lake) laptop with the integrated Intel Arc 140V (Xe2) GPU, via OpenVINO. All
inference runs on the iGPU; the NPU stays idle throughout, confirmed by the telemetry here.
This is reference characterization shared by a non-expert contributor — careful measurements
on one machine, offered so others can compare and correct, not… See the full description on the dataset page: https://huggingface.co/datasets/blairducrayoppat/openvino-arc140v-lunarlake.ARCHIVE-TEXT-URLS
Internet Archive English Text URLs Dataset
Dataset Description
This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials.
Dataset Summary
Total Rows: 11,151,637
Language: English
Source: Internet Archive
Format: CSV with metadata and direct text file URLs
Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.Munyarwanda-AI-AutoTrain
Munyarwanda AI - AutoTrain dataset
AutoTrain-ready version of arcange9/Munyarwanda-AI-Dataset v0.2.
Every example is pre-formatted in Qwen chat template as a single text column (5,542 train / 22 validation rows).
Intended recipe (Hugging Face AutoTrain, base model Qwen/Qwen3-0.6B):
LLM task, causal LM
text column: text
LoRA/PEFT + int4 quantization to fit free-tier GPUs
atcoder_arc_contests
Notification
Atcoder is selling this data now. If you are interested in accessing it please contact them.
Dataset Summary
This dataset aims to facilitate the creation of sophisticated, multi-turn dialogue datasets focused on coding
that could be used for training reasoning Large Language Models (LLMs), particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation techniques.
It also serves as a robust foundation for problem-solving in Large Language… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_arc_contests.ASRS-ChatGPT
Dataset Summary
The dataset contains a total of 9984 incident records and 9 columns. Some of the columns contain ground truth values whereas others contain information generated by ChatGPT based on the incident Narratives.
The creation of this dataset is aimed at providing researchers with columns generated by using ChatGPT API which is not freely available.
Dataset Structure
The column names present in the dataset and their descriptions are provided below:
Column… See the full description on the dataset page: https://huggingface.co/datasets/archanatikayatray/ASRS-ChatGPT.SNET_Archive
