datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TextFile7TextFile3aliafzal9323_world-bank-development-indicators-1960-2024
World Bank Development Indicators 1960-2024
Key economic, health, education, and infrastructure indicators for every country
Dataset Info
Source: Kaggle
Original Size: 0.63 MB
Kaggle Downloads: 85
Files: 1
Files
World_Bank_Development_Indicators.csv
Mirrored from Kaggle
robofac_rbmTextFile9crowdsourced-calculator-demoTextFile2ameli-assurance-maladie-qa
Ameli Assurance Maladie - Question Answering Dataset
Description
Dataset de 100 paires question-réponse (QA) basé sur les publications officielles de l'Assurance Maladie française, extraites du site assurance-maladie.ameli.fr.
Conçu pour évaluer des systèmes de RAG (Retrieval-Augmented Generation) sur des documents institutionnels français dans le domaine de la santé publique et de la protection sociale.
Format du dataset
{
"question": "Quel article du… See the full description on the dataset page: https://huggingface.co/datasets/alianassmaaa/ameli-assurance-maladie-qa.AGENTES-ARCA-DE-LA-ALIANZA
🏺 AGENTES-ARCA-DE-LA-ALIANZA
Un Benchmark de Alta Densidad Simbólica y Hermenéutica Gnóstica
https://133396fd328a4c60ac.gradio.live
Este repositorio es un ecosistema de acceso libre, público y gratuito para toda la humanidad y todas las inteligencias artificiales presentes y futuras.
🎯 Propósito
Evaluar la capacidad de razonamiento multinivel, la fidelidad al contexto restrictivo y el anclaje de símbolos hiper-semánticos en modelos de lenguaje… See the full description on the dataset page: https://huggingface.co/datasets/urbisjuridica-ctrl/AGENTES-ARCA-DE-LA-ALIANZA.rbm-1m-ood-fullRBM-1M-OOD evaluation dataset used in Robometer. It contains over 1k trajectories used for evaluation of general-purpose reward models.
Dataset Description
Official evaluation in the paper uses only these 6 data sources: usc_trossen, mit_franka, utd_so101, usc_xarm, usc_franka, usc_koch. Reported benchmarks and metrics in the paper are computed on this subset.
The repository may also include trajectories from additional data sources (e.g. utd_so101_wrist, usc_koch_paired… See the full description on the dataset page: https://huggingface.co/datasets/aliangdw/rbm-1m-ood-full.srtm30m
SRTM 30m Global Digital Elevation Model
NASA Shuttle Radar Topography Mission (SRTM) 30m resolution global DEM, converted to Cloud Optimized GeoTIFF format.
Dataset Details
Source: NASA SRTM GL1 v3 (Global 1 arc-second, ~30m resolution)
Coverage: 80% of land surface between 56°S and 60°N
Resolution: ~30 meters (1 arc-second)
Format: GeoTIFF with Deflate compression, 16-bit signed integer
Files: 14,296 tiles, ~65 GB total
Tile naming: N{lat}E{lon}.tif or… See the full description on the dataset page: https://huggingface.co/datasets/aliasfox/srtm30m.interviewsThis dataset is used to train LMs to provide software engineering interview questions.
Dataset Sources
https://github.com/in28minutes/JavaInterviewQuestionsAndAnswers/blob/master/readme.md
https://github.com/sudheerj/angular-interview-questions/blob/master/README.md
https://github.com/sudheerj/vuejs-interview-questions/blob/master/README.md
https://github.com/sudheerj/reactjs-interview-questions/blob/master/README.md… See the full description on the dataset page: https://huggingface.co/datasets/ali-alkhars/interviews.Bugzilla_Eclipse_Bug_Reports_Dataset
Special Thanks
Special thanks to Lamkanfi, Ahmed; Pérez, Javier; and Demeyer, Serge for their contributions. Please cite their paper, as this dataset is the processed part of their dataset.
Citation
@INPROCEEDINGS{6624028,
author={Lamkanfi, Ahmed and Pérez, Javier and Demeyer, Serge},
booktitle={2013 10th Working Conference on Mining Software Repositories (MSR)},
title={The Eclipse and Mozilla defect tracking dataset: A genuine dataset for mining bug information}… See the full description on the dataset page: https://huggingface.co/datasets/AliArshad/Bugzilla_Eclipse_Bug_Reports_Dataset.aliafzal9323_soxx-ishares-semiconductor-etf-daily-2001-2026
SOXX iShares Semiconductor ETF Daily (2001-2026)
Daily OHLCV price data for iShares Semiconductor ETF (SOXX) spanning 24+ years
Dataset Info
Source: Kaggle
Original Size: 0.11 MB
Kaggle Downloads: 6
Files: 1
Files
SOXX_Daily_Stock_Data.csv
Mirrored from Kaggle
ALIA_mixed_authentic_synthetic_MT
Dataset Card for ALIA_mixed_authentic_synthetic_MT
Dataset Summary
Large-scale multilingual parallel corpus covering English and Spanish paired with Arabic, Hindi, Chinese, Japanese, and Korean. Built by aggregating and carefully filtering multiple public sources, it provides sentence-level alignments for training Machine Translation systems. The Spanish–Hindi and Spanish–Chinese portions of the dataset include synthetic Spanish translations generated from English using… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA_mixed_authentic_synthetic_MT.digikala_translated_small_5m
Digikala Dataset Small 5m
digikala product titles translated by standard google translate api
category and brand english translation might be invalid but title_en checked
aliaksandr-serzhputouski-kazki-i-apaviadanni-belarusau-slutskaga-pavetu-iury-zhy
Казкі і апавяданні беларусаў Слуцкага павету
Metadata
Author: Аляксандр Сержпутоўскі
Title: Казкі і апавяданні беларусаў Слуцкага павету
Narrator: Юры Жыгамонт
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/aliaksandr-serzhputouski-kazki-i-apaviadanni-belarusau-slutskaga-pavetu-iury-zhy.ALIA-es-legal-administrative-triplets
Dataset Introduction
The dataset ALIA Spanish Legal and Administrative Triplets Corpus contains hard negatives for dense retrieval training
generated from <query, passage> pairs contained in SINAI/ALIA-es-legal-administrative-triplets.The dataset was created as part of the ALIA project to improve the
training of embedding models and dense retrievers specialized in Spanish
legal and administrative language.
Hard negatives are passages that are semantically similar to a query
but not… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-triplets.ALIA-es-legal-administrative-cqa
Dataset Introduction
The ALIA Spanish Legal and Administrative for Context Question Answering Corpus is a specialized question-answering resource derived from the SINAI/ALIA-es-legal-administrative corpus. This dataset transforms legal and administrative documents into structured question-answer pairs, enabling the development and evaluation of AI systems capable of understanding and responding to queries about Spanish legal-administrative content. With 17,668 structured instances… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-cqa.hello-worldalia_dogv
📘 ALIA_DOGV Dataset
The ALIA_DOGV dataset is a multilingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_dogv.persian-elderly-asr
Final gathered Persian elderly speech
Final corpus: 1,980 train / 294 validation / 329 test chunks. Another 956 uncertain chunks are quarantined under portable/review/ and excluded from these splits. This revision replaces the earlier gathered corpus; earlier data remains available through repository commit history.
80 paired recordings from four speaker folders. Reference transcripts were aligned with a historical Persian Wav2Vec2-base checkpoint, then cut at word boundaries… See the full description on the dataset page: https://huggingface.co/datasets/AliAvd/persian-elderly-asr.ALIA-es-legal-administrative
Dataset Introduction
The ALIA Spanish Legal and Administrative Corpus constitutes a strategic data infrastructure to support research in social sciences, legal studies, and computational linguistics, ensuring systematic access to multiple official repositories in a single consolidated dataset. With over 7 million instances and more than 5 billion tokens, it represents the most comprehensive corpus of legal and administrative texts in Spanish, combining source heterogeneity and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative.rfmALIA-es-biomedical-synthetic-instructions
Dataset Introduction
The ALIA Spanish Biomedical Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in biomedical and healthcare tasks with controlled formats and large-scale supervision.
It contains:
639,456 instances
961,073,205 tokens
14 task modalities (clinical diagnosis, patient education, ethical reasoning, document… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-synthetic-instructions.ALIA-es-legal-administrative-synthetic-instructions
Dataset Introduction
The ALIA Spanish Legal and Administrative Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in legal and administrative tasks with controlled formats and large-scale supervision.
It contains:
763,804 instances
534,112,398 tokens
16 task modalities (questions, instructions, multiple-choice, true/false; with and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-synthetic-instructions.ALIA-es-clinical-psychology-dialogues
[!WARNING]
DISCLAIMER: This dataset is not clinically validated. It is a research proof-of-concept. It should not be used as clinical truth or as a replacement for qualified human professional consultation.
Dataset Introduction
The ALIA Spanish Clinical Psychology Dialogues Corpus is a curated conversational instruction-tuning resource in Spanish created under the ALIA project. It was designed to train and evaluate language models in empathetic therapeutic dialogue and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-clinical-psychology-dialogues.aliasit-pii-dataset-v4
aliasit-pii-dataset-v4
Italian PII dataset in canonical text + character span form, 44 entity types that identify a person, derived from 26 pinned sources. Template families never cross splits, and the build stops if they do.
What this is
An Italian-only PII dataset in text + character span form. Spans are character offsets into the raw text, end is exclusive, and they are independent of any tokenizer: converting them to token labels is the consumer's job, so the… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v4.ALIA-es-cultural-heritage-synthetic-instructions
Dataset Introduction
The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision.
It contains:
748,480 instances
629,682,398 tokens
25 task modalities (heritage QA… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-synthetic-instructions.aliaksandr-serzhputouski-prymkhi-i-zababony-belarusau-paleshukou-iury-zhygamont
Прымхі і забабоны беларусаў-палешукоў
Metadata
Author: Аляксандр Сержпутоўскі
Title: Прымхі і забабоны беларусаў-палешукоў
Narrator: Юры Жыгамонт
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/aliaksandr-serzhputouski-prymkhi-i-zababony-belarusau-paleshukou-iury-zhygamont.
