datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aliafzal9323_world-bank-development-indicators-1960-2024
World Bank Development Indicators 1960-2024
Key economic, health, education, and infrastructure indicators for every country
Dataset Info
Source: Kaggle
Original Size: 0.63 MB
Kaggle Downloads: 85
Files: 1
Files
World_Bank_Development_Indicators.csv
Mirrored from Kaggle
crowdsourced-calculator-demoaliafzal9323_soxx-ishares-semiconductor-etf-daily-2001-2026
SOXX iShares Semiconductor ETF Daily (2001-2026)
Daily OHLCV price data for iShares Semiconductor ETF (SOXX) spanning 24+ years
Dataset Info
Source: Kaggle
Original Size: 0.11 MB
Kaggle Downloads: 6
Files: 1
Files
SOXX_Daily_Stock_Data.csv
Mirrored from Kaggle
robocasa_cue_stack_aliasingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda_omron",
"total_episodes": 40,
"total_frames": 48319,
"total_tasks": 1,
"total_videos": 80,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Leejungwook/robocasa_cue_stack_aliasing.stage_aliasingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_piper_follower",
"total_episodes": 20,
"total_frames": 21170,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Leejungwook/stage_aliasing.aliasit-pii-dataset-v4
aliasit-pii-dataset-v4
Italian PII dataset in canonical text + character span form, 44 entity types that identify a person, derived from 26 pinned sources. Template families never cross splits, and the build stops if they do.
What this is
An Italian-only PII dataset in text + character span form. Spans are character offsets into the raw text, end is exclusive, and they are independent of any tokenizer: converting them to token labels is the consumer's job, so the… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v4.aliasit-pii-dataset-v4-small
aliasit-pii-dataset-v4-small
One fifth of the training split of aliasit-pii-dataset-v4, sampled uniformly on ids so the category distribution survives, with validation and test kept whole. For training runs that fail fast and still evaluate on the real thing.
What this is
An Italian-only PII dataset in text + character span form. Spans are character offsets into the raw text, end is exclusive, and they are independent of any tokenizer: converting them to token… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v4-small.ALIA-es-discriminative-hate-speech
Dataset Introduction
The ALIA Spanish Discriminative Hate Speech Corpus is a large-scale Spanish dataset for hate-speech detection built from curated social-media comments and automatically annotated using a multi-expert LLM pipeline with Fusion of Experts (FoE)[1].
The release contains:
228,708 instances
Spanish comments from YouTube and TikTok
Per-expert predictions and explanations from three LLM experts
Final fused outputs (foe_class, foe_score) for discriminative… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-discriminative-hate-speech.ALIA_syntethic_MT_V2
ALIA Synthetic MT
ALIA Synthetic MT V2 is a parallel corpus derived from Berria news articles and legal/administrative domains BOPV and Parlamento, comprising content published in 2025 as well as archived material from 2023.
The dataset provides synthetic translations into English and Spanish, generated using two distinct Large Language Models: Qwen3.5-27B and LatxaQ.
Model Details
This dataset utilizes the following models for translation generation:… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/ALIA_syntethic_MT_V2.TR-DataAnalystBench
TR-DataAnalystBench
A Turkish-language benchmark for evaluating whether language models can perform
data-analyst style reasoning over tables and charts: reading a value,
finding the maximum/minimum, comparing two years, computing an average or a
(signed) percentage change, ranking, summarizing a trend, and — importantly —
abstaining when the data does not contain the answer.
Gold answers are computed and verified with Python (not produced by a language
model), so the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/alialp207/TR-DataAnalystBench.scikit-learn-issuesaliasit-pii-dataset-v5
aliasit-pii-dataset-v5
Italian PII dataset in canonical text + character span form, the 42 entity types of v4, with the form variety and the document length that a gold set of real Italian documents showed were missing: perturbed surface forms, composed multi-section documents with anaphoric surname references, and numeric material deliberately labelled O.
Read this first: the text is rewritten
Every dataset in this family before v5 kept the source text byte for… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v5.aliasit-pii-dataset-v2
aliasit-pii-dataset-v2
Dataset PII multilingue in formato canonico testo + character span, 82 categorie, 6 lingue, derivato da quattro sorgenti pinnate a commit.
documenti
170.448
entita'
3.127.084
categorie
82 (48 marcate critiche)
lingue
de, en, es, fr, it, pt
tassonomia
v1.1.0 (taxonomy.yaml nel repo)
versione
2.0.0
costruito il
2026-08-26T17:19:46+00:00
Formato
La rappresentazione canonica e' testo + character span, non BIO: gli… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v2.clickhouse-server-imagenahj-al-balaghaنهج البلاغة وهو مجموع ما اختاره الشريف الرضي من كلام سيدنا أمير المؤمنين علي بن أبي طالب عليه السلام
شرح الأستاذ الإمام الشيخ محمد عبدة مفتي الديار المصرية سابقا الجزء الأول
الناشر دار المعرفة للطباعة والنشر بيروت لبنان
Overview
The Nahj al-Balagha dataset is a structured CSV file containing textual data from the renowned collection of sermons, letters, and maxims attributed to Imam Ali ibn Abi Talib (AS). Compiled by Sharif Razi in the 10th century, this dataset provides a… See the full description on the dataset page: https://huggingface.co/datasets/aliahabeeb/nahj-al-balagha.aliasit-pii-dataset-v3
aliasit-pii-dataset-v3
Multilingual PII span-annotation dataset, 79 entity types, 30 pinned sources, Italian-first. Text plus character spans, single-label BIO.
Documents
182,862
Entities
3,172,535
Entity types
79 (46 marked critical)
Languages
ar, de, el, en, es, fr, it, nl, pt, sl, tr
Taxonomy
v3.0.0 — taxonomy.yaml ships in this repo
Revision
6.0.0
Representation
text + character spans, end exclusive, non-overlapping
Distinct sources
27, each pinned… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v3.aliasit-pii-dataset-v3-small
aliasit-pii-dataset-v3-small
Half the training split of aliasit-pii-dataset-v3, with validation and test kept whole. For cheap training runs that still evaluate on the real thing.
This is a subsample of mapo80/aliasit-pii-dataset-v3. The training split is halved by uniform sampling on document id; validation and test are the full splits of the parent dataset, unchanged. It exists to make a training run cheap without making the evaluation a different question. Every category… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v3-small.rbm-metaworld-metaworld-train-lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 100,
"total_frames": 3198,
"total_tasks": 20,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aliangdw/rbm-metaworld-metaworld-train-lerobot.rbm-mwThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 100,
"total_frames": 3198,
"total_tasks": 20,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aliangdw/rbm-mw.aliafzal9323_psi-invesco-semiconductors-etf-daily-2005-2026
PSI Invesco Semiconductors ETF Daily (2005-2026)
Daily OHLCV price data for Invesco Dynamic Semiconductors ETF (PSI) spanning 20+
Dataset Info
Source: Kaggle
Original Size: 0.08 MB
Kaggle Downloads: 13
Files: 1
Files
PSI_Daily_Stock_Data.csv
Mirrored from Kaggle
metaworld-eval-lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unknown",
"total_episodes": 151,
"total_frames": 4829,
"total_tasks": 17,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 1,
"splits": {
"train": "0:151"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aliangdw/metaworld-eval-lerobot.aliasit-pii-dataset-v2-small
aliasit-pii-dataset-v2-small
Sottoinsieme stratificato di aliasit-pii-dataset-v2 per le prove di training: stesse categorie e stesso formato, dimensioni da smoke test.
documenti
8.620
entita'
124.644
categorie
82 (48 marcate critiche)
lingue
de, en, es, fr, it, pt
tassonomia
v1.1.0 (taxonomy.yaml nel repo)
versione
1.0.0
costruito il
2026-08-26T17:36:35+00:00
Formato
La rappresentazione canonica e' testo + character span, non BIO: gli… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v2-small.notification-timing-dataset
Notification Timing Dataset
100K synthetic samples for training notification bad-timing prediction models.
21 features covering time context, battery state, user activity, and notification history.
Based on feature engineering from C-3PO (Cheetah Mobile, 600M MAU).
See model: alianassmaaa/notification-bad-timing-detector
usc-xarm-policy-ranking-v30-sampleThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 2,
"total_frames": 32,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aliangdw/usc-xarm-policy-ranking-v30-sample.usul-alkafi
Usul Al-Kafi Dataset
Overview
This dataset contains the full text of الأصول من الكافي, a foundational Islamic book authored by ثقة الإسلام أبي جعفر محمد بن يعقوب بن إسحاق الكليني الرازي. Each row in the dataset represents a single page from the book, preserving the original Arabic text structure. The book is divided into eight parts and includes beneficial annotations derived from various commentaries.
Dataset Details
Version: 1.0
Source: Digitized from Dar… See the full description on the dataset page: https://huggingface.co/datasets/aliahabeeb/usul-alkafi.rbm-1m-oodThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 782,
"total_frames": 25024,
"total_tasks": 95,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:782"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aliangdw/rbm-1m-ood.github-issuesmulti-lingual-qac-alias-graph
Multi-lingual chemical QAC — Alias-Graph Retrieval benchmark
Given a chemistry concept (named in several languages), can a retriever find the documents that genuinely talk about it, across languages, without being fooled by documents about chemically similar look-alike concepts?
Configs: corpus (gold + hard-negative documents), queries (technical questions about each concept, plus the concept's multilingual name_set and the source_publication each query was generated from)… See the full description on the dataset page: https://huggingface.co/datasets/MehdiAstaraki/multi-lingual-qac-alias-graph.aliafzal9323_us-cdc-places-county-health-data-2024
US CDC PLACES County Health Data 2024
40 health measures across all US counties from the CDC PLACES project 2024
Dataset Info
Source: Kaggle
Original Size: 9.24 MB
Kaggle Downloads: 9
Files: 1
Files
PLACES__Local_Data_for_Better_Health__County_Data__2025_release.csv
Mirrored from Kaggle
pointer-aliasing-v5
