datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ISEAR-dataset-completeILDC_35k_COMPLETEvn-provinces-housing-floor-area-completed
Vietnam provinces housing floor area completed
Floor area of housing completed during the year (thousand square metres). Coverage 2010-2023. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (882 rows)
data/provinces.csv
data/provinces.dta… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-housing-floor-area-completed.completeRXN
CompleteRxn Benchmark
206,423 reactions for the task of reaction completion: given an atom-unbalanced
USPTO reaction SMILES, predict the missing molecules to produce a balanced equation.
Ground-truth targets come from FlowER
mechanistic steps mapped to USPTO records. Three split types (random, group OOD, extreme
OOD) with 5 repetitions each test generalization across structural novelty levels.
Code: r/CompleteRxn-Benchmarking-2280
Files
benchmark_final.csv ← 206… See the full description on the dataset page: https://huggingface.co/datasets/completeRXN-benchmark-26/completeRXN.aayushmishra1512_faang-complete-stock-data
FAANG- Complete Stock Data
It contains data of Stock of the FAANG companies from when they began trading.
Dataset Info
Source: Kaggle
Original Size: 0.51 MB
Kaggle Downloads: 8,685
Files: 5
Files
Amazon.csv
Apple.csv
Facebook.csv
Google.csv
Netflix.csv
Mirrored from Kaggle
completeRXN
CompleteRxn Benchmark
206,423 reactions for the task of reaction completion: given an atom-unbalanced
USPTO reaction SMILES, predict the missing molecules to produce a balanced equation.
Ground-truth targets come from FlowER
mechanistic steps mapped to USPTO records. Three split types (random, group OOD, extreme
OOD) with 5 repetitions each test generalization across structural novelty levels.
Code: r/CompleteRxn-Benchmarking-2280
Files
benchmark_final.csv ←… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/completeRXN.complete-effect-9a435b
complete-effect-9a435b
Synthetic sensors test data: 37 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Velvet-Andrew/complete-effect-9a435b.complete-goal-35f45b
complete-goal-35f45b
Synthetic products test data: 51 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/yuiyamada/complete-goal-35f45b.wonbias-complete-dataset
WoNBias: A Bengali Dataset for Gender Bias Detection
Dataset Details
Overview
A manually annotated corpus of Bengali designed to identify gender-based biases, stereotypes, and harmful language against women. Supports research in ethical NLP, content moderation, and computational social science.
Dataset Description
Basic Info
Purpose: Detect gender bias and harmful language against women in Bengali text
Language: Bengali (Bangla)
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/gender-bias-bengali/wonbias-complete-dataset.msa-darja-pairs-completeNASA_Completeanime_rating_completeNYT_complete_datasetThe Dataset contains 1500 New York Times article title and their actual articles, along with responses of 3 others (LLama3.1, Gemma2-9b, Mixtral-8-7B) models with temperature 0.0,
0.2, 0.4, 0.6, 0.8, 1.0
For any further query, drop a mail at dibakarghosh1868@gmail.com
aayushmishra1512_fifa-2021-complete-player-data
FIFA 2021 Complete Player Dataset
This Data set contains data of the players in FIFA-2021
Dataset Info
Source: Kaggle
Original Size: 0.40 MB
Kaggle Downloads: 5,443
Files: 1
Files
FIFA-21 Complete.csv
Mirrored from Kaggle
prompt_response_1K_PIIS_completecomplete_nepal_share_market_datacomplete_foodie_datasetclinical-quad-negative-signal-rate-reporting-lag-bias-summary-completeness-omission-event-v0.1What this repo does
This dataset models negative evidence suppression in clinical trial narratives. It predicts when the interaction between negative signal rate, reporting lag, author bias, and low summary completeness indicates that adverse or null findings are likely omitted from the narrative.
Core quad
negative_signal_rate_index
reporting_lag_days
author_bias_index
summary_completeness_index
Prediction target
label_omission_event
Row structure
Each row represents a results-to-summary… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-negative-signal-rate-reporting-lag-bias-summary-completeness-omission-event-v0.1.batch1_2_complete_cleanedlegal-hearing-bundle-index-completeness-coherence-v0.1What this dataset does
You receive
hearing issues
core docs list
bundle index
pagination logic
version and duplicates
missing flags
You decide
coherent
or
incoherent
Daily use
bundle QC
missing document detection
wrong version detection
overstuffed bundle detection
pagination index fix flag
card_completenessbatch1cleaned_completejerrybase-songs-complete
jerrybase-songs-complete
Description
Complete song data from JerryBase
License
cc-by-4.0
Usage
from datasets import load_dataset
dataset = load_dataset("jaysooner/jerrybase-songs-complete")
Source
This dataset is part of the GratefulGPT project - an AI system focused on Grateful Dead knowledge and culture.
complete_credit_risk_analysis_platform-logs_auditfyp-complete-datasetComplete_Coca-Cola_Stocks_datasets
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Atif-067/Complete_Coca-Cola_Stocks_datasets.EUS_Complete_Classification
