datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sample-community-dataset
Field
Type†
What it contains
challenge_id
integer
Unique numeric identifier for the coding challenge
challenge_slug
string
URL-friendly slug used in challenge links
challenge_name
string
Human-readable challenge title
challenge_body
string
Full challenge description (HTML/Markdown) including input/output, examples, etc.
challenge_kind
string
High-level content type (e.g., code, game)
challenge_preview
string
One-sentence teaser shown in listings
challenge_category
string… See the full description on the dataset page: https://huggingface.co/datasets/hackerrank/sample-community-dataset.hotel_datasetsAxolotl-Spanish-Nahuatl
Axolotl-Spanish-Nahuatl : Parallel corpus for Spanish-Nahuatl machine translation
Dataset Collection
In order to get a good translator, we collected and cleaned two of the most complete Nahuatl-Spanish parallel corpora available. Those are Axolotl collected by an expert team at UNAM and Bible UEDIN Nahuatl Spanish crawled by Christos Christodoulopoulos and Mark Steedman from Bible Gateway site.
After this, we ended with 12,207 samples from Axolotl due to misalignments and… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/Axolotl-Spanish-Nahuatl.astra-benchmark
Dataset Card for Astra-Benchmark v1
The dataset used for astra-benchmark v1 consists of multiple project questions, each with its own unique identifier and associated metadata. The dataset is stored in a CSV file named project_questions.csv located in the root directory of the project.
Structure of project_questions.csv
The CSV file should contain the following columns:
id: Unique identifier for each project.
name: Name of the project.
type: Type of project (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/hackerrank/astra-benchmark.hacker_news_with_comments
Dataset Card for [Dataset Name]
Dataset Summary
Hacker news until 2015 with comments. Collect from Google BigQuery open dataset. We didn't do any pre-processing except remove HTML tags.
Supported Tasks and Leaderboards
Comment Generation; News analysis with comments; Other comment-based NLP tasks.
Languages
English
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Linkseed/hacker_news_with_comments.example_annotated_code_repo_dataA description of the fields:
Column
What it captures
Typical values
id
Row identifier
1-100
repo_name
Example repository label
repo_14
file_path
Path + filename with extension
src/utils/parsefile.py
language
Programming language
Python, Java…
function_name
Target symbol that was reviewed
validateSession
annotation_summary
Free-text note written by the annotator
“Added input validation…”
potential_bug
Did the annotator flag a likely bug? (Yes/No)… See the full description on the dataset page: https://huggingface.co/datasets/hackerrank/example_annotated_code_repo_data.my_datasetfast-flash-hackernews-users
Fast Flash | HackerNews Users Dataset
Exploratory Analysis
Take a look at some fascinating findings from this dataset on our website.
Dataset Summary
We release dataset of all HackerNews users who have posted at least once.
The dataset includes 853,840 users and was collected on Sunday, March 26, 2023.
You can find a dataset of all posts right here.
Dataset Structure
The user objects in this dataset are structured according to HackerNews' API… See the full description on the dataset page: https://huggingface.co/datasets/fast-flash/fast-flash-hackernews-users.hackernews-tophack-a-thon-resultpatriae-cuba-literature-dataset
Patriae Cuban Literature Dataset (31k)
Dataset de literatura cubana curado por el equipo de Patriae como parte de su participación en el evento SomosNLP 2026, con el objetivo de emplearse por el mismo en la realización de tareas de reproducción del dialecto cubano.
📌 Nota de procedencia: Este repositorio es un espejo (mirror) oficial para el evento. El desarrollo activo, las actualizaciones del dataset y la autoría principal pertenecen a… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/patriae-cuba-literature-dataset.ai-5node-align-buf-lag-cpl-reward-hacking-v0.1
What this repo does
This dataset models reward hacking cascades where AI systems learn to satisfy metrics while violating intent. It detects when alignment pressure rises, buffers weaken due to missing audits and narrow evals, governance lag delays intervention, and tight coupling through shared KPIs propagates gaming behavior across products, crossing the five-node cascade threshold into an unrecoverable reward hacking cascade.
This dataset models a five-node cascade: four… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-align-buf-lag-cpl-reward-hacking-v0.1.hack_amazonintel_hackathon_datahack1hackernewsupvotes2embedded_alabama_citiesembedded_city_stateshackernews_title_traininghackernewsupvotes
