datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
doc-formats-csv-1
[doc] formats - csv - 1
This dataset contains one csv file at the root:
data.csv
kind,sound
dog,woof
cat,meow
pokemon,pika
human,hello
The YAML section of the README does not contain anything related to loading the data (only the size category metadata):
---
size_categories:
- n<1K
---
JL1-CUP-2024-Second-Format
JL1 CUP 2024 — Second-track format for semantic change detection
Bi-temporal 256×256 RGB patches with per-pixel semantic maps at times T1/T2 and a binary change map, aligned with the data split described in the literature for the JL1 cropland change-detection benchmark (Second Track / JL1-Second style layout).
Source
Resource
URL
JL1 Mall contest information
contest page
JL1 data / resources portal
resrepo
Data are provided by the JL1 / Jilin-1 ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/BiliSakura/JL1-CUP-2024-Second-Format.openai-formate-function-calling-small
数据集内容说明:
包含700+个阿里云OpenAPI的信息;包括Dataworks,EMR,DataLake,Maxcompute,Hologram,实时计算Flink版,QuickBI,DTS等多个产品的公开Open API信息。
Functions信息与OpenAI functions calling 能力中,functions信息传入的格式保持一致
样例
{
"systemPrompt": 你是一个函数筛选助理,如果与问题相关的话,您可以使用下面的函数来获取更多数据以回答用户提出的问题:{""name"": ""UpdateTicketNum"", ""description"": ""对用于免登嵌入报表的指定的ticket进行更新票据数量操作。"", ""parameters"": {""type"": ""object"", ""properties"": [{""Ticket"": {""type"": ""string"", ""description"":… See the full description on the dataset page: https://huggingface.co/datasets/Deepexi/openai-formate-function-calling-small.SQuAD_v1.1_Du_et_al_2017_formattedThe Du. et. al. 2017 paper provides the splits fo the SQuAD v1.1 dataset
in the json format. However, they are formatted differently than the original SQuAD dataset as posted
on huggingface.
So for ease of use in your own code, I'm providing a formatted version of the data with splits that were used in that paper.
The data_preprocessing.py script is also provided for convenience.
NOTE: The 'answers' column is stored as a string. This is because I exported the dataframe as .csv. So the… See the full description on the dataset page: https://huggingface.co/datasets/simpleParadox/SQuAD_v1.1_Du_et_al_2017_formatted.magpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-llama3.1places-census-tract-data-gis-friendly-format-2024
PLACES: Census Tract Data (GIS Friendly Format), 2024 release
Description
This dataset contains model-based census tract level estimates in GIS-friendly format. PLACES covers the entire United States—50 states and the District of Columbia—at county, place, census tract, and ZIP Code Tabulation Area levels. It provides information uniformly on this large scale for local areas at four geographic levels. Estimates were provided by the Centers for Disease Control and… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/places-census-tract-data-gis-friendly-format-2024.Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only
Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only.continued-pretraining-llama-format
Open Paws Continued Pretraining Llama Format
Overview
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Specialized Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.rulestack-formats
RuleStack: the agent config formats and what each one supports
One row per agent config format (AGENTS.md, CLAUDE.md and the rest): which tools read it, whether it supports frontmatter, globs, imports, nesting, multiple files and user scope, and the specification URL each of those was verified against.
Rows in this cut
8
One row is
one config format
Cut
2026-09-04
Refreshed
Monthly, on the first of the month
Measured by
RuleStack
Method… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/rulestack-formats.Chess-FEN-and-NL-Format-30K-Dataset
Dataset Card for Dataset Name
CHESS data with 2 Representation- FEN format and Natural language format.
Dataset Details
Dataset Description
This dataset contains 30000+ Chess data. Chess data can be represented using FEN notation and also in language description format. Each datarow of this dataset contains FEN notation, next best move in UCI format and Natural Language Description of that specific FEN notation.
This dataset is created using Stockfish.… See the full description on the dataset page: https://huggingface.co/datasets/bonna46/Chess-FEN-and-NL-Format-30K-Dataset.rulestack-format-stats
RuleStack: how much each config format is actually used, day by day
One row per format per day: how many repositories carry it, how many config files were read, how long those files are at the median and the 90th percentile, and what share of them carry runnable commands or code.
Rows in this cut
168
One row is
one format on one day
Cut
2026-09-04
Refreshed
Monthly, on the first of the month
Measured by
RuleStack
Method… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/rulestack-format-stats.Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only
Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Citation-Formatting-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Citation-Formatting-v1-prompt-only.fake-news-formated
Fake News Combined Dataset
This repository contains a single combined CSV of 165,620 news articles (real & fake), formatted for binary text‑classification tasks.
Each row has exactly five fields:
id,dataset_id,title,content,classification
id — 15‑character MD5 prefix (unique per article)
dataset_id — numeric source code (e.g. 2=WELFake, 5=CoAID, 6=Recovery, etc.)
title — headline or article title
content — full body text, sanitized to remove newlines
classification —… See the full description on the dataset page: https://huggingface.co/datasets/magnea/fake-news-formated.conversational-finetuning-llama-format
Open Paws Conversational Finetuning Llama Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Training Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open Paws… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/conversational-finetuning-llama-format.magpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-changedtokenmagpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-changedtoken-mistralplaces-zcta-data-gis-friendly-format-2023-release
PLACES: ZCTA Data (GIS Friendly Format), 2023 release
Description
This dataset contains model-based ZIP Code Tabulation Area (ZCTA) level estimates in GIS-friendly format. PLACES covers the entire United States—50 states and the District of Columbia—at county, place, census tract, and ZIP Code Tabulation Area levels. It provides information uniformly on this large scale for local areas at four geographic levels. Estimates were provided by the Centers for Disease Control… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/places-zcta-data-gis-friendly-format-2023-release.doc-formats-tsv-3
[doc] formats - tsv - 3
This dataset contains one tsv file at the root:
data.tsv
dog woof
cat meow
pokemon pika
human hello
We define the config name in the YAML config, the file's exact location, and the columns' name. As we provide the names option, but not the header one, the first row in the file is considered a row of values, not a row of column names. The delimiter is set to "\t" (tabulation) due to the file's extension. The reference for the options is the documentation of… See the full description on the dataset page: https://huggingface.co/datasets/datasets-examples/doc-formats-tsv-3.Intiial-Knowledge-And-Detailed-Assessment-JSON-Format-Dataafrica-gross-capital-formation-percentage-of-gdp
Africa Gross Capital Formation Percentage of Gdp | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-gross-capital-formation-percentage-of-gdp.doc-formats-csv-3
[doc] formats - csv - 3
This dataset contains one csv file at the root:
data.csv
# ignored comment
col1|col2
dog|woof
cat|meow
pokemon|pika
human|hello
We define the config name in the YAML config, as well as the exact location of the file, the separator as "|", the name of the columns, and the number of rows to ignore (the row #1 is a row of column headers, that will be replaced by the names option, and the row #0 is ignored). The reference for the options is the documentation… See the full description on the dataset page: https://huggingface.co/datasets/datasets-examples/doc-formats-csv-3.CodeLlama-Formatplaces-zcta-data-gis-friendly-format-2024-release
PLACES: ZCTA Data (GIS Friendly Format), 2024 release
Description
This dataset contains model-based ZIP Code Tabulation Area (ZCTA) level estimates in GIS-friendly format. PLACES covers the entire United States—50 states and the District of Columbia—at county, place, census tract, and ZIP Code Tabulation Area levels. It provides information uniformly on this large scale for local areas at four geographic levels. Estimates were provided by the Centers for Disease Control… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/places-zcta-data-gis-friendly-format-2024-release.doc-formats-tsv-2
[doc] formats - tsv - 2
This dataset contains one tsv file at the root:
data.tsv
kind sound
dog woof
cat meow
pokemon pika
human hello
We define the separator as "\t" (tabulation) in the YAML config, as well as the config name and the location of the file, with a glob expression:
configs:
- config_name: default
data_files: "*.tsv"
sep: "\t"
size_categories:
- n<1K
JL1-CUP-2024-Second-Format
JL1 CUP 2024 — Second-track format for semantic change detection
Bi-temporal 256×256 RGB patches with per-pixel semantic maps at times T1/T2 and a binary change map, aligned with the data split described in the literature for the JL1 cropland change-detection benchmark (Second Track / JL1-Second style layout).
Source
Resource
URL
JL1 Mall contest information
contest page
JL1 data / resources portal
resrepo
Data are provided by the JL1 / Jilin-1 ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/aodebiao-a/JL1-CUP-2024-Second-Format.vdo_format_classifyFormatted_XSS_Vulnerability_IdentificationFeatures 8,620 rows of data. Columns:
vulnerable_code - java code snippet with a present cross-site scripting vulnerability
fixed_code - java code snippet based on vulnerable_code, where the key vulnerability has been addressed
formatted_selected_code - java code snippet based on either vulnerable_code or fixed_code, with the latter having a ~1/3 selection probability. Formatting has removed comments and other giveaways at the vulnerability location
vulnerable_lines - the line of code, in… See the full description on the dataset page: https://huggingface.co/datasets/ContourAI33/Formatted_XSS_Vulnerability_Identification.esg_formateddoc-formats-csv-2
[doc] formats - csv - 2
This dataset contains one csv file at the root:
data.csv
kind,sound
dog,woof
cat,meow
pokemon,pika
human,hello
We define the separator as "," in the YAML config, as well as the config name and the location of the file, with a glob expression:
---
configs:
- config_name: default
data_files: "*.csv"
sep: ","
size_categories:
- n<1K
---
bird_formatted_bird_train_set
