datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ARCHIVE-TEXT-URLS
Internet Archive English Text URLs Dataset
Dataset Description
This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials.
Dataset Summary
Total Rows: 11,151,637
Language: English
Source: Internet Archive
Format: CSV with metadata and direct text file URLs
Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.TextCodeDepot
Dataset description:
The Python Code Chatbot dataset is a collection of Python code snippets extracted from various publicly available datasets and platforms. It is designed to facilitate training conversational AI models that can understand and generate Python code. The dataset consists of a total of 1,37,183 prompts, each representing a dialogue between a human and an AI Scientist.
Prompt Card:
Each prompt in the dataset follows a specific format known as the "Prompt… See the full description on the dataset page: https://huggingface.co/datasets/anujsahani01/TextCodeDepot.DomesticNames_AllStates_Text Domestic Names from the Federal Government's repository of official geographic names [CSV dataset]
This Dataset includes 980,065 geographic names as of September 10, 2023.
It is apparent that no currently released LLMs are pretrained on datasets with many of these geographic names (i.e., features), descriptions, and histories.
Example: feature_name: Abercrombie Gulch
GPT-3.5 responds "I'm not aware of a specific location called Abercrombie Gulch in my training data,..." when prompted… See the full description on the dataset page: https://huggingface.co/datasets/cellos/DomesticNames_AllStates_Text.ph_en_text_detoxedPhEnText Detoxed is a large-scale and multi-domain lexical data written in Philippine English and Taglish text. The news articles, religious articles and court decisions collated by the original researchers were filtered for toxicity and special characters were further preprocessed. This dataset has been configured to easily fine-tune LLaMA-based models (Alpaca, Guanaco, Vicuna, LLaMA 2, etc.) In total, this dataset contains 6.29 million rows of training data and 2.7 million rows of testing… See the full description on the dataset page: https://huggingface.co/datasets/NLPinas/ph_en_text_detoxed.synthetic_text_to_sql_th
Synthetic Text-to-SQL Thai Dataset
Thai translation of the gretelai/synthetic_text_to_sql dataset.
Dataset Description
This dataset contains Thai translations of synthetic text-to-SQL examples covering various domains and SQL patterns.
Source
Original Dataset: gretelai/synthetic_text_to_sql
Created by: Gretel.ai
Statistics
Split
Rows
Train
100,000
Test
5,851
Total
105,851
Columns
Column
Description… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/synthetic_text_to_sql_th.text-to-mongodb-queries-llm
Dataset Description
This dataset contains 10,000+ complex SQL-style analytical questions
mapped to MongoDB queries and aggregation pipelines.
Features
Multiple schemas
$group, $sum, $avg, $lookup
Nested documents
Long analytical questions
Use Cases
Fine-tuning small LLMs (Qwen, Mistral, LLaMA 3B)
Text-to-Mongo query generation
Data analytics agents
ocr_textsEncyclopedia_Sindhiana_text_corpus
Encyclopedia Sindhiana Dataset
Overview
The Encyclopedia Sindhiana Dataset is a collection of encyclopedia articles created for Natural Language Processing (NLP) research.
Each entry contains:
Category — Topic label of the article (14 unique classes)
Title — Title of the article
Content — Full text of the article
The dataset is structured in CSV format (UTF-8 encoding).
Use Cases
Text Classification (predict article categories)
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/DanishMahdi/Encyclopedia_Sindhiana_text_corpus.TextGraphs17-shared-task-datasetWe present a dataset for graph-based question answering. The dataset consists of <question; candidate answer> pairs. For each candidate, we present a graph that is obtained by finding the shortest path between named entities mentioned in a question and a candidate answer. As a knowledge graph, we adopted Wikidata. Our dataset has the following fields:
sample_id - an identifier for <question, candidate answer>;
question - question text;
questionEntity - comma-separated list of names (textual… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/TextGraphs17-shared-task-dataset.texthumanizer-raw-datatext
