CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Navanjana /ARCHIVE-TEXT-URLS Internet Archive English Text URLs Dataset Dataset Description This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials. Dataset Summary Total Rows: 11,151,637 Language: English Source: Internet Archive Format: CSV with metadata and direct text file URLs Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.texttext-generation1M<n<10M1 likes297 downloads10mo agoHugging Face02anujsahani01 /TextCodeDepot Dataset description: The Python Code Chatbot dataset is a collection of Python code snippets extracted from various publicly available datasets and platforms. It is designed to facilitate training conversational AI models that can understand and generate Python code. The dataset consists of a total of 1,37,183 prompts, each representing a dialogue between a human and an AI Scientist. Prompt Card: Each prompt in the dataset follows a specific format known as the "Prompt… See the full description on the dataset page: https://huggingface.co/datasets/anujsahani01/TextCodeDepot.textquestion-answering10K<n<100K2 likes57 downloads3y agoHugging Face03cellos /DomesticNames_AllStates_Text Domestic Names from the Federal Government's repository of official geographic names [CSV dataset] This Dataset includes 980,065 geographic names as of September 10, 2023. It is apparent that no currently released LLMs are pretrained on datasets with many of these geographic names (i.e., features), descriptions, and histories. Example: feature_name: Abercrombie Gulch GPT-3.5 responds "I'm not aware of a specific location called Abercrombie Gulch in my training data,..." when prompted… See the full description on the dataset page: https://huggingface.co/datasets/cellos/DomesticNames_AllStates_Text.tabularquestion-answering100K<n<1M2 likes53 downloads3y agoHugging Face04NLPinas /ph_en_text_detoxedPhEnText Detoxed is a large-scale and multi-domain lexical data written in Philippine English and Taglish text. The news articles, religious articles and court decisions collated by the original researchers were filtered for toxicity and special characters were further preprocessed. This dataset has been configured to easily fine-tune LLaMA-based models (Alpaca, Guanaco, Vicuna, LLaMA 2, etc.) In total, this dataset contains 6.29 million rows of training data and 2.7 million rows of testing… See the full description on the dataset page: https://huggingface.co/datasets/NLPinas/ph_en_text_detoxed.texttext-generation1M<n<10M2 likes51 downloads3y agoHugging Face05Porameht /synthetic_text_to_sql_th Synthetic Text-to-SQL Thai Dataset Thai translation of the gretelai/synthetic_text_to_sql dataset. Dataset Description This dataset contains Thai translations of synthetic text-to-SQL examples covering various domains and SQL patterns. Source Original Dataset: gretelai/synthetic_text_to_sql Created by: Gretel.ai Statistics Split Rows Train 100,000 Test 5,851 Total 105,851 Columns Column Description… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/synthetic_text_to_sql_th.texttable-question-answering100K<n<1M0 likes34 downloads8mo agoHugging Face06channudambal /text-to-mongodb-queries-llm Dataset Description This dataset contains 10,000+ complex SQL-style analytical questions mapped to MongoDB queries and aggregation pipelines. Features Multiple schemas $group, $sum, $avg, $lookup Nested documents Long analytical questions Use Cases Fine-tuning small LLMs (Qwen, Mistral, LLaMA 3B) Text-to-Mongo query generation Data analytics agents texttext-generation10K<n<100K2 likes31 downloads9mo agoHugging Face07SantiagoPG /ocr_textstabularquestion-answering10K<n<100K0 likes13 downloads3y agoHugging Face08DanishMahdi /Encyclopedia_Sindhiana_text_corpusgated Encyclopedia Sindhiana Dataset Overview The Encyclopedia Sindhiana Dataset is a collection of encyclopedia articles created for Natural Language Processing (NLP) research. Each entry contains: Category — Topic label of the article (14 unique classes) Title — Title of the article Content — Full text of the article The dataset is structured in CSV format (UTF-8 encoding). Use Cases Text Classification (predict article categories) Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/DanishMahdi/Encyclopedia_Sindhiana_text_corpus.texttext-classification1K<n<10K0 likes9 downloads4mo agoHugging Face09s-nlp /TextGraphs17-shared-task-datasetWe present a dataset for graph-based question answering. The dataset consists of <question; candidate answer> pairs. For each candidate, we present a graph that is obtained by finding the shortest path between named entities mentioned in a question and a candidate answer. As a knowledge graph, we adopted Wikidata. Our dataset has the following fields: sample_id - an identifier for <question, candidate answer>; question - question text; questionEntity - comma-separated list of names (textual… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/TextGraphs17-shared-task-dataset.textquestion-answering10K<n<100K0 likes8 downloads3y agoHugging Face10vishal-adithya /texthumanizer-raw-datatextquestion-answering10K<n<100K1 likes7 downloads1y agoHugging Face115digit /texttexttext-classificationn<1K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.