CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01neo4j /text2cypher-2024v1 Neo4j-Text2Cypher (2024) Dataset The Neo4j-Text2Cypher (2024) Dataset brings together instances from publicly available datasets, cleaning and organizing them for smoother use. Each entry includes a “question, schema, cypher” triplet at minimum, with a total of 44,387 instances — 39,554 for training and 4,833 for testing. An overview of the dataset is shared at Link Have ideas or insights? Contact us: Neo4j/Team-GenAI Fields Fields and their descriptions are as… See the full description on the dataset page: https://huggingface.co/datasets/neo4j/text2cypher-2024v1.texttext-generation10K<n<100K55 likes313 downloads1y agoHugging Face02FBK-MT /Neo-GATE Dataset card for Neo-GATE Homepage: https://mt.fbk.eu/neo-gate/ Dataset summary Neo-GATE is a bilingual corpus designed to benchmark the ability of machine translation (MT) systems to translate from English into Italian using gender-inclusive neomorphemes. It is built upon GATE (Rarrick et al., 2023), a benchmark for the evaluation of gender rewriters and gender bias in MT. Neo-GATE includes 841 test entries (Neo-GATE.tsv) and 100 dev entries (Neo-GATE-dev.tsv). Each… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Neo-GATE.texttranslation1K<n<10K11 likes212 downloads2y agoHugging Face03neonforestmist /smolgpt-markdown-stories SmolGPT-Fables Stories A deterministic, text-only corpus of 96,000 original English Markdown stories built for SmolGPT-Fables. Every row is one complete supervised story example with an exact prompt / completion boundary, a requested scene count from one to six, and plain-language conditioning fields. No model, API, browser, or network service was used to create this dataset. Dataset summary 96,000 stories across 96,000 isolated story families 25 genres and all… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/smolgpt-markdown-stories.tabulartext-generation10K<n<100K0 likes132 downloads3mo agoHugging Face04neoai-inc /LIT-RAGBench LIT-RAGBench LIT-RAGBench is a benchmark for evaluating generator capabilities in Retrieval-Augmented Generation (RAG). It focuses on whether a model can answer questions correctly given retrieved documents, independent of retrieval quality. The benchmark covers five categories: Integration, Reasoning, Logic, Table, and Abstention. Dataset Summary LIT-RAGBench contains: 114 human-constructed Japanese questions An English version generated by machine translation with… See the full description on the dataset page: https://huggingface.co/datasets/neoai-inc/LIT-RAGBench.textquestion-answeringn<1K0 likes101 downloads6mo agoHugging Face05emgena /omnimcp_graphrag_neo4j_cypher_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_neo4j_cypher_teaser.texttext-generationn<1K0 likes79 downloads10d agoHugging Face06neomoon50 /Nemotron-Personas-Korea Nemotron-Personas-Korea 우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템 A compound AI approach to personas grounded in real-world distributions 데이터셋 개요 (Overview) Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 통계청(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다. Nemotron-Personas-Korea는… See the full description on the dataset page: https://huggingface.co/datasets/neomoon50/Nemotron-Personas-Korea.imagetext-generation1M<n<10M0 likes71 downloads5mo agoHugging Face07neogenesislab /korean-llm-citation-baseline-2026 DOI This dataset is citable via DataCite DOI 10.5281/zenodo.20018479 (Zenodo record). Cite as: @dataset{neogenesis_20018479, author = {Heo, Yesol and Neo Genesis Lab}, title = {Korean LLM Citation Baseline 2026 (Neo Genesis GEO Measurement)}, year = 2026, publisher = {Zenodo}, doi = {10.5281/zenodo.20018479}, url = {https://doi.org/10.5281/zenodo.20018479} } Korean LLM Citation Baseline 2026 (Neo Genesis GEO… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/korean-llm-citation-baseline-2026.tabulartext-generationn<1K0 likes52 downloads5mo agoHugging Face08neogenesislab /cross-agent-review-queue-2026 DOI This dataset is citable via DataCite DOI 10.5281/zenodo.20018477 (Zenodo record). Cite as: @dataset{neogenesis_20018477, author = {Heo, Yesol and Neo Genesis Lab}, title = {Cross-Agent Code Review Queue (Codex <-> Claude, Neo Genesis 2026)}, year = 2026, publisher = {Zenodo}, doi = {10.5281/zenodo.20018477}, url = {https://doi.org/10.5281/zenodo.20018477} } Cross-Agent Code Review Queue (Codex <-> Claude, Neo… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/cross-agent-review-queue-2026.texttext-generationn<1K0 likes49 downloads5mo agoHugging Face09neoluigi /MELD-MPCA Dataset Information the dataset is a jsonl file containing each dialogue (context) per line. Field Amount Dialogue (context/line) 1022 Diff User 260 each dialogue context messages of a conversation, with those informations: user content emotion type n_turn summary traits distanglement ref_speaker ref_utterance tar_speaker selected_speaker The user of the message The content of the message The emotion of the user Either a positive or negative emotion The… See the full description on the dataset page: https://huggingface.co/datasets/neoluigi/MELD-MPCA.textsummarization1K<n<10K0 likes43 downloads1y agoHugging Face10Dimitrex93 /neo1.0-benchmark Neo — LLM-Halluzinations-Benchmark (Deutsch) 134 Fragen (Multiple Choice + offene Fragen) in Deutsch, gebaut, um typische Halluzinationen kleiner lokaler LLMs zu provozieren — nicht um Weltwissen abzufragen. Genutzt für das eigene Fine-Tune Dimitrex93/neo1.0-3b und als SFT-Ground-Truth. English: 134 German-language questions (MC + open) designed to trigger the hallucination patterns of small local LLMs. Used as benchmark and SFT ground truth for neo1.0-3b. Evaluation harness:… See the full description on the dataset page: https://huggingface.co/datasets/Dimitrex93/neo1.0-benchmark.textquestion-answeringn<1K0 likes40 downloads2d agoHugging Face11REILX /neo_sft_phase2_conversations 1. The original dataset can be found at: https://huggingface.co/datasets/m-a-p/neo_sft_phase2 2. Split multi-turn conversations into individual single-turn samples Approach: Treat each round of dialogue as an independent question-and-answer pair, and construct the sample using contextual information. Specific operations: For each "conversations", iterate through each round of dialogue. Concatenate the "value" of the current "human" round with the dialogue from all… See the full description on the dataset page: https://huggingface.co/datasets/REILX/neo_sft_phase2_conversations.texttext-generation100K<n<1M0 likes38 downloads2y agoHugging Face12dotwee /structured-stern-neon-articles Structured Stern NEON Community Articles This repository contains approximately 20k user written texts, articles, and poetry pulled from archives of the Stern NEON website. Stern NEON was a community platform where users could write and publish their own articles. Many of the articles are personal stories, poems, or opinion pieces. The articles are structured in a way that they can be used for further analysis. Dataset Details Uses This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.tabulartext-classification10K<n<100K0 likes38 downloads8mo agoHugging Face13giseldo /neo_ara_v2tabulartext-generation10K<n<100K1 likes37 downloads1y agoHugging Face14neogenesislab /sora-multi-device-orchestration-2026 DOI This dataset is citable via DataCite DOI 10.5281/zenodo.20018481 (Zenodo record). Cite as: @dataset{neogenesis_20018481, author = {Heo, Yesol and Neo Genesis Lab}, title = {Sora Multi-Device Orchestration Architecture 2026}, year = 2026, publisher = {Zenodo}, doi = {10.5281/zenodo.20018481}, url = {https://doi.org/10.5281/zenodo.20018481} } Sora Multi-Device Orchestration Architecture 2026 A reference… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/sora-multi-device-orchestration-2026.texttext-generationn<1K0 likes35 downloads5mo agoHugging Face15Shumatsurontek /neo-sql-reasoning-combined Neo SQL + Reasoning Combined Dataset Combined SFT dataset for fine-tuning SQL, reasoning, and math models. Built for the neo-deep-agent-lab project. Sources & Proportions Source Proportion Records Focus gretelai/synthetic_text_to_sql 50% ~5,000 SQL generation nohurry/Opus-4.6-Reasoning-3000x-filtered 30% ~2,100 Reasoning openai/gsm8k 20% ~1,400 Math problems Format All samples are normalized to SFT chat format: { "messages": [… See the full description on the dataset page: https://huggingface.co/datasets/Shumatsurontek/neo-sql-reasoning-combined.texttext-generation10K<n<100K0 likes34 downloads6mo agoHugging Face16Alex01837178373 /neolurk-dataset Neolurk.org Memes Dataset Этот датасет содержит очищенный текстовый корпус русской интернет-энциклопедии мемов Neolurk (современный преемник Lurkmore). Описание Данные предназначены для обучения и файнтюнинга больших языковых моделей (LLM), помогая им лучше понимать русскоязычный интернет-фольклор, сленг, мемы и исторический контекст субкультур. Количество статей: 48 234 статьи Формат: JSON Lines (.jsonl), где каждая строка содержит: pageid (int): уникальный… See the full description on the dataset page: https://huggingface.co/datasets/Alex01837178373/neolurk-dataset.texttext-generation10K<n<100K0 likes33 downloads3mo agoHugging Face17neovalle /H4rmony_dpo Citation Information @article{neovalle2024h4rmony, author = {Vallego, Jorge}, title = {H4rmony DPO Dataset}, howpublished = {Hugging Face Hub}, year = {2024}, url = {https://huggingface.co/datasets/neovalle/H4rmony_dpo} } This dataset is based on neovalle/H4rmony, and optimised to the format required by DPOTrainer from the trl library. textquestion-answering1K<n<10K11 likes28 downloads8mo agoHugging Face18neogenesislab /whylab-gemini-2-5-docker-validation 🛈 Anonymity Notice (2026-05-12): The associated manuscript is currently under peer review at a double-blind venue. Author identity and venue-specific identifiers have been withheld throughout this README, the BibTeX templates, and the CITATION.cff block. The dataset itself remains CC-BY-4.0 and is independently citable via its Zenodo DOI 10.5281/zenodo.20018468. The author byline will be restored after the review outcome is announced. DOI This dataset is citable via DataCite DOI… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/whylab-gemini-2-5-docker-validation.tabularothern<1K0 likes28 downloads5mo agoHugging Face19neogenesislab /quant-v11-ensemble-6alpha-specs-2026 DOI This dataset is citable via DataCite DOI 10.5281/zenodo.20018487 (Zenodo record). Cite as: @dataset{neogenesis_20018487, author = {Heo, Yesol and Neo Genesis Lab}, title = {Quant v11 Ensemble 6-Alpha Specs & Risk Engineering 2026}, year = 2026, publisher = {Zenodo}, doi = {10.5281/zenodo.20018487}, url = {https://doi.org/10.5281/zenodo.20018487} } Quant v11 Ensemble - 6-Alpha Specs & Risk Engineering 2026 A… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/quant-v11-ensemble-6alpha-specs-2026.texttext-generationn<1K0 likes22 downloads5mo agoHugging Face20REILX /neo_sft_phase2_single dataset The original dataset can be found at: https://huggingface.co/datasets/m-a-p/neo_sft_phase2 Use the following code to select two-turn conversations for your SFT dataset. code import json def process_conversations(input_file, output_file): with open(input_file, 'r', encoding='utf-8') as f_in, \ open(output_file, 'w', encoding='utf-8') as f_out: data = json.load(f_in) for item in data: conversations =… See the full description on the dataset page: https://huggingface.co/datasets/REILX/neo_sft_phase2_single.texttext-generation10K<n<100K0 likes18 downloads2y agoHugging Face21abhi26 /research-papers-gpt-neox abhi26/research-papers-gpt-neox This dataset contains processed research papers optimized for GPT-NeoX-20B training. The text has been cleaned, chunked to 2048 tokens, and formatted for causal language modeling. Dataset Details Total Samples: 9993 Unique Papers: 1017 Average Tokens per Sample: 1965.4 Token Range: 10 - 91659 Max Token Limit: 2048 Source Subdirectories: 1 Dataset Structure Each sample contains: text: The processed research paper text or chunk… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/research-papers-gpt-neox.texttext-generation1K<n<10K0 likes14 downloads1y agoHugging Face22REILX /neo_sft_phase2_multi 1. The original dataset can be found at: https://huggingface.co/datasets/m-a-p/neo_sft_phase2 2. Split multi-turn conversations into individual single-turn samples Approach: Treat each round of dialogue as a separate question-and-answer pair, and construct the sample by leveraging the contextual information. Specific Operations: For each "conversation," iterate through all the dialogue rounds. Concatenate the "value" of all "human" turns within each "conversation" to… See the full description on the dataset page: https://huggingface.co/datasets/REILX/neo_sft_phase2_multi.texttext-generation100K<n<1M1 likes8 downloads2y agoHugging Face23giseldo /neo_ara_v1tabulartext-generation10K<n<100K1 likes8 downloads1y agoHugging Face24leo20240112 /pile-neox-uint16-partsTokenized uint16 shard parts for language-model pretraining. Original source: The Pile / NeoX-style preprocessing. tabulartext-generationn<1K0 likes6 downloads3mo agoHugging Face25dlewicki /neocortirrhea-lexicon Dataset Card: Neocortirrhea Lexicon Entry Summary This dataset entry defines and contextualizes the psychological, neurological, and somatic neologism Neocortirrhea. Dataset Structure JSON Lines Representation (data.jsonl) { "term": "Neocortirrhea", "part_of_speech": "noun", "phonetic": "/ˌniː.oʊˌkɔːr.tɪˈriː.ə/", "etymology": "Neocortex (higher-order cognitive processing) + -rrhea (Greek rhoia: abnormal/excessive flow or… See the full description on the dataset page: https://huggingface.co/datasets/dlewicki/neocortirrhea-lexicon.texttext-generationn<1K0 likes6 downloads1mo agoHugging Face26NeoAI-Official /Moon-1-Datatexttext-generationn<1K0 likes3 downloads6mo agoHugging Face27NeoAI-Official /Moon-2-Datatexttext-generation1K<n<10K0 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.