CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Dk587 /arctic Arctic Shift Reddit Archive Every Reddit comment and submission since 2005, organized as monthly Parquet shards What is it? The full Reddit archive from Arctic Shift, converted to Parquet and hosted here for easy access. Covers every public subreddit from 2005-12 through 2026-02. Right now the archive has 1.6B items (362.1M comments, 1.2B submissions) in 181.4 GB of compressed Parquet. Comments and submissions are stored as separate datasets, split into monthly shards… See the full description on the dataset page: https://huggingface.co/datasets/Dk587/arctic.tabulartext-generation1B<n<10B1 likes1.1k downloads6mo agoHugging Face02blairducrayoppat /openvino-arc140v-lunarlake OpenVINO local-inference on an Intel Arc 140V (Lunar Lake) iGPU Reference performance data for running local models on a single Intel Core Ultra 7 258V (Lunar Lake) laptop with the integrated Intel Arc 140V (Xe2) GPU, via OpenVINO. All inference runs on the iGPU; the NPU stays idle throughout, confirmed by the telemetry here. This is reference characterization shared by a non-expert contributor — careful measurements on one machine, offered so others can compare and correct, not… See the full description on the dataset page: https://huggingface.co/datasets/blairducrayoppat/openvino-arc140v-lunarlake.tabulartext-generationn<1K0 likes675 downloads8d agoHugging Face03ajibawa-2023 /Technical-Architectures-Large Technical Architectures Large (294k Samples) Overview Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 293,000+ distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8. Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Technical-Architectures-Large.tabulartext-generation100K<n<1M8 likes248 downloads2mo agoHugging Face04mcp-tool-shop /jam-rollout-arc-evals Rollout arc — raw generations Every model generation behind the write-ups in mcp-tool-shop-org/ai-jam-sessions under experiments/rollout-arc/p4/. Two things you can do with this. Check our arithmetic. The repo has the readout scripts, the preregistrations and the intervals — but the generations they were computed from are ~51 MB and were never committed, so a clone got the conclusions and no way to recompute them. These are those files, unfiltered. Or run the loop yourself. The… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-rollout-arc-evals.tabulartext-generation1K<n<10K0 likes108 downloads9d agoHugging Face05akumalondon /Rail_Freight_Logistics_Company_Email_Archive_Sample Ukrainian Rail-Freight Correspondence Corpus (Sample) Real operational correspondence from a working freight forwarding business, and the documents attached to it — consignment notes, service acts, invoices, wagon manifests. Not scraped, not synthetic, and never published anywhere before. This is a de-identified sample released for evaluation. It is drawn from a larger private archive; see Full archive below. Published by Akuma London · akumalondon.com Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.tabulartext-generation1K<n<10K0 likes89 downloads12d agoHugging Face06audreyt /sayit-archive-tw Audrey Tang Transcript Corpus Public transcripts of Audrey Tang (唐鳳) — 2025 Right Livelihood Laureate, civic hacker, and Taiwan's Cyber Ambassador. Tang is co-author of Plurality: The Future of Collaborative Technology and Democracy and an inaugural Senior Accelerator Fellow at the Oxford Institute for Ethics in AI. She served as Taiwan's first Digital Minister (2016–2024) and the world's first nonbinary cabinet minister, awarded the Right Livelihood Award for "advancing the social… See the full description on the dataset page: https://huggingface.co/datasets/audreyt/sayit-archive-tw.tabulartext-generation100K<n<1M0 likes88 downloads7mo agoHugging Face07Banodoco /discord-archive Discord Archive This is an archive of messages from the Banodoco Discord community, where technical and artistic practitioners have been discussing open source AI art for the past three years. The archive captures a long-running community record of people learning, training, evaluating, and using open source AI art models in practice. It contains discussion around model releases, workflows, tooling, troubleshooting, creative experiments, training details, and the many small… See the full description on the dataset page: https://huggingface.co/datasets/Banodoco/discord-archive.tabulartext-generation1M<n<10M4 likes68 downloads4mo agoHugging Face08G4KMU /review_arcade ACL ARR Reviews - Review Arcade Project This dataset contains paper reviews from the ACL ARR (Association for Computational Linguistics - Annual Review of Research) program. The reviews are organized into splits corresponding to papers that were accepted or rejected for publication, as well as specific subsets used for the research analysis. This dataset was introduced in the paper Review Arcade: On the Human Alignment and Gameability of LLM Reviews. Code:… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/review_arcade.tabulartext-classification1K<n<10K2 likes60 downloads4mo agoHugging Face09Nabidnur /arc-agi-2-grids ARC-AGI-2 Grids — training + analysis corpus (NVARC-compatible) Companion dataset for the Kaggle ARC Prize 2026 (ARC-AGI-2) solver built on sorokin/qwen3_4b_grids15_sft139 + per-task rank-256 LoRA (NVARC lineage). Everything here is generated from public canonical data only (1,000 training / 120 evaluation tasks); no hidden competition data is included. Contents Path Rows Description train/train_tasks.jsonl 1,000 canonical training tasks (full I/O)… See the full description on the dataset page: https://huggingface.co/datasets/Nabidnur/arc-agi-2-grids.tabulartext-generation100K<n<1M0 likes34 downloads2d agoHugging Face10nuprl-staging /agent-archive Agent Archive: PR Interactions This dataset contains pull request interactions from GitHub repositories where AI coding agents (Claude Code, Copilot, etc.) contributed code. The data was sourced from the AgentPack dataset and enriched with full PR timeline events from the GitHub API. Dataset Description Each row represents a single pull request with its complete interaction history, including comments, reviews, commits, and other timeline events. Column… See the full description on the dataset page: https://huggingface.co/datasets/nuprl-staging/agent-archive.tabulartext-generation10K<n<100K0 likes22 downloads8mo agoHugging Face11autoshift /Technical-Architectures-Large Technical Architectures Large (210k+ Samples) Overview Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 210,000 distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8. Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/autoshift/Technical-Architectures-Large.tabulartext-generation100K<n<1M0 likes20 downloads2mo agoHugging Face12KotshinZ /arc-agi-augmented-100 ARC-AGI Augmented Dataset This dataset is an augmented version of the Abstraction and Reasoning Corpus (ARC-AGI), processed for training neural networks (such as Transformers or Neural Cellular Automata). Dataset Details Original Source: ARC-AGI Benchmark License: MIT Augmentation Method: Dihedral Transformations: 8 symmetries (rotations/flips). Color Permutation: Random permutation of colors 1-9 (0 is fixed as background). Translational Padding: Randomly positioning the… See the full description on the dataset page: https://huggingface.co/datasets/KotshinZ/arc-agi-augmented-100.tabularimage-to-image100K<n<1M1 likes18 downloads9mo agoHugging Face13louisbrulenaudet /code-deontologie-architectes Code de déontologie des architectes, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-deontologie-architectes.tabulartext-generationn<1K0 likes16 downloads1y agoHugging Face14hunterbown /bell-labs-technical-archive Bell Labs Documents and Stuff This is a conservative public-release subset of the internal BELLA continued-pretraining corpus. It keeps the Bell-system technical material that survived a stricter final pass for public dataset hosting and removes records that still looked risky, off-scope, or too low-signal for a Hugging Face corpus listing. What is in the release Split Documents train 1220 validation 29 test 42 The release contains 1291 documents out… See the full description on the dataset page: https://huggingface.co/datasets/hunterbown/bell-labs-technical-archive.tabulartext-generation1K<n<10K0 likes16 downloads6mo agoHugging Face15Nan-Do /atcoder_arc_contestsgated Notification Atcoder is selling this data now. If you are interested in accessing it please contact them. Dataset Summary This dataset aims to facilitate the creation of sophisticated, multi-turn dialogue datasets focused on coding that could be used for training reasoning Large Language Models (LLMs), particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation techniques. It also serves as a robust foundation for problem-solving in Large Language… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_arc_contests.tabulartext-classification1M<n<10M1 likes14 downloads2mo agoHugging Face16fineset-io /neural-architecture-search-papers Neural Architecture Search (NAS) Papers — FineSet A research-paper dataset on Neural Architecture Search (NAS) Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Neural Architecture Search (NAS) Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/neural-architecture-search-papers.tabulartext-classificationn<1K0 likes14 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.