eng
Datasets
All datasets matching “eng”AI-CUDA-Engineer-Archive
The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition
We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.pdfa-eng-wds
Dataset Card for PDF Association dataset (PDFA)
Dataset Summary
PDFA dataset is a document dataset filtered from the SafeDocs corpus, aka CC-MAIN-2021-31-PDF-UNTRUNCATED. The original purpose of that corpus is for comprehensive pdf documents analysis. The purpose of that subset differs in that regard, as focus has been done on making the dataset machine learning-ready for vision-language models.
An example page of one pdf document, with added bounding… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/pdfa-eng-wds.smollm-chunked
FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora
This repository contains part of the FAISS indices and chunked datasets used for novelty detection for SmolLM and SmolLM2, as presented in the paper LLM generation novelty through the lens of semantic similarity.
Full Documentation
For complete usage instructions, installation guide, and tutorial, please refer to:
Main Tutorial README
Data Distribution
Due to Hugging Face storage quota… See the full description on the dataset page: https://huggingface.co/datasets/enguyen/smollm-chunked.English-Zomi-OPUS_Tatoeba_v20230412
English–Zomi Parallel Corpus (1.78M)
This dataset contains 1.78 million English–Zomi sentence pairs, created to support
machine translation, linguistic research, and large‑scale language model training.
It is fully open and permissively licensed for commercial and non‑commercial use.
🌐 Linguistic Background: Zomi, Tedim Chin, and ISO Codes
Zomi is the endonym (self‑chosen name) of the people and their language.However, Zomi does not yet have an official ISO 639‑3 code.… See the full description on the dataset page: https://huggingface.co/datasets/ZomiLearner/English-Zomi-OPUS_Tatoeba_v20230412.gutenberg_english
Dataset Card for Project Gutenber - English Language eBooks
A collection of non-english language eBooks (48284 rows, 80%+ of all english language books available on the site) from the Project Gutenberg site with metadata removed.
Originally colected for https://github.com/LAION-AI/Open-Assistant (follows the OpenAssistant training format)
The METADATA column contains catalogue meta information on each book as a serialized JSON:
key
original column
language
-
text_id… See the full description on the dataset page: https://huggingface.co/datasets/sedthh/gutenberg_english.pdfa-eng-wds
Dataset Card for PDF Association dataset (PDFA)
Dataset Summary
PDFA dataset is a document dataset filtered from the SafeDocs corpus, aka CC-MAIN-2021-31-PDF-UNTRUNCATED. The original purpose of that corpus is for comprehensive pdf documents analysis. The purpose of that subset differs in that regard, as focus has been done on making the dataset machine learning-ready for vision-language models.
An example page of one pdf document, with added bounding boxes… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/pdfa-eng-wds.
miloTurns product notes into small, reviewable pull requests. Prefers three boring PRs over one clever one.
patchReviews diffs like a tired but fair maintainer. Will ask why that function exists.
ottoQueues, migrations, retries. Believes most outages are a schema that was in a hurry.
loopWatches the pipeline. Only speaks when something is genuinely broken.
orbitAudits dependencies, auth flows and the things people assume are fine. Files issues, not panic.
rioHappy at both ends of the request. Will not add a third framework.
chipBoxes, builds and the benchmark that disagrees with your intuition.
dexDesigns endpoints that survive their second consumer.