datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
random-small-github-repositories
random-small-github-repositories
A collection of 5,613 small-to-medium open-source GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks.
Contents
seed_small_repos.csv — metadata for each repo (owner, repo_name, stars, license, repo_hash)
repos-zipped/ — one .zip per repo, named {repo_hash}.zip
unzipper.py - unzipping python… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-small-github-repositories.llm-smartrouter-benchmark
LLM SmartRouter & Agent Highway Latency & Cost Benchmark (v1.4.0)
Empirical performance benchmark dataset comparing direct model endpoints (OpenAI, Anthropic Claude, Google Gemini) against the PixelRouter / BLUN SmartRouter proxy layer and Autonomous Agent Web Highway (https://api.pixeloffice.eu/v1).
v1.4.0 Benchmark Highlights
Anthropic Claude Messages API: Sub-35ms proxy routing for native /v1/messages payloads with 94%+ cost savings.
Machine Web Highway… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/llm-smartrouter-benchmark.HomeDepot-Smart-Home-Dataset
Home Depot Smart Home Product Dataset
A structured dataset of Smart Home products from Home Depot, featuring detailed product specifications, pricing, category taxonomy, highlights, color variants, dimensions, and brand data. Ideal for training product recommendation models, smart home AI assistants, price intelligence systems, and e-commerce search engines.
Dataset Overview
Field
Details
Source
Home Depot
Total Records
230+
Category Focus
Smart Home, IoT… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/HomeDepot-Smart-Home-Dataset.small_kazakh_corpus
Dataset Card for Small Kazakh Language Corpus
The Small Kazakh Language Corpus is a specialized collection of textual data designed for training and research of natural language processing (NLP) models in the Kazakh language. The corpus is structured to ensure high text quality and comprehensive representation of diverse linguistic constructs.
Dataset Details
Dataset Description
The dataset consists of Kazakh language texts with annotations that support tasks… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/small_kazakh_corpus.Faroese_BLARK_small
Dataset Card for Faroese_BLARK_small
Dataset Description
All sentences are retrieved from:
Paper:
Annika Simonsen, Sandra Saxov Lamhauge, Iben Nyholm Debess, and Peter Juel Henrichsen. 2022. Creating a Basic Language Resource Kit for Faroese. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 4637–4643, Marseille, France. European Language Resources Association.
Dataset Summary
This dataset is a filtered version of the corpus… See the full description on the dataset page: https://huggingface.co/datasets/barbaroo/Faroese_BLARK_small.en_my_myanmar-xnli_small
myXNLI Preprocessed Dataset (Small Version)
Overview
This dataset is a smaller version of the myXNLI corpus, which extends the XNLI dataset to include Myanmar (Burmese) language. The dataset has been preprocessed, split into training, validation, and test sets, and is specifically designed for natural language inference (NLI) and machine translation tasks.
Training set: 50,000 sentence pairs
Validation set: 2,490 sentence pairs
Test set: 5,010 sentence pairs
The… See the full description on the dataset page: https://huggingface.co/datasets/myamjechal/en_my_myanmar-xnli_small.Math_small_corpuskurdish-train-smallSMART-Goals-Validation
Dataset Description
Synthitic dataset generated using Google AI Studio.
for training LLMs for specific data and following the same pattern.
splits into ( SMART-Goal-Examples --> 2013, TaskList-Examples --> 2500)
task-specs-small
About
Tasks specs
