datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ENADE_Brazilian_national_university_examination_MCQ_483Sci2Pol-BenchSci2Pol-Bench
Data, scripts, and recipes for the benchmark Sci2Pol-Bench, a comprehensive benchmark for evaluating large language models.
About •
Usage•
Authors
About
The data consists of policy briefs obtained from Nature Energy, Nature Climate, Nature Cities, and Journal of Health and Social Behavior Policy Briefs.
Policy briefs originally were introduced in the Nature Energy journal with the goal of:
This format aims to provide… See the full description on the dataset page: https://huggingface.co/datasets/Northwestern-CSSI/Sci2Pol-Bench.northwind_opinion_mining_corpus
Opinion Mining Text Corpus
A labeled text corpus for opinion mining and sentiment analysis tasks, compiled from an open product review text corpus dataset publicly hosted on this Hub. The source corpus was assembled by a university research center.
This card does not yet list the source dataset or the applicable usage terms.
northwind_sales_anonymized_2023
Sales Transactions (Anonymized)
Anonymized sales transaction records generated internally by Northwind Analytics. No external source.
License: MIT.
CodeBench-30
CodeBench-30
A focused benchmark of 30 coding problems designed to evaluate AI model coding ability across languages, domains, and difficulty levels. Inspired by the evaluation philosophy of SWE-Bench — real problems, clear expected outputs, no ambiguity.
Dataset Summary
Property
Value
Total problems
30
Difficulty split
10 Easy / 10 Medium / 10 Hard
Languages
Python, JavaScript, SQL
Format
Code generation from description
Evaluation
Expected output… See the full description on the dataset page: https://huggingface.co/datasets/North-ML1/CodeBench-30.Northwind-Curated-Catalog
Northwind Library — Curated Catalog
This card contains catalog decisions for the Northwind Library digitization desk. The table below is the authoritative source for the current catalog-repair request.
Catalog publication: 4 accessions repaired
Accession register
accession_id
collection
received_on
catalog_state
folio_count
NW-104
Maritime Maps
2025-02-03
verified
18
NW-207
Orchard Letters
2025-02-09
verified
8
NW-318
Foundry Plans
2025-02-12… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/Northwind-Curated-Catalog.tisus_mcq_example_examhh-rlhf-harmlessInternal copy of https://huggingface.co/datasets/Anthropic/hh-rlhf.
mm_homework_datasetwind-edge-1.6-sft
Wind Lite SFT
Custom supervised fine-tuning dataset for Wind Lite 1.6 by North AI.
Dataset Summary
20,000 high-quality instruction-response pairs covering identity grounding, math reasoning, coding, general knowledge, and multi-turn conversations.
Data Composition
Category
Count
Description
Math & Reasoning
~7,000
Arithmetic, algebra, percentages, unit conversions — with step-by-step working
Coding
~4,000
Python, JavaScript, SQL, systems — with… See the full description on the dataset page: https://huggingface.co/datasets/North-ML1/wind-edge-1.6-sft.hh-rlhf-helpful-and-harmlessInternal copy of https://huggingface.co/datasets/Anthropic/hh-rlhf.
hh-rlhf-helpfulInternal copy of https://huggingface.co/datasets/Anthropic/hh-rlhf.
Northpakistandatanorth-maluku-local-languagewind-arc-1.7-sftThis dataset is Wind Arc 1.7's SFT. It includes 30 rows of Identity, and Christianity.
maasai-translation-corpus
Maasai-English Translation Corpus
Parallel English↔Maasai translation pairs for low-resource MT, language preservation, and culturally grounded tooling.
Overview
Total pairs: 9,910
Splits: 8,434 train / 738 valid / 738 test
Directions: 4,955 en→mas and 4,955 mas→en
Quality tiers: 8,444 gold and 1,466 silver
Main sources: 8,444 Bible-derived pairs, 680 cultural manual pairs, 70 knowledge-driven cultural pairs, 132 public-domain Hollis proverb pairs, 504 public-domain… See the full description on the dataset page: https://huggingface.co/datasets/NorthernTribe-Research/maasai-translation-corpus.North-Welsh-CEFR
