ain
Datasets
All datasets matching “ain”CATalog
Dataset Summary
CATalog is a diverse, open-source Catalan corpus for language modelling. It consists of text documents from 26 different sources, including web crawling, news, forums, digital libraries and public institutions, totaling in 17.45 billion words.
Supported Tasks and Leaderboards
Fill-Mask
Text Generation
other:Language-Modelling: The dataset is suitable for training a model in Language Modelling, predicting the next word in a given context. Success is… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CATalog.synthetic_dem
Dataset Card for synthetic_dem
Dataset Summary
The Synthetic DEM Corpus is the result of the first phase of a collaboration between El Colegio de México (COLMEX) and the Barcelona Supercomputing Center (BSC).
It all began when COLMEX was looking for a way to have its Diccionario del Español de México (DEM), which can be accessed online, include the option to play each of its words with a Mexican accent through synthetic speech files. On the other hand, BSC is always on… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/synthetic_dem.ainoidenshi
Bangumi Image Base of Ai No Idenshi
This is the image base of bangumi AI no Idenshi, we detected 70 characters, 4221 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/ainoidenshi.ISSEAINativeBench
AI-NativeBench Data
This data/ directory contains the dataset for "AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems".
It is split into two parts:
raw/: original, unmodified experimental data organized by model and task/architecture.
processed/: derived run artifacts, aggregated tables, figures, and analysis scripts built on top of the raw traces.
Directory overview
data/
├── raw/ # Raw experimental data (per-model, per… See the full description on the dataset page: https://huggingface.co/datasets/AINativeOps/AINativeBench.carbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences
A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences
and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
Feature
Type
Description
record_id
string
NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.
