CoolFace
26 results

ain

projecte-aina /CATalog Dataset Summary CATalog is a diverse, open-source Catalan corpus for language modelling. It consists of text documents from 26 different sources, including web crawling, news, forums, digital libraries and public institutions, totaling in 17.45 billion words. Supported Tasks and Leaderboards Fill-Mask Text Generation other:Language-Modelling: The dataset is suitable for training a model in Language Modelling, predicting the next word in a given context. Success is… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CATalog.textfill-mask10M<n<100M8 likes3.9k downloads1y agoHugging Faceprojecte-aina /synthetic_dem Dataset Card for synthetic_dem Dataset Summary The Synthetic DEM Corpus is the result of the first phase of a collaboration between El Colegio de México (COLMEX) and the Barcelona Supercomputing Center (BSC). It all began when COLMEX was looking for a way to have its Diccionario del Español de México (DEM), which can be accessed online, include the option to play each of its words with a Mexican accent through synthetic speech files. On the other hand, BSC is always on… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/synthetic_dem.audioautomatic-speech-recognition100K<n<1M2 likes1.5k downloads1y agoHugging FaceBangumiBase /ainoidenshi Bangumi Image Base of Ai No Idenshi This is the image base of bangumi AI no Idenshi, we detected 70 characters, 4221 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/ainoidenshi.1K<n<10K0 likes1.2k downloads2y agoHugging FaceAinncy /ISSEtext100K<n<1M1 likes1k downloads1y agoHugging FaceAINativeOps /AINativeBench AI-NativeBench Data This data/ directory contains the dataset for "AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems". It is split into two parts: raw/: original, unmodified experimental data organized by model and task/architecture. processed/: derived run artifacts, aggregated tables, figures, and analysis scripts built on top of the raw traces. Directory overview data/ ├── raw/ # Raw experimental data (per-model, per… See the full description on the dataset page: https://huggingface.co/datasets/AINativeOps/AINativeBench.0 likes885 downloads6mo agoHugging FaceAINovice2005 /carbon-cpu-enriched-sequences carbon-cpu-enriched-sequences A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation. Information of Features Feature Type Description record_id string NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity. begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.tabulartext-generation10M<n<100M0 likes707 downloads11d agoHugging Face