CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01singletongue /wikipedia-paragraphs wikipedia-paragraphs wikipedia-paragraphs is a dataset generated from Wikipedia, designed for natural language processing (NLP) research. Each entry contains cleaned paragraph text and Wikilink information extracted from a Wikipedia page, along with useful metadata such as categories, templates, and the associated Wikidata QID. Dataset structure Configurations The dataset is organized into multiple configurations, such as enwiki-20260607-v1.2.1.… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/wikipedia-paragraphs.tabular100M<n<1B2 likes10k downloads3mo agoHugging Face02mrlbenchmarks /global-piqa-parallel Global PIQA Parallel Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems. In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.imagequestion-answering10K<n<100K10 likes3.4k downloads4mo agoHugging Face03cloverx-id /xone-repository-parallel-en-id-corpusWe are currently developing new version of LMSE translation scoring model and processing additional data sources. We estimate the dataset will expand, with significantly improved quality.(Delayed..) A score of 55% and above indicates high-quality translation pairs, even if the first version of the model we developed gave them such a score. We will try to release a newer model in the future with better quality and consistently fast scoring speeds, and release it to the public once we decide… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/xone-repository-parallel-en-id-corpus.tabulartranslation100M<n<1B2 likes2k downloads1d agoHugging Face04Samarth-27 /Crop-Recommendation-Parameters 🌱 Crop Recommendation Dataset A machine learning dataset for crop recommendation based on soil properties and environmental conditions. The dataset contains measurements of essential soil nutrients and climatic parameters, along with the crop label that is suitable for those conditions. This dataset can be used for machine learning classification, agricultural analytics, decision-support systems, and smart farming applications. 📌 Dataset Overview Property… See the full description on the dataset page: https://huggingface.co/datasets/Samarth-27/Crop-Recommendation-Parameters.tabular1K<n<10K0 likes532 downloads2mo agoHugging Face05JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the cut — and answer with a single unconstrained paragraph of dense reasoning, focused on the exact state at the cut and what the local grammar, notation, or argument forces next. Both the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.tabular10M<n<100M0 likes496 downloads20d agoHugging Face06para-zhou /CDial-BiasOfficial release of CDial-Bias dataset. Notation: Before downloading the dataset, please be aware that: The CDial-Bias Dataset is released for research purpose only and other usages require further permission. Please ensure the usage contributes to improving the safety and fairness of AI technologies. No malicious usage is allowed. Paper: https://aclanthology.org/2022.findings-emnlp.262/ Github Repo: https://github.com/para-zhou/CDial-Bias Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/para-zhou/CDial-Bias.tabulartext-classification10K<n<100K8 likes465 downloads2y agoHugging Face07Sudehsna /Romansh_German_Parallel_Data Romansh–German Parallel Dataset (FineWeb-Based) This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction. Description This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.tabular10K<n<100K2 likes464 downloads1y agoHugging Face08giggseh /kepler-lc-star-params Kepler Stellar Lightcurves Complete Kepler mission lightcurves for ~190,000 stars, paired with stellar parameters. Each shard contains 500 stars. Fields Field Type Description target_id int Kepler Input Catalog (KIC) ID flux list<float32> Brightness measurements, normalised to median = 1.0 time list<float64> Timestamps (BJD - 2454833.0, days) n_points int Number of measurements duration_days float Observation span Teff float Effective temperature (K)… See the full description on the dataset page: https://huggingface.co/datasets/giggseh/kepler-lc-star-params.tabular100K<n<1M0 likes450 downloads6mo agoHugging Face09justintiensmith /Ordering_Constrained_ParaphrasesThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/Ordering_Constrained_Paraphrases.tabularrobotics100K<n<1M0 likes420 downloads3mo agoHugging Face10JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids. The source thoughts are two parts: a <think> block, then a single paragraph of dense reasoning about the immediate continuation. Only the part after </think> — the paragraph — becomes the VALUE. The reasoning inside the think block is dropped. The VALUE is capped at 512 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabular10M<n<100M0 likes388 downloads20d agoHugging Face11togethercomputer /ParallelKernelBench_Problems ParallelKernelBench (benchmark) Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels. This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Inputs are deterministic — reproduce them with create_input_tensor(rank, world_size, problem_id, base_shape, dtype, trial) from that file; you do not need stored .pt files. Files Path Description… See the full description on the dataset page: https://huggingface.co/datasets/togethercomputer/ParallelKernelBench_Problems.tabulartext-generationn<1K0 likes327 downloads3mo agoHugging Face12ajd12342 /paraspeechcaps ParaSpeechCaps We release ParaSpeechCaps (Paralinguistic Speech Captions), a large-scale dataset that annotates speech utterances with rich style captions ('A male speaker with a husky, raspy voice delivers happy and admiring remarks at a slow speed in a very noisy American environment. His speech is enthusiastic and confident, with occasional high-pitched inflections.'). It supports 59 style tags covering styles like pitch, rhythm, emotion, and more, spanning speaker-level… See the full description on the dataset page: https://huggingface.co/datasets/ajd12342/paraspeechcaps.tabular1M<n<10M25 likes312 downloads10mo agoHugging Face13dartbrains /paranoia Paranoia Naturalistic fMRI dataset: 22 subjects listened to a three-part ambiguous social narrative (~22 minutes total) designed to elicit varying levels of paranoid interpretation. TR = 1.0 s. 3-second fixation before each run. This repo mirrors the fmriprep-preprocessed dataset originally distributed via DataLad at https://gin.g-node.org/ljchang/Paranoia. fmriprep version 1.2.6-1. Layout derivatives/fmriprep/sub-tbXXXX/ anat/ func/ figures/ participants.tsv… See the full description on the dataset page: https://huggingface.co/datasets/dartbrains/paranoia.imagefeature-extractionn<1K0 likes283 downloads3mo agoHugging Face14justintiensmith /Reorient_Block_ParaphrasesThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/Reorient_Block_Paraphrases.tabularrobotics10K<n<100K0 likes268 downloads3mo agoHugging Face15willychan21 /ParallelKernelBench_Problems ParallelKernelBench (benchmark) Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels. This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Files Path Description data/problems.parquet One row per problem (tabular access) reference/*.py Reference solution() implementations utils/input_output_tensors.py Input/output tensor… See the full description on the dataset page: https://huggingface.co/datasets/willychan21/ParallelKernelBench_Problems.tabulartext-generationn<1K0 likes255 downloads4mo agoHugging Face16ParallaxData /snack_transfer_clean_trimmed ParallaxData/snack_transfer_clean_trimmed Edit trims Frame-aligned trim of willlyb/snack_transfer. 7 episodes retained; 3684 frames at 50 FPS. Source was not modified. Videos, numeric rows, and supported recovery sidecars use identical retained frame intervals. Each episode starts at zero. RGB cuts may be re-encoded. Existing tracking loss and reconstructed holds remain; trimming does not make them valid training labels. See meta/trim_provenance.json for bounds and original… See the full description on the dataset page: https://huggingface.co/datasets/ParallaxData/snack_transfer_clean_trimmed.tabular1K<n<10K0 likes252 downloads3d agoHugging Face17mlfoundations-dev /d1_code_long_paragraphstabular10K<n<100K0 likes246 downloads1y agoHugging Face18ParallaxData /snack_transferThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "grabette", "total_episodes": 9, "total_frames": 5528, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 50, "splits": { "train": "0:9" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ParallaxData/snack_transfer.tabularrobotics1K<n<10K0 likes235 downloads4d agoHugging Face19Faless /harvest_apples_with_agilex_piper_sim_ee_paraphrases20This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 25, "features": { "observation.state": { "dtype": "float32", "fps": 25, "shape": [ 8 ], "names": [ "ee.x", "ee.y", "ee.z", "ee.roll", "ee.pitch", "ee.yaw"… See the full description on the dataset page: https://huggingface.co/datasets/Faless/harvest_apples_with_agilex_piper_sim_ee_paraphrases20.tabularrobotics1M<n<10M0 likes231 downloads4mo agoHugging Face20ilsp /scipar_parallel_docs SciPar Parallel Documents Dataset Description This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts. In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories. This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences. To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.tabulartext-generation1K<n<10K2 likes223 downloads3y agoHugging Face21bowang0911 /paraphrasing-french Attribution MTEB-format derivative of ismailiismail/paraphrasing_french. Query = phrase; corpus = paraphrase. tabulartext-retrieval1K<n<10K0 likes222 downloads3mo agoHugging Face22justicedao /ipfs_paraguay_laws_ir Paraguay legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_paraguay_laws (revision 2e8736d28819fbb4eb711f364e8343ae8b12b4c3) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Paraguay prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_paraguay_laws_ir.tabulartext-retrieval100K<n<1M1 likes207 downloads2d agoHugging Face23ParallaxData /door_open_clean_trimmed ParallaxData/door_open_clean_trimmed Edit trims Frame-aligned trim of ParallaxData/door_open. 10 episodes retained; 4796 frames at 50 FPS. Source was not modified. Videos, numeric rows, and supported recovery sidecars use identical retained frame intervals. Each episode starts at zero. RGB cuts may be re-encoded. Existing tracking loss and reconstructed holds remain; trimming does not make them valid training labels. See meta/trim_provenance.json for bounds and original… See the full description on the dataset page: https://huggingface.co/datasets/ParallaxData/door_open_clean_trimmed.tabular1K<n<10K0 likes202 downloads3d agoHugging Face24tt1225 /robocasa-100demos-5chosen-tasks_params_virtual_views_splattedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "panda", "total_episodes": 364, "total_frames": 92882, "total_tasks": 96, "total_videos": 5096, "total_chunks": 1, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:364" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tt1225/robocasa-100demos-5chosen-tasks_params_virtual_views_splatted.tabularrobotics10K<n<100K0 likes198 downloads8mo agoHugging Face25juliensimon /omni-solar-wind-parameters OMNI Hourly Solar Wind Parameters Credit: NASA Part of a dataset collection on Hugging Face. Dataset description Merged hourly near-Earth solar wind magnetic field, plasma, energetic particle parameters combined with geomagnetic and solar activity indices from NASA's OMNI dataset. The master bridge dataset for space weather analysis -- it time-aligns IMF, solar wind, and geomagnetic response in a single file. The OMNI dataset from NASA's Goddard Space… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/omni-solar-wind-parameters.tabulartabular-regression100K<n<1M0 likes165 downloads5d agoHugging Face26browndw /human-ai-parallel-corpus-biber Citation If you use the corpus as part of your research, please cite: Do LLMs write like humans? Variation in grammatical and rhetorical styles @misc{reinhart2024llmswritelikehumans, title={Do LLMs write like humans? Variation in grammatical and rhetorical styles}, author={Alex Reinhart and David West Brown and Ben Markey and Michael Laudenbach and Kachatad Pantusen and Ronald Yurko and Gordon Weinberg}, year={2024}, eprint={2410.16107}… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-biber.tabular10K<n<100K0 likes150 downloads2y agoHugging Face27failed09 /bashkir-wikipedia-parallel Bashkir-Russian Wikipedia Parallel Corpus Sentence-level Bashkir-Russian parallel text from Wikipedia, scored and filtered for machine translation. Overview Sentence-level Bashkir-Russian parallel dataset extracted from the corresponding Bashkir and Russian Wikipedia dumps dated 2026-08-01. Candidate pairs are scored for semantic alignment with multilingual sentence encoders (Meta LASER3, Google LaBSE) and the in-domain Bashkir-Russian Pair Scorer. The filtered… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-parallel.tabulartranslation100K<n<1M0 likes147 downloads9d agoHugging Face28ParallaxData /sliding_door_clean_trimmed ParallaxData/sliding_door_clean_trimmed Edit trims Frame-aligned trim of ParallaxData/sliding_door. 3 episodes retained; 1364 frames at 50 FPS. Source was not modified. Videos, numeric rows, and supported recovery sidecars use identical retained frame intervals. Each episode starts at zero. RGB cuts may be re-encoded. Existing tracking loss and reconstructed holds remain; trimming does not make them valid training labels. See meta/trim_provenance.json for bounds and… See the full description on the dataset page: https://huggingface.co/datasets/ParallaxData/sliding_door_clean_trimmed.tabular1K<n<10K0 likes146 downloads3d agoHugging Face29pere /wiki_paragraphs_norwegian WIKI Paragraphs Norwegian A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format. Features Multiple splits for different use cases Random shuffle with Fisher-Yates algorithm Structured format with text and metadata Size-varied validation/test sets (100 to 10k samples) Splits Overview Split Name Samples Typical Usage train 1,000,000 Primary training data validation 10,000 Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.tabulartext-generation1M<n<10M0 likes142 downloads2y agoHugging Face30JDhruv14 /Brihat_Parashara_Hora_Shastratabular1K<n<10K0 likes137 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.