CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Stage-jh-monitor /total-300-random-jh-epoch4 total-300-random-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3890625 Action score: 0.440625 Valid samples: 320/320 tabularn<1K0 likes4.6k downloads16d agoHugging Face02cambridgeltl /vsr_random VSR: Visual Spatial Reasoning This is the random set of VSR: Visual Spatial Reasoning (TACL 2023) [paper]. Usage from datasets import load_dataset data_files = {"train": "train.jsonl", "dev": "dev.jsonl", "test": "test.jsonl"} dataset = load_dataset("cambridgeltl/vsr_random", data_files=data_files) Note that the image files still need to be downloaded separately. See data/ for details. Go to our github repo for more introductions. Citation If you find VSR… See the full description on the dataset page: https://huggingface.co/datasets/cambridgeltl/vsr_random.imagetext-classification10K<n<100K4 likes1.6k downloads4y agoHugging Face03marin-dna /vertebrate-v1-issue473-fullwindow-cds-random-val marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val CDS full-window vertebrate projection sequences for the issue #473 random validation control. The source is the immutable issue #417 accepted-sequence table. The split uniformly samples 16,384 original-orientation CDS rows without replacement using seed 42. Sampling occurs before reverse-complement augmentation. Selected rows are removed from training; reverse complements are then added only to the remaining training… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val.tabular10M<n<100M0 likes883 downloads1mo agoHugging Face04ranausmans /synthetic-social-networks Synthetic Social Networks (Dataset) Raw experimental outputs from the Synthetic Social Networks study: 59,776 in-character LLM-agent posts from 528 production trials, and 64,562 posts total when the original pipeline-verification runs are included. The artifact combines an exploratory stage with a separately frozen, preregistered 448-trial matched-exposure confirmation. Each production trial includes peer-vote traces from in-character voting by other agents.… See the full description on the dataset page: https://huggingface.co/datasets/ranausmans/synthetic-social-networks.tabularother10K<n<100K1 likes334 downloads1mo agoHugging Face05ryankim17920 /open-bottleneck-ranklong27b-slurm-364982-rollouts Open Bottleneck RankLong 27B — Slurm array 364982 Compact rollout evidence archived from completed Slurm array 364982. Config Files / steps Records JSONL bytes Note rank_a40 60 (1–60) 15,360 66,445,670 Complete local rollout evidence rank_a80 54 (1–54) 13,824 59,386,451 Includes the cancelled arm's final dumped step (54.jsonl) Each JSONL record contains input, output, gts, score, acc, response_length, grouprel_reward, and step. Only rollout evidence is archived… See the full description on the dataset page: https://huggingface.co/datasets/ryankim17920/open-bottleneck-ranklong27b-slurm-364982-rollouts.tabular10K<n<100K0 likes94 downloads2mo agoHugging Face06threatcluster /ransomware-leak-site-victims Ransomware leak-site victims Every victim listing ThreatCluster has collected first-hand from ransomware and extortion leak sites: the group, the named organisation, when it appeared, and the sector and country where known. Built from the ThreatCluster corpus. 20,627 rows, snapshot generated 2026-09-06. Fields Field Description group_name Ransomware or extortion group that published the listing victim_name Organisation named by the group country… See the full description on the dataset page: https://huggingface.co/datasets/threatcluster/ransomware-leak-site-victims.tabulartabular-classification10K<n<100K0 likes65 downloads19d agoHugging Face07Kiria-Nozan /TRIM-gpt-oss-120b-separate-neighbors-only-para-random-feature-num TRIM Agent Reasoning Messages (HF Public Export) This directory is a Hugging Face-friendly public export of the TRIM agent reasoning SFT data. What Is Included Provider: vllm Model: gpt-oss-120b SFT mode: local_neighbor_only Splits present: train Records in this export manifest: 10056 Tasks in this split: AMES, BBB_Martins, Bioavailability_Ma, CYP2C9_Substrate_CarbonMangels, CYP2D6_Substrate_CarbonMangels, CYP3A4_Substrate_CarbonMangels, Carcinogens_Lagunin, ClinTox… See the full description on the dataset page: https://huggingface.co/datasets/Kiria-Nozan/TRIM-gpt-oss-120b-separate-neighbors-only-para-random-feature-num.tabular10K<n<100K0 likes47 downloads5mo agoHugging Face08BITnene465 /TVR_rankingtabular100K<n<1M0 likes46 downloads8mo agoHugging Face09SicariusSicariiStuff /RolePlay_Collection_random_ShareGPT RolePlay_Collection_random_ShareGPT This collection of random datasets for roleplay has been gathered from various sources. It has been processed, cleaned, and grammar-checked. However, it still requires significant work to be fully usable. Additional cleaning is necessary. I hope this proves helpful to someone. tabular1K<n<10K6 likes45 downloads2y agoHugging Face10randomyellowdude /repro-understanding-lora-as-knowledge-memory-an-empirical-analysis-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes38 downloads2mo agoHugging Face11SabaPivot /repro-randomized-feasibility-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes34 downloads2mo agoHugging Face12HBB-Community /randomizer-5Mэто случайные цифорки Нехнаю зачем я это создал tabular1M<n<10M0 likes31 downloads18d agoHugging Face13EleutherAI /lm-eval-EleutherAI_tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748 Dataset Card for Evaluation run of EleutherAI/tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748 Dataset automatically created during the evaluation run of model EleutherAI/tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748 The dataset is composed of 2 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/lm-eval-EleutherAI_tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748.tabular1K<n<10K0 likes25 downloads11mo agoHugging Face14ranork /turkish-question-departmenttabularzero-shot-classificationn<1K0 likes24 downloads2y agoHugging Face15flavianv /musical-instruments-ranker-unseen-resampled-20260924-v1 Musical Instruments original ranker protocol: resampled v1 A training-only expansion of the original ranking dataset. Same 1,648 training queries, same original negative-generation rules. Four distinct orderings of each complete reference bundle give 6,592 positive rows. Two independent negatives per strategy are requested: random set, role collision, wrong item, wrong query. There are 13,150 negatives/comparisons (exactly twice the original 6,575): random set 3,296; role… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/musical-instruments-ranker-unseen-resampled-20260924-v1.tabulartext-classification10K<n<100K0 likes19 downloads1d agoHugging Face16suraj-ranganath /tell-human-detectors TELL Human Detectors Split This dataset is a prompt-disjoint validation/test split of human_detectors.json from Jenna Russell, Marzena Karpinska, and Mohit Iyyer, "People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text" (arXiv:2501.15654). The split is intended as a research artifact for evaluating AI-vs-human writing detectors and explanation-quality methods. It preserves the upstream fields, with one documented issue type… See the full description on the dataset page: https://huggingface.co/datasets/suraj-ranganath/tell-human-detectors.tabulartext-classificationn<1K0 likes18 downloads4mo agoHugging Face17nberkowitz /gpn_combined_random_uniform_10Mb_256tabular10M<n<100M0 likes17 downloads1y agoHugging Face18anis-mselmi /random-user-profiles Random User Profiles 2,000 synthetic user profiles with random ids, ages, plans and scores. For testing only. Note: This is randomly generated synthetic data with no real-world meaning. Generated for testing and demonstration purposes. tabulartabular-classification1K<n<10K0 likes16 downloads2mo agoHugging Face19novastar111 /pacman_eval_easy_braidthin_random100_v1_20260820 Pacman Easy braided/thinned — random 100 v1 This is the exact 100-episode Pacman Easy shard used by the published UWM trajectories. Source dataset: pacman_2d_easy_braidthin_oracle_success_v1_20260815 (500 episodes) Selection: MT19937 random sample without replacement (seed 2026082001) Records: 100 JSONL SHA-256: 788d71aeb295e3a54ddfca870d8c7c720f59185c43ff21a3105cdc6ecf8d52b0 Full model trajectories:… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pacman_eval_easy_braidthin_random100_v1_20260820.tabularn<1K0 likes16 downloads1mo agoHugging Face20groupfairnessllm /random_effect_exampletabular1K<n<10K0 likes15 downloads5mo agoHugging Face21dispatchAI /speed-ranking Speed Ranking All 31 working dispatchAI models ranked by CPU inference speed. 🚀 dispatchAI tabularn<1K0 likes15 downloads3mo agoHugging Face22novastar111 /pacman_eval_forward_stress_g12_random100_20260818 Forward stress G12 — deterministic random 100 This is the exact 100-episode Pacman evaluation shard used for the final UWM trajectories. Source dataset: pacman_2d_easy_g12_f2_greedyfood_trap_forward_eval_v1_20260818 Sampling: Python MT19937 without replacement, seed 2026081812 Records: 100 JSONL SHA-256: 8118f7b275d345a1a6558df4f53c6de3382b23d8c663ebd112660e52219a6051 Trajectories: https://huggingface.co/datasets/novastar111/uwm_paper_eval_trajectories sample_manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pacman_eval_forward_stress_g12_random100_20260818.tabularn<1K0 likes14 downloads1mo agoHugging Face23Weyl09 /mosim-humanoid-walk-random-ar1-dt005 MoSim Humanoid Walk Random AR(1), dt=0.005 Humanoid Walk state transition dataset collected from DMC with random AR(1)-style actions. Format Each .npz contains: data: float32 array with shape (num_trajectories, 1000, state_dim + action_dim + state_dim). metadata: object with state_dim, action_dim, and DMC qpos/qvel sizes when available. Each transition row is [s_t, a_t, s_(t+1)]. For Humanoid Walk: state_dim = 55 action_dim = 21 row dimension = 55 + 21 + 55 = 131… See the full description on the dataset page: https://huggingface.co/datasets/Weyl09/mosim-humanoid-walk-random-ar1-dt005.tabularreinforcement-learningn<1K1 likes12 downloads4mo agoHugging Face24open-llm-leaderboard /Nexesenex__Llama_3.2_1b_RandomLego_RP_R1_0.1-detailsgated Dataset Card for Evaluation run of Nexesenex/Llama_3.2_1b_RandomLego_RP_R1_0.1 Dataset automatically created during the evaluation run of model Nexesenex/Llama_3.2_1b_RandomLego_RP_R1_0.1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nexesenex__Llama_3.2_1b_RandomLego_RP_R1_0.1-details.tabular10K<n<100K0 likes11 downloads2y agoHugging Face25novastar111 /pacman_eval_two_ghost_braidthin_random100_20260817 HF Two-Ghost braided/thinned — deterministic random 100 This is the exact 100-episode Pacman evaluation shard used for the final UWM trajectories. Source dataset: pacman_2d_easy_g9f8_two_ghost_braidthin_oracle_success_test_v1_20260817 Sampling: Python MT19937 without replacement, seed 2026081702 Records: 100 JSONL SHA-256: fbaf50a9a9799384b2cc1daf5c4c5c7e81c3fed8a23b01cc3b5b53747f3ed034 Trajectories: https://huggingface.co/datasets/novastar111/uwm_paper_eval_trajectories… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pacman_eval_two_ghost_braidthin_random100_20260817.tabularn<1K0 likes10 downloads1mo agoHugging Face26EleutherAI /lm-eval-EleutherAI_deep-ignorance-random-init Dataset Card for Evaluation run of EleutherAI/deep-ignorance-random-init Dataset automatically created during the evaluation run of model EleutherAI/deep-ignorance-random-init The dataset is composed of 0 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/lm-eval-EleutherAI_deep-ignorance-random-init.tabular1K<n<10K0 likes9 downloads11mo agoHugging Face27novastar111 /pacman_eval_forward_stress_g8_random100_20260818 Forward stress G8 — deterministic random 100 This is the exact 100-episode Pacman evaluation shard used for the final UWM trajectories. Source dataset: pacman_2d_easy_g8_f2_greedyfood_trap_forward_eval_v1_20260818 Sampling: Python MT19937 without replacement, seed 2026081808 Records: 100 JSONL SHA-256: 5044700e1e317981e889a34936593e8d95f431fa34e0e95705b1cc758168a4c6 Trajectories: https://huggingface.co/datasets/novastar111/uwm_paper_eval_trajectories sample_manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pacman_eval_forward_stress_g8_random100_20260818.tabularn<1K0 likes9 downloads1mo agoHugging Face28RandomOscillations /autotrain-data-amdal-mining-llama2-7b-cleantabularn<1K0 likes8 downloads1y agoHugging Face29ClickNoow /5k-dataset-geogpt-fineweb-randomtabular1K<n<10K0 likes8 downloads4mo agoHugging Face30novastar112 /visgym_pacman_2d_random VisGym Pacman2D Random Policy World-Model Trajectories This dataset contains random-policy interaction trajectories for the custom VisGym Pacman2D environment. Each row is one environment trajectory. The history entries include image_prev, image, and image_next base64 JPEG frames, the prompt shown at the step, the sampled action, reward, and environment info. The policy samples move actions uniformly and samples the valid stop action with the configured random stop probability after… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/visgym_pacman_2d_random.tabularimage-to-text100K<n<1M0 likes7 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.