datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
X-Atlas-Orion
X-Atlas/Orion
X-Atlas: Orion edition (X-Atlas/Orion) is a Perturb-seq atlas containing two genome-wide Fix-Cryopreserve-ScRNAseq (FiCS) Perturb-seq screens that target all human
protein-coding genes (n = 18,903 genes). The dataset is comprised of eight million HCT116 and HEK293T cells, each deeply sequenced to a median of 16,000 unique molecular
identifiers (UMIs) per cell. The median on-target knockdown efficiency is 75.4% in HCT116 cells and 51.5% in HEK293T cells, with a median… See the full description on the dataset page: https://huggingface.co/datasets/Xaira-Therapeutics/X-Atlas-Orion.X-Atlas-Orion
X-Atlas Orion Dataset (SLAF Format)
Attribution
This is a re-release of data originally generated by Xaira Therapeutics.
Original Dataset: Xaira-Therapeutics/X-Atlas-Orion
Original Format: Parquet files
This Release: Same data in SLAF (Sparse Lazy Array Format)
License: CC-BY-NC-SA-4.0 (Creative Commons Attribution-NonCommercial-ShareAlike 4.0)
Original Citation:
@article{huang2025xatlasorion,
title={X-Atlas/Orion: Genome-wide Perturb-seq Datasets via a Scalable… See the full description on the dataset page: https://huggingface.co/datasets/slaf-project/X-Atlas-Orion.AlgoTune
Website |
Paper |
Code
How good are language models at coming up with new algorithms? To try to answer this, we built a benchmark, AlgoTune, comprised of 154 widely used math, physics, and computer science functions. For each function, the goal is to write code that produces the same outputs as the original function, while being faster. In addition to the benchmark, we also provide an agent, AlgoTuner, which allows language models to easily optimize code.… See the full description on the dataset page: https://huggingface.co/datasets/oripress/AlgoTune.mualem-recitations-original
المصاحف القرآنية
مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية
البيانات الوصفية للمصاحف
ds = load_dataset('obadx/mualem-recitations-original', name='moshaf_metadata')['train']
وصف أوجه حفص
Attribute Name
Arabic Name
Values
Default Value
More Info
rewaya
الرواية
- hafs (حفص)
The type of the quran Rewaya.
recitation_speed
سرعة التلاوة
- mujawad (مجود)-… See the full description on the dataset page: https://huggingface.co/datasets/obadx/mualem-recitations-original.game-recordings-v3
OriginLab Game Recordings v0.3.0
Human gameplay captured in-engine under per-title licenses at 1080p / 60 FPS CFR on one shared frame clock: every stream starts at frame 0 and frame k matches frame k across pre-HUD and post-HUD RGB, surface normals, metric depth, audio, camera telemetry, keyboard and mouse inputs, in-engine action events and game state, and world telemetry — plus per-frame training tables.
Watch full playable previews of every modality, side by side and in… See the full description on the dataset page: https://huggingface.co/datasets/originlab/game-recordings-v3.dolma_20bn_wiki_upsampleframe-synced-multiplayer
Origin Lab Frame-Synced Multiplayer: Eight Players, One Frame Clock
Six to eight players, each on their own PC on residential internet, each
recording their own live view of one match, and frame k on every machine is
the same server instant, verified four independent ways, with the
verification script in this repo. Every player ships the full engine stack
at 1080p / 60 FPS on one shared frame grid. Because the alignment is measured
rather than assumed, the release also… See the full description on the dataset page: https://huggingface.co/datasets/originlab/frame-synced-multiplayer.dolma_20bn_instruct_upsampleasap-7-originaldolma_18bn_stratified_sampledolma_20bn_cc_high_qualitydolma_20bn_prop_stratified_sampledolma_20bn_no_instructo-ring_removeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_joint_0.pos",
"left_joint_1.pos",
"left_joint_2.pos",
"left_joint_3.pos",
"left_joint_4.pos",
"left_joint_5.pos"… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/o-ring_remove.clear_square_origThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 19189,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lirislab/clear_square_orig.dolma_18bn_prop_stratified_sampleasap-8-originalasappp-1-2-originalpick_tblock_mp_original_controllerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "DualPanda",
"total_episodes": 1000,
"total_frames": 173000,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/younghyopark/pick_tblock_mp_original_controller.contrastive-pretraining
Contrastive Pretraining
Per-language query/document pairs produced by the retrieval-common-crawl pipeline.
Each config corresponds to a single language or source with identical LightOn-style schema.
Config overview
Configs available are: fw-edu, fw2-arb_Arab, fw2-ces_Latn, fw2-cmn_Hani, fw2-dan_Latn, fw2-deu_Latn, fw2-ell_Grek, fw2-fas_Arab, fw2-fra_Latn, fw2-hun_Latn, fw2-ind_Latn, fw2-ita_Latn, fw2-jpn_Jpan, fw2-nld_Latn, fw2-pol_Latn, fw2-por_Latn, fw2-rus_Cyrl… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/contrastive-pretraining.eval1_mix_orig93_booster55_difficult30x2_h264This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-group47/eval1_mix_orig93_booster55_difficult30x2_h264.asappp-3-6-originaljetson_orin_nano_super_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 1,
"total_frames": 590,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Ruth011/jetson_orin_nano_super_1.dumpling_pick_originalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 89795,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hesh0629/dumpling_pick_original.game-scenes-posed-rgbd
Origin Lab Game Scenes: Posed RGB-D Flythroughs of Game Worlds
Every frame carries the camera that rendered it and the depth the engine computed for it. Ten game worlds, with the camera released from the player for 60% of the footage: metric depth, world-space normals, 4x4 pose, and per-frame intrinsics on one frame index, plus hundreds of full in-place turns and long stretches in which the world is frozen and only the camera moves. Two trajectories per world and one whole… See the full description on the dataset page: https://huggingface.co/datasets/originlab/game-scenes-posed-rgbd.iers-earth-orientation
IERS Earth Orientation Parameters
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Earth Orientation Parameters (EOP) from the IERS finals2000A series. Includes polar motion, UT1-UTC, length of day, and nutation offsets. Updated daily.
Earth Orientation Parameters describe the irregularities in Earth's rotation and the motion of its poles. These parameters are essential for transforming between celestial and terrestrial… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/iers-earth-orientation.train-dataset-original-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 210,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hngchris/train-dataset-original-test.Original-alpha-suppression-task-boosto-ring_remove_20260831_115728This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_joint_0.pos",
"left_joint_1.pos",
"left_joint_2.pos",
"left_joint_3.pos",
"left_joint_4.pos",
"left_joint_5.pos"… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/o-ring_remove_20260831_115728.grab_brush_diff_bin_orientThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 40,
"total_frames": 35726,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/guanfengliu/grab_brush_diff_bin_orient.
