datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
icl-dataset-end-effector-space
icl-dataset-fixed-obs
Derived from adityx23/icl-dataset
(lerobot v2.1 format). Every existing column, task, episode flag
(success/valid/keep), and episode_uid is carried through unchanged.
What's added
Two new features, observation.left_ee / observation.right_ee (float32,
shape [7], names qw, qx, qy, qz, x, y, z): the Cartesian end-effector
pose of each arm, forward-kinematics'd from that frame's recorded
observation.state (real joint encoders) through the same… See the full description on the dataset page: https://huggingface.co/datasets/Hannibal52Barca/icl-dataset-end-effector-space.nyu_door_opening_surprising_effectiveness_rawnyu_door_opening_surprising_effectiveness_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "hello_stretch",
"total_episodes": 435,
"total_frames": 18196,
"total_tasks": 1,
"total_videos": 435,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 3,
"splits": {
"train": "0:435"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/nyu_door_opening_surprising_effectiveness_lerobot.demo-end-effector-space
demo_action_space
Derived from adityx23/icl-demo-dataset
(lerobot v2.1 format, 285 episodes / 254,171 frames / 27 tasks). Every
existing column, task, episode flag (success/valid/keep), and
episode_uid is carried through unchanged.
Sibling dataset: demo_joint_space
adds the same episodes' joint-space IK targets instead of Cartesian poses.
Same source, same episode indices, same pipeline.
What's added
Two new features, observation.left_ee / observation.right_ee… See the full description on the dataset page: https://huggingface.co/datasets/Hannibal52Barca/demo-end-effector-space.EffectData
EffectData
This repository contains the dataset released with the paper "EffectMaker: Unifying Reasoning and Generation for Customized Visual Effect Creation".
Shiyuan Yang1,2,†,*, Ruihuang Li1,†, Jiale Tao1, Shuai Shao1,‡, Qinglin Lu1,✉, Jing Liao2,✉1Tencent Hunyuan 2City University of Hong Kong †Equal Contribution *Work Done During Internship at Tencent Hunyuan ‡Project Lead ✉Corresponding Authors
We introduce EffectData, the largest and high-quality… See the full description on the dataset page: https://huggingface.co/datasets/ysy31415926/EffectData.EffectErase
EffectErase Dataset
🚨 License Update Notice (Important)We have updated the dataset license and access terms.If you submitted an access request before March 20, 2026, 20:23 (AoE),please resubmit your request to acknowledge the updated license.
⚠️ This dataset is for non-commercial research purposes only.Commercial use is strictly prohibited without explicit permission from the authors.
Overview
The EffectErase dataset is designed for video object removal and… See the full description on the dataset page: https://huggingface.co/datasets/FudanCVL/EffectErase.china-effective-laws-regulations
全国现行法律法规合集
现行有效的中华人民共和国法律、行政法规、监察法规、地方性法规、司法解释结构化文本。一部法规一行,一条法条一行,供查阅、检索、RAG 和法律 NLP 使用。
数据来自全国人大常委会办公厅 国家法律法规数据库,下载口径为官网的 「有效及尚未生效」。正文由 Word 原文用脚本抽取,未经大模型改写。
这不是官方汇编,不能替代公报或标准文本,也不能作为法律意见。 电子文本与标准文本不一致时,以法律规定的标准文本为准。
快照日期:2026-08-26
效力说明
本数据集 以现行有效法律法规为主体:
效力 status
法规份数
说明
有效
17,649
现行有效,默认应使用这一部分
尚未生效
7
已公布、施行日晚于快照日
失效
45
文件名含「失效」,多为已到期的全国人大常委会试点授权决定
使用时请筛选 status == "有效",即可得到现行有效文本。同一部法若有修正前后多个版本,均予保留,用 filename_date 区分,采用最新日期即可。… See the full description on the dataset page: https://huggingface.co/datasets/senry5433/china-effective-laws-regulations.task828_copa_commonsense_cause_effect
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task828_copa_commonsense_cause_effect
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task828_copa_commonsense_cause_effect.variant-effect-prediction
Updates
[2025-09-09] We have added ClinVar variant effect prediction results to the repository. The evaluation dataset was sourced from SongLab. The benchmark includes comparisons of GENERator against Evo2, NT, NT-v2, HyenaDNA, GPN-MSA, CADD, phyloP, and phastCons.
Abouts
The human reference genome data is sourced from the NCBI website.
We have applied minor formatting adjustments to the dataset to facilitate streamlined data analysis.
How to use
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/GenerTeam/variant-effect-prediction.epidemic_sound_effectsnyu_door_opening_surprising_effectivenessThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 484,
"total_frames": 20405,
"total_tasks": 1,
"total_videos": 484,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 3,
"splits": {
"train": "0:484"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/nyu_door_opening_surprising_effectiveness.variant_effect_coding
🧬 BioReasonIncentivizing Multimodal Biological Reasoning within a DNA-LLM Model
Variant Effect Coding Dataset
50,083 core variant entries from GPN-MSA study using ClinVar pathogenic variants and gnomAD benign variants (MAF>5%), split by chromosome (Chr 1-7,9-22,X,Y for train, Chr 8 for test) for pathogenic/benign classification.
Usage
from datasets import load_dataset
dataset = load_dataset("wanglab/variant_effect_coding")
example = dataset["train"][0]… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/variant_effect_coding.epidemic_sound_effects_t5_debiasedsonniss_game_effectsplate_effects
Multi-Source Domain Adaptation for Bioimaging Data (MSCDA-BioIm)
MSCDA-BioIm is a biomedical microscopy benchmark for evaluating test-time and in-context domain adaptation under realistic batch effects. Built from the large-scale JUMP-CP dataset, it targets mechanism-of-action (MoA) classification using five-channel images of compounds associated with eight well-defined MoA classes. The dataset is organized by experimental batches and imaging… See the full description on the dataset page: https://huggingface.co/datasets/anasanchezf/plate_effects.IC-Effectsae-ts-effects
SAE-TS Effects Dataset
This dataset contains pre-computed feature effects for the SAE-TS (SAE-Targeted Steering) method described in our paper Improving Steering Vectors by Targeting Sparse Autoencoder Features.
Contents
The dataset contains two files:
effects_2b.pt: Pre-computed effects for Gemma-2B model
effects_9b.pt: Pre-computed effects for Gemma-9B model
Each file is a PyTorch saved dictionary containing:
features: The steering vectors used (shape: [num_features… See the full description on the dataset page: https://huggingface.co/datasets/schalnev/sae-ts-effects.cagi-variant-effect-glm-tang
GLM-Tang Task 3: CAGI Regulatory Variant Effects
This dataset packages the saturation-mutagenesis MPRA variants used for
Task 3 of Tang et al. The task is zero-shot variant-effect prediction:
compare a reference sequence with a matched single-nucleotide alternate
sequence and test whether the model score tracks the measured regulatory
effect.
Choosing a configuration
Config
Rows
Sequence length
Intended use
paper-230
5,056
230 nt
Official… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/cagi-variant-effect-glm-tang.variant_effect_non_snv
🧬 BioReasonIncentivizing Multimodal Biological Reasoning within a DNA-LLM Model
Variant Effect Coding Non-SNVs Dataset
36,088 core non-SNV entries from ClinVar 2024-02-28 release, filtered for coding variants with ≥2-star review status, using stratified train/test splits for balanced disease representation in pathogenic/benign classification.
Usage
from datasets import load_dataset
dataset = load_dataset("wanglab/variant_effect_non_snv")
example =… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/variant_effect_non_snv.daily-paper-2026-09-04-redundancy-effect-skill-registry-routing
The Redundancy Effect: Decomposing Skill-Registry Size from Duplicate Mass in Router Accuracy of a 2,200-Skill Agent Harness
TL;DR — In a ~2,200-skill agent harness, much of the strict top-1 routing-accuracy loss seen as registries grow is a scoring artifact of near-duplicate skills rather than router degradation: a closed-form symmetry result shows that at fixed quality strict top-1 decays as 1/(m+1) in the number of redundant copies while Recall@k (k >= m+1) is unaffected -… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-09-04-redundancy-effect-skill-registry-routing.ArXiv-10
ArXiv-10
"ArXiv-10" dataset consists of titles and abstracts extracted from 100k scientific papers on ArXiv, covering ten distinct research categories.
These categories includes subfields of computer science, physics, and mathematics.
To ensure consistency and manageability, the dataset consists of precisely 10k samples per category.
This dataset provides a practical resource for researchers and practitioners interested in LLM.
What is different about this dataset is the high… See the full description on the dataset page: https://huggingface.co/datasets/effectiveML/ArXiv-10.dllm-effect-parents-w128-tau2-half-v2
Deprecated — do not use
The payload on this repository's main branch was removed on 2026-09-19.
It used the superseded pre-fix counterfactual dependency construction
(manifest.json SHA-256 678fc4069de11c941120e3bfe431857ea15d70b8a02bde34a1e5091ad82f0888) and is not the corrected
d1-marginal graph.
Use the corrected canonical B64 replacement:
zimplex/dllm-effect-parents-llada2-finemath-half-d1marginal-w128-tau2-b64-v3
(manifest… See the full description on the dataset page: https://huggingface.co/datasets/zimplex/dllm-effect-parents-w128-tau2-half-v2.effective-performance-measurement
Effective Performance Measurement: KPI Extraction Datasets
This dataset repository accompanies the ACL 2026 (Industry Track) paper: "Effective Performance Measurement: Challenges and Opportunities in KPI Extraction from Earnings Calls".
It contains three novel benchmarks and a prediction set designed to evaluate the extraction of Key Performance Indicators (KPIs) from unstructured financial texts, specifically comparing highly regulated SEC filings to conversational earnings calls.… See the full description on the dataset page: https://huggingface.co/datasets/AAU-NLP/effective-performance-measurement.Effective_prompts_for_scraping_AI_data_to_train_another_modelEffective prompts for scraping AI data to train another model !
Ripple_effects_of_prizewinning_across_knowledge_space_and_collaboration_networksnmt-pe-effects
Neural Machine Translation Quality and Post-Editing Performance
This is a repository for an experiment relating NMT quality and post-editing efforts, presented at EMNLP2021 (presentation recording).
Please cite the following paper when you use this research:
@inproceedings{zouhar2021neural,
title={Neural Machine Translation Quality and Post-Editing Performance},
author={Zouhar, Vil{\'e}m and Popel, Martin and Bojar, Ond{\v{r}}ej and Tamchyna, Ale{\v{s}}}… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/nmt-pe-effects.2D-animation-effectsdllm-effect-parents-dreamreasoner-finecode-d1marginal-approx-w128-tau2-b64-v6
DART FineCode deterministic sample — Dream-org/DreamReasoner-8B-Base dependency sidecar
This is the sealed counterfactual dependency release generated
with corrected d1-marginal, approximate window 128, tau 2. It is compatible with block size 64 and pairs only
with zimplex/dllm-dreamreasoner-finecode-b64-v1,
whose manifest SHA-256 is 4a0aece1752c3cd1b3fcd09c9e1690a2e4814ea7853f9a3ec5299c963f2ab9e9.
The sidecar manifest SHA-256 is… See the full description on the dataset page: https://huggingface.co/datasets/zimplex/dllm-effect-parents-dreamreasoner-finecode-d1marginal-approx-w128-tau2-b64-v6.batch-effects-leaderboard-resultsrollouts-molmoact2-500ep-eetc-wrong-taskThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/end-effector-trace-conditioning/rollouts-molmoact2-500ep-eetc-wrong-task.
