datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mala-monolingual-integration
MaLA Corpus: Massive Language Adaptation Corpus
This is the noisy version that integrates texts from different sources.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-integration.faiss-integration-testaxolotl-ota-sft-integration
Axolotl ⇄ OpenThoughts-Agent SFT-backend integration — overview
Status: COMPLETE + merged to penfever/working (merge 02d676d0, 2026-07-02).
Where it ran: TACC Vista (GH200, aarch64), conda env sft-axolotl.
One-line result: OpenThoughts-Agent can now run SFT through axolotl (--sft_backend axolotl)
as a drop-in alternative to LLaMA-Factory, with the delphi chat-template masking validated
on the real delphi path (jinja-as-ground-truth: train == serve).
1. What was done… See the full description on the dataset page: https://huggingface.co/datasets/laion/axolotl-ota-sft-integration.test-braindecode-integration
EEG Dataset
This dataset was created using braindecode, a library for deep learning with EEG/MEG/ECoG signals.
Dataset Information
Number of recordings: 1
Number of channels: 26
Sampling frequency: 250.0 Hz
Data type: Windowed (from Epochs object)
Number of windows: 48
Total size: 0.04 MB
Storage format: zarr
Usage
To load this dataset:
from braindecode.datasets import BaseConcatDataset
# Load dataset from Hugging Face Hub
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Kkuntal990/test-braindecode-integration.Renewable-Energy-Integration-Simscapeintegration-Data-Collection-Dataset
integration Data Collection Dataset
Merged LeRobot v3.0 dataset: 100 episodes, 26,953 frames, 30 FPS, robot type so_follower. Two 640×480 camera streams (front and wrist) and six-dimensional actions/state are preserved.
All episodes are in the train split. Episodes and global frame indices are continuous. Task labels are preserved exactly and mapped into one shared task table. No episodes were deduplicated or dropped. Original video files are copied without re-encoding;… See the full description on the dataset page: https://huggingface.co/datasets/Hailey-5-2026/integration-Data-Collection-Dataset.rollout-integration-20260918This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Hailey-5-2026/rollout-integration-20260918.cdg-tech-components-integration-dataset
Character Pool Dataset: 81 Characters - Generated by Conversation Dataset Generator
This dataset was generated using the Conversation Dataset Generator script available at https://cahlen.github.io/conversation-dataset-generator/.
Generation Parameters
Number of Conversations Requested: 1000
Number of Conversations Successfully Generated: 1000
Total Turns: 5132
Model ID: meta-llama/Meta-Llama-3-8B-Instruct
Generation Mode:
Topic & Scenario
Initial Topic:… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/cdg-tech-components-integration-dataset.llama_index_integration_dataGEM-Integration
GEM-Integration
This dataset is the redistribution-approved data package for the GEM-Atlas
research platform. It contains harmonized tables derived from Kochanowski et al.
(2021), explicit condition splits, schemas, checksums, and metadata-only source
manifests.
Included values
16 EV3/EV4 strain-effector conditions.
Source-derived 13C-MFA flux estimates from EV3.
Absolute intracellular metabolite measurements from EV4.
The separate EV2 protein series. It is not… See the full description on the dataset page: https://huggingface.co/datasets/lhallee/GEM-Integration.single-cell-integrationclinical-meaning-integration-fragmentation-analysis-v0.1What this dataset tests
Whether an intelligence system can evaluatethe integrity of a patient’s meaning-making processduring illness and recovery.
Required outputs
coherence integrity score
fragmentation markers
denial indicators
adaptive reframing presence
narrative stability index
meaning failure mode
Use case
Second layer of the Healing Narrative Coherence Corpus.
production-integration-recipes
SHAR Production Integration Recipes
Five runnable production-delivery recipes from SHAR Production — https://sharprod.com/
The dataset contains bilingual recipe metadata and synthetic examples only. Dataset records and synthetic fixtures are CC BY 4.0; linked source code and documentation are MIT. No client data or media is included. Codex assisted with implementation and validation; SHAR Production is the accountable publisher.
Charter_values_citizenship_integration
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/19058
Description
The integration agreement form prepared for signing the pact between foreign and state, in addition to providing the alien's commitments, indicates, the statement by the person concerned, to adhere to the Charter of the values of citizenship and integration of the decree of the Minister of 23 April 2007, pledging to respect its principles. The Charter of citizenship and integration… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Charter_values_citizenship_integration.chinese-dataset-integrationdepthai_integration_submissio33422343223nThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 2,
"total_frames": 163,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kma7h/depthai_integration_submissio33422343223n.depthai_integration_submissionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 5,
"total_frames": 2458,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kma7h/depthai_integration_submission.integration_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
26
],
"names": [
"dg5f_0.pos",
"dg5f_1.pos",
"dg5f_2.pos",
"dg5f_3.pos",
"dg5f_4.pos",
"dg5f_5.pos"… See the full description on the dataset page: https://huggingface.co/datasets/Kaz55/integration_test.depthai_integration_submissio3342233223nThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 2,
"total_frames": 331,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kma7h/depthai_integration_submissio3342233223n.depthai_integration_submissio22nThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 2,
"total_frames": 184,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kma7h/depthai_integration_submissio22n.depthai_integration_submissio2nThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 2,
"total_frames": 243,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kma7h/depthai_integration_submissio2n.langchain-python-integrationsZapier-Integrations-260825
Zapier Apps/Integrations
Extracted: Aug 26, 2025
simpleqa_verified-integration-tests
simpleqa_verified-integration-tests Evaluation Results
Eval created with evaljobs.
This dataset contains evaluation results for the model hf-inference-providers/openai/gpt-oss-20b:cheapest using the eval script simpleqa_verified_custom.py.
To browse the results interactively, visit this Space.
How to Run This Eval
pip install git+https://github.com/dvsrepo/evaljobs.git
export HF_TOKEN=your_token_here
evaljobs dvilasuero/simpleqa_verified-integration-tests \… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/simpleqa_verified-integration-tests.omg-game-api-integration-docs
OMG Game Aggregation Platform — API Integration Documentation
About OMG
OMG is a professional game aggregation platform that provides third-party merchants with a unified API to access a wide variety of game content including Slots, Fishing Games, Mass Table Games, and Mini Games.
Official documentation: docs.omgapi.cc
Dataset Description
This dataset contains the complete API integration documentation for the OMG game aggregation platform… See the full description on the dataset page: https://huggingface.co/datasets/omgapi/omg-game-api-integration-docs.test-integrationafrica-mauritius-budget-data-2016-2017-ministry-of-social-integration-and-e-5cf52e85
Budget Data 2016 2017 Ministry of Social Integration and E | Africa (MDPA)
87 rows - 1 Africa country/area - 2016-2017 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 87 rows from MDPA, covering Budget Data 2016 2017 Ministry of Social Integration and E. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-mauritius-budget-data-2016-2017-ministry-of-social-integration-and-e-5cf52e85.eval_LAVLA_S1_VIII_reverse_integrationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 2,
"total_frames": 159,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Alkatt/eval_LAVLA_S1_VIII_reverse_integration.crewai-nexaapi-integrationIyBDcmV3QUkgKyBOZXhhQVBJIEludGVncmF0aW9uCgo+IEJ1aWxkIG9ic2VydmFibGUgQUkgYWdlbnRzIHdpdGggJDAuMDAzIGltYWdlIGdlbmVyYXRpb24uIFBhaXIgQ3Jld0FJIHdpdGggTmV4YUFQSSBmb3IgdGhlIHVsdGltYXRlIEFJIGFnZW50IHN0YWNrLgoKWyFbTmV4YUFQSV0oaHR0cHM6Ly9pbWcuc2hpZWxkcy5pby9iYWRnZS9OZXhhQVBJLTU2JTJCJTIwTW9kZWxzLWJsdWUpXShodHRwczovL25leGEtYXBpLmNvbSkKWyFbUHlQSV0oaHR0cHM6Ly9pbWcuc2hpZWxkcy5pby9weXBpL3YvbmV4YWFwaSldKGh0dHBzOi8vcHlwaS5vcmcvcHJvamVjdC9uZXhhYXBpKQpbIVtucG1dKGh0dHBzOi8vaW1nLnNoaWVsZHMuaW8vbnBtL3YvbmV4YWFwaSldKGh… See the full description on the dataset page: https://huggingface.co/datasets/nickyni/crewai-nexaapi-integration.integration_test_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
26
],
"names": [
"dg5f_0.pos",
"dg5f_1.pos",
"dg5f_2.pos",
"dg5f_3.pos",
"dg5f_4.pos",
"dg5f_5.pos"… See the full description on the dataset page: https://huggingface.co/datasets/Kaz55/integration_test_2.
