integration
mala-monolingual-integration
MaLA Corpus: Massive Language Adaptation Corpus
This is the noisy version that integrates texts from different sources.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-integration.faiss-integration-testaxolotl-ota-sft-integration
Axolotl ⇄ OpenThoughts-Agent SFT-backend integration — overview
Status: COMPLETE + merged to penfever/working (merge 02d676d0, 2026-07-02).
Where it ran: TACC Vista (GH200, aarch64), conda env sft-axolotl.
One-line result: OpenThoughts-Agent can now run SFT through axolotl (--sft_backend axolotl)
as a drop-in alternative to LLaMA-Factory, with the delphi chat-template masking validated
on the real delphi path (jinja-as-ground-truth: train == serve).
1. What was done… See the full description on the dataset page: https://huggingface.co/datasets/laion/axolotl-ota-sft-integration.test-braindecode-integration
EEG Dataset
This dataset was created using braindecode, a library for deep learning with EEG/MEG/ECoG signals.
Dataset Information
Number of recordings: 1
Number of channels: 26
Sampling frequency: 250.0 Hz
Data type: Windowed (from Epochs object)
Number of windows: 48
Total size: 0.04 MB
Storage format: zarr
Usage
To load this dataset:
from braindecode.datasets import BaseConcatDataset
# Load dataset from Hugging Face Hub
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Kkuntal990/test-braindecode-integration.Renewable-Energy-Integration-Simscapeintegration-Data-Collection-Dataset
integration Data Collection Dataset
Merged LeRobot v3.0 dataset: 100 episodes, 26,953 frames, 30 FPS, robot type so_follower. Two 640×480 camera streams (front and wrist) and six-dimensional actions/state are preserved.
All episodes are in the train split. Episodes and global frame indices are continuous. Task labels are preserved exactly and mapped into one shared task table. No episodes were deduplicated or dropped. Original video files are copied without re-encoding;… See the full description on the dataset page: https://huggingface.co/datasets/Hailey-5-2026/integration-Data-Collection-Dataset.
