datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cav3_t-type_calcium_channels_butkiewicz
Dataset Details
Dataset Description
This dataset was initially curated from HTS data at the PubChem database.
The curation process is documented in Butkiewicz et al.
Primary screening with AID 449739 identified inhibitors of Cav3 T-type calcium channels.
Four follow-up screens were performed to confirm inhibitory effects on smaller sets of compounds
involving AID 493021, AID 493022, AID 493023, and AID 493041.
AID 489005 was performed as counter screen validating active… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/cav3_t-type_calcium_channels_butkiewicz.train_800_sparse__bbox__separate_channel__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__bbox__separate_channel__sim__all_cameras__live.telegram-public-channels-2026-W36
Telegram public channels: a 5,177-channel snapshot with topics and a recommendation graph
A single snapshot of 5,177 public Telegram channels, measured on 31 August 2026, together with the
recommendation graph Telegram itself exposes between them.
Public Telegram data is hard to get in tabular form. The libraries that read it need a phone number
and a user session, and the two Telegram datasets that rank on Kaggle today are both from 2021.
This is a current measurement… See the full description on the dataset page: https://huggingface.co/datasets/starnikovoleg/telegram-public-channels-2026-W36.Top_100_YouTube_ChannelsTop 100 YouTube Channels (Updated Monthly) - ytRank.com
Discover the latest ranking of the top 100 YouTube channels, updated every month. This dataset provides an up-to-date list of channels sorted by subscriber count, offering insights into the most popular creators and trending content on the platform. Stay informed about the biggest names in the YouTube community with our accurate and regularly refreshed data.
train_800_sparse__mask__separate_channel__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__separate_channel__sim__all_cameras__live.train_800_dense__point__separate_channel__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_dense__point__separate_channel__sim__all_cameras__live.zaduha-channel-history
Zaduha Chat — Telegram discussion archive (2022-08-18–2026-08-18)
Public release / Публічний реліз. This repository is intentionally public,
but it contains user-generated content and participant display names that may
constitute personal data. Public availability is not evidence that every
participant consented to research reuse, profiling, or republication.
Репозиторій навмисно опубліковано, але він містить користувацький контент та
імена учасників, які можуть бути… See the full description on the dataset page: https://huggingface.co/datasets/transhumanist-already-exists/zaduha-channel-history.train_800_complex__bbox__separate_channel__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_complex__bbox__separate_channel__sim__all_cameras__live.train_800_complex__point__separate_channel__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_complex__point__separate_channel__sim__all_cameras__live.telegram-channel-dataset
Telegram AI Image Dataset — Cleaned for VLM LoRA Training
A cleaned dataset of 1,050 AI-generated images with their generation prompts, collected from a Chinese Telegram channel focused on GPT-Image-2 prompt engineering.
Each image is paired with a structured generation prompt in Chinese/English. Prompts have been cleaned of channel boilerplate — no bot instructions, hashtags, source credits, model name prefixes, emoji title lines, or channel footer ads.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/GCStream/telegram-channel-dataset.train_800_sparse__point__separate_channel__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__point__separate_channel__sim__all_cameras__live.train_800_dense__bbox__separate_channel__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_dense__bbox__separate_channel__sim__all_cameras__live.train_800_dense__mask__separate_channel__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_dense__mask__separate_channel__sim__all_cameras__live.train_800_complex__mask__separate_channel__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_complex__mask__separate_channel__sim__all_cameras__live.2026.RA.NBS-Channel-Comparison
2026.RA.NBS-Channel-Comparison
A preregistered negative on representation-level conditioning: information that provably improves negotiation outcomes when delivered as prompt text has zero effect when delivered as distilled virtual tokens — even though the virtual-token channel closes about three quarters of the distributional gap it was trained to close.
What the experiment was
Six copies of Qwen/Qwen3-8B negotiate a multi-issue deal. Each seat holds a private… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.NBS-Channel-Comparison.africa-synth-retail-and-ecommerce-cross-channel-sales-data-nigeria
Cross Channel Sales Data | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-retail-and-ecommerce-cross-channel-sales-data-nigeria.channel-metadataDataset containing video metadata from a few tech channels, i.e.
James Briggs
Yannic Kilcher
sentdex
Daniel Bourke
AI Coffee Break with Letitia
Alex Ziskind
political-whatsapp-channels
Political party WhatsApp channels
An open dataset of European political parties' presence on WhatsApp, with the
public follower counts of their Channels tracked over time. Published under
CC BY 4.0 by WhaTools (https://wha.tools).
Parties tracked: 415 (80 with a WhatsApp Channel, 80 with at least one follower reading)
Follower readings: 967
Live source page: https://wha.tools/political-parties-whatsapp-channels
Canonical repo (JSON + geography + full history):… See the full description on the dataset page: https://huggingface.co/datasets/WhaTools/political-whatsapp-channels.Wireless-Channel-Parameter-Estimation-Datasetiras-chopped-photometric-channel-survey
IRAS Chopped Photometric Channel Survey
Raster-scan maps from the IRAS Chopped Photometric Channel at 50 and 100 microns, provided as RAW and CLEAN products.
Data structure
Each configuration contains one source-named fixed-size-list column. One Parquet row is one FITS axis-1 scanline; row order runs through axis 2 and then axis 3. Reshape the flattened column to (NAXIS3, NAXIS2, NAXIS1) to recover the two-plane image. Axis lengths vary by observation and are… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/iras-chopped-photometric-channel-survey.news_channel_ordinal
Dataset Card for "news_channel_ordinal"
More Information needed
hidden-channel-405b-activationsrutube-channels
Dataset Card for Rutube channels
Dataset Summary
This dataset was scraped from channel pages on the Russian video-sharing platform Rutube. It includes all information from the channel card. The dataset was collected by processing 36 million channels, starting from the first one. At the time the dataset was collected, it is assumed that these were all the channels available on this platform. Some fields may be empty, but the string is expected to contain some data, empty… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rutube-channels.potassium_ion_channel_kir2_1_butkiewicz
Dataset Details
Dataset Description
The Kir2.1 inward-rectifier potassium ion channel is
a target in the treatment of cardiovascular, neurological, renal and
metabolic disorders. Primary assay AID 1672. Validation screens AID
2032 and AID 463252. Counter screens AID 2105, AID 2345, AID 2236, and
AID 2329. The final set of 172 active compounds was constructed
subtracting the actives in AID 2105, AID 2345, AID 2236, and AID 2329
from the molecules found active in both, AID… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/potassium_ion_channel_kir2_1_butkiewicz.lerobot_teleop_aux_channel_max_100This dataset was created using LeRobot.
Dataset Description
This was a test of a minimal "robot": teleop and robot are each a single digital channel. This caused lerobot training (using ACT policy) to fail, giving an error about multiplying matrices with incorrect dimensions. Conclusion: I don't think lerobot supports single degree of freedom robots.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aux_channel_follower"… See the full description on the dataset page: https://huggingface.co/datasets/primordial-spork/lerobot_teleop_aux_channel_max_100.kcnq2_potassium_channel_butkiewicz
Dataset Details
Dataset Description
This dataset was initially curated from HTS data at
the PubChem database. Details are reported by Butkiewicz et al. (2013).
Primary screen AID 2239, AID 2287 validated active compounds to be
potentiators. Counter screens are AID 2282, AID 2283, and AID 2558.
Final set of 213 active compounds was acquired by removing the active
compounds of AID 2282, AID 2283 and AID 2558 from the confirmatory
screen active set of compounds (AID 2287).… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/kcnq2_potassium_channel_butkiewicz.2026-03-09-no-cap-channel-repaired-dev-checkThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 38,
"total_frames": 10210,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 200,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:38"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lyl472324464/2026-03-09-no-cap-channel-repaired-dev-check.channel-metadatadead-channel-results
dit-score results
Raw per-pair fidelity scores for quantized diffusion transformers, from the dit-score harness.
Study 1 (2026-07-16). Krea 2 Turbo, five quant formats vs BF16 reference, RTX 4090, 1024px, seed-locked, n=96 pairs per format. LPIPS (alex) + PSNR + ImageReward delta. Method, sample grids and analysis in the GitHub README.
mining_massive_data_N6_channel11
