datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Motion-o-MCoT-PLM-motion-keyframes
Motion-o-MCoT (PLM + motion keyframes)
Subset of STGR: STR_plm_rdcap rows with <motion in reasoning_process, plus sharded keyframes under videos/stgr/plm/kfs/.
Train split: 3,168 examples (see export_manifest.json in the repo for exact export stats).
Keyframes: JPEGs are stored under shard subfolders (e.g. videos/stgr/plm/kfs/plm_0150/…) so each directory stays under Hugging Face file-count limits. Each key_frames[].path in the JSON is relative to videos/stgr/plm/kfs/ (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/bishoygaloaa/Motion-o-MCoT-PLM-motion-keyframes.dataset-ohada-droit-commercial-general-echantillon
Dataset OHADA — Droit Commercial Général (AUDCG) — Échantillon
Description
Échantillon de 10 entrées extraites d'un dataset de fine-tuning juridique en cours de conception, portant sur l'Acte Uniforme relatif au Droit Commercial Général (AUDCG) — le texte fondamental du statut du commerçant, des actes de commerce, de la preuve et de la prescription en matière commerciale dans l'espace OHADA (Organisation pour l'Harmonisation en Afrique du Droit des Affaires — 17… See the full description on the dataset page: https://huggingface.co/datasets/Bisilivan/dataset-ohada-droit-commercial-general-echantillon.bist-klines
bist-klines
Borsa İstanbul OHLCV mumları — Kronos
fine-tune'u için düz dizi formatında hazırlanmış hâli.
Kaynak: Yahoo Finance (<SEMBOL>.IS, auto_adjust=True).
193 seri / 1.019.897 mum — BIST-100 evreninden ~97 sembol saatlik (son ~2 yıl)
ve ~96 sembol günlük (tam geçmiş).
Dosyalar
dosya
şekil
içerik
bars.npy
[N, 6] float32
open, high, low, close, volume, amount (= close × volume)
stamps.npy
[N, 5] float32
minute, hour, weekday, day, month
ts.npy… See the full description on the dataset page: https://huggingface.co/datasets/zubziretta/bist-klines.Corpus_biscegliese
📚 Corpus Biscegliese — Dialect Dataset (JSONL)
The Corpus Biscegliese is a JSONL dataset designed for training, evaluating, and studying language models specialized in the Biscegliese dialect, a local linguistic variety spoken in Bisceglie (Apulia, Southern Italy).
This dataset was used to train the model:👉 Biscegliese-Qwen2.5-GGUFhttps://huggingface.co/vamoruso/biscegliese-qwen2.5-gguf
📌 Dataset Structure
The dataset is provided as a JSON Lines (JSONL)… See the full description on the dataset page: https://huggingface.co/datasets/vamoruso/Corpus_biscegliese.repro-probabilistic-bisection-algorithm-provably-achieves-exponential-convergence-traces
Agent traces
Agent sessions published from a Trackio Logbook.
hard-negatives-traversal
If our work was helpful conside citing us ☺️
@misc{sinha2025bicaeffectivebiomedicaldense,
title={BiCA: Effective Biomedical Dense Retrieval with Citation-Aware Hard Negatives},
author={Aarush Sinha and Pavan Kumar S and Roshan Balaji and Nirav Pravinbhai Bhatt},
year={2025},
eprint={2511.08029},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/abs/2511.08029},
}
bishop-chess-dataset
Bishop Chess Concepts Dataset
Cleaned, training-ready chess text focused on bishop concepts and strategy, distilled from the
Waterhorse/chess_data dataset.
At a glance
15,677 records · ~18.59M tokens (measured with the GPT-4 cl100k BPE).
Game-level-disjoint train/test split (no game leaks across splits), seed 20260708, test fraction 0.02.
Split
Records
Tokens
train
15,363
18,152,676
test
314
433,026
Sources (origin):
Source
Train records… See the full description on the dataset page: https://huggingface.co/datasets/pkloats/bishop-chess-dataset.medical-instruction-120k
What is the Dataset About?🤷🏼♂️
The dataset is useful for training a Generative Language Model for the Medical application and instruction purposes, the dataset consists of various thoughs proposed by the people [mentioned as the Human ] and there responses including Medical Terminologies not limited to but including names of the drugs, prescriptions, yogic exercise suggessions, breathing exercise suggessions and few natural home made prescriptions.
How the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/bisonnetworking/medical-instruction-120k.bist-tool-sft
BIST Tool SFT
Fine-tuning data for market-data tool calling (Gemma 4 style), Turkish and English.
~1200 examples, 23 tools, BIST stock-flow heavy. Includes a large manual curated block of natural TR questions (bilanço/kârlılık/KAP/sermaye/teknik/tarama) with agent-style multi-tool plans; EUREN always multi-market search (never FX).
Files
File
Format
gemma4_sft.jsonl
Alpaca: instruction, input, output
gemma4_sft_text.jsonl
Full chat as { "text": "..."… See the full description on the dataset page: https://huggingface.co/datasets/ardakalayci/bist-tool-sft.bishop-tasks-v1
bishop-tasks-v1
43 exact pattern-recognition tasks for the
bishop-env RL
environment, on the topics of Pattern Recognition and Machine Learning (Bishop): probability
and Bayes, information theory, linear regression and ridge, naive Bayes, Bernoulli mixtures and
EM, k-means, conjugate priors, and d-separation in directed graphical models.
field
meaning
task_id
bi-000 … bi-042
category
probability / information-theory / regression / classification / bayesian-inference… See the full description on the dataset page: https://huggingface.co/datasets/eltociear/bishop-tasks-v1.dataset-ohada-droit-societes-echantillon
Dataset OHADA — Droit des sociétés commerciales (AUSCGIE) — Échantillon
Description
Échantillon de 10 entrées extraites d'un dataset de fine-tuning juridique de 40 entrées portant sur les dispositions générales de l'Acte Uniforme relatif au Droit des Sociétés Commerciales et du Groupement d'Intérêt Économique (AUSCGIE), le texte fondamental du droit des sociétés dans l'espace OHADA (Organisation pour l'Harmonisation en Afrique du Droit des Affaires — 17 pays… See the full description on the dataset page: https://huggingface.co/datasets/Bisilivan/dataset-ohada-droit-societes-echantillon.2hop-citation-graphs
If our work was helpful conside citing us ☺️
@misc{sinha2025bicaeffectivebiomedicaldense,
title={BiCA: Effective Biomedical Dense Retrieval with Citation-Aware Hard Negatives},
author={Aarush Sinha and Pavan Kumar S and Roshan Balaji and Nirav Pravinbhai Bhatt},
year={2025},
eprint={2511.08029},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/abs/2511.08029},
}
PMID_CITED_forKGDataset collected from PGB: A PubMed Graph Benchmark for Heterogeneous Network Representation Learning
Description :
inbound_citation: List List of PMID that cites the paper
outbound_citation: List References of the paper
PMID : Pubmed ID
bisac-topicsforbidden-bishop-effectPlease message or write for removal.
No infrigement is intended. All the content was scraped from the open-web and is structured to allow for the next generation of academic LLM data generation tasks.
lalainy__ECE-PRYMMAL-YL-0.5B-SLERP-BIS-V1-details
Dataset Card for Evaluation run of lalainy/ECE-PRYMMAL-YL-0.5B-SLERP-BIS-V1
Dataset automatically created during the evaluation run of model lalainy/ECE-PRYMMAL-YL-0.5B-SLERP-BIS-V1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lalainy__ECE-PRYMMAL-YL-0.5B-SLERP-BIS-V1-details.bismillaBispatialstructure-Bigraph-Modelbishedata_decret_marche_publique_062023_BISBisha_University_QA
