datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
job_assistantThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 109,
"total_frames": 51991,
"total_tasks": 1,
"total_videos": 218,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:109"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/wockd/job_assistant.so100_table_assistantThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 30,
"total_frames": 26243,
"total_tasks": 4,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masato-ka/so100_table_assistant.empowerment_v2_assistant_ds_1_100customer_assistant
Dataset Card for customer_assistant
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
This dataset contains:
A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/argilla/customer_assistant.job-search-assistant-agent-tracewhere-assistants-part-ways
Where Assistants Part Ways
This is the complete, offline-reproducible artifact for the paper Measuring
Moral Disagreement Between AI Assistants. It
contains the 630-item evaluation corpus, all final recorded annotations and
responses from seven models, collection metadata, corrected parse fields, and
code that reproduces the paper's full result object.
The artifact is a bounded record of one study run. Its labels are model-panel
labels, not independent human judgments, and its… See the full description on the dataset page: https://huggingface.co/datasets/ephipi/where-assistants-part-ways.assistant-axis-qwen-3-14b
Assistant Axis — Qwen3-14B
Pre-computed Assistant Axis and per-role persona vectors for Qwen/Qwen3-14B,
produced with the pipeline from safety-research/assistant-axis.
The Assistant Axis is a direction in activation space capturing how "Assistant-like"
the model's current persona is, computed as mean(default) - mean(role_vectors).
Contents
axis.pt — the final Assistant Axis, shape (n_layers, hidden_dim).
vectors/ — per-role mean activation vectors (one .pt per… See the full description on the dataset page: https://huggingface.co/datasets/dungnv/assistant-axis-qwen-3-14b.synthesized-coding-assistant-dataset
Synthesized Coding Assistant Dataset
Overview
Coding assistants are increasingly used for real-world software engineering workflows. However, there are relatively few datasets that closely resemble how such assistants operate in practice.
Many existing coding datasets are based on single-turn or single-iteration tasks, where a model receives one coding request and directly produces an answer or patch. In contrast, practical coding assistants often work through… See the full description on the dataset page: https://huggingface.co/datasets/squeezebits/synthesized-coding-assistant-dataset.job-search-assistant-agent-traceasia-who-pharmaceutical-technicians-and-assistants
Pharmaceutical Technicians and Assistants (number) | Asia (WHO GHO)
🌏 182 observations · 28 Asia countries · 1983–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 182 observations of Pharmaceutical Technicians and Assistants (number) data across 28 Asia countries, spanning 1983–2024, covering 1 distinct indicators.
About the source
Source: WHO Global Health Observatory
Publisher: World Health Organization
License:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-pharmaceutical-technicians-and-assistants.32b_swebv_26rollouts_patch_verifier_thought_action_wo_ASSISTANTAssistantz8
OpenAssistant Conversations Dataset (OASST1)
Dataset Summary
In an effort to democratize research on large-scale alignment, we release OpenAssistant
Conversations (OASST1), a human-generated, human-annotated assistant-style conversation
corpus consisting of 161,443 messages in 35 different languages, annotated with 461,292
quality ratings, resulting in over 10,000 fully annotated conversation trees. The corpus
is a product of a worldwide crowd-sourcing effort… See the full description on the dataset page: https://huggingface.co/datasets/Sellopale/Assistantz8.evalap-assistant-lasuite-tools-comparison-v2-101
Assistant LASuite Tools Comparison v2 (ID: 101)
Generated locally via notebook and pushed to EvalAP
Overview
This dataset contains 13 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Metrics: answer_relevancy, faithfulness, judge_exactness, judge_notator, judge_precision
Scores
MFS_questions_v01
model
answer_relevancy
judge_exactness
judge_notator
judge_precision
Mistral Medium (With WEB)(no sources)
0.95 ±… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-assistant-lasuite-tools-comparison-v2-101.home-assistant-local-llm-voice-benchmark
Home Assistant Local LLM Voice Benchmark
Per-model tool-call accuracy and component latency for running a Home Assistant voice assistant against local LLMs.
Measured per-model tool-call accuracy and component latency for running a Home Assistant voice assistant against local LLMs.
Broken out by pipeline component rather than reported as one opaque round trip, so you can tell whether your latency is wake-word, speech-to-text, the model, or text-to-speech before you go optimizing… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/home-assistant-local-llm-voice-benchmark.agreement_assistant_ds_fagreement_assistant_dsrag_audit_assistant_samples_clusters_analysiseurope-who-pharmaceutical-technicians-and-assistants
Pharmaceutical Technicians and Assistants (number) | Europe (WHO GHO)
🇪🇺 29 observations · 12 Europe countries · 2000–2024 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 29 observations of Pharmaceutical Technicians and Assistants (number) data across 12 Europe countries, spanning 2000–2024, covering 1 distinct indicators.
About the source
Source: WHO Global Health Observatory
Publisher: World Health Organization… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-who-pharmaceutical-technicians-and-assistants.eval_job_assistant2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 2,
"total_frames": 1639,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/wockd/eval_job_assistant2.scientific-literature-research-assistant-datajob_assistant_12_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 101,
"total_frames": 41734,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/wockd/job_assistant_12_1.eval_job_assistant_12_1_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 3,
"total_frames": 3612,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/wockd/eval_job_assistant_12_1_3.eval_job_assistant6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 749,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/wockd/eval_job_assistant6.eval_job_assistant_12_1_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 3,
"total_frames": 4223,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/wockd/eval_job_assistant_12_1_4.eval_job_assistant_12_1_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 774,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/wockd/eval_job_assistant_12_1_2.ru_virtual_assistant_chatgpt_distill
📊 Virtual Assistant Queries Dataset (Russian, Synthetic, 100K)
Описание
Этот датасет содержит 100,000 синтетически сгенерированных пользовательских запросов к виртуальному ассистенту на русском языке. Он предназначен для задач анализа пользовательского опыта, обработки естественного языка и предсказательного моделирования.
Каждая запись представляет собой реалистичный запрос пользователя, категорию запроса, устройство, с которого он был сделан, и оценку качества… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/ru_virtual_assistant_chatgpt_distill.eval_job_assistant_12_1_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 3,
"total_frames": 1830,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/wockd/eval_job_assistant_12_1_1.eval_job_assistant5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 3,
"total_frames": 2538,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/wockd/eval_job_assistant5.Assistantsellopale53oi
OpenAssistant Conversations Dataset (OASST1)
Dataset Summary
In an effort to democratize research on large-scale alignment, we release OpenAssistant
Conversations (OASST1), a human-generated, human-annotated assistant-style conversation
corpus consisting of 161,443 messages in 35 different languages, annotated with 461,292
quality ratings, resulting in over 10,000 fully annotated conversation trees. The corpus
is a product of a worldwide crowd-sourcing effort… See the full description on the dataset page: https://huggingface.co/datasets/Sellopale/Assistantsellopale53oi.posthog-assistant-isolation-results-001Normalized run events for dataset viewer compatibility.
