datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
soa-15.8-sid-11.5soa-65-sid-17soa-65-sid-15soa-15.8-sid-3.7medical-dialogue-to-soap-summary
Dataset Card for Synthetic Medical Dialogues and SOAP Summaries
Dataset Description
Abstract
This dataset consists of 10,000 synthetic dialogues between a patient and clinician, created using the GPT-4 dataset from NoteChat, based on PubMed Central (PMC) case-reports. Accompanying these dialogues are SOAP summaries generated through GPT-4. The dataset is split into 9250 training, 500 validation, and 250 test entries, each containing a dialogue column, a SOAP… See the full description on the dataset page: https://huggingface.co/datasets/omi-health/medical-dialogue-to-soap-summary.soa-21.9-sid-3.7soa-49.8-sid-46sid-39.4-soa-35.4soa-15.8-sid-26.1soa-49.8-sid-35.6soa-65-sid-25soa-21.9-sid-11.5sid-40.8-soa-49.8SoAyBench
SoAyBench
by WangYC
We've based SoAyBench creation on AMiner. To really understand how well LLMs can use SoAPI, we need to make AMiner's basic SoAPIs available for LLMs to use. We also need a test set made up of academic (question, solution, answer) triplets for checking how they're doing. The tricky part is, academic data keeps changing fast – stuff like info on scholars and their publications. So, keeping a test set with fixed answers is tough.
To tackle this, what we've done is… See the full description on the dataset page: https://huggingface.co/datasets/frederickwang99/SoAyBench.soa-fullThis dataset is a shuffled list of downloadable CC0 image titles and URLs from Smithsonian Open Access.
Some images may be omitted due to limitations or oversights in the preprocessing pipeline, but there's no deliberate curation.
This dataset only contains metadata; a tool like https://github.com/rom1504/img2dataset can be used to download the actual images:
img2dataset --url_list data --output_folder data_files \
--input_format "parquet" --output_format files \
--caption_col "text"… See the full description on the dataset page: https://huggingface.co/datasets/madebyollin/soa-full.soa-28.7-sod-11.5soarm100_dataThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 25,
"total_frames": 11178,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Dangvi/soarm100_data.soa-49.8-sid-30soarm100_data2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 25,
"total_frames": 11150,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Dangvi/soarm100_data2.so_arm101_grab_red_diceThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 100,
"total_frames": 116695,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tatsuyaaaaaaa/so_arm101_grab_red_dice.soa-15.8-sod-3.7so_arm_100_image_rnd_4This dataset was created using LeRobot.
Dataset Description
The SO ARM 100 test dataset series is an example dataset for testing purposes. It is a synthetic dataset from a
high-fidelity pick and place task recorded from an expert policy trained in IsaacLab. The dataset consists of image
observations, joint angle observations, and joint angle actions. There are five datasets in the series with varying
degrees of domain randomization, as shown below. In addition to the state… See the full description on the dataset page: https://huggingface.co/datasets/Nfiniteai/so_arm_100_image_rnd_4.soa-21.9-sod-11.5soa-15.8-sod-11.5soa-28.7-sid-3.7SO_ARM101_6000soa-full-florence2
Smithsonian Open Access Dataset with Florence-2 Caption
日本語はこちら
This dataset is made of soa-full.
soa-full is an CC-0 image dataset from Smithsonian Open Access. However, the dataset does not contain the image caption.
Therefore, we caption the images by Florence 2.
Usage
from datasets import load_dataset
dataset = load_dataset("aipicasso/soa-full-florence2")
Intended Use
Research Vision & Language
Develop text-to-image model or image-to-text model.… See the full description on the dataset page: https://huggingface.co/datasets/aipicasso/soa-full-florence2.soa-21.9-sod-3.7receiver_Lhandover_soapbottle_sceneC_subscene1_qi_V2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_tb4",
"total_episodes": 7,
"total_frames": 7350,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:7"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/MoAIBo/receiver_Lhandover_soapbottle_sceneC_subscene1_qi_V2.sid-31.6-soa-21.9
