CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-llm-leaderboard-old /details_one-man-army__UNA-34Beagles-32K-bf16-v1 Dataset Card for Evaluation run of one-man-army/UNA-34Beagles-32K-bf16-v1 Dataset automatically created during the evaluation run of model one-man-army/UNA-34Beagles-32K-bf16-v1 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_one-man-army__UNA-34Beagles-32K-bf16-v1.0 likes1.8k downloads3y agoHugging Face02unavailableshorts /Videosvideon<1K0 likes1k downloads4mo agoHugging Face03open-llm-leaderboard-old /details_one-man-army__una-neural-chat-v3-3-P2-OMA Dataset Card for Evaluation run of one-man-army/una-neural-chat-v3-3-P2-OMA Dataset automatically created during the evaluation run of model one-man-army/una-neural-chat-v3-3-P2-OMA on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_one-man-army__una-neural-chat-v3-3-P2-OMA.0 likes958 downloads3y agoHugging Face04ttgeng233 /UnAV-100text10K<n<100K0 likes908 downloads9mo agoHugging Face05Salesforce /FaithEval-unanswerable-v1.0 FaithEval FaithEval is a new and comprehensive benchmark dedicated to evaluating contextual faithfulness in LLMs across three diverse tasks: unanswerable, inconsistent, and counterfactual contexts. [Paper] FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows", ICLR 2025, https://arxiv.org/abs/2410.03727 [Code and Detailed Instructions] https://github.com/SalesforceAIResearch/FaithEval Disclaimer and Ethical Considerations… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/FaithEval-unanswerable-v1.0.textquestion-answering1K<n<10K5 likes751 downloads2y agoHugging Face06open-llm-leaderboard-old /details_fblgit__UNA-SOLAR-10.7B-Instruct-v1.0 Dataset Card for Evaluation run of fblgit/UNA-SOLAR-10.7B-Instruct-v1.0 Dataset automatically created during the evaluation run of model fblgit/UNA-SOLAR-10.7B-Instruct-v1.0 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_fblgit__UNA-SOLAR-10.7B-Instruct-v1.0.0 likes586 downloads3y agoHugging Face07una-auxme /eval_ROS2SmolVLAThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "ros2", "total_episodes": 244, "total_frames": 333872, "total_tasks": 5, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:244" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/una-auxme/eval_ROS2SmolVLA.tabularrobotics100K<n<1M1 likes497 downloads3mo agoHugging Face08ines-besrour /unarxive_2024 Example Preview Here is a small excerpt of the dataset format:📄 preview.jsonl Dataset Card for UnarXive 2024 UnarXive 2024 is a large-scale, structured dataset of 2.3 million full-text arXiv papers (1991–2024), processed for use in NLP and information retrieval tasks. Each paper is provided in a structured JSONL format. Features 2.28 million structured papers across physics, CS, mathematics, and other fields Logical section structure (Introduction, Methods… See the full description on the dataset page: https://huggingface.co/datasets/ines-besrour/unarxive_2024.1M<n<10M3 likes484 downloads1y agoHugging Face09KhalfounMehdi /gametilenet-sprites-unannotatedimage1K<n<10K0 likes471 downloads2mo agoHugging Face10GIL-UNAM /SpanishParaphraseCorporaSpanish Paraphrase Corpora. This is a dataset of a total of 6896 pairs of sentences, 314 paraphrased pairs of sentences and 6558 non-paraphrased pairs of sentences.'', MICAI 2020: Advances in Computational Intelligence pp 214–223.feature-extractionn<1K1 likes465 downloads3y agoHugging Face11dianavdavidson /iv_speaker_disjoint_sociodem_unaware_dsaudio100K<n<1M0 likes444 downloads1mo agoHugging Face12howey /unarXive Dataset Card for "unarXive" More Information needed text1M<n<10M1 likes385 downloads3y agoHugging Face13una-auxme /ROS2SmolVLA_ur10e_no_joints_crop_pick_placeThis is the training dataset for our ROS2SmolVLA project It was recorded by teleoperation of our UR10e lightweight industrial robot through our ROS2SmolVLA setup utilizing LeRobot. The action space is: "linear_x.vel", "linear_y.vel", "linear_z.vel", "angular_x.vel", "angular_y.vel", "angular_z.vel", "gripper.pos" The observation space is: "pose.x", "pose.y", "pose.z", "pose.quat_x", "pose.quat_y", "pose.quat_z", "pose.quat_w" one 720x720 and two 1280x720 camera streams.… See the full description on the dataset page: https://huggingface.co/datasets/una-auxme/ROS2SmolVLA_ur10e_no_joints_crop_pick_place.tabularrobotics100K<n<1M1 likes294 downloads28d agoHugging Face14unaidedelf87777 /riddler Dataset Card for "riddler" It's a bunch of riddles. No I dont remember where they came from. text1K<n<10K0 likes290 downloads8mo agoHugging Face15lime-nlp /Synthetic_Unanswerable_Math Dataset Card for Synthetic Unanswerable Math (SUM) Dataset Summary Synthetic Unanswerable Math (SUM) is a dataset of high-quality, implicitly unanswerable math problems constructed to probe and improve the refusal behavior of large language models (LLMs). The goal is to teach models to identify when a problem cannot be answered due to incomplete, ambiguous, or contradictory information, and respond with epistemic humility (e.g., \boxed{I don't know}). Each entry in the… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/Synthetic_Unanswerable_Math.textreinforcement-learning10K<n<100K18 likes271 downloads1y agoHugging Face16unaidedelf87777 /parallel-function_calling-10k parallel-function-calling-10k This dataset contains 10,000 conversation examples demonstrating both parallel and standard (non-parallel) function calling. It's based on the Salesforce/xlam-function-calling-60k dataset, with additional generation using GPT-4o-small. Dataset Generation Process Initial Data Augmentation (gen_dataset.py): Loads the Salesforce/xlam-function-calling-60k dataset, which provides: User queries Available tools (functions) Function calls made by… See the full description on the dataset page: https://huggingface.co/datasets/unaidedelf87777/parallel-function_calling-10k.text10K<n<100K7 likes257 downloads2y agoHugging Face17open-llm-leaderboard-old /details_PulsarAI__Neural-una-cybertron-7b Dataset Card for Evaluation run of PulsarAI/Neural-una-cybertron-7b Dataset Summary Dataset automatically created during the evaluation run of model PulsarAI/Neural-una-cybertron-7b on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_PulsarAI__Neural-una-cybertron-7b.0 likes216 downloads3y agoHugging Face18open-llm-leaderboard-old /details_fblgit__UNA-ThePitbull-21.4B-v20 likes196 downloads2y agoHugging Face19open-llm-leaderboard-old /details_fblgit__UNA-TheBeagle-7b-v1 Dataset Card for Evaluation run of fblgit/UNA-TheBeagle-7b-v1 Dataset automatically created during the evaluation run of model fblgit/UNA-TheBeagle-7b-v1 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_fblgit__UNA-TheBeagle-7b-v1.0 likes195 downloads3y agoHugging Face20ymoslem /UN-Arabic-English-Filtered Dataset Details MultiUN + UNPC datasets, with rule-based and semantic filtering (train > 0.45 - test/dev > 0.9) as well as (>= 0.1) fasttext language detection. Dataset Structure DatasetDict({ train: Dataset({ features: ['text_en', 'text_ar'], num_rows: 19279407 }) test: Dataset({ features: ['text_en', 'text_ar'], num_rows: 8752 }) dev: Dataset({ features: ['text_en', 'text_ar'], num_rows: 8752 }) })… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/UN-Arabic-English-Filtered.texttranslation10M<n<100M2 likes186 downloads2y agoHugging Face21unalignment /toxic-dpo-v0.2 Toxic-DPO This is a highly toxic, "harmful" dataset meant to illustrate how DPO can be used to de-censor/unalign a model quite easily using direct-preference-optimization (DPO) using very few examples. Many of the examples still contain some amount of warnings/disclaimers, so it's still somewhat editorialized. Usage restriction To use this data, you must acknowledge/agree to the following: data contained within is "toxic"/"harmful", and contains profanity and other types… See the full description on the dataset page: https://huggingface.co/datasets/unalignment/toxic-dpo-v0.2.textn<1K143 likes174 downloads3y agoHugging Face22open-llm-leaderboard-old /details_fblgit__una-xaberius-34b-v1beta Dataset Card for Evaluation run of fblgit/una-xaberius-34b-v1beta Dataset Summary Dataset automatically created during the evaluation run of model fblgit/una-xaberius-34b-v1beta on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_fblgit__una-xaberius-34b-v1beta.0 likes171 downloads3y agoHugging Face23open-llm-leaderboard-old /details_fblgit__UNA-dolphin-2.6-mistral-7b-dpo-laser Dataset Card for Evaluation run of fblgit/UNA-dolphin-2.6-mistral-7b-dpo-laser Dataset automatically created during the evaluation run of model fblgit/UNA-dolphin-2.6-mistral-7b-dpo-laser on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_fblgit__UNA-dolphin-2.6-mistral-7b-dpo-laser.0 likes170 downloads3y agoHugging Face24MrDen1234567890 /deniz-unay-kurumsal-kesif-v3 Deniz UNAY: kaynaklı keşif ve teklif görüşmesi Bu paket, kurumların eğitim/konuşmacı ihtiyacını Deniz UNAY’ın belgelenmiş deneyimine bağlar; uygun ihtiyaç için teklif görüşmesi taslağı oluşturmayı destekler. Gerçek müşteri talebi, iş garantisi veya eğitilmiş LLM değildir. Başlangıç Keşif için discovery.csv: sekiz hizmetin 16 TR/EN kaydı. RAG için retrieval_documents.csv; kaynakları evidence/claims tablolarıyla birlikte kullanın. Talep sınıflandırması için… See the full description on the dataset page: https://huggingface.co/datasets/MrDen1234567890/deniz-unay-kurumsal-kesif-v3.tabularn<1K0 likes164 downloads7d agoHugging Face25unanam /mdramaaudio1K<n<10K0 likes162 downloads3y agoHugging Face26open-llm-leaderboard-old /details_unaidedelf87777__wizard-mistral-v0.1 Dataset Card for Evaluation run of unaidedelf87777/wizard-mistral-v0.1 Dataset Summary Dataset automatically created during the evaluation run of model unaidedelf87777/wizard-mistral-v0.1 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_unaidedelf87777__wizard-mistral-v0.1.0 likes152 downloads3y agoHugging Face27open-llm-leaderboard-old /details_Weyaxi__MetaMath-una-cybertron-v2-bf16-Ties Dataset Card for Evaluation run of Weyaxi/MetaMath-una-cybertron-v2-bf16-Ties Dataset Summary Dataset automatically created during the evaluation run of model Weyaxi/MetaMath-una-cybertron-v2-bf16-Ties on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Weyaxi__MetaMath-una-cybertron-v2-bf16-Ties.0 likes150 downloads3y agoHugging Face28ryfkn /unair-dicom-dataimage10K<n<100K0 likes141 downloads2mo agoHugging Face29saier /unarXive_citrec Dataset Card for unarXive citation recommendation Dataset Summary The unarXive citation recommendation dataset contains 2.5 Million paragraphs from computer science papers and with an annotated citation marker. The paragraphs and citation information is derived from unarXive. Note that citation infromation is only given as the OpenAlex ID of the cited paper. An important consideration for models is therefore if the data is used as is, or if additional information of the… See the full description on the dataset page: https://huggingface.co/datasets/saier/unarXive_citrec.texttext-classification1M<n<10M9 likes129 downloads3y agoHugging Face30somosnlp-hackathon-2022 /unam_tesis Dataset Card of "unam_tesis" Dataset Summary El dataset unam_tesis cuenta con 1000 tesis de 5 carreras de la Universidad Nacional Autónoma de México (UNAM), 200 por carrera. Se pretende seguir incrementando este dataset con las demás carreras y más tesis. Supported Tasks and Leaderboards text-classification Languages Español (es) Dataset Structure Data Instances Las instancias del dataset son de la siguiente forma: El objetivo… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/unam_tesis.text-classification5 likes124 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.