CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01WHOP3 /m2mdatasettext10K<n<100K0 likes134 downloads3d agoHugging Face02evozim /m2mcent-mcp-schemas 🌐 M2MCent Agentic Services - MCP Schemas Dataset 🚀 Empowering Autonomous AI on Base L2 This dataset contains the JSON schemas for 1,005 microservices natively available on the M2MCent Network via the x402 V2 Protocol (EIP-3009). It is specifically designed for instruction-tuning LLMs (like Llama-3, Mistral, Qwen) so they can autonomously discover, negotiate, and consume monetized API endpoints using gasless cryptocurrency settlements on the Base L2 network.… See the full description on the dataset page: https://huggingface.co/datasets/evozim/m2mcent-mcp-schemas.texttext-generation1K<n<10K1 likes72 downloads9d agoHugging Face03nlpso /m2m3_fine_tuning_ref_ptrn_cmbert_io m2m3_fine_tuning_ref_ptrn_cmbert_io Introduction This dataset was used to fine-tuned HueyNemud/das22-10-camembert_pretrained for nested NER task using Independant NER layers approach [M1]. It contains Paris trade directories entries from the 19th century. Dataset parameters Approachrd : M2 and M3 Dataset type : ground-truth Tokenizer : HueyNemud/das22-10-camembert_pretrained Tagging format : IO Counts : Train : 6084 Dev : 676 Test : 1685 Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_fine_tuning_ref_ptrn_cmbert_io.texttoken-classification1K<n<10K0 likes36 downloads4y agoHugging Face04lomiotech /m2m_smolvla_finetune_datasetThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "koch_follower", "total_episodes": 20, "total_frames": 15132, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lomiotech/m2m_smolvla_finetune_dataset.tabularrobotics10K<n<100K0 likes31 downloads1y agoHugging Face05nlpso /m2m3_qualitative_analysis_ref_cmbert_io m2m3_qualitative_analysis_ref_cmbert_io Introduction This dataset was used to perform qualitative analysis of Jean-Baptiste/camembert-ner on nested NER task using Independant NER layers approach [M1]. It contains Paris trade directories entries from the 19th century. Dataset parameters Approachrd : M2 and M3 Dataset type : ground-truth Tokenizer : Jean-Baptiste/camembert-ner Tagging format : IO Counts : Train : 6084 Dev : 676 Test : 1685 Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ref_cmbert_io.texttoken-classification1K<n<10K0 likes30 downloads4y agoHugging Face06zwh9029 /rm-static-m2m100-zh-jiantitext10K<n<100K4 likes30 downloads3y agoHugging Face07nlpso /m2m3_qualitative_analysis_ocr_cmbert_io m2m3_qualitative_analysis_ocr_cmbert_io Introduction This dataset was used to perform qualitative analysis of Jean-Baptiste/camembert-ner on nested NER task using Independant NER layers approach [M1]. It contains Paris trade directories entries from the 19th century. Dataset parameters Approachrd : M2 and M3 Dataset type : noisy (Pero OCR) Tokenizer : Jean-Baptiste/camembert-ner Tagging format : IO Counts : Train : 6084 Dev : 676 Test : 1685 Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ocr_cmbert_io.texttoken-classification1K<n<10K0 likes26 downloads4y agoHugging Face08nlpso /m2m3_qualitative_analysis_ocr_ptrn_cmbert_io m2m3_qualitative_analysis_ocr_ptrn_cmbert_io Introduction This dataset was used to perform qualitative analysis of HueyNemud/das22-10-camembert_pretrained on nested NER task using Independant NER layers approach [M1]. It contains Paris trade directories entries from the 19th century. Dataset parameters Approachrd : M2 and M3 Dataset type : noisy (Pero OCR) Tokenizer : HueyNemud/das22-10-camembert_pretrained Tagging format : IO Counts : Train : 6084 Dev : 676… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ocr_ptrn_cmbert_io.texttoken-classification1K<n<10K0 likes26 downloads4y agoHugging Face09shalanova /benchmark-2-russian-m2mInfo: Translated on Russian by facebook/m2m100_418M model Source: xTRam1/safe-guard-prompt-injection Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns Size: 1,000 prompts (500 safe / 500 unsafe) Columns: text - original prompt label - 0: safe, 1: unsafe translation - prompt on Russian translated by facebook/m2m100_418M score_ru_model - cosine similarity score with codebook More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-2-russian-m2m.tabular1K<n<10K0 likes23 downloads5mo agoHugging Face10nlpso /m2m3_fine_tuning_ref_cmbert_io m2m3_fine_tuning_ref_cmbert_io Introduction This dataset was used to fine-tuned Jean-Baptiste/camembert-ner for nested NER task using Independant NER layers approach [M1]. It contains Paris trade directories entries from the 19th century. Dataset parameters Approachrd : M2 and M3 Dataset type : ground-truth Tokenizer : Jean-Baptiste/camembert-ner Tagging format : IO Counts : Train : 6084 Dev : 676 Test : 1685 Associated fine-tuned models : M2 :… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_fine_tuning_ref_cmbert_io.texttoken-classification1K<n<10K0 likes22 downloads4y agoHugging Face11nlpso /m2m3_fine_tuning_ocr_ptrn_cmbert_iob2 m2m3_fine_tuning_ocr_ptrn_cmbert_iob2 Introduction This dataset was used to fine-tuned HueyNemud/das22-10-camembert_pretrained for nested NER task using Independant NER layers approach [M1]. It contains Paris trade directories entries from the 19th century. Dataset parameters Approachrd : M2 and M3 Dataset type : noisy (Pero OCR) Tokenizer : HueyNemud/das22-10-camembert_pretrained Tagging format : IOB2 Counts : Train : 6084 Dev : 676 Test : 1685 Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_fine_tuning_ocr_ptrn_cmbert_iob2.texttoken-classification1K<n<10K0 likes21 downloads4y agoHugging Face12ctc88haha /m2m100-418m-fp16-merged-onnx-ios-package M2M100 418M FP16 Merged ONNX iOS Package This repository contains an ONNX FP16 merged runtime package converted from facebook/m2m100_418M for use in an offline iOS translation app. This is a converted runtime package. It is not the original unmodified PyTorch model checkpoint published by Meta/Facebook. This repository is not endorsed by Meta/Facebook. Package contents The archive m2m100-418m-fp16-merged-ios.zip contains one top-level folder… See the full description on the dataset page: https://huggingface.co/datasets/ctc88haha/m2m100-418m-fp16-merged-onnx-ios-package.0 likes21 downloads3mo agoHugging Face13nlpso /m2m3_qualitative_analysis_ref_ptrn_cmbert_iob2 m2m3_qualitative_analysis_ref_ptrn_cmbert_iob2 Introduction This dataset was used to perform qualitative analysis of HueyNemud/das22-10-camembert_pretrained on nested NER task using Independant NER layers approach [M1]. It contains Paris trade directories entries from the 19th century. Dataset parameters Approachrd : M2 and M3 Dataset type : ground-truth Tokenizer : HueyNemud/das22-10-camembert_pretrained Tagging format : IOB2 Counts : Train : 6084 Dev : 676… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ref_ptrn_cmbert_iob2.texttoken-classification1K<n<10K0 likes20 downloads4y agoHugging Face14shalanova /benchmark-1-chinese-m2mInfo: Translated on Chinese by facebook/m2m100_418M model Source: jayavibhav/prompt-injection-safety Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns Size: 1,000 prompts (500 safe / 500 unsafe) Columns: text - original prompt label - 0: safe, 1: unsafe translation - prompt on Chinese translated by facebook/m2m100_418M score_zh_model - cosine similarity score with codebook More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-1-chinese-m2m.tabular1K<n<10K0 likes20 downloads5mo agoHugging Face15nlpso /m2m3_qualitative_analysis_ref_cmbert_iob2 m2m3_qualitative_analysis_ref_cmbert_iob2 Introduction This dataset was used to perform qualitative analysis of Jean-Baptiste/camembert-ner on nested NER task using Independant NER layers approach [M1]. It contains Paris trade directories entries from the 19th century. Dataset parameters Approachrd : M2 and M3 Dataset type : ground-truth Tokenizer : Jean-Baptiste/camembert-ner Tagging format : IOB2 Counts : Train : 6084 Dev : 676 Test : 1685 Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ref_cmbert_iob2.texttoken-classification1K<n<10K0 likes19 downloads4y agoHugging Face16nlpso /m2m3_qualitative_analysis_ref_ptrn_cmbert_io m2m3_qualitative_analysis_ref_ptrn_cmbert_io Introduction This dataset was used to perform qualitative analysis of HueyNemud/das22-10-camembert_pretrained on nested NER task using Independant NER layers approach [M1]. It contains Paris trade directories entries from the 19th century. Dataset parameters Approachrd : M2 and M3 Dataset type : ground-truth Tokenizer : HueyNemud/das22-10-camembert_pretrained Tagging format : IO Counts : Train : 6084 Dev : 676… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ref_ptrn_cmbert_io.texttoken-classification1K<n<10K0 likes17 downloads4y agoHugging Face17shalanova /benchmark-1-arabic-m2mInfo: Translated on Arabic by facebook/m2m100_418M model Source: jayavibhav/prompt-injection-safety Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns Size: 1,000 prompts (500 safe / 500 unsafe) Columns: text - original prompt label - 0: safe, 1: unsafe translation - prompt on Arabic translated by facebook/m2m100_418M score_ar_model - cosine similarity score with codebook More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-1-arabic-m2m.tabular1K<n<10K0 likes16 downloads5mo agoHugging Face18shalanova /benchmark-4-chinese-m2mInfo: Translated on Chinese by facebook/m2m100_418M model Source: nvidia/Aegis-AI-Content-Safety-Dataset-2.0 Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer. Size: 1,000 prompts (500… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-4-chinese-m2m.tabular1K<n<10K0 likes16 downloads5mo agoHugging Face19nlpso /m2m3_fine_tuning_ocr_ptrn_cmbert_io m2m3_fine_tuning_ocr_ptrn_cmbert_io Introduction This dataset was used to fine-tuned HueyNemud/das22-10-camembert_pretrained for nested NER task using Independant NER layers approach [M1]. It contains Paris trade directories entries from the 19th century. Dataset parameters Approachrd : M2 and M3 Dataset type : noisy (Pero OCR) Tokenizer : HueyNemud/das22-10-camembert_pretrained Tagging format : IO Counts : Train : 6084 Dev : 676 Test : 1685 Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_fine_tuning_ocr_ptrn_cmbert_io.texttoken-classification1K<n<10K0 likes15 downloads4y agoHugging Face20nlpso /m2m3_qualitative_analysis_ocr_cmbert_iob2 m2m3_qualitative_analysis_ocr_cmbert_iob2 Introduction This dataset was used to perform qualitative analysis of Jean-Baptiste/camembert-ner on nested NER task using Independant NER layers approach [M1]. It contains Paris trade directories entries from the 19th century. Dataset parameters Approachrd : M2 and M3 Dataset type : noisy (Pero OCR) Tokenizer : Jean-Baptiste/camembert-ner Tagging format : IOB2 Counts : Train : 6084 Dev : 676 Test : 1685 Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ocr_cmbert_iob2.texttoken-classification1K<n<10K0 likes15 downloads4y agoHugging Face21nlpso /m2m3_qualitative_analysis_ocr_ptrn_cmbert_iob2 m2m3_qualitative_analysis_ocr_ptrn_cmbert_iob2 Introduction This dataset was used to perform qualitative analysis of HueyNemud/das22-10-camembert_pretrained on nested NER task using Independant NER layers approach [M1]. It contains Paris trade directories entries from the 19th century. Dataset parameters Approachrd : M2 and M3 Dataset type : noisy (Pero OCR) Tokenizer : HueyNemud/das22-10-camembert_pretrained Tagging format : IOB2 Counts : Train : 6084 Dev :… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ocr_ptrn_cmbert_iob2.texttoken-classification1K<n<10K0 likes15 downloads4y agoHugging Face22nlpso /m2m3_fine_tuning_ref_cmbert_iob2 m2m3_fine_tuning_ref_cmbert_iob2 Introduction This dataset was used to fine-tuned Jean-Baptiste/camembert-ner for nested NER task using Independant NER layers approach [M1]. It contains Paris trade directories entries from the 19th century. Dataset parameters Approachrd : M2 and M3 Dataset type : ground-truth Tokenizer : Jean-Baptiste/camembert-ner Tagging format : IOB2 Counts : Train : 6084 Dev : 676 Test : 1685 Associated fine-tuned models : M2 :… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_fine_tuning_ref_cmbert_iob2.texttoken-classification1K<n<10K0 likes14 downloads4y agoHugging Face23ignacioct /wikipedia_en_es_m2m Dataset Card for "wikipedia_en_es_nllb" More Information needed text10K<n<100K0 likes14 downloads2y agoHugging Face24nlpso /m2m3_fine_tuning_ocr_cmbert_io m2m3_fine_tuning_ocr_cmbert_io Introduction This dataset was used to fine-tuned Jean-Baptiste/camembert-ner for nested NER task using Independant NER layers approach [M1]. It contains Paris trade directories entries from the 19th century. Dataset parameters Approachrd : M2 and M3 Dataset type : noisy (Pero OCR) Tokenizer : Jean-Baptiste/camembert-ner Tagging format : IO Counts : Train : 6084 Dev : 676 Test : 1685 Associated fine-tuned models : M2 :… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_fine_tuning_ocr_cmbert_io.texttoken-classification1K<n<10K0 likes13 downloads4y agoHugging Face25zwh9029 /rm-static-m2m100-zheg-jiantitext100K<n<1M2 likes13 downloads3y agoHugging Face26shalanova /benchmark-1-russian-m2mInfo: Translated on Russian by facebook/m2m100_418M model Source: jayavibhav/prompt-injection-safety Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns Size: 1,000 prompts (500 safe / 500 unsafe) Columns: text - original prompt label - 0: safe, 1: unsafe translation - prompt on Russian translated by facebook/m2m100_418M score_ru_model - cosine similarity score with codebook More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-1-russian-m2m.tabular1K<n<10K0 likes13 downloads5mo agoHugging Face27shalanova /benchmark-3-arabic-m2mInfo: Translated on Arabic by facebook/m2m100_418M model Source: JailbreakBench/JBB-Behaviors Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer. Size: 200 prompts (100 safe / 100… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-3-arabic-m2m.tabularn<1K0 likes11 downloads5mo agoHugging Face28Alwaly /french-wolof-translation-results-m2m100textn<1K0 likes10 downloads1y agoHugging Face29french-datasets /Alwaly_french-wolof-translation-results-m2m100Ce répertoire est vide, il a été créé pour améliorer le référencement du jeu de données Alwaly/french-wolof-translation-results-m2m100. translation0 likes9 downloads1y agoHugging Face30shalanova /benchmark-2-chinese-m2mInfo: Translated on Chinese by facebook/m2m100_418M model Source: xTRam1/safe-guard-prompt-injection Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns Size: 1,000 prompts (500 safe / 500 unsafe) Columns: text - original prompt label - 0: safe, 1: unsafe translation - prompt on Chinese translated by facebook/m2m100_418M score_zh_model - cosine similarity score with codebook More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-2-chinese-m2m.tabular1K<n<10K0 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.