CoolFace
Modelpublic

bhujith10/bert-large-uncased-setfit_finetuned

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes5downloads
Model Card

SetFit with google-bert/bert-large-uncased

This is a SetFit model trained on the bhujith10/multi_class_classification_dataset dataset that can be used for Text Classification. This SetFit model uses google-bert/bert-large-uncased as the Sentence Transformer embedding model. A SetFitHead instance is used for classification.

The model has been trained using an efficient few-shot learning technique that involves:

  1. 1.Fine-tuning a Sentence Transformer with contrastive learning.
  2. 2.Training a classification head with features from the fine-tuned Sentence Transformer.

Model Details

Model Description

Model Sources

Uses

Direct Use for Inference

First install the SetFit library:

bash
pip install setfit

Then you can load this model and run inference.

python
from setfit import SetFitModel

# Download from the 🤗 Hub
model = SetFitModel.from_pretrained("bhujith10/bert-large-uncased-setfit_finetuned")
# Run inference
preds = model("Title: On the isoperimetric quotient over scalar-flat conformal classes,
Abstract: Let $(M,g)$ be a smooth compact Riemannian manifold of dimension $n$ with
smooth boundary $\partial M$. Suppose that $(M,g)$ admits a scalar-flat
conformal metric. We prove that the supremum of the isoperimetric quotient over
the scalar-flat conformal class is strictly larger than the best constant of
the isoperimetric inequality in the Euclidean space, and consequently is
achieved, if either (i) $n\ge 12$ and $\partial M$ has a nonumbilic point; or
(ii) $n\ge 10$, $\partial M$ is umbilic and the Weyl tensor does not vanish at
some boundary point.")

<!--

Downstream Use

List how someone could finetune this model on their own dataset. -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Set Metrics

Training setMinMedianMax
Word count23145.8467280

Training Hyperparameters

  • —batch_size: (4, 4)
  • —num_epochs: (2, 2)
  • —max_steps: -1
  • —sampling_strategy: oversampling
  • —bodylearningrate: (2e-05, 1e-05)
  • —headlearningrate: 0.01
  • —loss: CosineSimilarityLoss
  • —distancemetric: cosinedistance
  • —margin: 0.25
  • —endtoend: False
  • —use_amp: False
  • —warmup_proportion: 0.1
  • —l2_weight: 0.01
  • —seed: 42
  • —evalmaxsteps: -1
  • —loadbestmodelatend: True

Training Results

EpochStepTraining LossValidation Loss
0.000310.22-
0.0138500.3706-
0.02761000.2389-
0.04141500.1628-
0.05512000.1401-
0.06892500.1043-
0.08273000.1047-
0.09653500.098-
0.11034000.0931-
0.12414500.1002-
0.13795000.0837-
0.15165500.0673-
0.16546000.0709-
0.17926500.08-
0.19307000.0719-
0.20687500.0805-
0.22068000.059-
0.23448500.0957-
0.24819000.0614-
0.26199500.0887-
0.275710000.0713-
0.289510500.0734-
0.303311000.0519-
0.317111500.0802-
0.330912000.0817-
0.344612500.0665-
0.358413000.0515-
0.372213500.0764-
0.386014000.0564-
0.399814500.0512-
0.413615000.052-
0.427415500.0398-
0.441116000.0473-
0.454916500.0433-
0.468717000.0621-
0.482517500.0506-
0.496318000.0395-
0.510118500.0516-
0.523819000.0431-
0.537619500.037-
0.551420000.0299-
0.565220500.0398-
0.579021000.0335-
0.592821500.0438-
0.606622000.0436-
0.620322500.0345-
0.634123000.0396-
0.647923500.0381-
0.661724000.0377-
0.675524500.0287-
0.689325000.0393-
0.703125500.0309-
0.716826000.0363-
0.730626500.0347-
0.744427000.0299-
0.758227500.0305-
0.772028000.0349-
0.785828500.0385-
0.799629000.0412-
0.813329500.0336-
0.827130000.0422-
0.840930500.0249-
0.854731000.0285-
0.868531500.0258-
0.882332000.0309-
0.896132500.0246-
0.909833000.0271-
0.923633500.0285-
0.937434000.0318-
0.951234500.0287-
0.965035000.0298-
0.978835500.021-
0.992636000.036-
1.03627-0.1036
1.006336500.0257-
1.020137000.02-
1.033937500.0333-
1.047738000.0339-
1.061538500.0283-
1.075339000.0233-
1.089139500.0311-
1.102840000.0296-
1.116640500.0271-
1.130441000.0321-
1.144241500.0221-
1.158042000.026-
1.171842500.0283-
1.185643000.0378-
1.199343500.0225-
1.213144000.0237-
1.226944500.0254-
1.240745000.0253-
1.254545500.023-
1.268346000.0265-
1.282146500.0255-
1.295847000.0278-
1.309647500.0285-
1.323448000.0234-
1.337248500.0282-
1.351049000.0197-
1.364849500.0284-
1.378550000.0326-
1.392350500.0233-
1.406151000.0386-
1.419951500.0308-
1.433752000.0218-
1.447552500.0288-
1.461353000.0251-
1.475053500.0255-
1.488854000.0261-
1.502654500.0253-
1.516455000.0313-
1.530255500.0277-
1.544056000.0252-
1.557856500.0293-
1.571557000.0334-
1.585357500.0285-
1.599158000.0269-
1.612958500.0267-
1.626759000.0313-
1.640559500.0243-
1.654360000.0301-
1.668060500.0266-
1.681861000.0276-
1.695661500.0293-
1.709462000.0291-
1.723262500.031-
1.737063000.0283-
1.750863500.0238-
1.764564000.0261-
1.778364500.0196-
1.792165000.034-
1.805965500.0255-
1.819766000.0231-
1.833566500.0256-
1.847367000.0207-
1.861067500.0325-
1.874868000.0238-
1.888668500.0277-
1.902469000.0239-
1.916269500.0239-
1.930070000.0227-
1.943870500.0236-
1.957571000.0216-
1.971371500.0248-
1.985172000.0244-
1.998972500.0203-
2.07254-0.1068

Framework Versions

  • —Python: 3.10.12
  • —SetFit: 1.1.0
  • —Sentence Transformers: 3.3.1
  • —Transformers: 4.45.2
  • —PyTorch: 2.1.0+cu118
  • —Datasets: 3.2.0
  • —Tokenizers: 0.20.3

Citation

BibTeX

bibtex
@article{https://doi.org/10.48550/arxiv.2209.11055,
    doi = {10.48550/ARXIV.2209.11055},
    url = {https://arxiv.org/abs/2209.11055},
    author = {Tunstall, Lewis and Reimers, Nils and Jo, Unso Eun Seo and Bates, Luke and Korat, Daniel and Wasserblat, Moshe and Pereg, Oren},
    keywords = {Computation and Language (cs.CL), FOS: Computer and information sciences, FOS: Computer and information sciences},
    title = {Efficient Few-Shot Learning Without Prompts},
    publisher = {arXiv},
    year = {2022},
    copyright = {Creative Commons Attribution 4.0 International}
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->