HarishMaths's picture
Update README.md
3c9f67d verified
|
Raw History Blame Contribute Delete
15.6 kB
metadata
tags:
  - sentence-transformers
  - sentence-similarity
  - feature-extraction
  - dense
  - generated_from_trainer
  - dataset_size:3872
  - loss:MultipleNegativesRankingLoss
widget:
  - source_sentence: >-
      (g) If a Member wishes to extend their stay and has enough Nightly Upgrade
      Award(s) to cover the extension, the Member must book a separate
      reservation for the additional nights and request to use Nightly Upgrade
      Awards on Marriott Websites or by calling Member Support; the Nightly
      Upgrade Award request cannot be processed at the Participating Property.
    sentences:
      - Flexible rates cancel up to a deadline the property sets.
      - >-
        DONT book a non-refundable hotel without reading the cancellation policy
        because you must understand the exact penalty structurewhether you
        forfeit one night, the full amount, or a percentageto determine your
        coverage needs and ensure your insurance limit is adequate.
      - >-
        As for semi-flexible plans, they might require notice at least five days
        before check-in.
  - source_sentence: >-
      Checking out late at a hotel isnt guaranteed, especially when it comes to
      complimentary late check-out.
    sentences:
      - >-
        However, the late check-out policy will vary based on the specific
        hotels policy.
      - >-
        The hotel guest damage clause is a crucial aspect of your reservation
        agreement that aims to protect both the hotels property and the guests
        interests.
      - Yes, if a clean room is available.
  - source_sentence: refund terms and conditions
    sentences:
      - CANCELLATION OR MODIFICATION OF A SERVICE RESERVATION
      - >-
        Elite status doesn't always change the written policy, but it can give
        you leverage with customer service if you need an exception.
      - 'Trick #2  Resell the nonrefundable hotel room'
  - source_sentence: occupancy rules guidelines for guests
    sentences:
      - >-
        Most hotel insurance policies provide coverage for theft, damage, or
        loss of personal property under certain conditions.
      - Choose designated smoking areas outside the hotel.
      - >-
        Semi-Flexible Rates : Some properties offer rates that allow
        cancellation with a fee (e.g., $50) or partial refund up to a certain
        point.
  - source_sentence: Smoking or vaping is allowed in designated rooms.
    sentences:
      - >-
        Some upgrades to Premium Rooms require payment in local currency and
        cannot be purchased with Points.
      - >-
        Marriott is committed to providing its guests and associates with a
        smoke-free environment, and is proud to boast one of the most
        comprehensive smoke-free hotel policies in the industry.
      - >-
        Participating Properties outside the United States may provide
        alternative services and benefits to the Elite membership benefits set
        forth in these Program Rules, depending on local law and policy.
pipeline_tag: sentence-similarity
library_name: sentence-transformers
metrics:
  - pearson_cosine
  - spearman_cosine
model-index:
  - name: SentenceTransformer
    results:
      - task:
          type: semantic-similarity
          name: Semantic Similarity
        dataset:
          name: val
          type: val
        metrics:
          - type: pearson_cosine
            value: 0.6244156998181909
            name: Pearson Cosine
          - type: spearman_cosine
            value: 0.6463453957364179
            name: Spearman Cosine

SentenceTransformer

This is a sentence-transformers model trained for semantic text understanding. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, classification, clustering, and more.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Maximum Sequence Length: 128 tokens
  • Output Dimensionality: 384 dimensions
  • Similarity Function: Cosine Similarity
  • Supported Modality: Text

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
  (1): Pooling({'embedding_dimension': 384, 'pooling_mode': 'mean', 'include_prompt': True})
  (2): Normalize({})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("HarishMaths/Hotel-Policy-Embedding")
# Run inference
sentences = [
    'Smoking or vaping is allowed in designated rooms.',
    'Marriott is committed to providing its guests and associates with a smoke-free environment, and is proud to boast one of the most comprehensive smoke-free hotel policies in the industry.',
    'Some upgrades to Premium Rooms require payment in local currency and cannot be purchased with Points.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.9151, 0.2602],
#         [0.9151, 1.0000, 0.3726],
#         [0.2602, 0.3726, 1.0000]])

Evaluation

Metrics

Semantic Similarity

Metric Value
pearson_cosine 0.6244
spearman_cosine 0.6463

Training Details

Training Dataset

Unnamed Dataset

  • Size: 3,872 training samples
  • Columns: sentence_0 and sentence_1
  • Approximate statistics based on the first 100 samples:
    sentence_0 sentence_1
    type string string
    modality text text
    details
    • min: 6 tokens
    • mean: 24.39 tokens
    • max: 117 tokens
    • min: 9 tokens
    • mean: 27.12 tokens
    • max: 103 tokens
  • Samples:
    sentence_0 sentence_1
    Three months after booking and 20 days after purchasing insurance, the tour operator files bankruptcy and ceases all operations. California SB 644 requires hotels and third-party booking sites to give a full refund when a guest cancels within 24 hours of booking, as long as the reservation was made at least 72 hours before check-in.
    One of the easiest ways to avoid resort fees is by booking an award stay. For a typical domestic Hilton hotel, it will show you the points options and cash rates all on one screen.
    The General Contractor went out to several Home Depot locations around the City to find over 250 battery-operated smoke detectors. Modern sensors detect vapour, and the resulting charge is identical to a cigarette violation.
  • Loss: MultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim",
        "gather_across_devices": false,
        "directions": [
            "query_to_doc"
        ],
        "partition_mode": "joint",
        "hardness_mode": null,
        "hardness_strength": 0.0
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 32
  • num_train_epochs: 25
  • fp16: True
  • per_device_eval_batch_size: 32
  • multi_dataset_batch_sampler: round_robin

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 32
  • num_train_epochs: 25
  • max_steps: -1
  • learning_rate: 5e-05
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_steps: 0
  • optim: adamw_torch_fused
  • optim_args: None
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 1
  • average_tokens_across_devices: True
  • max_grad_norm: 1
  • label_smoothing_factor: 0.0
  • bf16: False
  • fp16: True
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: None
  • trackio_bucket_id: None
  • trackio_static_space_id: None
  • per_device_eval_batch_size: 32
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • dataloader_multiprocessing_context: None
  • dataloader_in_order: True
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_static_graph: None
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: None
  • fsdp_config: None
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: round_robin
  • router_mapping: {}
  • learning_rate_mapping: {}
  • warmup_ratio: None

Training Logs

Epoch Step Training Loss val_spearman_cosine
-1 -1 - 0.1722
0.8264 100 - 0.2700
1.0 121 - 0.2967
1.6529 200 - 0.4105
2.0 242 - 0.4552
2.4793 300 - 0.4991
3.0 363 - 0.5376
3.3058 400 - 0.5566
4.0 484 - 0.5790
4.1322 500 3.5721 0.5825
4.9587 600 - 0.6156
5.0 605 - 0.6162
5.7851 700 - 0.6144
6.0 726 - 0.6166
6.6116 800 - 0.6266
7.0 847 - 0.6269
7.4380 900 - 0.6348
8.0 968 - 0.6276
8.2645 1000 2.4308 0.6329
9.0 1089 - 0.6336
9.0909 1100 - 0.6332
9.9174 1200 - 0.6399
10.0 1210 - 0.6381
10.7438 1300 - 0.6390
11.0 1331 - 0.6397
11.5702 1400 - 0.6446
12.0 1452 - 0.6418
12.3967 1500 2.1294 0.6463

Framework Versions

  • Python: 3.13.15
  • Sentence Transformers: 5.7.0
  • Transformers: 5.16.1
  • PyTorch: 2.11.0+cu128
  • Accelerate: 1.14.0
  • Datasets: 4.8.5
  • Tokenizers: 0.23.1