OceanBEATs โ€” corrected September 2026 models

OceanBEATs adapts BEATs to underwater acoustic recordings using domain-adaptive pretraining (DAPT). The corrected Stage 1 encoder and matching 56-class SED head below are the model pair used for Table 1 of the September 2026 minor revision of Noda et al., Scientific Reports (revision in review). The Stage 2 encoder listed after them is used for the FRDR and HICEAS analyses of the final revision.

Corrected model pair

File Bytes SHA-256
BEATs_DAPT_MAM_fixed_step127641.pt 361345049 2a2d1d93f53ec29227bdd52da087fd0abcf0ce797c3c4a8629cd1435a314a6f9
sed_head_fixed_s42_ep7.pt 17963283 9b2b202ab3e52b0d1efe4cd3479ee479db7646b0f42ab5b0e32f1e3ca551f119

Use these two files together. The encoder contains cfg and a model state dictionary; the head contains a head state dictionary, the ordered 56-class labels, and training hparams. The files are PyTorch checkpoints, not a Transformers from_pretrained bundle. Only load checkpoints from a trusted source and verify their hashes before loading.

The head configuration retains generic original training/validation and log path strings for provenance; the referenced CSV contents are not included. No raw audio or individual reference embeddings are distributed in this pair.

Stage 2 encoder (final revision)

File Bytes SHA-256
BEATs_DAPT_MAM_fixed_palaoa_step6385.pt 361346274 4f7869751d7f15e3a806fb062902654597ca5566be610fedc1762c440d5c2a89

The manuscript keeps its two-stage DAPT design. Stage 2 continues from the Stage 1 checkpoint above on the 2021 PALAOA subset (approximately 287 h) with the same input-mask implementation and frozen teacher: 102,168 training windows (1,032 of 103,200 held out by random_split with seed 42), batch size 16, drop_last=True, 6,385 optimiser steps, encoder learning rate 1e-5 and predictor/mask-token learning rate 1e-4. The final checkpoint is used; no downstream selection was applied. Stage 2 changes the embeddings relative to Stage 1: for the same 1,623 validation clips the mean cosine similarity between Stage 1 and Stage 2 embeddings is 0.913 (minimum 0.793), so the CCED2 reference of the final revision was refitted on Stage 2 embeddings. The 56-class SED head was trained on the Stage 1 encoder and is not a matching head for the Stage 2 encoder.

Training and evaluation identity

  • BEATs AS-2M (iter3+) initialisation; 16 kHz mono input.
  • World-DAPT: 2,042,268 non-overlapping 10-s windows (approximately 5,673 h).
  • Frozen teacher and fixed 1,024-cluster targets; 75% of student patch embeddings replaced by a trainable mask token before transformer encoding.
  • Cross-entropy on masked positions; AdamW with encoder learning rate 1e-4 and predictor/mask-token learning rate 1e-3, 5% warm-up then cosine decay, and bfloat16 autocast.
  • One shuffled pass, batch size 16, drop_last=True: 127,641 optimiser steps; the incomplete final batch of 12 windows is unused. No DAPT validation split or downstream checkpoint selection was used for this endpoint.
  • Matching supervised head: seed 42, epoch 7. With this head, Event/Clip/2-s-segment F1 is 0.493/0.749/0.518. Table 1 of the final revision reports the mean of eight SED-head seeds (0.493/0.739/0.523), of which this head is one; the per-seed values are in the code and results release. These are frozen aggregate results, not an evaluation on a publicly reproducible 56-class dataset.

The retained BEATs+DAPT analyses use corrected window-aware extraction: start_sec when supplied, otherwise max(0, center_sec - 5), with boundary padding as needed. Reusing a legacy extractor can change results even when the correct checkpoint is loaded.

Version history โ€” do not substitute legacy files

  • December 2025 SimCLR-style DAPT was invalid because AMP fp16 prevented encoder updates. That analysis is superseded.
  • May 2026 step-120,000 MAM is also superseded: masking was applied after unmasked audio passed through the encoder, rather than before transformer encoding. It does not implement the stated masked-input objective.
  • September 2026 step-127,641 (Stage 1) and its PALAOA continuation at step 6,385 (Stage 2) use the corrected input-mask implementation.

The two May files remain available unchanged for historical provenance:

Legacy file SHA-256
beats_dapt_mam_step120000.pt 0fe9f7dd92780c2e564f1df06a192482dbcb9a56bdab4202f4d94862b9168f89
sed_head_56_fulldata_ep8.pt 135d11738a6619a57769955468ce5cb6eee3f07044fa45e6c950bf25ac4f8f60

Neither is a substitute for the corrected pair. The original model-card history remains available at repository revision dbb29a3dfc4fe1605c9fdd87079723db12903849; its May-era numerical and availability claims are not the current minor-revision record.

Download and code

The corrected pair is fixed by tag v3.0.3-sr-minor-2026-09-13; tag v3.1.6-sr-minor-2026-09-19 fixes the pair together with the Stage 2 encoder and corresponds to the code and results release of the final revision (GitHub commit 9596fd72285413acfb3ccf208683c2c2d62470d2; model files unchanged since tag v3.1.0-sr-minor-2026-09-19). To download the pair:

from huggingface_hub import hf_hub_download

for name in ["BEATs_DAPT_MAM_fixed_step127641.pt", "sed_head_fixed_s42_ep7.pt"]:
    path = hf_hub_download(
        repo_id="BiologgingSolutions/OceanBEATs",
        revision="v3.0.3-sr-minor-2026-09-13",
        filename=name,
        local_dir="weights",
    )
    print(path)

Check each downloaded file with shasum -a 256 against the table above. The corrected code and aggregate-results release is GitHub commit fe3cc9f1a696e6814329120fdf0a64bc6327ed5c. Its artifact map defines the verification and full-rerun boundaries. That GitHub release predates this model upload; its statement that corrected weights were not yet public describes the release-time state. This model revision supplies the two matching files, but does not remove the remaining data limitations.

Data and reproducibility limits

The internal 56-class audio, clip-level metadata and split membership remain non-public because of sensitive location/operational information and the permissions of the original collaborating organisations. The complete label taxonomy and aggregate statistics are reported in the manuscript materials; exact retraining and Table 1 evaluation cannot be reproduced publicly.

The corrected fitted CCED2 kNN reference contains individual restricted reference embeddings and is not distributed. This model pair does not supply the fitted CCED2 reference models, all evaluation embeddings, or window-level score arrays. Public aggregate checks are not equivalent to complete raw-data reproduction. Results are research diagnostics, not validated deployment performance or evidence that the complete DGPU loop caused the Promoter gain.

Licence and acknowledgement

The model weights retain CC BY 4.0, including commercial use with attribution; the companion source code is MIT-licensed. The DGPU framework and CCED2 score are subject to patent applications filed by Biologging Solutions Inc.; these copyright licences do not grant patent rights.

We acknowledge Microsoft BEATs and the World-DAPT source collections: NOAA SanctSound, US Navy USWTR, NOAA NRS/ONMS, ICListen / ONC, NPS Glacier Bay, and PALAOA Ekstrรถm Ice Shelf.

Citation

Noda, T. and Koizumi, T. Discovery and promotion of unknown sounds into operational detection targets for underwater passive acoustic monitoring under false alarm constraints. Scientific Reports (revision in review, 2026).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support