OceanBEATs โ corrected September 2026 models
OceanBEATs adapts BEATs to underwater acoustic recordings using domain-adaptive pretraining (DAPT). The corrected Stage 1 encoder and matching 56-class SED head below are the model pair used for Table 1 of the September 2026 minor revision of Noda et al., Scientific Reports (revision in review). The Stage 2 encoder listed after them is used for the FRDR and HICEAS analyses of the final revision.
Corrected model pair
| File | Bytes | SHA-256 |
|---|---|---|
BEATs_DAPT_MAM_fixed_step127641.pt |
361345049 | 2a2d1d93f53ec29227bdd52da087fd0abcf0ce797c3c4a8629cd1435a314a6f9 |
sed_head_fixed_s42_ep7.pt |
17963283 | 9b2b202ab3e52b0d1efe4cd3479ee479db7646b0f42ab5b0e32f1e3ca551f119 |
Use these two files together. The encoder contains cfg and a model
state dictionary; the head contains a head state dictionary, the ordered
56-class labels, and training hparams. The files are PyTorch checkpoints,
not a Transformers from_pretrained bundle. Only load checkpoints from a
trusted source and verify their hashes before loading.
The head configuration retains generic original training/validation and log path strings for provenance; the referenced CSV contents are not included. No raw audio or individual reference embeddings are distributed in this pair.
Stage 2 encoder (final revision)
| File | Bytes | SHA-256 |
|---|---|---|
BEATs_DAPT_MAM_fixed_palaoa_step6385.pt |
361346274 | 4f7869751d7f15e3a806fb062902654597ca5566be610fedc1762c440d5c2a89 |
The manuscript keeps its two-stage DAPT design. Stage 2 continues from the
Stage 1 checkpoint above on the 2021 PALAOA subset (approximately 287 h) with
the same input-mask implementation and frozen teacher: 102,168 training windows
(1,032 of 103,200 held out by random_split with seed 42), batch size 16,
drop_last=True, 6,385 optimiser steps, encoder learning rate 1e-5 and
predictor/mask-token learning rate 1e-4. The final checkpoint is used; no
downstream selection was applied. Stage 2 changes the embeddings relative to
Stage 1: for the same 1,623 validation clips the mean cosine similarity between
Stage 1 and Stage 2 embeddings is 0.913 (minimum 0.793), so the CCED2 reference
of the final revision was refitted on Stage 2 embeddings. The 56-class SED head was trained on the Stage 1 encoder and
is not a matching head for the Stage 2 encoder.
Training and evaluation identity
- BEATs AS-2M (iter3+) initialisation; 16 kHz mono input.
- World-DAPT: 2,042,268 non-overlapping 10-s windows (approximately 5,673 h).
- Frozen teacher and fixed 1,024-cluster targets; 75% of student patch embeddings replaced by a trainable mask token before transformer encoding.
- Cross-entropy on masked positions; AdamW with encoder learning rate 1e-4 and predictor/mask-token learning rate 1e-3, 5% warm-up then cosine decay, and bfloat16 autocast.
- One shuffled pass, batch size 16,
drop_last=True: 127,641 optimiser steps; the incomplete final batch of 12 windows is unused. No DAPT validation split or downstream checkpoint selection was used for this endpoint. - Matching supervised head: seed 42, epoch 7. With this head, Event/Clip/2-s-segment F1 is 0.493/0.749/0.518. Table 1 of the final revision reports the mean of eight SED-head seeds (0.493/0.739/0.523), of which this head is one; the per-seed values are in the code and results release. These are frozen aggregate results, not an evaluation on a publicly reproducible 56-class dataset.
The retained BEATs+DAPT analyses use corrected window-aware extraction:
start_sec when supplied, otherwise max(0, center_sec - 5), with boundary
padding as needed. Reusing a legacy extractor can change results even when
the correct checkpoint is loaded.
Version history โ do not substitute legacy files
- December 2025 SimCLR-style DAPT was invalid because AMP fp16 prevented encoder updates. That analysis is superseded.
- May 2026 step-120,000 MAM is also superseded: masking was applied after unmasked audio passed through the encoder, rather than before transformer encoding. It does not implement the stated masked-input objective.
- September 2026 step-127,641 (Stage 1) and its PALAOA continuation at step 6,385 (Stage 2) use the corrected input-mask implementation.
The two May files remain available unchanged for historical provenance:
| Legacy file | SHA-256 |
|---|---|
beats_dapt_mam_step120000.pt |
0fe9f7dd92780c2e564f1df06a192482dbcb9a56bdab4202f4d94862b9168f89 |
sed_head_56_fulldata_ep8.pt |
135d11738a6619a57769955468ce5cb6eee3f07044fa45e6c950bf25ac4f8f60 |
Neither is a substitute for the corrected pair. The original model-card
history remains available at repository revision
dbb29a3dfc4fe1605c9fdd87079723db12903849; its May-era numerical and
availability claims are not the current minor-revision record.
Download and code
The corrected pair is fixed by tag v3.0.3-sr-minor-2026-09-13; tag
v3.1.6-sr-minor-2026-09-19 fixes the pair together with the Stage 2 encoder and
corresponds to the
code and results release of the final revision
(GitHub commit 9596fd72285413acfb3ccf208683c2c2d62470d2; model files unchanged since tag
v3.1.0-sr-minor-2026-09-19). To download the pair:
from huggingface_hub import hf_hub_download
for name in ["BEATs_DAPT_MAM_fixed_step127641.pt", "sed_head_fixed_s42_ep7.pt"]:
path = hf_hub_download(
repo_id="BiologgingSolutions/OceanBEATs",
revision="v3.0.3-sr-minor-2026-09-13",
filename=name,
local_dir="weights",
)
print(path)
Check each downloaded file with shasum -a 256 against the table above.
The corrected code and aggregate-results release
is GitHub commit fe3cc9f1a696e6814329120fdf0a64bc6327ed5c.
Its artifact map
defines the verification and full-rerun boundaries. That GitHub release
predates this model upload; its statement that corrected weights were not yet
public describes the release-time state. This model revision supplies the
two matching files, but does not remove the remaining data limitations.
Data and reproducibility limits
The internal 56-class audio, clip-level metadata and split membership remain non-public because of sensitive location/operational information and the permissions of the original collaborating organisations. The complete label taxonomy and aggregate statistics are reported in the manuscript materials; exact retraining and Table 1 evaluation cannot be reproduced publicly.
The corrected fitted CCED2 kNN reference contains individual restricted reference embeddings and is not distributed. This model pair does not supply the fitted CCED2 reference models, all evaluation embeddings, or window-level score arrays. Public aggregate checks are not equivalent to complete raw-data reproduction. Results are research diagnostics, not validated deployment performance or evidence that the complete DGPU loop caused the Promoter gain.
Licence and acknowledgement
The model weights retain CC BY 4.0, including commercial use with attribution; the companion source code is MIT-licensed. The DGPU framework and CCED2 score are subject to patent applications filed by Biologging Solutions Inc.; these copyright licences do not grant patent rights.
We acknowledge Microsoft BEATs and the World-DAPT source collections: NOAA SanctSound, US Navy USWTR, NOAA NRS/ONMS, ICListen / ONC, NPS Glacier Bay, and PALAOA Ekstrรถm Ice Shelf.
Citation
Noda, T. and Koizumi, T. Discovery and promotion of unknown sounds into operational detection targets for underwater passive acoustic monitoring under false alarm constraints. Scientific Reports (revision in review, 2026).