colaBertFt checkpoint archive

This repository is a preservation archive for checkpoints produced by the colaBertFt research code. The models use a BERT-large-shaped architecture with tensorized attention and feed-forward layers. One pair uses low-rank activations throughout; the other restricts the low-rank activation behavior according to only_lr_act in its config.

This is an archival release, not a polished inference package. Loading the exports may require the custom modeling code in the source repository.

Contents

Remote path Description
bak/chkpt_not_onlylr.pt Epoch-2 PyTorch training snapshot with model, optimizer, gradient-scaler, and epoch state.
bak/chkpt_onlylr.pt Epoch-2 only_lr_act training snapshot with model, optimizer, gradient-scaler, and epoch state.
checkpoints/cola_bert_llama_mlp_large_0.7-004/ FP16 model state dict and architecture config corresponding exactly to the model tensors in bak/chkpt_not_onlylr.pt.
checkpoints/cola_bert_llama_mlp_large_0.7_only_lr-001/ FP16 model state dict and architecture config corresponding exactly to the model tensors in bak/chkpt_onlylr.pt.
model_checkpoints/rte_cola_bert_llama_mlp_large_0.7_only_lr-001/ PyTorch Lightning RTE checkpoint from epoch 1/global step 68. Its filename records validation accuracy 0.7329; this archive did not independently reproduce that evaluation.
model_checkpoints/colaBert_rte_mlpId_0.7329_2e2e91ae/ Converted FP16 safetensors RTE export, config, and the original generated model card.

Exact SHA-256 checksums are stored in ARTIFACT_CHECKSUMS.sha256 here and in the source repository.

Provenance and training

  • Base model named by the training configuration: bert-large-uncased.
  • Downstream task: GLUE RTE.
  • RTE Lightning checkpoint metadata: 2 labels, 2,490 training samples, learning rate 2e-5, Adam epsilon 1e-8, warmup proportion 0.1, weight decay 0.01, and 117 planned training steps.
  • Archive source files were created in August and September 2025 and archived in July 2026.

The two 2.4 GB training snapshots contain optimizer state and are retained for exact training-resume preservation. Their 615 model tensors match their corresponding smaller model.pt exports exactly.

Loading and safety

The .pt and .ckpt files use Python/PyTorch serialization. Only load them when you trust this archive, and prefer a current PyTorch release with weights_only=True where compatible. The .safetensors export avoids pickle deserialization but still requires the matching custom architecture.

These artifacts are research outputs and have not been audited for production use, bias, robustness, or generalization beyond the recorded RTE experiment.

Licensing

The upstream bert-large-uncased model is distributed under Apache-2.0. The source project does not currently declare a license of its own, so this archive does not make an additional licensing grant for the custom code or derived artifacts. Users are responsible for checking the applicable upstream terms.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support