|
Download README.md from mooreanthony/tmp-multitask: direct link, hf CLI and curl.
- Browser
- Download file 2.45 kB
-
https://huggingface.co/mooreanthony/tmp-multitask/resolve/main/README.md
- Command line
-
hf download hf://mooreanthony/tmp-multitask/README.md
-
curl -L -o README.md https://huggingface.co/mooreanthony/tmp-multitask/resolve/main/README.md
2.45 kB
| license: bsd-3-clause | |
| tags: | |
| - pytorch | |
| - coca | |
| - multitask | |
| # Coca for Multitask | |
| ## Overview | |
| A small **Coca** implementation for **Multitask**, packaged with an explicit configuration and an initialization checkpoint. The **huge** variant is a reproducible starting point, not a trained model release. | |
| ## Repository status | |
| - The Python file contains the model and runnable example or training entry point. | |
| - `config.json` records the generated architecture settings. | |
| - `training_args.json` records the default experiment recipe. | |
| - `model.safetensors` is a valid initialization checkpoint for smoke tests; it is **not** presented as a trained benchmark checkpoint. | |
| - No benchmark score is claimed in this repository. | |
| ## Architecture | |
| | Item | Value | | |
| |---|---| | |
| | Architecture | Coca | | |
| | Scale | huge | | |
| | Attention | sparse | | |
| | Fusion | tucker | | |
| | Activation | mish | | |
| | Normalization | batchnorm | | |
| ## Default experiment recipe | |
| The included configuration uses **adafactor** with a **exponential** schedule. These are starting values in the script, not evidence of a completed run. For a meaningful evaluation, train all baselines with the same data exposure, tuning budget, and random seeds. | |
| ## Quick check | |
| ```bash | |
| python predict.py --help | |
| ``` | |
| Inspect the script's `__main__` block for its generated smoke-test example. Because this is a custom implementation, generic automatic loading APIs require an explicit adapter before use. | |
| ## Evaluation guidance | |
| A useful first evaluation would use **a task-specific held-out set**, report the task metric across at least three seeds, and include a matched-capacity baseline. Keep training logs and environment versions with any published result. | |
| ## Limitations | |
| The initialization checkpoint has not been trained or audited for robustness, fairness, or domain transfer. The implementation should be treated as an experimental starting point. Results from a future trained checkpoint must be documented separately from the defaults shipped here. | |
| ## Files | |
| - `predict.py` — primary artifact | |
| - `README.md` — this documentation | |
| - `config.json` — architecture configuration | |
| - `training_args.json` — default experiment settings | |
| - `model.safetensors` — initialization checkpoint | |
| ## License | |
| Released under **bsd-3-clause**. Review the source-data terms separately when this repository is used with external datasets. | |