|
Download lecture_6/data/README.md from ChatterjeeLab/CIS6270: direct link, hf CLI and curl.
- Browser
- Download file 1.68 kB
-
https://huggingface.co/ChatterjeeLab/CIS6270/resolve/main/lecture_6/data/README.md
- Command line
-
hf download hf://ChatterjeeLab/CIS6270/lecture_6/data/README.md
-
curl -L -o README.md https://huggingface.co/ChatterjeeLab/CIS6270/resolve/main/lecture_6/data/README.md
1.68 kB
| # Data for Lecture 6 | |
| `phrases.txt` and `variable_text.txt` were generated deterministically with Python | |
| `random.Random(6270)`. Each contains 640 examples. Their vocabulary describes | |
| colored shapes and directions; no external dataset is needed. | |
| | Color | Shape | Direction | | |
| | --- | --- | --- | | |
| | red | circle | left | | |
| | blue | square | right | | |
| | green | triangle | up | | |
| | gold | star | down | | |
| The fixed-length corpus has four tokens per line. The variable-length corpus | |
| uses `COLOR SHAPE`, `COLOR SHAPE moves DIRECTION`, or the latter followed by | |
| `slowly` or `twice`. The color, shape, and direction remain correlated. This | |
| makes invalid independent combinations easy to inspect. | |
| Training uses a seeded shuffle and 80/20 example split. The grammar has a small | |
| support, so examples repeat and both partitions contain the same grammar. | |
| Training-support fraction measures membership in this small support; it is | |
| not a test of open-ended language generalization. | |
| For a custom corpus, supply a UTF-8 text file with one whitespace-tokenized | |
| sequence per line. Fixed-length methods require equal token counts. The | |
| expanding method accepts variable lengths up to 32. Vocabulary and maximum | |
| length are saved in the checkpoint, so sampling does not need the original file. | |
| The continuous dataset is sampled from four equally weighted Gaussians centered | |
| at `(−1.5,−1.5)`, `(−1.5,1.5)`, `(1.5,−1.5)`, and `(1.5,1.5)`, with coordinate | |
| standard deviation 0.22. The latent example embeds these points on | |
| `(x, y, 0.3(x²−y²))` before learning its autoencoder. SSFM uses analytic | |
| Ornstein-Uhlenbeck marginals with initial variance one, drift `−x`, and diffusion | |
| coefficient 0.7. | |