Download lecture_6/data/README.md from ChatterjeeLab/CIS6270: direct link, hf CLI and curl.
- Browser
- Download file 1.68 kB
-
https://huggingface.co/ChatterjeeLab/CIS6270/resolve/main/lecture_6/data/README.md
- Command line
-
hf download hf://ChatterjeeLab/CIS6270/lecture_6/data/README.md
-
curl -L -o README.md https://huggingface.co/ChatterjeeLab/CIS6270/resolve/main/lecture_6/data/README.md
Data for Lecture 6
phrases.txt and variable_text.txt were generated deterministically with Python
random.Random(6270). Each contains 640 examples. Their vocabulary describes
colored shapes and directions; no external dataset is needed.
| Color | Shape | Direction |
|---|---|---|
| red | circle | left |
| blue | square | right |
| green | triangle | up |
| gold | star | down |
The fixed-length corpus has four tokens per line. The variable-length corpus
uses COLOR SHAPE, COLOR SHAPE moves DIRECTION, or the latter followed by
slowly or twice. The color, shape, and direction remain correlated. This
makes invalid independent combinations easy to inspect.
Training uses a seeded shuffle and 80/20 example split. The grammar has a small support, so examples repeat and both partitions contain the same grammar. Training-support fraction measures membership in this small support; it is not a test of open-ended language generalization.
For a custom corpus, supply a UTF-8 text file with one whitespace-tokenized sequence per line. Fixed-length methods require equal token counts. The expanding method accepts variable lengths up to 32. Vocabulary and maximum length are saved in the checkpoint, so sampling does not need the original file.
The continuous dataset is sampled from four equally weighted Gaussians centered
at (−1.5,−1.5), (−1.5,1.5), (1.5,−1.5), and (1.5,1.5), with coordinate
standard deviation 0.22. The latent example embeds these points on
(x, y, 0.3(x²−y²)) before learning its autoencoder. SSFM uses analytic
Ornstein-Uhlenbeck marginals with initial variance one, drift −x, and diffusion
coefficient 0.7.