# Data for Lecture 6 `phrases.txt` and `variable_text.txt` were generated deterministically with Python `random.Random(6270)`. Each contains 640 examples. Their vocabulary describes colored shapes and directions; no external dataset is needed. | Color | Shape | Direction | | --- | --- | --- | | red | circle | left | | blue | square | right | | green | triangle | up | | gold | star | down | The fixed-length corpus has four tokens per line. The variable-length corpus uses `COLOR SHAPE`, `COLOR SHAPE moves DIRECTION`, or the latter followed by `slowly` or `twice`. The color, shape, and direction remain correlated. This makes invalid independent combinations easy to inspect. Training uses a seeded shuffle and 80/20 example split. The grammar has a small support, so examples repeat and both partitions contain the same grammar. Training-support fraction measures membership in this small support; it is not a test of open-ended language generalization. For a custom corpus, supply a UTF-8 text file with one whitespace-tokenized sequence per line. Fixed-length methods require equal token counts. The expanding method accepts variable lengths up to 32. Vocabulary and maximum length are saved in the checkpoint, so sampling does not need the original file. The continuous dataset is sampled from four equally weighted Gaussians centered at `(−1.5,−1.5)`, `(−1.5,1.5)`, `(1.5,−1.5)`, and `(1.5,1.5)`, with coordinate standard deviation 0.22. The latent example embeds these points on `(x, y, 0.3(x²−y²))` before learning its autoencoder. SSFM uses analytic Ornstein-Uhlenbeck marginals with initial variance one, drift `−x`, and diffusion coefficient 0.7.