CIS6270 / lecture_6 /data /README.md
pranamanam's picture
Add Lecture 6 flow maps with tested training, sampling, and course navigation
da3babb verified
|
Raw History Blame Contribute Delete
1.68 kB
# Data for Lecture 6
`phrases.txt` and `variable_text.txt` were generated deterministically with Python
`random.Random(6270)`. Each contains 640 examples. Their vocabulary describes
colored shapes and directions; no external dataset is needed.
| Color | Shape | Direction |
| --- | --- | --- |
| red | circle | left |
| blue | square | right |
| green | triangle | up |
| gold | star | down |
The fixed-length corpus has four tokens per line. The variable-length corpus
uses `COLOR SHAPE`, `COLOR SHAPE moves DIRECTION`, or the latter followed by
`slowly` or `twice`. The color, shape, and direction remain correlated. This
makes invalid independent combinations easy to inspect.
Training uses a seeded shuffle and 80/20 example split. The grammar has a small
support, so examples repeat and both partitions contain the same grammar.
Training-support fraction measures membership in this small support; it is
not a test of open-ended language generalization.
For a custom corpus, supply a UTF-8 text file with one whitespace-tokenized
sequence per line. Fixed-length methods require equal token counts. The
expanding method accepts variable lengths up to 32. Vocabulary and maximum
length are saved in the checkpoint, so sampling does not need the original file.
The continuous dataset is sampled from four equally weighted Gaussians centered
at `(−1.5,−1.5)`, `(−1.5,1.5)`, `(1.5,−1.5)`, and `(1.5,1.5)`, with coordinate
standard deviation 0.22. The latent example embeds these points on
`(x, y, 0.3(x²−y²))` before learning its autoencoder. SSFM uses analytic
Ornstein-Uhlenbeck marginals with initial variance one, drift `−x`, and diffusion
coefficient 0.7.