|
Download lecture_6/README.md from ChatterjeeLab/CIS6270: direct link, hf CLI and curl.
- Browser
- Download file 12.5 kB
-
https://huggingface.co/ChatterjeeLab/CIS6270/resolve/main/lecture_6/README.md
- Command line
-
hf download hf://ChatterjeeLab/CIS6270/lecture_6/README.md
-
curl -L -o README.md https://huggingface.co/ChatterjeeLab/CIS6270/resolve/main/lecture_6/README.md
12.5 kB
| # Lecture 6: Flow Maps | |
| Chapter 6 of the notes, `book/chapters/ch6_flow_maps.tex`. | |
| Chapters 2 to 5 learned a local change and applied it again and again, so the number of steps set | |
| the cost of a sample. In this lecture we learn the map between two times directly. The map | |
| `X_{s,t}(x)` is fixed by three axioms, identity, cocycle and tangent, and every file writes it the | |
| same way, as the state plus the elapsed time times an average velocity the network produces, which | |
| makes the identity axiom hold for any weights at all. We differentiate the cocycle in the end time | |
| to get the Lagrangian identity, in the start time to get the Eulerian identity, and we leave it | |
| finite to get the composition identity. Each one becomes a squared residual, we pair each with a | |
| flow matching anchor on the diagonal, and from there we carry the construction into five settings: a | |
| point in space, a distribution over tokens, a Brownian path, a dimension that grows, and a | |
| conditioning state. | |
| Every file is standalone, runs offline on a laptop CPU, and prints the numbers the chapter works | |
| by hand. | |
| ## The files | |
| | File | Command | What it builds | | |
| | --- | --- | --- | | |
| | `flow_map.py` | `python lecture_6/flow_map.py` | No learning. The chapter's velocity, the map by RK4, the three axioms, the inverse, the Jacobian and its differential equation, Liouville's formula, density transport and the generator | | |
| | `residuals.py` | `python lecture_6/residuals.py` | The Lagrangian, Eulerian and composition residuals and the diagonal anchor, from the chapter's listing, with four models trained on the four Gaussian blobs | | |
| | `meanflow.py` | `python lecture_6/meanflow.py` | MeanFlow: the Eulerian identity in the average-velocity parameterization, with the total derivative from one forward-mode pass | | |
| | `shortcut.py` | `python lecture_6/shortcut.py` | Shortcut models: the composition identity indexed by step size, on a dyadic ladder from 1/128 to 1 | | |
| | `consistency.py` | `python lecture_6/consistency.py` | Consistency models on the edge `t = 1` with the boundary coefficients, and consistency trajectory models on the whole triangle | | |
| | `fmlm.py` | `python lecture_6/fmlm.py` | Flow Map Language Models: the two-time denoiser, the convex form, the cocycle in denoiser coordinates, the two gradients a softmax head sees, two KL terms and the decoding clock, on DNA | | |
| | `categorical_map.py` | `python lecture_6/categorical_map.py` | Categorical Flow Maps: the Lagrangian identity with a softmax head, the teacher gap and drift, and the bound that replaces the residual, on phrases | | |
| | `discrete_map.py` | `python lecture_6/discrete_map.py` | Discrete Flow Maps: why the derivative targets leave the simplex, and the repair in logit space, on DNA | | |
| | `stochastic_map.py` | `python lecture_6/stochastic_map.py` | Strong stochastic flow maps: the solution map of an SDE, the weak map's floor, the Legendre coefficients and the Chen relations | | |
| | `expanding.py` | `python lecture_6/expanding.py` | Expanding flow maps: a state whose length grows, local clocks, the composition identity under growth, and one draw of the insertion noise | | |
| | `posterior_map.py` | `python lecture_6/posterior_map.py` | Diamond Maps and Meta Flow Maps: the inner clock, posterior samples, the value function from K samples, its gradient and the importance weight | | |
| Every file accepts the shared flags of `cis6270.runner`, and `--steps 20 --quiet` finishes in a | |
| few seconds. In `flow_map.py`, `stochastic_map.py` and `posterior_map.py` every quantity comes | |
| from the chapter's own formulas, so there `--steps` sets the Runge-Kutta substep count, the Monte | |
| Carlo draw count and the inner substep count, which each file's docstring says. | |
| ## What a default run prints | |
| The chapter's worked numbers come out exactly, and the test suite asserts them. | |
| | Quantity | Chapter | Printed | | |
| | --- | --- | --- | | |
| | `X_{0,1}(0.7)`, the cocycle through 0.5, the inverse, `det J_{0,1}` | 4.1, 4.1, 0.7, 3 | 4.100000, 4.100000, 0.700000, 3.000000 | | |
| | `log det J_{0,1}` by Liouville, and the density ratio | 1.0986, 3.000 | 1.0986, 3.0000 | | |
| | Stay-still residuals at `x = 1`, `s = 0`, `t = 1` | 1, 1, 0 | 1.0000, 1.0000, 0.0000 | | |
| | MeanFlow on `dx/dt = x`: `1.718x = x + 0.718x` | 1.718, 1, 0.718 | 1.7183, 1.0000, 0.7183 | | |
| | Shortcut doubling: two steps of 0.25 against one of 0.5 | 1.297443 | 1.297443 | | |
| | Consistency distillation residual at `h = 0.01` | 8.189e-05 | 8.189e-05 | | |
| | FMLM composed target, its four KL terms and their sum | (0.06, 0.12, 0.76, 0.06), 0.030747 | the same, 0.030747 | | |
| | The two gradients in the logits at that target and student | (0.04, 0.03, -0.11, 0.04) and (0.0099, 0.01335, -0.03315, 0.0099) | the same | | |
| | The same two once the student saturates to `q_1 = 1e-06` | -0.0600 and -4.62e-08 | -0.0600 and -4.6171e-08 | | |
| | The decoding clock at 0.60, 0.75, 0.80, 0.85 | 0.006, 0.141, 0.456, 0.931 | 0.0057, 0.1411, 0.4561, 0.9313 | | |
| | Four steps on the clock | 0.773490, 0.804481, 0.828095 | the same | | |
| | Gamma on the last three times, and over the long first leg | 0.494620, 0.830144 | the same | | |
| | Categorical gap and drift, and the bound around a zero residual | antiparallel, 0.199 | 1e-34 residual, 0.1989 | | |
| | The same two terms averaged over three thousand states | 0.393, 0.095 | 0.3955, 0.0963 | | |
| | Discrete derivative target, three entries negative | (-0.3333, -0.6667, 1.3333, -0.3333) | the same | | |
| | The logit repair of (-0.284, 0.828, 0.456) | (0.0194, 0.6259, 0.3547) | the same | | |
| | The weak map's floor `2K_1` | 6.389 | 6.3886 by Monte Carlo | | |
| | Chen fusion of four half-interval draws | 0.50, -0.20 | 0.50, -0.20 | | |
| | Local clocks at `t = 0.8`, and omega at (0.3, 0.6, 0.9) | 0.800, 0.692, 0.333 and 0.125 | the same | | |
| | Jensen gap, and `E[V_K]` at K = 1, 2, 4, 8 | 0.434 and 1.000, 1.217, 1.340, 1.393 | the same | | |
| | Meta Flow Maps fixed point against the per-sample ratio | 0.625 and 0.540 | the same | | |
| For the saturated pair in that table, `fmlm.py` works the full expression of Section 6.27.4, the | |
| gradient of half the squared error through the softmax Jacobian, | |
| d/dz_i (1/2) sum_j (q_j - p_j)^2 = q_i (q_i - p_i) - q_i sum_j q_j (q_j - p_j) , | |
| at `p = (0.06, 0.12, 0.76, 0.06)` with the student moved to `q_1 = 1e-06` and the other three | |
| renormalized. We also differentiate the two losses automatically and compare, and the two routes | |
| agree to four figures on every coordinate. As you can see from the pair, `-0.0600` against | |
| `-4.6171e-08` is the six orders of magnitude the chapter draws its point from. | |
| We take the trained numbers from one default run on a laptop CPU, with seed 0. | |
| | Run | Headline | | |
| | --- | --- | | |
| | `residuals.py`, 600 steps each | distance to its own ODE 0.900 for the diagonal alone, 0.481, 0.376 and 0.541 with the Lagrangian, Eulerian and composition residuals; fraction on a mode 0.44 against 0.74, 0.66 and 0.70 | | |
| | `meanflow.py`, 2000 steps | nearest mode 0.627 at one step, 0.522 at two, 0.505 at four, all four modes found | | |
| | `shortcut.py`, 2000 steps | nearest mode 0.605 at one step and 0.508 at eight; the doubling residual rises from below 1e-04 at `d = 1/128` to 0.034 at `d = 1/2` | | |
| | `consistency.py`, 1500 steps each | nearest mode 0.513 for the consistency model at one step, 0.515, 0.420 and 0.470 for the trajectory model at one, two and four | | |
| | `fmlm.py`, 1500 steps | GC 0.586 at one step against 0.581 in the training data, motif 0.63, token entropy 1.371 of a possible 1.386, every sequence distinct | | |
| | `categorical_map.py`, 1000 steps each | every phrase in the language at both step counts; trained on the residual, 0.56 of one-step phrases distinct at entropy 2.75; trained on the bound, 0.33 at 2.54, and 0.12 at 2.05 by four steps | | |
| | `discrete_map.py`, 1500 steps | GC 0.609 at one step, motif 0.81, entropy 1.362, every sequence distinct; the unrepaired additive target is negative on 0.05% of its entries | | |
| | `stochastic_map.py` | the strong map's terminal cloud sits 8.8e-04 from the finest crossing at two hops, 1.1e-04 at four, 1.4e-05 at eight and 2e-06 at sixteen, the weak map's 3.06 at every hop count | | |
| | `expanding.py`, 1500 steps | GC 0.617 at one step, 0.621 on coordinates born at the source and 0.613 on coordinates born later; motif 0.75 | | |
| | `posterior_map.py` | on the blobs, the exact value -0.1127, the 16-sample estimate -0.1088, the single endpoint -0.1628 | | |
| The DNA files draw 1024 training sequences of the default length 16, which have GC fraction 0.581 | |
| and motif fraction 0.749. They read 64 samples, as the phrase file does, so a reported fraction | |
| has about 0.06 of sampling noise and its second digit moves between seeds. | |
| Longer runs, where they change the picture: | |
| - `residuals.py --steps 2400` moves the four distances to 1.033, 0.171, 0.247 and 0.511 and the | |
| four mode fractions to 0.27, 0.82, 0.71 and 0.79, against Table 6.2's 0.896, 0.292, 0.242 and | |
| 0.325 at 0.12, 0.84, 0.82 and 0.84. The gap between the anchor alone and the anchor plus any | |
| residual widens with the budget, and the three residuals stay inside each other's spread. | |
| - `meanflow.py --steps 8000` brings the one-step nearest mode from 0.627 to 0.437, below its own | |
| two-step and four-step figures, so the one-step map overtakes the grid. | |
| - `categorical_map.py --steps 3000` widens the gap between the two students, to 0.60 of one-step | |
| phrases distinct for the residual against 0.27 for the bound, and 0.54 against 0.09 at four | |
| steps. The bounded student loses variety as its budget grows, as Section 6.28.4 predicts, since | |
| the bound is minimized by a prediction that is constant in the end time. | |
| - The DNA runs are at their plateau by their default budgets. At `--steps 6000` the one-step motif | |
| fraction of `fmlm.py` reads 0.61 and of `discrete_map.py` 0.69, inside the sampling noise of the | |
| default run. | |
| At the default flags on one core, we get three or four seconds for `flow_map.py`, | |
| `stochastic_map.py` and `posterior_map.py`, which train nothing, and twelve to twenty-seven | |
| seconds for the eight that do, including `residuals.py` with its four models and | |
| `categorical_map.py` and `consistency.py` with their two. | |
| ## What is simplified | |
| - The four blob models of `residuals.py` train for 600 steps on a four-layer network, so the | |
| distances are larger than Table 6.2 of the chapter, which trains longer. The ordering the | |
| chapter draws, the anchor alone against the anchor plus any residual, comes out at this budget. | |
| - `shortcut.py` evaluates the zero-step term and the doubling term on the whole batch and weights | |
| them three to one, in place of the chapter's split of the batch in that proportion. The two have | |
| the same expectation and the weighted form costs one network call more per step. | |
| - `consistency.py` distils from the exact marginal velocity of the training set, which is a | |
| softmax over the data points, in place of a separately trained flow matching teacher. The | |
| teacher's discretization error is then the only error the student inherits, as Section 6.20 sets | |
| out. | |
| - `posterior_map.py` builds the inner map from the exact posterior over the training points, in | |
| place of training a conditional family, so the inner velocity is a closed form and the value, | |
| its gradient and the Jensen gap are all available to compare against. | |
| - We compute the decoding clock of `fmlm.py` by quadrature on the one-position integral, which is | |
| exact for the Gaussian interpolant and treats each position on its own, so the clock stops where | |
| correlations between positions begin. | |
| - `categorical_map.py` and `discrete_map.py` keep the end time below one, because the drift term | |
| of each loss divides by the remaining time. `fmlm.py` has no such term, so it draws a quarter of | |
| each batch at `t = 1`, the slice Section 6.26 asks for by name. | |
| - In `stochastic_map.py` we carry three Legendre coefficients and fuse them with the Chen | |
| relations written out for `n = 0, 1, 2`. We stop at three, because the chapter lists the | |
| constants `c_{3,m}` only up to that order. | |
| - Every network is `cis6270.nets.TwoTimeMLP`. The discrete files flatten the one-hot lift into its | |
| input and read its output as per-position logits, so one architecture covers both halves of the | |
| lecture and any difference in the samples comes from the loss. | |
| - Each file sets `torch.set_num_threads(1)`. These networks are small enough that splitting one | |
| matrix product across cores costs more than it saves. `tests/conftest.py` sets the same thing | |
| for the test session, and the line in each file acts when a reader runs the script directly. | |