CIS6270 / lecture_6 /README.md
pranamanam's picture
Upload 87 files
0ba9d09 verified
|
Raw History Blame Contribute Delete
12.5 kB
# Lecture 6: Flow Maps
Chapter 6 of the notes, `book/chapters/ch6_flow_maps.tex`.
Chapters 2 to 5 learned a local change and applied it again and again, so the number of steps set
the cost of a sample. In this lecture we learn the map between two times directly. The map
`X_{s,t}(x)` is fixed by three axioms, identity, cocycle and tangent, and every file writes it the
same way, as the state plus the elapsed time times an average velocity the network produces, which
makes the identity axiom hold for any weights at all. We differentiate the cocycle in the end time
to get the Lagrangian identity, in the start time to get the Eulerian identity, and we leave it
finite to get the composition identity. Each one becomes a squared residual, we pair each with a
flow matching anchor on the diagonal, and from there we carry the construction into five settings: a
point in space, a distribution over tokens, a Brownian path, a dimension that grows, and a
conditioning state.
Every file is standalone, runs offline on a laptop CPU, and prints the numbers the chapter works
by hand.
## The files
| File | Command | What it builds |
| --- | --- | --- |
| `flow_map.py` | `python lecture_6/flow_map.py` | No learning. The chapter's velocity, the map by RK4, the three axioms, the inverse, the Jacobian and its differential equation, Liouville's formula, density transport and the generator |
| `residuals.py` | `python lecture_6/residuals.py` | The Lagrangian, Eulerian and composition residuals and the diagonal anchor, from the chapter's listing, with four models trained on the four Gaussian blobs |
| `meanflow.py` | `python lecture_6/meanflow.py` | MeanFlow: the Eulerian identity in the average-velocity parameterization, with the total derivative from one forward-mode pass |
| `shortcut.py` | `python lecture_6/shortcut.py` | Shortcut models: the composition identity indexed by step size, on a dyadic ladder from 1/128 to 1 |
| `consistency.py` | `python lecture_6/consistency.py` | Consistency models on the edge `t = 1` with the boundary coefficients, and consistency trajectory models on the whole triangle |
| `fmlm.py` | `python lecture_6/fmlm.py` | Flow Map Language Models: the two-time denoiser, the convex form, the cocycle in denoiser coordinates, the two gradients a softmax head sees, two KL terms and the decoding clock, on DNA |
| `categorical_map.py` | `python lecture_6/categorical_map.py` | Categorical Flow Maps: the Lagrangian identity with a softmax head, the teacher gap and drift, and the bound that replaces the residual, on phrases |
| `discrete_map.py` | `python lecture_6/discrete_map.py` | Discrete Flow Maps: why the derivative targets leave the simplex, and the repair in logit space, on DNA |
| `stochastic_map.py` | `python lecture_6/stochastic_map.py` | Strong stochastic flow maps: the solution map of an SDE, the weak map's floor, the Legendre coefficients and the Chen relations |
| `expanding.py` | `python lecture_6/expanding.py` | Expanding flow maps: a state whose length grows, local clocks, the composition identity under growth, and one draw of the insertion noise |
| `posterior_map.py` | `python lecture_6/posterior_map.py` | Diamond Maps and Meta Flow Maps: the inner clock, posterior samples, the value function from K samples, its gradient and the importance weight |
Every file accepts the shared flags of `cis6270.runner`, and `--steps 20 --quiet` finishes in a
few seconds. In `flow_map.py`, `stochastic_map.py` and `posterior_map.py` every quantity comes
from the chapter's own formulas, so there `--steps` sets the Runge-Kutta substep count, the Monte
Carlo draw count and the inner substep count, which each file's docstring says.
## What a default run prints
The chapter's worked numbers come out exactly, and the test suite asserts them.
| Quantity | Chapter | Printed |
| --- | --- | --- |
| `X_{0,1}(0.7)`, the cocycle through 0.5, the inverse, `det J_{0,1}` | 4.1, 4.1, 0.7, 3 | 4.100000, 4.100000, 0.700000, 3.000000 |
| `log det J_{0,1}` by Liouville, and the density ratio | 1.0986, 3.000 | 1.0986, 3.0000 |
| Stay-still residuals at `x = 1`, `s = 0`, `t = 1` | 1, 1, 0 | 1.0000, 1.0000, 0.0000 |
| MeanFlow on `dx/dt = x`: `1.718x = x + 0.718x` | 1.718, 1, 0.718 | 1.7183, 1.0000, 0.7183 |
| Shortcut doubling: two steps of 0.25 against one of 0.5 | 1.297443 | 1.297443 |
| Consistency distillation residual at `h = 0.01` | 8.189e-05 | 8.189e-05 |
| FMLM composed target, its four KL terms and their sum | (0.06, 0.12, 0.76, 0.06), 0.030747 | the same, 0.030747 |
| The two gradients in the logits at that target and student | (0.04, 0.03, -0.11, 0.04) and (0.0099, 0.01335, -0.03315, 0.0099) | the same |
| The same two once the student saturates to `q_1 = 1e-06` | -0.0600 and -4.62e-08 | -0.0600 and -4.6171e-08 |
| The decoding clock at 0.60, 0.75, 0.80, 0.85 | 0.006, 0.141, 0.456, 0.931 | 0.0057, 0.1411, 0.4561, 0.9313 |
| Four steps on the clock | 0.773490, 0.804481, 0.828095 | the same |
| Gamma on the last three times, and over the long first leg | 0.494620, 0.830144 | the same |
| Categorical gap and drift, and the bound around a zero residual | antiparallel, 0.199 | 1e-34 residual, 0.1989 |
| The same two terms averaged over three thousand states | 0.393, 0.095 | 0.3955, 0.0963 |
| Discrete derivative target, three entries negative | (-0.3333, -0.6667, 1.3333, -0.3333) | the same |
| The logit repair of (-0.284, 0.828, 0.456) | (0.0194, 0.6259, 0.3547) | the same |
| The weak map's floor `2K_1` | 6.389 | 6.3886 by Monte Carlo |
| Chen fusion of four half-interval draws | 0.50, -0.20 | 0.50, -0.20 |
| Local clocks at `t = 0.8`, and omega at (0.3, 0.6, 0.9) | 0.800, 0.692, 0.333 and 0.125 | the same |
| Jensen gap, and `E[V_K]` at K = 1, 2, 4, 8 | 0.434 and 1.000, 1.217, 1.340, 1.393 | the same |
| Meta Flow Maps fixed point against the per-sample ratio | 0.625 and 0.540 | the same |
For the saturated pair in that table, `fmlm.py` works the full expression of Section 6.27.4, the
gradient of half the squared error through the softmax Jacobian,
d/dz_i (1/2) sum_j (q_j - p_j)^2 = q_i (q_i - p_i) - q_i sum_j q_j (q_j - p_j) ,
at `p = (0.06, 0.12, 0.76, 0.06)` with the student moved to `q_1 = 1e-06` and the other three
renormalized. We also differentiate the two losses automatically and compare, and the two routes
agree to four figures on every coordinate. As you can see from the pair, `-0.0600` against
`-4.6171e-08` is the six orders of magnitude the chapter draws its point from.
We take the trained numbers from one default run on a laptop CPU, with seed 0.
| Run | Headline |
| --- | --- |
| `residuals.py`, 600 steps each | distance to its own ODE 0.900 for the diagonal alone, 0.481, 0.376 and 0.541 with the Lagrangian, Eulerian and composition residuals; fraction on a mode 0.44 against 0.74, 0.66 and 0.70 |
| `meanflow.py`, 2000 steps | nearest mode 0.627 at one step, 0.522 at two, 0.505 at four, all four modes found |
| `shortcut.py`, 2000 steps | nearest mode 0.605 at one step and 0.508 at eight; the doubling residual rises from below 1e-04 at `d = 1/128` to 0.034 at `d = 1/2` |
| `consistency.py`, 1500 steps each | nearest mode 0.513 for the consistency model at one step, 0.515, 0.420 and 0.470 for the trajectory model at one, two and four |
| `fmlm.py`, 1500 steps | GC 0.586 at one step against 0.581 in the training data, motif 0.63, token entropy 1.371 of a possible 1.386, every sequence distinct |
| `categorical_map.py`, 1000 steps each | every phrase in the language at both step counts; trained on the residual, 0.56 of one-step phrases distinct at entropy 2.75; trained on the bound, 0.33 at 2.54, and 0.12 at 2.05 by four steps |
| `discrete_map.py`, 1500 steps | GC 0.609 at one step, motif 0.81, entropy 1.362, every sequence distinct; the unrepaired additive target is negative on 0.05% of its entries |
| `stochastic_map.py` | the strong map's terminal cloud sits 8.8e-04 from the finest crossing at two hops, 1.1e-04 at four, 1.4e-05 at eight and 2e-06 at sixteen, the weak map's 3.06 at every hop count |
| `expanding.py`, 1500 steps | GC 0.617 at one step, 0.621 on coordinates born at the source and 0.613 on coordinates born later; motif 0.75 |
| `posterior_map.py` | on the blobs, the exact value -0.1127, the 16-sample estimate -0.1088, the single endpoint -0.1628 |
The DNA files draw 1024 training sequences of the default length 16, which have GC fraction 0.581
and motif fraction 0.749. They read 64 samples, as the phrase file does, so a reported fraction
has about 0.06 of sampling noise and its second digit moves between seeds.
Longer runs, where they change the picture:
- `residuals.py --steps 2400` moves the four distances to 1.033, 0.171, 0.247 and 0.511 and the
four mode fractions to 0.27, 0.82, 0.71 and 0.79, against Table 6.2's 0.896, 0.292, 0.242 and
0.325 at 0.12, 0.84, 0.82 and 0.84. The gap between the anchor alone and the anchor plus any
residual widens with the budget, and the three residuals stay inside each other's spread.
- `meanflow.py --steps 8000` brings the one-step nearest mode from 0.627 to 0.437, below its own
two-step and four-step figures, so the one-step map overtakes the grid.
- `categorical_map.py --steps 3000` widens the gap between the two students, to 0.60 of one-step
phrases distinct for the residual against 0.27 for the bound, and 0.54 against 0.09 at four
steps. The bounded student loses variety as its budget grows, as Section 6.28.4 predicts, since
the bound is minimized by a prediction that is constant in the end time.
- The DNA runs are at their plateau by their default budgets. At `--steps 6000` the one-step motif
fraction of `fmlm.py` reads 0.61 and of `discrete_map.py` 0.69, inside the sampling noise of the
default run.
At the default flags on one core, we get three or four seconds for `flow_map.py`,
`stochastic_map.py` and `posterior_map.py`, which train nothing, and twelve to twenty-seven
seconds for the eight that do, including `residuals.py` with its four models and
`categorical_map.py` and `consistency.py` with their two.
## What is simplified
- The four blob models of `residuals.py` train for 600 steps on a four-layer network, so the
distances are larger than Table 6.2 of the chapter, which trains longer. The ordering the
chapter draws, the anchor alone against the anchor plus any residual, comes out at this budget.
- `shortcut.py` evaluates the zero-step term and the doubling term on the whole batch and weights
them three to one, in place of the chapter's split of the batch in that proportion. The two have
the same expectation and the weighted form costs one network call more per step.
- `consistency.py` distils from the exact marginal velocity of the training set, which is a
softmax over the data points, in place of a separately trained flow matching teacher. The
teacher's discretization error is then the only error the student inherits, as Section 6.20 sets
out.
- `posterior_map.py` builds the inner map from the exact posterior over the training points, in
place of training a conditional family, so the inner velocity is a closed form and the value,
its gradient and the Jensen gap are all available to compare against.
- We compute the decoding clock of `fmlm.py` by quadrature on the one-position integral, which is
exact for the Gaussian interpolant and treats each position on its own, so the clock stops where
correlations between positions begin.
- `categorical_map.py` and `discrete_map.py` keep the end time below one, because the drift term
of each loss divides by the remaining time. `fmlm.py` has no such term, so it draws a quarter of
each batch at `t = 1`, the slice Section 6.26 asks for by name.
- In `stochastic_map.py` we carry three Legendre coefficients and fuse them with the Chen
relations written out for `n = 0, 1, 2`. We stop at three, because the chapter lists the
constants `c_{3,m}` only up to that order.
- Every network is `cis6270.nets.TwoTimeMLP`. The discrete files flatten the one-hot lift into its
input and read its output as per-position logits, so one architecture covers both halves of the
lecture and any difference in the samples comes from the loss.
- Each file sets `torch.set_num_threads(1)`. These networks are small enough that splitting one
matrix product across cores costs more than it saves. `tests/conftest.py` sets the same thing
for the test session, and the line in each file acts when a reader runs the script directly.