CIS6270 / lecture_6 /README.md
pranamanam's picture
Upload 87 files
0ba9d09 verified
|
Raw History Blame Contribute Delete
12.5 kB

Lecture 6: Flow Maps

Chapter 6 of the notes, book/chapters/ch6_flow_maps.tex.

Chapters 2 to 5 learned a local change and applied it again and again, so the number of steps set the cost of a sample. In this lecture we learn the map between two times directly. The map X_{s,t}(x) is fixed by three axioms, identity, cocycle and tangent, and every file writes it the same way, as the state plus the elapsed time times an average velocity the network produces, which makes the identity axiom hold for any weights at all. We differentiate the cocycle in the end time to get the Lagrangian identity, in the start time to get the Eulerian identity, and we leave it finite to get the composition identity. Each one becomes a squared residual, we pair each with a flow matching anchor on the diagonal, and from there we carry the construction into five settings: a point in space, a distribution over tokens, a Brownian path, a dimension that grows, and a conditioning state.

Every file is standalone, runs offline on a laptop CPU, and prints the numbers the chapter works by hand.

The files

File Command What it builds
flow_map.py python lecture_6/flow_map.py No learning. The chapter's velocity, the map by RK4, the three axioms, the inverse, the Jacobian and its differential equation, Liouville's formula, density transport and the generator
residuals.py python lecture_6/residuals.py The Lagrangian, Eulerian and composition residuals and the diagonal anchor, from the chapter's listing, with four models trained on the four Gaussian blobs
meanflow.py python lecture_6/meanflow.py MeanFlow: the Eulerian identity in the average-velocity parameterization, with the total derivative from one forward-mode pass
shortcut.py python lecture_6/shortcut.py Shortcut models: the composition identity indexed by step size, on a dyadic ladder from 1/128 to 1
consistency.py python lecture_6/consistency.py Consistency models on the edge t = 1 with the boundary coefficients, and consistency trajectory models on the whole triangle
fmlm.py python lecture_6/fmlm.py Flow Map Language Models: the two-time denoiser, the convex form, the cocycle in denoiser coordinates, the two gradients a softmax head sees, two KL terms and the decoding clock, on DNA
categorical_map.py python lecture_6/categorical_map.py Categorical Flow Maps: the Lagrangian identity with a softmax head, the teacher gap and drift, and the bound that replaces the residual, on phrases
discrete_map.py python lecture_6/discrete_map.py Discrete Flow Maps: why the derivative targets leave the simplex, and the repair in logit space, on DNA
stochastic_map.py python lecture_6/stochastic_map.py Strong stochastic flow maps: the solution map of an SDE, the weak map's floor, the Legendre coefficients and the Chen relations
expanding.py python lecture_6/expanding.py Expanding flow maps: a state whose length grows, local clocks, the composition identity under growth, and one draw of the insertion noise
posterior_map.py python lecture_6/posterior_map.py Diamond Maps and Meta Flow Maps: the inner clock, posterior samples, the value function from K samples, its gradient and the importance weight

Every file accepts the shared flags of cis6270.runner, and --steps 20 --quiet finishes in a few seconds. In flow_map.py, stochastic_map.py and posterior_map.py every quantity comes from the chapter's own formulas, so there --steps sets the Runge-Kutta substep count, the Monte Carlo draw count and the inner substep count, which each file's docstring says.

What a default run prints

The chapter's worked numbers come out exactly, and the test suite asserts them.

Quantity Chapter Printed
X_{0,1}(0.7), the cocycle through 0.5, the inverse, det J_{0,1} 4.1, 4.1, 0.7, 3 4.100000, 4.100000, 0.700000, 3.000000
log det J_{0,1} by Liouville, and the density ratio 1.0986, 3.000 1.0986, 3.0000
Stay-still residuals at x = 1, s = 0, t = 1 1, 1, 0 1.0000, 1.0000, 0.0000
MeanFlow on dx/dt = x: 1.718x = x + 0.718x 1.718, 1, 0.718 1.7183, 1.0000, 0.7183
Shortcut doubling: two steps of 0.25 against one of 0.5 1.297443 1.297443
Consistency distillation residual at h = 0.01 8.189e-05 8.189e-05
FMLM composed target, its four KL terms and their sum (0.06, 0.12, 0.76, 0.06), 0.030747 the same, 0.030747
The two gradients in the logits at that target and student (0.04, 0.03, -0.11, 0.04) and (0.0099, 0.01335, -0.03315, 0.0099) the same
The same two once the student saturates to q_1 = 1e-06 -0.0600 and -4.62e-08 -0.0600 and -4.6171e-08
The decoding clock at 0.60, 0.75, 0.80, 0.85 0.006, 0.141, 0.456, 0.931 0.0057, 0.1411, 0.4561, 0.9313
Four steps on the clock 0.773490, 0.804481, 0.828095 the same
Gamma on the last three times, and over the long first leg 0.494620, 0.830144 the same
Categorical gap and drift, and the bound around a zero residual antiparallel, 0.199 1e-34 residual, 0.1989
The same two terms averaged over three thousand states 0.393, 0.095 0.3955, 0.0963
Discrete derivative target, three entries negative (-0.3333, -0.6667, 1.3333, -0.3333) the same
The logit repair of (-0.284, 0.828, 0.456) (0.0194, 0.6259, 0.3547) the same
The weak map's floor 2K_1 6.389 6.3886 by Monte Carlo
Chen fusion of four half-interval draws 0.50, -0.20 0.50, -0.20
Local clocks at t = 0.8, and omega at (0.3, 0.6, 0.9) 0.800, 0.692, 0.333 and 0.125 the same
Jensen gap, and E[V_K] at K = 1, 2, 4, 8 0.434 and 1.000, 1.217, 1.340, 1.393 the same
Meta Flow Maps fixed point against the per-sample ratio 0.625 and 0.540 the same

For the saturated pair in that table, fmlm.py works the full expression of Section 6.27.4, the gradient of half the squared error through the softmax Jacobian,

d/dz_i (1/2) sum_j (q_j - p_j)^2  =  q_i (q_i - p_i)  -  q_i sum_j q_j (q_j - p_j) ,

at p = (0.06, 0.12, 0.76, 0.06) with the student moved to q_1 = 1e-06 and the other three renormalized. We also differentiate the two losses automatically and compare, and the two routes agree to four figures on every coordinate. As you can see from the pair, -0.0600 against -4.6171e-08 is the six orders of magnitude the chapter draws its point from.

We take the trained numbers from one default run on a laptop CPU, with seed 0.

Run Headline
residuals.py, 600 steps each distance to its own ODE 0.900 for the diagonal alone, 0.481, 0.376 and 0.541 with the Lagrangian, Eulerian and composition residuals; fraction on a mode 0.44 against 0.74, 0.66 and 0.70
meanflow.py, 2000 steps nearest mode 0.627 at one step, 0.522 at two, 0.505 at four, all four modes found
shortcut.py, 2000 steps nearest mode 0.605 at one step and 0.508 at eight; the doubling residual rises from below 1e-04 at d = 1/128 to 0.034 at d = 1/2
consistency.py, 1500 steps each nearest mode 0.513 for the consistency model at one step, 0.515, 0.420 and 0.470 for the trajectory model at one, two and four
fmlm.py, 1500 steps GC 0.586 at one step against 0.581 in the training data, motif 0.63, token entropy 1.371 of a possible 1.386, every sequence distinct
categorical_map.py, 1000 steps each every phrase in the language at both step counts; trained on the residual, 0.56 of one-step phrases distinct at entropy 2.75; trained on the bound, 0.33 at 2.54, and 0.12 at 2.05 by four steps
discrete_map.py, 1500 steps GC 0.609 at one step, motif 0.81, entropy 1.362, every sequence distinct; the unrepaired additive target is negative on 0.05% of its entries
stochastic_map.py the strong map's terminal cloud sits 8.8e-04 from the finest crossing at two hops, 1.1e-04 at four, 1.4e-05 at eight and 2e-06 at sixteen, the weak map's 3.06 at every hop count
expanding.py, 1500 steps GC 0.617 at one step, 0.621 on coordinates born at the source and 0.613 on coordinates born later; motif 0.75
posterior_map.py on the blobs, the exact value -0.1127, the 16-sample estimate -0.1088, the single endpoint -0.1628

The DNA files draw 1024 training sequences of the default length 16, which have GC fraction 0.581 and motif fraction 0.749. They read 64 samples, as the phrase file does, so a reported fraction has about 0.06 of sampling noise and its second digit moves between seeds.

Longer runs, where they change the picture:

  • residuals.py --steps 2400 moves the four distances to 1.033, 0.171, 0.247 and 0.511 and the four mode fractions to 0.27, 0.82, 0.71 and 0.79, against Table 6.2's 0.896, 0.292, 0.242 and 0.325 at 0.12, 0.84, 0.82 and 0.84. The gap between the anchor alone and the anchor plus any residual widens with the budget, and the three residuals stay inside each other's spread.
  • meanflow.py --steps 8000 brings the one-step nearest mode from 0.627 to 0.437, below its own two-step and four-step figures, so the one-step map overtakes the grid.
  • categorical_map.py --steps 3000 widens the gap between the two students, to 0.60 of one-step phrases distinct for the residual against 0.27 for the bound, and 0.54 against 0.09 at four steps. The bounded student loses variety as its budget grows, as Section 6.28.4 predicts, since the bound is minimized by a prediction that is constant in the end time.
  • The DNA runs are at their plateau by their default budgets. At --steps 6000 the one-step motif fraction of fmlm.py reads 0.61 and of discrete_map.py 0.69, inside the sampling noise of the default run.

At the default flags on one core, we get three or four seconds for flow_map.py, stochastic_map.py and posterior_map.py, which train nothing, and twelve to twenty-seven seconds for the eight that do, including residuals.py with its four models and categorical_map.py and consistency.py with their two.

What is simplified

  • The four blob models of residuals.py train for 600 steps on a four-layer network, so the distances are larger than Table 6.2 of the chapter, which trains longer. The ordering the chapter draws, the anchor alone against the anchor plus any residual, comes out at this budget.
  • shortcut.py evaluates the zero-step term and the doubling term on the whole batch and weights them three to one, in place of the chapter's split of the batch in that proportion. The two have the same expectation and the weighted form costs one network call more per step.
  • consistency.py distils from the exact marginal velocity of the training set, which is a softmax over the data points, in place of a separately trained flow matching teacher. The teacher's discretization error is then the only error the student inherits, as Section 6.20 sets out.
  • posterior_map.py builds the inner map from the exact posterior over the training points, in place of training a conditional family, so the inner velocity is a closed form and the value, its gradient and the Jensen gap are all available to compare against.
  • We compute the decoding clock of fmlm.py by quadrature on the one-position integral, which is exact for the Gaussian interpolant and treats each position on its own, so the clock stops where correlations between positions begin.
  • categorical_map.py and discrete_map.py keep the end time below one, because the drift term of each loss divides by the remaining time. fmlm.py has no such term, so it draws a quarter of each batch at t = 1, the slice Section 6.26 asks for by name.
  • In stochastic_map.py we carry three Legendre coefficients and fuse them with the Chen relations written out for n = 0, 1, 2. We stop at three, because the chapter lists the constants c_{3,m} only up to that order.
  • Every network is cis6270.nets.TwoTimeMLP. The discrete files flatten the one-hot lift into its input and read its output as per-position logits, so one architecture covers both halves of the lecture and any difference in the samples comes from the loss.
  • Each file sets torch.set_num_threads(1). These networks are small enough that splitting one matrix product across cores costs more than it saves. tests/conftest.py sets the same thing for the test session, and the line in each file acts when a reader runs the script directly.