Download lecture_6/README.md from ChatterjeeLab/CIS6270: direct link, hf CLI and curl.
- Browser
- Download file 12.5 kB
-
https://huggingface.co/ChatterjeeLab/CIS6270/resolve/main/lecture_6/README.md
- Command line
-
hf download hf://ChatterjeeLab/CIS6270/lecture_6/README.md
-
curl -L -o README.md https://huggingface.co/ChatterjeeLab/CIS6270/resolve/main/lecture_6/README.md
Lecture 6: Flow Maps
Chapter 6 of the notes, book/chapters/ch6_flow_maps.tex.
Chapters 2 to 5 learned a local change and applied it again and again, so the number of steps set
the cost of a sample. In this lecture we learn the map between two times directly. The map
X_{s,t}(x) is fixed by three axioms, identity, cocycle and tangent, and every file writes it the
same way, as the state plus the elapsed time times an average velocity the network produces, which
makes the identity axiom hold for any weights at all. We differentiate the cocycle in the end time
to get the Lagrangian identity, in the start time to get the Eulerian identity, and we leave it
finite to get the composition identity. Each one becomes a squared residual, we pair each with a
flow matching anchor on the diagonal, and from there we carry the construction into five settings: a
point in space, a distribution over tokens, a Brownian path, a dimension that grows, and a
conditioning state.
Every file is standalone, runs offline on a laptop CPU, and prints the numbers the chapter works by hand.
The files
| File | Command | What it builds |
|---|---|---|
flow_map.py |
python lecture_6/flow_map.py |
No learning. The chapter's velocity, the map by RK4, the three axioms, the inverse, the Jacobian and its differential equation, Liouville's formula, density transport and the generator |
residuals.py |
python lecture_6/residuals.py |
The Lagrangian, Eulerian and composition residuals and the diagonal anchor, from the chapter's listing, with four models trained on the four Gaussian blobs |
meanflow.py |
python lecture_6/meanflow.py |
MeanFlow: the Eulerian identity in the average-velocity parameterization, with the total derivative from one forward-mode pass |
shortcut.py |
python lecture_6/shortcut.py |
Shortcut models: the composition identity indexed by step size, on a dyadic ladder from 1/128 to 1 |
consistency.py |
python lecture_6/consistency.py |
Consistency models on the edge t = 1 with the boundary coefficients, and consistency trajectory models on the whole triangle |
fmlm.py |
python lecture_6/fmlm.py |
Flow Map Language Models: the two-time denoiser, the convex form, the cocycle in denoiser coordinates, the two gradients a softmax head sees, two KL terms and the decoding clock, on DNA |
categorical_map.py |
python lecture_6/categorical_map.py |
Categorical Flow Maps: the Lagrangian identity with a softmax head, the teacher gap and drift, and the bound that replaces the residual, on phrases |
discrete_map.py |
python lecture_6/discrete_map.py |
Discrete Flow Maps: why the derivative targets leave the simplex, and the repair in logit space, on DNA |
stochastic_map.py |
python lecture_6/stochastic_map.py |
Strong stochastic flow maps: the solution map of an SDE, the weak map's floor, the Legendre coefficients and the Chen relations |
expanding.py |
python lecture_6/expanding.py |
Expanding flow maps: a state whose length grows, local clocks, the composition identity under growth, and one draw of the insertion noise |
posterior_map.py |
python lecture_6/posterior_map.py |
Diamond Maps and Meta Flow Maps: the inner clock, posterior samples, the value function from K samples, its gradient and the importance weight |
Every file accepts the shared flags of cis6270.runner, and --steps 20 --quiet finishes in a
few seconds. In flow_map.py, stochastic_map.py and posterior_map.py every quantity comes
from the chapter's own formulas, so there --steps sets the Runge-Kutta substep count, the Monte
Carlo draw count and the inner substep count, which each file's docstring says.
What a default run prints
The chapter's worked numbers come out exactly, and the test suite asserts them.
| Quantity | Chapter | Printed |
|---|---|---|
X_{0,1}(0.7), the cocycle through 0.5, the inverse, det J_{0,1} |
4.1, 4.1, 0.7, 3 | 4.100000, 4.100000, 0.700000, 3.000000 |
log det J_{0,1} by Liouville, and the density ratio |
1.0986, 3.000 | 1.0986, 3.0000 |
Stay-still residuals at x = 1, s = 0, t = 1 |
1, 1, 0 | 1.0000, 1.0000, 0.0000 |
MeanFlow on dx/dt = x: 1.718x = x + 0.718x |
1.718, 1, 0.718 | 1.7183, 1.0000, 0.7183 |
| Shortcut doubling: two steps of 0.25 against one of 0.5 | 1.297443 | 1.297443 |
Consistency distillation residual at h = 0.01 |
8.189e-05 | 8.189e-05 |
| FMLM composed target, its four KL terms and their sum | (0.06, 0.12, 0.76, 0.06), 0.030747 | the same, 0.030747 |
| The two gradients in the logits at that target and student | (0.04, 0.03, -0.11, 0.04) and (0.0099, 0.01335, -0.03315, 0.0099) | the same |
The same two once the student saturates to q_1 = 1e-06 |
-0.0600 and -4.62e-08 | -0.0600 and -4.6171e-08 |
| The decoding clock at 0.60, 0.75, 0.80, 0.85 | 0.006, 0.141, 0.456, 0.931 | 0.0057, 0.1411, 0.4561, 0.9313 |
| Four steps on the clock | 0.773490, 0.804481, 0.828095 | the same |
| Gamma on the last three times, and over the long first leg | 0.494620, 0.830144 | the same |
| Categorical gap and drift, and the bound around a zero residual | antiparallel, 0.199 | 1e-34 residual, 0.1989 |
| The same two terms averaged over three thousand states | 0.393, 0.095 | 0.3955, 0.0963 |
| Discrete derivative target, three entries negative | (-0.3333, -0.6667, 1.3333, -0.3333) | the same |
| The logit repair of (-0.284, 0.828, 0.456) | (0.0194, 0.6259, 0.3547) | the same |
The weak map's floor 2K_1 |
6.389 | 6.3886 by Monte Carlo |
| Chen fusion of four half-interval draws | 0.50, -0.20 | 0.50, -0.20 |
Local clocks at t = 0.8, and omega at (0.3, 0.6, 0.9) |
0.800, 0.692, 0.333 and 0.125 | the same |
Jensen gap, and E[V_K] at K = 1, 2, 4, 8 |
0.434 and 1.000, 1.217, 1.340, 1.393 | the same |
| Meta Flow Maps fixed point against the per-sample ratio | 0.625 and 0.540 | the same |
For the saturated pair in that table, fmlm.py works the full expression of Section 6.27.4, the
gradient of half the squared error through the softmax Jacobian,
d/dz_i (1/2) sum_j (q_j - p_j)^2 = q_i (q_i - p_i) - q_i sum_j q_j (q_j - p_j) ,
at p = (0.06, 0.12, 0.76, 0.06) with the student moved to q_1 = 1e-06 and the other three
renormalized. We also differentiate the two losses automatically and compare, and the two routes
agree to four figures on every coordinate. As you can see from the pair, -0.0600 against
-4.6171e-08 is the six orders of magnitude the chapter draws its point from.
We take the trained numbers from one default run on a laptop CPU, with seed 0.
| Run | Headline |
|---|---|
residuals.py, 600 steps each |
distance to its own ODE 0.900 for the diagonal alone, 0.481, 0.376 and 0.541 with the Lagrangian, Eulerian and composition residuals; fraction on a mode 0.44 against 0.74, 0.66 and 0.70 |
meanflow.py, 2000 steps |
nearest mode 0.627 at one step, 0.522 at two, 0.505 at four, all four modes found |
shortcut.py, 2000 steps |
nearest mode 0.605 at one step and 0.508 at eight; the doubling residual rises from below 1e-04 at d = 1/128 to 0.034 at d = 1/2 |
consistency.py, 1500 steps each |
nearest mode 0.513 for the consistency model at one step, 0.515, 0.420 and 0.470 for the trajectory model at one, two and four |
fmlm.py, 1500 steps |
GC 0.586 at one step against 0.581 in the training data, motif 0.63, token entropy 1.371 of a possible 1.386, every sequence distinct |
categorical_map.py, 1000 steps each |
every phrase in the language at both step counts; trained on the residual, 0.56 of one-step phrases distinct at entropy 2.75; trained on the bound, 0.33 at 2.54, and 0.12 at 2.05 by four steps |
discrete_map.py, 1500 steps |
GC 0.609 at one step, motif 0.81, entropy 1.362, every sequence distinct; the unrepaired additive target is negative on 0.05% of its entries |
stochastic_map.py |
the strong map's terminal cloud sits 8.8e-04 from the finest crossing at two hops, 1.1e-04 at four, 1.4e-05 at eight and 2e-06 at sixteen, the weak map's 3.06 at every hop count |
expanding.py, 1500 steps |
GC 0.617 at one step, 0.621 on coordinates born at the source and 0.613 on coordinates born later; motif 0.75 |
posterior_map.py |
on the blobs, the exact value -0.1127, the 16-sample estimate -0.1088, the single endpoint -0.1628 |
The DNA files draw 1024 training sequences of the default length 16, which have GC fraction 0.581 and motif fraction 0.749. They read 64 samples, as the phrase file does, so a reported fraction has about 0.06 of sampling noise and its second digit moves between seeds.
Longer runs, where they change the picture:
residuals.py --steps 2400moves the four distances to 1.033, 0.171, 0.247 and 0.511 and the four mode fractions to 0.27, 0.82, 0.71 and 0.79, against Table 6.2's 0.896, 0.292, 0.242 and 0.325 at 0.12, 0.84, 0.82 and 0.84. The gap between the anchor alone and the anchor plus any residual widens with the budget, and the three residuals stay inside each other's spread.meanflow.py --steps 8000brings the one-step nearest mode from 0.627 to 0.437, below its own two-step and four-step figures, so the one-step map overtakes the grid.categorical_map.py --steps 3000widens the gap between the two students, to 0.60 of one-step phrases distinct for the residual against 0.27 for the bound, and 0.54 against 0.09 at four steps. The bounded student loses variety as its budget grows, as Section 6.28.4 predicts, since the bound is minimized by a prediction that is constant in the end time.- The DNA runs are at their plateau by their default budgets. At
--steps 6000the one-step motif fraction offmlm.pyreads 0.61 and ofdiscrete_map.py0.69, inside the sampling noise of the default run.
At the default flags on one core, we get three or four seconds for flow_map.py,
stochastic_map.py and posterior_map.py, which train nothing, and twelve to twenty-seven
seconds for the eight that do, including residuals.py with its four models and
categorical_map.py and consistency.py with their two.
What is simplified
- The four blob models of
residuals.pytrain for 600 steps on a four-layer network, so the distances are larger than Table 6.2 of the chapter, which trains longer. The ordering the chapter draws, the anchor alone against the anchor plus any residual, comes out at this budget. shortcut.pyevaluates the zero-step term and the doubling term on the whole batch and weights them three to one, in place of the chapter's split of the batch in that proportion. The two have the same expectation and the weighted form costs one network call more per step.consistency.pydistils from the exact marginal velocity of the training set, which is a softmax over the data points, in place of a separately trained flow matching teacher. The teacher's discretization error is then the only error the student inherits, as Section 6.20 sets out.posterior_map.pybuilds the inner map from the exact posterior over the training points, in place of training a conditional family, so the inner velocity is a closed form and the value, its gradient and the Jensen gap are all available to compare against.- We compute the decoding clock of
fmlm.pyby quadrature on the one-position integral, which is exact for the Gaussian interpolant and treats each position on its own, so the clock stops where correlations between positions begin. categorical_map.pyanddiscrete_map.pykeep the end time below one, because the drift term of each loss divides by the remaining time.fmlm.pyhas no such term, so it draws a quarter of each batch att = 1, the slice Section 6.26 asks for by name.- In
stochastic_map.pywe carry three Legendre coefficients and fuse them with the Chen relations written out forn = 0, 1, 2. We stop at three, because the chapter lists the constantsc_{3,m}only up to that order. - Every network is
cis6270.nets.TwoTimeMLP. The discrete files flatten the one-hot lift into its input and read its output as per-position logits, so one architecture covers both halves of the lecture and any difference in the samples comes from the loss. - Each file sets
torch.set_num_threads(1). These networks are small enough that splitting one matrix product across cores costs more than it saves.tests/conftest.pysets the same thing for the test session, and the line in each file acts when a reader runs the script directly.