CIS6270 / lecture_5 /README.md
pranamanam's picture
Upload 87 files
0ba9d09 verified
|
Raw History Blame Contribute Delete
9.83 kB

Lecture 5: Discrete Flow Matching (Chapter 5)

Chapter 4 fixed a corruption process and solved for its reverse. In this lecture we reverse that order, picking the probability path first, writing down a velocity that generates it, and reading the sampler off the velocity. On tokens the velocity is a rate matrix read from the destination's side, and the master equation is the condition under which it generates a path. The mixture path (1 - kappa_t) delta_{x_0} + kappa_t delta_{x_1} at each position has a conditional velocity, which a posterior average turns into the schedule factor times a denoiser, and we recover that denoiser with cross-entropy on the data token. Everything after that is a choice made inside the same construction: a schedule, a corrector, a kinetic-optimal path, edits that change the length, a relaxation of the token into a point of the simplex or of a sphere, a guidance tilt, or a different coupling.

Two conventions hold in every file, and we state them at the top of each. Time runs from the source at t = 0 to the data at t = 1, so kappa_t here equals the survival probability alpha_{1-t} of Chapter 4. The probability velocity u_t(y, x) is the same number Chapter 4 writes as the rate R_t(x, y), with the two arguments exchanged.

The files

Every script here runs on the synthetic DNA that cis6270.data generates, 512 sequences of length 16 over A, C, G, T, with a GC-rich background and the motif ACGT planted in three quarters of them at one of three positions. Those sequences have GC fraction 0.577, motif fraction 0.777 and token entropy 1.374 of log 4 = 1.386, and we measure a trained sampler against those three numbers.

File Command What it does
dfm.py python lecture_5/dfm.py The coupling, the mixture path, the conditional velocity, the posterior average, the cross-entropy loss and the two-part Euler rule. Trains a denoiser and generates DNA from either source, --source uniform or --source mask.
velocity.py python lecture_5/velocity.py Works the mathematics with nothing learned: a jump rate written as a velocity, the two rules, the master equation, the worked DNA velocity, survival and hazard, and the continuous and discrete velocities side by side.
corrector.py python lecture_5/corrector.py Velocities with zero net flux, the three-state example, the corrector for the mask source, the jump counts at three corrector strengths, and the parallel-step error.
kinetic_optimal.py python lecture_5/kinetic_optimal.py The flux and its energy, the velocity of least energy, the kinetic optimality of the mixture path, the transition distance on DNA, the worked metric path, and the energy of the linear and geodesic schedules.
edit_flows.py python lecture_5/edit_flows.py Insertions, deletions and substitutions; alignments and the column-wise path; the posterior over two alignments; the rate loss; one sampler step; and the length statistics of a run on DNA pairs of different lengths.
dirichlet.py python lecture_5/dirichlet.py The Dirichlet family as a conditional path, the speed from the flux across a cut, the two-letter closed form, the worked speeds on a DNA base, and the posterior average. Trains a denoiser on the simplex and integrates to t = 8.
fisher.py python lecture_5/fisher.py The Fisher information as a metric, the square-root map to the sphere, geodesics, the tangent projection and the regression loss. Trains a velocity field and samples with exponential-map steps.
gumbel.py python lecture_5/gumbel.py The Gumbel-max trick, the cooling temperature, the annealed path, the replicator equation, the inference-time velocity, and training against the marginal velocity.
guidance.py python lecture_5/guidance.py Guidance as a tilt of the rates, the h-transform, the reachable set, the guidance strength, the classifier-free form, the Taylor expansion of the log ratio, and the guided sampler. Steers toward the GC label with --gamma.
multi_objective.py python lecture_5/multi_objective.py The Pareto front, MOG-DFM's rank, direction and cone with the controller that widens it, and pCoMole's worst-objective utility with a rollout of the base flow on DNA.
rectify.py python lecture_5/rectify.py Total correlation on two bases, ReDi's recoupling rounds run on the exact sampler, and AReUReDi's Tchebycheff tilt with its Metropolis acceptance step.

Every file takes the shared flags of cis6270.runner. --steps 20 --quiet runs each one in a few seconds, and --steps controls the training steps in the five files that train, the grid in the five that integrate, and the number of rollouts in multi_objective.py.

The numbers a default run produces

At seed 0 on one core, the five files that train take under half a minute each, and the other six take one to seven seconds. The worked numbers are exact and repeat on any machine; the trained numbers move a little with the hardware's floating point.

File Headline numbers
dfm.py loss 0.876 nats, GC fraction 0.578, motif fraction 0.672, token entropy 1.374 of log 4 = 1.386
velocity.py p_dot = (-0.1, -0.2, 0.7, -0.4), Euler step (0.09, 0.18, 0.37, 0.36), hazard 2 at t = 0.5 and 10 at t = 0.9, Euler error 0.0002 against the closed form
corrector.py masked fraction 0.500 at t = 0.5 for every eta, jumps per position 1.000, 2.991 and 10.934 at eta = 0, 2, 10, pair survival 0.133 measured against the predicted 0.135 at eta = 2
kinetic_optimal.py rate out of C 5.780, p_dot = (-0.3997, -0.4774, 1.3546, -0.4774) from the velocity and from the path, energy 27.63 and 41.45 for the linear schedule, pi^2 = 9.8696 for the geodesic one
edit_flows.py averaged edit rates 0.667, 1.333 and 0.667 for a total of 2.667, example loss 3.446, untouched 0.8452 against the first-order 0.8350, DNA length 13.945 at t = 0.5 against the predicted 13.941
dirichlet.py speeds 0.232, 0.199, 0.159, 0.104 at t = 1, continuity residual 2.5e-08, loss 1.027 nats, GC fraction 0.532, motif fraction 0.547
fisher.py Fisher-Rao distance 1.485 against the Euclidean 0.141 and 0.013 on the two moves, arc 0.3713 from the midpoint, loss 0.278, GC fraction 0.576, motif fraction 0.203
gumbel.py path (0.0528, 0.0214, 0.8781, 0.0477) at t = 1, replicator velocity (-0.136, -0.217, 0.500, -0.147), loss 0.672 nats, GC fraction 0.572, motif fraction 0.516
guidance.py guided destination (0.029, 0.203, 0.478, 0.290), exit rate 2.50 at the exact h, GC fraction 0.531 unguided, 0.657 at gamma = 2 and 0.698 at gamma = 5
multi_objective.py combined score (-1.511, 1.541, 2.297, -2.327), cone angles 47.7, 37.9 and 107.7 degrees, selection (0.371, 0.629, 0) after the chapter's rollouts, and on DNA a base selection of (0.417, 0.583) that the rollout value turns into (0.962, 0.038)
rectify.py total correlation 1.386, 1.101, 0.843, 0.696 and 0.616 over rounds 0 to 4 against the limit 0.520; one-step success 25.0%, 40.3%, 52.5%, 58.0% and 60.4% against the limit 62.5%

With four thousand training steps, about two minutes a file, we get a sampler that reaches the planted motif more often: dfm.py goes from 0.672 to 0.734 against the data's 0.777, and fisher.py from 0.203 to 0.453, while the other numbers in each report stay where the table has them.

tests/test_lecture_5.py asserts the chapter's worked numbers against these files, including the posterior 0.625 and 0.375 at a masked position, the entropy 0.6616 that bounds the loss there, the velocity (0.2, 1.2, -1.6, 0.2) out of G, the metric-path rates 0.898 and 4.882 out of C, the Dirichlet speeds, the Fisher-Rao arc, the Gumbel-softmax path at four times, the guided destination distribution, the Pareto front, and the ReDi rounds.

What is simplified

  • The corrector simulation uses 40,000 positions where the chapter uses 200,000. Pass --positions 200000 to reproduce the chapter's figures exactly.
  • The parallel-step section measures the survival probability of a jointly revealed pair, which is the quantity the chapter computes in closed form. At a fine grid the end-to-end error of the corrector stays at or above its value at eta = 0, because the admissible step shrinks with eta.
  • edit_flows.py runs the column-wise path on one sampled alignment per pair, which the chapter's marginalization theorem permits, and trains nothing, so the edit intensities of its sampler section are the chapter's own numbers.
  • dirichlet.py evaluates the derivative of the regularized incomplete beta function by a central difference on scipy.special.betainc, in double precision whatever the ambient dtype, since the two values agree to six decimals. For the continuity check we differentiate the flux with a five-point stencil.
  • rectify.py runs ReDi with the exact marginal velocity of the current coupling on a two-position state space, so it trains nothing, and its AReUReDi section works one Metropolis iteration and stops there.
  • guidance.py trains the denoiser and the property model under one optimizer, and the property is the GC label of cis6270.data, not a biological measurement.
  • multi_objective.py trains nothing. Its two DNA objectives are the GC fraction and the AT fraction of a finished sequence, its feasible set is the sequences carrying the planted motif, and the base flow a rollout continues is the mask-source sampler with the data's letter frequency as its denoiser.
  • The relaxed paths replace the token lookup of cis6270.nets.SequenceDenoiser with a linear map on the relaxed token, and leave the rest of the network untouched.