Download lecture_5/README.md from ChatterjeeLab/CIS6270: direct link, hf CLI and curl.
- Browser
- Download file 9.83 kB
-
https://huggingface.co/ChatterjeeLab/CIS6270/resolve/main/lecture_5/README.md
- Command line
-
hf download hf://ChatterjeeLab/CIS6270/lecture_5/README.md
-
curl -L -o README.md https://huggingface.co/ChatterjeeLab/CIS6270/resolve/main/lecture_5/README.md
Lecture 5: Discrete Flow Matching (Chapter 5)
Chapter 4 fixed a corruption process and solved for its reverse. In this lecture we reverse that
order, picking the probability path first, writing down a velocity that generates it, and reading
the sampler off the velocity. On tokens the velocity is a rate matrix read from the destination's
side, and the master equation is the condition under which it generates a path. The mixture path
(1 - kappa_t) delta_{x_0} + kappa_t delta_{x_1} at each position has a conditional velocity, which
a posterior average turns into the schedule factor times a denoiser, and we recover that denoiser
with cross-entropy on the data token. Everything after that is a choice made inside the same
construction: a schedule, a corrector, a kinetic-optimal path, edits that change the length, a
relaxation of the token into a point of the simplex or of a sphere, a guidance tilt, or a different
coupling.
Two conventions hold in every file, and we state them at the top of each. Time runs from the
source at t = 0 to the data at t = 1, so kappa_t here equals the survival probability
alpha_{1-t} of Chapter 4. The probability velocity u_t(y, x) is the same number Chapter 4
writes as the rate R_t(x, y), with the two arguments exchanged.
The files
Every script here runs on the synthetic DNA that cis6270.data generates, 512 sequences of length
16 over A, C, G, T, with a GC-rich background and the motif ACGT planted in three quarters of them
at one of three positions. Those sequences have GC fraction 0.577, motif fraction 0.777 and token
entropy 1.374 of log 4 = 1.386, and we measure a trained sampler against those three numbers.
| File | Command | What it does |
|---|---|---|
dfm.py |
python lecture_5/dfm.py |
The coupling, the mixture path, the conditional velocity, the posterior average, the cross-entropy loss and the two-part Euler rule. Trains a denoiser and generates DNA from either source, --source uniform or --source mask. |
velocity.py |
python lecture_5/velocity.py |
Works the mathematics with nothing learned: a jump rate written as a velocity, the two rules, the master equation, the worked DNA velocity, survival and hazard, and the continuous and discrete velocities side by side. |
corrector.py |
python lecture_5/corrector.py |
Velocities with zero net flux, the three-state example, the corrector for the mask source, the jump counts at three corrector strengths, and the parallel-step error. |
kinetic_optimal.py |
python lecture_5/kinetic_optimal.py |
The flux and its energy, the velocity of least energy, the kinetic optimality of the mixture path, the transition distance on DNA, the worked metric path, and the energy of the linear and geodesic schedules. |
edit_flows.py |
python lecture_5/edit_flows.py |
Insertions, deletions and substitutions; alignments and the column-wise path; the posterior over two alignments; the rate loss; one sampler step; and the length statistics of a run on DNA pairs of different lengths. |
dirichlet.py |
python lecture_5/dirichlet.py |
The Dirichlet family as a conditional path, the speed from the flux across a cut, the two-letter closed form, the worked speeds on a DNA base, and the posterior average. Trains a denoiser on the simplex and integrates to t = 8. |
fisher.py |
python lecture_5/fisher.py |
The Fisher information as a metric, the square-root map to the sphere, geodesics, the tangent projection and the regression loss. Trains a velocity field and samples with exponential-map steps. |
gumbel.py |
python lecture_5/gumbel.py |
The Gumbel-max trick, the cooling temperature, the annealed path, the replicator equation, the inference-time velocity, and training against the marginal velocity. |
guidance.py |
python lecture_5/guidance.py |
Guidance as a tilt of the rates, the h-transform, the reachable set, the guidance strength, the classifier-free form, the Taylor expansion of the log ratio, and the guided sampler. Steers toward the GC label with --gamma. |
multi_objective.py |
python lecture_5/multi_objective.py |
The Pareto front, MOG-DFM's rank, direction and cone with the controller that widens it, and pCoMole's worst-objective utility with a rollout of the base flow on DNA. |
rectify.py |
python lecture_5/rectify.py |
Total correlation on two bases, ReDi's recoupling rounds run on the exact sampler, and AReUReDi's Tchebycheff tilt with its Metropolis acceptance step. |
Every file takes the shared flags of cis6270.runner. --steps 20 --quiet runs each one in a few
seconds, and --steps controls the training steps in the five files that train, the grid in the
five that integrate, and the number of rollouts in multi_objective.py.
The numbers a default run produces
At seed 0 on one core, the five files that train take under half a minute each, and the other six take one to seven seconds. The worked numbers are exact and repeat on any machine; the trained numbers move a little with the hardware's floating point.
| File | Headline numbers |
|---|---|
dfm.py |
loss 0.876 nats, GC fraction 0.578, motif fraction 0.672, token entropy 1.374 of log 4 = 1.386 |
velocity.py |
p_dot = (-0.1, -0.2, 0.7, -0.4), Euler step (0.09, 0.18, 0.37, 0.36), hazard 2 at t = 0.5 and 10 at t = 0.9, Euler error 0.0002 against the closed form |
corrector.py |
masked fraction 0.500 at t = 0.5 for every eta, jumps per position 1.000, 2.991 and 10.934 at eta = 0, 2, 10, pair survival 0.133 measured against the predicted 0.135 at eta = 2 |
kinetic_optimal.py |
rate out of C 5.780, p_dot = (-0.3997, -0.4774, 1.3546, -0.4774) from the velocity and from the path, energy 27.63 and 41.45 for the linear schedule, pi^2 = 9.8696 for the geodesic one |
edit_flows.py |
averaged edit rates 0.667, 1.333 and 0.667 for a total of 2.667, example loss 3.446, untouched 0.8452 against the first-order 0.8350, DNA length 13.945 at t = 0.5 against the predicted 13.941 |
dirichlet.py |
speeds 0.232, 0.199, 0.159, 0.104 at t = 1, continuity residual 2.5e-08, loss 1.027 nats, GC fraction 0.532, motif fraction 0.547 |
fisher.py |
Fisher-Rao distance 1.485 against the Euclidean 0.141 and 0.013 on the two moves, arc 0.3713 from the midpoint, loss 0.278, GC fraction 0.576, motif fraction 0.203 |
gumbel.py |
path (0.0528, 0.0214, 0.8781, 0.0477) at t = 1, replicator velocity (-0.136, -0.217, 0.500, -0.147), loss 0.672 nats, GC fraction 0.572, motif fraction 0.516 |
guidance.py |
guided destination (0.029, 0.203, 0.478, 0.290), exit rate 2.50 at the exact h, GC fraction 0.531 unguided, 0.657 at gamma = 2 and 0.698 at gamma = 5 |
multi_objective.py |
combined score (-1.511, 1.541, 2.297, -2.327), cone angles 47.7, 37.9 and 107.7 degrees, selection (0.371, 0.629, 0) after the chapter's rollouts, and on DNA a base selection of (0.417, 0.583) that the rollout value turns into (0.962, 0.038) |
rectify.py |
total correlation 1.386, 1.101, 0.843, 0.696 and 0.616 over rounds 0 to 4 against the limit 0.520; one-step success 25.0%, 40.3%, 52.5%, 58.0% and 60.4% against the limit 62.5% |
With four thousand training steps, about two minutes a file, we get a sampler that reaches the
planted motif more often: dfm.py goes from 0.672 to 0.734 against the data's 0.777, and
fisher.py from 0.203 to 0.453, while the other numbers in each report stay where the table has
them.
tests/test_lecture_5.py asserts the chapter's worked numbers against these files, including the
posterior 0.625 and 0.375 at a masked position, the entropy 0.6616 that bounds the loss there, the
velocity (0.2, 1.2, -1.6, 0.2) out of G, the metric-path rates 0.898 and 4.882 out of C, the
Dirichlet speeds, the Fisher-Rao arc, the Gumbel-softmax path at four times, the guided destination
distribution, the Pareto front, and the ReDi rounds.
What is simplified
- The corrector simulation uses 40,000 positions where the chapter uses 200,000. Pass
--positions 200000to reproduce the chapter's figures exactly. - The parallel-step section measures the survival probability of a jointly revealed pair, which is
the quantity the chapter computes in closed form. At a fine grid the end-to-end error of the
corrector stays at or above its value at
eta = 0, because the admissible step shrinks witheta. edit_flows.pyruns the column-wise path on one sampled alignment per pair, which the chapter's marginalization theorem permits, and trains nothing, so the edit intensities of its sampler section are the chapter's own numbers.dirichlet.pyevaluates the derivative of the regularized incomplete beta function by a central difference onscipy.special.betainc, in double precision whatever the ambient dtype, since the two values agree to six decimals. For the continuity check we differentiate the flux with a five-point stencil.rectify.pyruns ReDi with the exact marginal velocity of the current coupling on a two-position state space, so it trains nothing, and its AReUReDi section works one Metropolis iteration and stops there.guidance.pytrains the denoiser and the property model under one optimizer, and the property is the GC label ofcis6270.data, not a biological measurement.multi_objective.pytrains nothing. Its two DNA objectives are the GC fraction and the AT fraction of a finished sequence, its feasible set is the sequences carrying the planted motif, and the base flow a rollout continues is the mask-source sampler with the data's letter frequency as its denoiser.- The relaxed paths replace the token lookup of
cis6270.nets.SequenceDenoiserwith a linear map on the relaxed token, and leave the rest of the network untouched.