CIS6270 / lecture_5 /README.md
pranamanam's picture
Upload 87 files
0ba9d09 verified
|
Raw History Blame Contribute Delete
9.83 kB
# Lecture 5: Discrete Flow Matching (Chapter 5)
Chapter 4 fixed a corruption process and solved for its reverse. In this lecture we reverse that
order, picking the probability path first, writing down a velocity that generates it, and reading
the sampler off the velocity. On tokens the velocity is a rate matrix read from the destination's
side, and the master equation is the condition under which it generates a path. The mixture path
`(1 - kappa_t) delta_{x_0} + kappa_t delta_{x_1}` at each position has a conditional velocity, which
a posterior average turns into the schedule factor times a denoiser, and we recover that denoiser
with cross-entropy on the data token. Everything after that is a choice made inside the same
construction: a schedule, a corrector, a kinetic-optimal path, edits that change the length, a
relaxation of the token into a point of the simplex or of a sphere, a guidance tilt, or a different
coupling.
Two conventions hold in every file, and we state them at the top of each. Time runs from the
source at `t = 0` to the data at `t = 1`, so `kappa_t` here equals the survival probability
`alpha_{1-t}` of Chapter 4. The probability velocity `u_t(y, x)` is the same number Chapter 4
writes as the rate `R_t(x, y)`, with the two arguments exchanged.
## The files
Every script here runs on the synthetic DNA that `cis6270.data` generates, 512 sequences of length
16 over A, C, G, T, with a GC-rich background and the motif ACGT planted in three quarters of them
at one of three positions. Those sequences have GC fraction 0.577, motif fraction 0.777 and token
entropy 1.374 of `log 4 = 1.386`, and we measure a trained sampler against those three numbers.
| File | Command | What it does |
| --- | --- | --- |
| `dfm.py` | `python lecture_5/dfm.py` | The coupling, the mixture path, the conditional velocity, the posterior average, the cross-entropy loss and the two-part Euler rule. Trains a denoiser and generates DNA from either source, `--source uniform` or `--source mask`. |
| `velocity.py` | `python lecture_5/velocity.py` | Works the mathematics with nothing learned: a jump rate written as a velocity, the two rules, the master equation, the worked DNA velocity, survival and hazard, and the continuous and discrete velocities side by side. |
| `corrector.py` | `python lecture_5/corrector.py` | Velocities with zero net flux, the three-state example, the corrector for the mask source, the jump counts at three corrector strengths, and the parallel-step error. |
| `kinetic_optimal.py` | `python lecture_5/kinetic_optimal.py` | The flux and its energy, the velocity of least energy, the kinetic optimality of the mixture path, the transition distance on DNA, the worked metric path, and the energy of the linear and geodesic schedules. |
| `edit_flows.py` | `python lecture_5/edit_flows.py` | Insertions, deletions and substitutions; alignments and the column-wise path; the posterior over two alignments; the rate loss; one sampler step; and the length statistics of a run on DNA pairs of different lengths. |
| `dirichlet.py` | `python lecture_5/dirichlet.py` | The Dirichlet family as a conditional path, the speed from the flux across a cut, the two-letter closed form, the worked speeds on a DNA base, and the posterior average. Trains a denoiser on the simplex and integrates to `t = 8`. |
| `fisher.py` | `python lecture_5/fisher.py` | The Fisher information as a metric, the square-root map to the sphere, geodesics, the tangent projection and the regression loss. Trains a velocity field and samples with exponential-map steps. |
| `gumbel.py` | `python lecture_5/gumbel.py` | The Gumbel-max trick, the cooling temperature, the annealed path, the replicator equation, the inference-time velocity, and training against the marginal velocity. |
| `guidance.py` | `python lecture_5/guidance.py` | Guidance as a tilt of the rates, the h-transform, the reachable set, the guidance strength, the classifier-free form, the Taylor expansion of the log ratio, and the guided sampler. Steers toward the GC label with `--gamma`. |
| `multi_objective.py` | `python lecture_5/multi_objective.py` | The Pareto front, MOG-DFM's rank, direction and cone with the controller that widens it, and pCoMole's worst-objective utility with a rollout of the base flow on DNA. |
| `rectify.py` | `python lecture_5/rectify.py` | Total correlation on two bases, ReDi's recoupling rounds run on the exact sampler, and AReUReDi's Tchebycheff tilt with its Metropolis acceptance step. |
Every file takes the shared flags of `cis6270.runner`. `--steps 20 --quiet` runs each one in a few
seconds, and `--steps` controls the training steps in the five files that train, the grid in the
five that integrate, and the number of rollouts in `multi_objective.py`.
## The numbers a default run produces
At seed 0 on one core, the five files that train take under half a minute each, and the other six
take one to seven seconds. The worked numbers are exact and repeat on any machine; the trained
numbers move a little with the hardware's floating point.
| File | Headline numbers |
| --- | --- |
| `dfm.py` | loss 0.876 nats, GC fraction 0.578, motif fraction 0.672, token entropy 1.374 of `log 4 = 1.386` |
| `velocity.py` | `p_dot = (-0.1, -0.2, 0.7, -0.4)`, Euler step `(0.09, 0.18, 0.37, 0.36)`, hazard 2 at `t = 0.5` and 10 at `t = 0.9`, Euler error 0.0002 against the closed form |
| `corrector.py` | masked fraction 0.500 at `t = 0.5` for every `eta`, jumps per position 1.000, 2.991 and 10.934 at `eta = 0, 2, 10`, pair survival 0.133 measured against the predicted 0.135 at `eta = 2` |
| `kinetic_optimal.py` | rate out of C 5.780, `p_dot = (-0.3997, -0.4774, 1.3546, -0.4774)` from the velocity and from the path, energy 27.63 and 41.45 for the linear schedule, `pi^2 = 9.8696` for the geodesic one |
| `edit_flows.py` | averaged edit rates 0.667, 1.333 and 0.667 for a total of 2.667, example loss 3.446, untouched 0.8452 against the first-order 0.8350, DNA length 13.945 at `t = 0.5` against the predicted 13.941 |
| `dirichlet.py` | speeds 0.232, 0.199, 0.159, 0.104 at `t = 1`, continuity residual `2.5e-08`, loss 1.027 nats, GC fraction 0.532, motif fraction 0.547 |
| `fisher.py` | Fisher-Rao distance 1.485 against the Euclidean 0.141 and 0.013 on the two moves, arc 0.3713 from the midpoint, loss 0.278, GC fraction 0.576, motif fraction 0.203 |
| `gumbel.py` | path `(0.0528, 0.0214, 0.8781, 0.0477)` at `t = 1`, replicator velocity `(-0.136, -0.217, 0.500, -0.147)`, loss 0.672 nats, GC fraction 0.572, motif fraction 0.516 |
| `guidance.py` | guided destination `(0.029, 0.203, 0.478, 0.290)`, exit rate 2.50 at the exact `h`, GC fraction 0.531 unguided, 0.657 at `gamma = 2` and 0.698 at `gamma = 5` |
| `multi_objective.py` | combined score `(-1.511, 1.541, 2.297, -2.327)`, cone angles 47.7, 37.9 and 107.7 degrees, selection `(0.371, 0.629, 0)` after the chapter's rollouts, and on DNA a base selection of `(0.417, 0.583)` that the rollout value turns into `(0.962, 0.038)` |
| `rectify.py` | total correlation 1.386, 1.101, 0.843, 0.696 and 0.616 over rounds 0 to 4 against the limit 0.520; one-step success 25.0%, 40.3%, 52.5%, 58.0% and 60.4% against the limit 62.5% |
With four thousand training steps, about two minutes a file, we get a sampler that reaches the
planted motif more often: `dfm.py` goes from 0.672 to 0.734 against the data's 0.777, and
`fisher.py` from 0.203 to 0.453, while the other numbers in each report stay where the table has
them.
`tests/test_lecture_5.py` asserts the chapter's worked numbers against these files, including the
posterior 0.625 and 0.375 at a masked position, the entropy 0.6616 that bounds the loss there, the
velocity `(0.2, 1.2, -1.6, 0.2)` out of G, the metric-path rates 0.898 and 4.882 out of C, the
Dirichlet speeds, the Fisher-Rao arc, the Gumbel-softmax path at four times, the guided destination
distribution, the Pareto front, and the ReDi rounds.
## What is simplified
- The corrector simulation uses 40,000 positions where the chapter uses 200,000. Pass
`--positions 200000` to reproduce the chapter's figures exactly.
- The parallel-step section measures the survival probability of a jointly revealed pair, which is
the quantity the chapter computes in closed form. At a fine grid the end-to-end error of the
corrector stays at or above its value at `eta = 0`, because the admissible step shrinks with
`eta`.
- `edit_flows.py` runs the column-wise path on one sampled alignment per pair, which the chapter's
marginalization theorem permits, and trains nothing, so the edit intensities of its sampler
section are the chapter's own numbers.
- `dirichlet.py` evaluates the derivative of the regularized incomplete beta function by a central
difference on `scipy.special.betainc`, in double precision whatever the ambient dtype, since the
two values agree to six decimals. For the continuity check we differentiate the flux with a
five-point stencil.
- `rectify.py` runs ReDi with the exact marginal velocity of the current coupling on a
two-position state space, so it trains nothing, and its AReUReDi section works one Metropolis
iteration and stops there.
- `guidance.py` trains the denoiser and the property model under one optimizer, and the property is
the GC label of `cis6270.data`, not a biological measurement.
- `multi_objective.py` trains nothing. Its two DNA objectives are the GC fraction and the AT
fraction of a finished sequence, its feasible set is the sequences carrying the planted motif,
and the base flow a rollout continues is the mask-source sampler with the data's letter frequency
as its denoiser.
- The relaxed paths replace the token lookup of `cis6270.nets.SequenceDenoiser` with a linear map on
the relaxed token, and leave the rest of the network untouched.