Download composition_rules.yaml from graziul/differance-engine: direct link, hf CLI and curl.
- Browser
- Download file 20 kB
-
https://huggingface.co/graziul/differance-engine/resolve/main/composition_rules.yaml
- Command line
-
hf download hf://graziul/differance-engine/composition_rules.yaml
-
curl -L -o composition_rules.yaml https://huggingface.co/graziul/differance-engine/resolve/main/composition_rules.yaml
20 kB
| # Composition Rules | |
| # Each rule specifies how formalisms compose. Type-checked: input and output | |
| # signatures must match for a composition to be valid. The matching engine uses | |
| # these rules to decompose a paper's claimed-novel concept into a composition of | |
| # known formalisms. | |
| # | |
| # Schema: | |
| # id: stable identifier | |
| # name: human-readable name (lowercase, underscorized) | |
| # description: what the rule does, in plain language | |
| # signature: | |
| # operation: the mathematical operation type(s) accepted or produced | |
| # domain: input domain type(s) | |
| # codomain: output codomain type(s) | |
| # objective_family: objective type(s) accepted or produced | |
| # input_constraints: | |
| # formalism_ids: optional list of specific formalisms this rule applies to | |
| # meso_types: optional list of meso types this rule accepts | |
| # macro_types: optional list of macro types this rule accepts | |
| # output_signature: | |
| # operation: the resulting operation type | |
| # codomain: the resulting codomain type | |
| # objective_family: the resulting objective family (if modified) | |
| # meso_type: the resulting meso type | |
| # macro_type: the resulting macro type | |
| # preserves: properties that survive the transformation | |
| # introduces: new properties the transformation adds | |
| # examples: | |
| # - concrete example of the composition | |
| # status: seed | verified | deprecated | |
| # | |
| # Composition rule families: | |
| # 1. Neuralization primitives (~5) | |
| # 2. Kernelization primitives (~3) | |
| # 3. Structural wrappers (~4) | |
| # 4. Objective transforms (~3) | |
| # TOTAL: ~15 rules | |
| composition_rules: | |
| # ============================================================================= | |
| # 1. NEURALIZATION PRIMITIVES | |
| # "Neuralize": replace a fixed function/operator with a learned parameterized map | |
| # ============================================================================= | |
| - id: neuralize | |
| name: neuralize | |
| decomposes_to: [gradient_descent, sgd] | |
| description: > | |
| Replace a fixed function (kernel, distance metric, projection matrix) with a | |
| learned parameterized map φ_θ typically implemented as a deep neural network. | |
| The structural operation stays the same; the function implementing it becomes | |
| a learnable composition of linear+nonlinear transforms optimized by SGD. | |
| input_constraints: | |
| meso_types: | |
| - kernel_method | |
| - linear_projection | |
| - spectral_method | |
| - optimal_transport | |
| macro_types: | |
| - eigenvalue_problem | |
| - optimization | |
| output_signature: | |
| meso_type: none # neural network is implementation detail, not meso type | |
| preserves: true # the underlying mathematical operation | |
| preserves: | |
| - "the mathematical operation (project, decompose, match, transform)" | |
| - "the objective family (correlation, divergence, energy, etc.)" | |
| - "the domain and codomain types" | |
| introduces: | |
| - "differentiable parameterization: φ_θ replaces fixed function" | |
| - "SGD-based fitting instead of closed-form solution" | |
| - "non-convex optimization landscape" | |
| - "scale: can handle higher-dimensional inputs" | |
| examples: | |
| - "Kernel CCA → Deep CCA: replace kernel k(x,y) with learned encoders f_θ(x), g_φ(y)" | |
| - "Kernel PCA → Autoencoder: replace kernel with encoder-decoder pair" | |
| - "MMD → Deep MMD: replace fixed kernel with learned feature extractor" | |
| status: seed | |
| - id: attention_wrap | |
| name: attention_wrap | |
| decomposes_to: [softmax_attention, boltzmann_distribution] | |
| description: > | |
| Wrap a pairwise operation in a softmax-normalized weighted aggregation. | |
| For any function s(x_i, x_j) that scores compatibility between elements, | |
| produce output as Σ_j softmax(s(x_i, x_j))·v(x_j). This is the transformer's | |
| core operation: a Boltzmann-weighted sum over a set of values. | |
| input_constraints: | |
| formalisms: [] # any scoring function can be attention-wrapped | |
| output_signature: | |
| operation: aggregate | |
| codomain: vector | |
| objective_family: none | |
| preserves: | |
| - "the scoring function s(x_i, x_j) as the core computation" | |
| - "the value function v(x_j)" | |
| introduces: | |
| - "Boltzmann / softmax partition function as normalizer: Z_i = Σ_j exp(s(q_i, k_j))" | |
| - "set-to-vector aggregation that is permutation-equivariant" | |
| - "variable-length input handling" | |
| examples: | |
| - "dot-product attention: s(q,k) = q·k/√d → softmax → weighted sum of values" | |
| - "kernel attention: s(q,k) = k(q,k) → a kernel smoother with learnable parameters" | |
| - "self-attention = attention_wrap(dot-product) over same sequence" | |
| status: seed | |
| - id: residualize | |
| name: residualize | |
| decomposes_to: [residual_connection] | |
| description: > | |
| Add identity skip connection: y = F(x) + x. This is a structural | |
| transformation that makes the function learn perturbations around identity | |
| rather than the full mapping. Mathematically equivalent to applying an | |
| Euler discretization of an ODE with step size 1. | |
| input_constraints: {} # any vector-to-vector transform | |
| output_signature: | |
| operation: transform | |
| codomain: vector | |
| preserves: true | |
| preserves: | |
| - "the wrapped function F(x)" | |
| - "dimensionality: output dim = input dim" | |
| introduces: | |
| - "identity skip path: gradient can bypass F, mitigating vanishing gradients" | |
| - "ODE interpretation: x_{t+1} = x_t + F(x_t) as Euler step" | |
| - "the function learns F(x) = H(x) - x rather than H(x) directly" | |
| examples: | |
| - "ResNet block = residualize(convolution + batch_norm + ReLU)" | |
| - "Transformer sublayer = residualize(multi_head_attention) then residualize(FFN)" | |
| - "Residual flow in normalizing flows" | |
| status: seed | |
| - id: normalize | |
| name: normalize | |
| decomposes_to: [layer_normalization, batch_normalization] | |
| description: > | |
| Apply standardization: x' = γ·(x - μ)/σ + β. Removes first and second | |
| moment variation across a specified axis (batch, layer, instance, group). | |
| This is a whitening operation restricted to the first two moments. | |
| input_constraints: {} # any numeric tensor | |
| output_signature: | |
| operation: transform | |
| codomain: vector | |
| preserves: false # moments change | |
| preserves: | |
| - "dimensionality" | |
| - "the subsequent computation's functional form" | |
| introduces: | |
| - "zero mean, unit variance along normalization axis (before γ,β)" | |
| - "learnable affine parameters γ, β" | |
| - "stabilized gradient flow during training" | |
| examples: | |
| - "LayerNorm(x) before self-attention in Transformer" | |
| - "BatchNorm between conv layers in ResNet" | |
| status: seed | |
| - id: encode_decode | |
| name: encode_decode | |
| decomposes_to: [pca, vae] | |
| description: > | |
| Bottleneck compression: x → encode → z (latent) → decode → x̂. | |
| The encoder maps to a lower-dimensional latent; the decoder reconstructs. | |
| This is the universal autoencoder pattern. Information-theoretically, | |
| it's rate-distortion with λ controlling the bottleneck width. | |
| input_constraints: {} # any domain type | |
| output_signature: | |
| operation: transform | |
| codomain: vector | |
| objective_family: reconstruction # ||x - x̂||² or cross-entropy | |
| preserves: | |
| - "the identity map through the bottleneck (x ≈ decode(encode(x)))" | |
| introduces: | |
| - "latent representation z in R^d (typically d < input dim)" | |
| - "information bottleneck: the latent discards everything not needed for reconstruction" | |
| - "if stochastic: variational bound (ELBO) on log p(x)" | |
| examples: | |
| - "Autoencoder = neuralize(encode_decode)" | |
| - "VAE = encode_decode + KL regularizer on latent" | |
| - "U-Net = encode_decode with skip connections between corresponding resolutions" | |
| status: seed | |
| # ============================================================================= | |
| # 2. KERNELIZATION PRIMITIVES | |
| # ============================================================================= | |
| - id: kernelize | |
| name: kernelize | |
| decomposes_to: [kernel_pca, kernel_cca, kernel_ridge_regression] | |
| description: > | |
| Lift a linear method to a nonlinear one by replacing inner products ⟨x, y⟩ | |
| with a positive-definite kernel k(x, y). This is the kernel trick: the | |
| method stays algebraically identical but now operates in a reproducing | |
| kernel Hilbert space (RKHS) implicitly defined by the feature map φ. | |
| Equivalent to: apply linear method to φ(x) without ever computing φ(x). | |
| input_constraints: | |
| meso_types: | |
| - linear_projection | |
| - spectral_method | |
| macro_types: | |
| - eigenvalue_problem | |
| - optimization | |
| output_signature: | |
| meso_type: kernel_method | |
| macro_type: eigenvalue_problem # kernel matrices are eigen-decomposed | |
| preserves: true | |
| preserves: | |
| - "the algebraic form of the method" | |
| - "the objective family" | |
| - "convexity (when the original method is convex)" | |
| introduces: | |
| - "implicit feature map φ: X → H (RKHS)" | |
| - "kernel Gram matrix K_{ij} = k(x_i, x_j) as sufficient statistic" | |
| - "O(N³) or O(N²) complexity (unless Nyström / random features)" | |
| - "Mercer condition: k must be positive-definite" | |
| examples: | |
| - "PCA → Kernel PCA: K = k(x_i, x_j), eigen-decompose centered K" | |
| - "CCA → Kernel CCA: eigenproblem on K_x^{-1/2} K_x K_y K_y^{-1/2}" | |
| - "Ridge regression → Kernel ridge regression: α = (K + λI)^{-1} y" | |
| - "Fisher LDA → Kernel FDA" | |
| status: seed | |
| - id: random_fourier_features | |
| name: random_fourier_features | |
| decomposes_to: [fourier_transform, kernel_pca] | |
| description: > | |
| Approximate a shift-invariant kernel k(x,y) = k(x-y) by sampling random | |
| Fourier features: z(x) = √(2/D) · [cos(ω₁·x + b₁), ..., cos(ω_D·x + b_D)] | |
| where ω_d ~ p(ω) (the kernel's spectral density). Then k(x,y) ≈ z(x)·z(y). | |
| This linearizes the kernel method: kernelized methods become linear methods | |
| in the random feature space, recovering O(ND) complexity. | |
| input_constraints: | |
| meso_types: | |
| - kernel_method | |
| formalisms: [] # any shift-invariant kernel | |
| output_signature: | |
| meso_type: linear_projection # method is now linear in z(x) | |
| preserves: true | |
| preserves: | |
| - "the method's algebraic form (now linear in z(x))" | |
| - "the objective family" | |
| introduces: | |
| - "explicit D-dimensional feature map approximating RKHS" | |
| - "O(ND) complexity instead of O(N²) or O(N³)" | |
| - "approximation error O(D^{-1/2}) by Bochner's theorem + Hoeffding" | |
| examples: | |
| - "RFF for RBF kernel: ω ~ N(0, σ^{-2} I)" | |
| - "Deep sets / point cloud methods that use RFF as positional encoding" | |
| - "Transformer sinusoidal position encoding (closely related)" | |
| status: seed | |
| - id: nystrom | |
| name: nystrom_approximation | |
| decomposes_to: [svd, kernel_pca] | |
| description: > | |
| Low-rank approximation of a kernel Gram matrix by subsampling m landmark | |
| points: K ≈ K_{nm} K_{mm}^{-1} K_{mn}. Reduces complexity from O(N³) to | |
| O(Nm² + m³). The Nyström method is the quadrature-based numerical | |
| approximation of the integral eigenproblem underlying the kernel expansion. | |
| input_constraints: | |
| meso_types: | |
| - kernel_method | |
| output_signature: | |
| preserves: true | |
| preserves: | |
| - "the kernel method's structure" | |
| - "the objective family" | |
| introduces: | |
| - "low-rank approximation: m landmarks, rank at most m" | |
| - "O(Nm² + m³) complexity" | |
| - "approximation quality depends on landmark selection" | |
| examples: | |
| - "Nyström kernel PCA" | |
| - "Nyström kernel ridge regression" | |
| - "Landmark-based spectral clustering" | |
| status: seed | |
| # ============================================================================= | |
| # 3. STRUCTURAL WRAPPERS | |
| # ============================================================================= | |
| - id: diffuse | |
| name: diffuse | |
| decomposes_to: [diffusion_sde, langevin_dynamics] | |
| description: > | |
| Wrap a sampling/generative process in a forward noising + reverse denoising | |
| SDE. Forward: dx = f(x,t)dt + g(t)dW incrementally destroys structure. | |
| Reverse: dx = [f(x,t) - g(t)²∇_x log p_t(x)]dt + g(t)dW reconstructs. | |
| The core component is score matching: learn s_θ(x,t) ≈ ∇_x log p_t(x). | |
| The generative model = reverse-time SDE driven by learned score. | |
| input_constraints: {} # any data distribution | |
| output_signature: | |
| operation: sample | |
| meso_type: diffusion_process | |
| macro_type: stochastic_process | |
| preserves: false | |
| introduces: | |
| - "continuous-time stochastic process: Itô SDE" | |
| - "score function s(x,t) = ∇_x log p_t(x) as central object" | |
| - "denoising score matching: ||s_θ(x_t,t) - ∇ log p(x_t|x_0)||²" | |
| - "probability flow ODE for deterministic sampling (same marginals)" | |
| - "ancestral sampling via SDE discretization (Euler-Maruyama, etc.)" | |
| examples: | |
| - "DDPM = diffuse(Gaussian forward + learned reverse)" | |
| - "Score-based SDE = diffuse with VP/VE/sub-VP SDEs" | |
| - "Cold diffusion = diffuse with arbitrary degradation (not just Gaussian)" | |
| status: seed | |
| - id: contrastivize | |
| name: contrastivize | |
| decomposes_to: [infonce, contrastive_learning, mutual_info_max] | |
| description: > | |
| Convert a generative/similarity objective into a contrastive one: | |
| pull positive pairs together, push negative pairs apart. The core | |
| operation is maximize I(x; y) or its lower bound via noise-contrastive | |
| estimation. Any embedding method can be contrastivized by defining | |
| positive pairs (e.g., augmentations of same instance) and negative | |
| pairs (other instances). | |
| input_constraints: | |
| meso_types: | |
| - joint_embedding | |
| - linear_projection | |
| output_signature: | |
| meso_type: joint_embedding | |
| objective_family: information # mutual information lower bound | |
| preserves: false | |
| preserves: | |
| - "the encoder architecture" | |
| introduces: | |
| - "InfoNCE loss: -log(exp(sim(z_i, z_i^+)/τ) / Σ_j exp(sim(z_i, z_j)/τ))" | |
| - "temperature parameter τ controlling hardness" | |
| - "negative sampling strategy" | |
| - "uniformity + alignment decomposition (Wang & Isola 2020)" | |
| examples: | |
| - "SimCLR = contrastivize(ResNet encoder) with image augmentations" | |
| - "CLIP = contrastivize(dual encoder) with image-text pairs" | |
| - "SimSiam = contrastivize without negatives (stop-gradient trick)" | |
| status: seed | |
| - id: adversarize | |
| name: adversarize | |
| decomposes_to: [gan, wasserstein_distance, js_divergence] | |
| description: > | |
| Convert an optimization (typically generative) into a two-player | |
| minimax game: min_θ max_φ V(θ, φ). Player G (generator) minimizes; | |
| Player D (discriminator/critic) maximizes. At Nash equilibrium (if | |
| reached), G's distribution matches the target. This is functionally | |
| equivalent to minimizing a divergence (JS, Wasserstein, etc.) but | |
| the adversarial formulation replaces the explicit density with a | |
| learned critic. | |
| input_constraints: | |
| formalisms: [] # any generative model | |
| output_signature: | |
| meso_type: game_theoretic | |
| objective_family: adversarial | |
| preserves: false | |
| preserves: | |
| - "the generator architecture" | |
| introduces: | |
| - "minimax objective: min_G max_D V(D,G)" | |
| - "learned divergence: the critic D implicitly defines the loss" | |
| - "mode collapse risk (especially with JS-GAN)" | |
| - "training instability from non-stationary objectives" | |
| examples: | |
| - "GAN = adversarize(Gaussian latent generator + CNN discriminator)" | |
| - "WGAN = adversarize with Wasserstein critic + gradient penalty" | |
| - "Adversarial autoencoder = adversarize(AE latent regularizer)" | |
| status: seed | |
| - id: regularize | |
| name: regularize | |
| decomposes_to: [ridge_regression, lasso, elastic_net] | |
| description: > | |
| Add a penalty term to the objective: L_total = L_task + λ·R(θ). | |
| Common R: L2 (weight decay — Gaussian prior), L1 (sparsity — Laplace prior), | |
| dropout (stochastic regularization — approximate Bayesian model averaging), | |
| spectral norm (Lipschitz constraint). | |
| input_constraints: {} # any optimization | |
| output_signature: | |
| preserves: false | |
| preserves: | |
| - "the task objective L_task" | |
| - "the optimization algorithm" | |
| introduces: | |
| - "regularization penalty λ·R(θ)" | |
| - "bias-variance tradeoff via λ" | |
| - "Bayesian interpretation: R(θ) = -log p(θ)" | |
| examples: | |
| - "Weight decay = regularize(SGD, L2)" | |
| - "LASSO = regularize(OLS, L1)" | |
| - "Elastic Net = regularize(OLS, L1+L2)" | |
| - "Dropout = regularize with stochastic binary mask" | |
| status: seed | |
| # ============================================================================= | |
| # 4. OBJECTIVE TRANSFORMS | |
| # ============================================================================= | |
| - id: predict_in_codomain | |
| name: predict_in_codomain | |
| decomposes_to: [cca, kernel_cca] | |
| description: > | |
| Instead of predicting in the original data space, predict in a transformed | |
| (typically lower-dimensional or structured) latent space. The prediction | |
| target is not x_{t+1} but z_{t+1} = f(x_{t+1}) where f is an encoder. | |
| This is the core operation in JEPA and related architectures. It is | |
| equivalent to applying a projection before a prediction objective. | |
| input_constraints: {} # any predictive model | |
| output_signature: | |
| preserves: true | |
| preserves: | |
| - "the predictive architecture" | |
| - "the objective form (MSE, contrastive, etc.)" | |
| introduces: | |
| - "encoder f mapping to latent space" | |
| - "prediction target is f(x_target) not x_target" | |
| - "the encoder acts as a regularizer: irrelevant variation is projected out" | |
| - "connection to CCA: if encoders maximize correlation and predictor is linear" | |
| examples: | |
| - "JEPA = predict_in_codomain(joint_embedding) = CCA in latent space" | |
| - "BYOL = predict_in_codomain(contrastive without negatives)" | |
| - "World models / Dreamer: predict in latent dynamics space" | |
| status: seed | |
| - id: minimize_energy | |
| name: minimize_energy | |
| decomposes_to: [free_energy_min, hopfield_network, boltzmann_distribution] | |
| description: > | |
| Reframe the problem as energy minimization over a scalar field E(x). | |
| The solution is x* = argmin E(x). Training = shape the energy landscape | |
| so that desired configurations are low-energy and undesired ones are | |
| high-energy. This is the universal statistical physics framing of | |
| learning: any loss function is an energy function, and any optimizer | |
| is doing energy minimization. | |
| input_constraints: {} # any optimization | |
| output_signature: | |
| objective_family: energy | |
| meso_type: energy_model | |
| preserves: false # reframes the objective | |
| preserves: | |
| - "the optimal point x*" | |
| - "the gradient field ∇E(x) (visible in Langevin sampling)" | |
| introduces: | |
| - "energy landscape E(x) as central object" | |
| - "Boltzmann distribution: p(x) ∝ exp(-βE(x))" | |
| - "free energy: F = -β^{-1} log ∫ exp(-βE(x)) dx" | |
| - "sampling = Langevin dynamics on E(x)" | |
| examples: | |
| - "Hopfield network = minimize_energy(associative memory)" | |
| - "Energy-Based Models (EBM): directly parameterize E_θ(x)" | |
| - "Score-based models: s_θ(x) = -∇_x E_θ(x)" | |
| status: seed | |
| - id: variational_bound | |
| name: variational_bound | |
| decomposes_to: [elbo, variational_inference, kl_divergence_min] | |
| description: > | |
| Replace an intractable marginal likelihood log p(x) with a tractable | |
| lower bound (ELBO): log p(x) ≥ E_q[log p(x|z)] - KL(q(z|x) || p(z)). | |
| The bound is tight when q(z|x) = p(z|x). The gap is exactly KL(q||p(z|x)). | |
| This is the fundamental operation behind VAEs, variational inference, | |
| and any method that replaces exact inference with amortized inference. | |
| input_constraints: | |
| meso_types: | |
| - variational | |
| - probabilistic_inference | |
| output_signature: | |
| objective_family: divergence | |
| preserves: false | |
| preserves: | |
| - "the generative model p(x|z) and prior p(z)" | |
| introduces: | |
| - "inference network q(z|x) (amortized inference)" | |
| - "ELBO as surrogate objective" | |
| - "reparameterization trick for gradient estimation" | |
| - "KL gap: tightness depends on q's expressiveness" | |
| examples: | |
| - "VAE = variational_bound(encode_decode with stochastic latent)" | |
| - "IWAE: tighter bound with importance weighting" | |
| - "β-VAE: β-weighted KL term for disentanglement" | |
| status: seed | |